Fifteen systems · eighty adjudicated claims · one instrument that failed its own test
The map, and the method that audits it
How the world's systems couple, written so that every claim about them carries the verdict an independent panel gave it, and the evidence it does or does not rest on.
Most explanations of how the world works arrive with no method attached, which makes them indistinguishable from the confident commentary they were written to replace. This edition ships both. The map is the substance: fifteen systems, their couplings, six named transmission chains. The method is the annotation: every claim that touches a system carries the verdict three independent raters gave it, and a visible mark for whether its evidence can be retrieved at all.
The subject
A prophet with a scoreboard
In May 2024, in a Beijing classroom, a former high-school teacher named Jiang Xueqin made three predictions on camera. Trump would win in November. Trump would start a war with Iran. And the United States would lose it. He sketched the invasion scenario in detail, put the force requirement in the millions, and described what would happen to the troops already in theatre.
On 28 February 2026 the United States and Israel began the war. Within three days the video had earned him a hundred thousand subscribers, an epithet, and an audience that now numbers in the millions. By July he had left teaching and was touring: Tokyo, Los Angeles, New York, three-hour sessions in packed rooms.
The standard reaction is one of two things and both are useless. The first is that he is a prophet. The second is that a stopped clock is right twice a day. Neither is a finding. Both are ways of not doing the work.
The work is available, because he made it available. He treats his predictions as the test of his model and says so repeatedly. He names three events that would break it outright. He publishes his sources. He tells his audience which of his claims are intuition and which are documents. Whatever else is true of him, he has built something that can be scored, which is more than most of the people who dismiss him have done.
So it was scored. 80 claims, one at a time, against the record, by three raters blind to this project. What comes back is not a verdict on a man. It is a shape, and the shape is the reason this document exists.
He is far more reliable about what a document says than about why anything happened. Claims asserting a checkable predicate confirm at 58%. Claims asserting a conclusion built on one confirm at 12%. On causes, which are 49% of everything he says, the confirmed rate is 9%, or roughly right about one time in eleven.
Those two rates are the finding. The five-row ordering by subject that this project originally led with is reported further down, together with the reason it should not be read as a rank order.
Four reading paths run over the same source. The map is the world. The method is how any claim about it was judged. The failed test is what happened when the instrument was pointed at itself. Practice is the only one where you can be wrong.
The map
Fifteen systems, and what moves between them
Systems are arranged by layer: substrate at the top, coordination and decision at the bottom. An edge means one system transmits into another through a named mechanism, not that they are vaguely related. Select a system to light its couplings, see the mechanism on every edge, and read the adjudicated claims that touch it.
The tint is the method annotating the map. Each system is shaded by how often the claims filed against it were confirmed, on the resolved denominator. The substrate systems hold their colour. State capacity and geopolitics, which carries more claims than any other system, confirms at 5%.
And the map shows a coverage gap the rates cannot. The eighty claims touch only 9 of 15 systems. 6 are drawn hatched because no claim was filed against them at all: Water, Critical minerals, Physical networks, Debt, fiscal & state capacity, Public health & biosecurity and Climate & the energy transition. A confident account of how the world works said nothing measurable about any of them.
Coupling direction is drawn in neutral ink on purpose. Amber and blue are reserved for the distinction between what a source asserted and what the record shows, and are not reused for anything else.
148 coupling edges across 15 systems, each naming its mechanism, bundled toward a vertical spine so the graph resolves at rest. Bundling is geometry only: no coupling weight is invented, because the data does not carry one. The declared cross-cutting key cross_system is not a system and is not drawn. Node counts read confirmed over resolved; an asterisk marks a system with fewer than 3 resolved claims, where the rate is shown but should not be read as a measurement.
The map
Six chains worth knowing cold
A chain is a sequence of couplings with a mechanism at each step. Chains are falsifiable; vibes are not. Each is stated with the link that would break it. Highlight any chain on the map above.
The map
Three couplings people assert that do not hold
These are marked distinctly because asserting a coupling that is not there is the most common way a confident map goes wrong. The struck line is the assertion; beneath it is what the evidence supports instead.
The join
How the method annotates the map
Selecting a system above lists the claims filed against it with their verdicts, so the map is not a set of assertions but a set of adjudicated assertions. The join is enforced in the data: every claim carries a system field, and the validator refuses to build if one fails to resolve against the registry.
The single most useful thing to carry from the method into the map: claims that assert a checkable predicate confirm at 58%; claims that assert a conclusion built on one confirm at 12%. On this map, a mechanism is a predicate. An intent is a conclusion.
The subject
What licence the subject gives
A body of commentary is not a claim. It is a few hundred claims of five different kinds, delivered in one voice, and its truth value is not a single number. The first move is therefore mechanical: break the corpus into separable assertions, tag each by how it was obtained, check it, and rate it. 80 claims survived that filter as load-bearing, meaning the framework changes shape if they fail. Occult and metaphysical material was set aside rather than rated, on the principle that a document should refuse where it cannot judge.
The second move is to check whether the subject consents to being scored this way. He does, explicitly, naming three events that would refute his model without hedging any of them. All three remain open. None has triggered. The adjudication that follows is fair rather than hostile precisely because he set the terms.
That is not a small thing. Most commentators of his reach never state a condition under which they would be wrong, and the ones who do rarely pick conditions that could plausibly occur. It is also what makes the corpus usable as a test bed for the method, which is the other half of what this edition is for.
Confidence is a stylistic property, not an epistemic one. In this corpus a verbatim quotation from a published policy brief and an unsourced assertion about elite ritual arrive in the same voice, in the same paragraph rhythm, with the same certainty. 56 of 80 load-bearing claims, 70%, are conclusion-bearing rather than directly checkable. A reader who tracks tone instead of provenance will be wrong in both directions: credulous where the tone is confident, and dismissive where it is confident about something true.
The method
The finding, in its narrow form
The instinct when auditing a confident source is to sort claims by subject. That instinct is wrong. What predicts whether a claim survives checking is not what it is about, but whether it asserts something a reader could look up, or something inferred from that.
The result that survived every control
Does confirmation track what a claim is about, or how it is built?
- Source
- 80 adjudicated claims. Structure labelled by two classifiers blind to every verdict; verdicts by three raters blind to the structure labels.
- Encoding
- Bar length is the share of resolved claims rated CONFIRMED.
- Reading
- The two variables came from parties that could not see each other, so the gap is not one judgement correlating with itself.
Data behind this figure
| Claim structure | Resolved n | Confirmed % |
|---|---|---|
| Predicate only | 24 | 58% |
| Conclusion bearing | 49 | 12% |
The gap is roughly 46 points and holds across every rater in both experimental arms, including the arm that never saw this project's rubric. It is not a definitional trick: conclusion-bearing claims are not automatically unconfirmable, and several here are confirmed outright.
The subject
The case at full strength
A verdict that has not met the best version of the argument is not worth having, so here is the best version. Start with the thing he gets most right, which is also the thing least discussed about him: he is a competent macro-financial analyst. His account of the dollar system is standard economic history told accurately, from Bretton Woods through the closing of the gold window to the petrodollar, and the diagnosis he attaches to it is one many orthodox economists would recognise.
Then he does something more specific. He argues that the political neutrality reserve status depended on was destroyed by sanctions policy, and that the observable consequence would be a rotation out of Treasuries and into gold. That is a mechanism, it is checkable, and the reserve and central-bank-purchase data run with him.
The strongest single passage in the corpus is not a prediction at all. It is a transmission chain described in early August that required no coordinating actor to be right: Japan supplies marginal liquidity to the world through the carry trade and imports most of its energy from the Gulf, so a Gulf disruption raises Japan's import bill, weakens the yen, and eventually forces the Bank of Japan to recall liquidity from abroad to defend it, pulling support out from under the American asset prices carrying the economy. Every link is public. It is drawn on the map above as a named chain, with the step that would break it.
His most repeated rhetorical move is to tell his audience to stop speculating and go download the document, and he does it about a defence strategy, a 1996 strategy memo, and a policy brief on a Spanish enclave. In all three cases the document exists and says substantially what he says it says. One of the strongest confirmations in the ledger is a pillar of the published defence strategy that he reported word for word.
This is what the predicate rate is made of. Where the claim is "the document says X," he is reliable at 58%. It is the next sentence, the one that says what the document means, that fails.
The subject
The signature the failures share
The failures are not scattered. They have one shape, and naming it is more useful than counting them.
Where a structural, legal or bureaucratic cause is available, he reaches past it for a coordinating actor.
At the enclave the available cause was two rulings of the Spanish Supreme Court; he reached for Washington and Tel Aviv. On conscription the available cause was a registration provision in an appropriations bill; he reached for a war plan requiring bodies. On the defence strategy four pillars were on the page; he added a fifth describing an intention. On missile attrition the available cause was that the launchers were being destroyed; he reached for a deliberate feint.
That last one is instructive because it may be right. The record does not adjudicate between attrition and strategy, and both fit. What is diagnostic is not that he chose the strategic reading but that he chose it without noting the other, in a corpus where he is otherwise careful to enumerate possibilities.
A second, related signature runs through the collapse material. He imports a mechanism from a field he has read about, reproduces the framework accurately, and then leans his conclusion on the one component that field has specifically discarded. The systems-collapse framing is faithful to its source and is a real contribution to how his audience thinks about fragility. The agent he then hangs the forecast on is the part that source exists to demote.
The enclave session is the most instructive document in the corpus because it contains his best and worst work seventy-two hours apart. In the same three hours he surfaced a policy brief the Western press had almost entirely missed and quoted it accurately, and asserted an orchestration for which no evidence exists and a permanent territorial loss that was already reversing while he said it.
Both halves are true, and a ruling that drops either one is not a ruling, it is a preference. That is the whole reason this project rates 80 claims separately rather than rating a man.
The method
The descent, and what it depends on
The project's original headline was a monotone accuracy descent across five claim types. Under the independent panel that ordering survives, at 60 → 45 → 36 → 22 → 9. Survival is not the whole answer, because the question is whether the ordering is a property of the corpus or of the instrument reading it.
The five-row descent depends on the instrument that measures it
Does the accuracy ordering across claim types reproduce for raters who never saw the rubric?
- Source
- Same 80 claims under four adjudications: the three-rater rubric panel and the two-rater control arm, each split by claim structure.
- Encoding
- Line per adjudication. Solid, raters using the rubric. Dashed, the control arm mean.
- Reading
- Full-set monotonic in 3 of 3 rubric raters and 0 of 2 control raters. Predicate-only monotonic in 3 of 3 and 2 of 2.
Data behind this figure
| Claim type | All · rubric | All · control | Predicate-only · rubric | Predicate-only · control |
|---|---|---|---|---|
| Dates | 60 | 70 | 100 | 92 |
| Institutions | 45 | 50 | 80 | 80 |
| Documents | 36 | 64 | 50 | 75 |
| Quantities | 22 | 28 | 29 | 36 |
| Causes | 9 | 12 | 0 | 0 |
How precise are any of these numbers?
The argument above, and the control below it, are conducted on point estimates drawn from samples of nine to thirty-two. Before asking whether the ordering holds under one adjudication or another, it is worth asking whether it is resolved at all.
Every adjacent pair in the descent overlaps
The ordering is argued at length. Is it resolved between any two neighbouring rows?
- Source
- data/claims.json. 95% Wilson intervals on the same k and n the descent is computed from.
- Encoding
- Band is the interval; an open dot marks a sample under ten; the bracket is the overlap with the row above.
- Reading
- 4 of 4 adjacent pairs overlap. Only the extremes separate, and only because the causal row has n=32.
Data behind this figure
| Claim type | k / n | Point | 95% interval | Width |
|---|---|---|---|---|
| Dates | 6 / 10 | 60% | 31–83% | 52 pts |
| Institutions | 5 / 11 | 45% | 21–72% | 51 pts |
| Documents | 4 / 11 | 36% | 15–65% | 49 pts |
| Quantities | 2 / 9 | 22% | 6–55% | 48 pts |
| Causes | 3 / 32 | 9% | 3–24% | 21 pts |
It is not. All 4 adjacent pairs overlap at 95%. The rows are separated by fifteen to twenty-three points and their intervals are around fifty points wide. On this evidence no two neighbouring rows are distinguishable, and the five-row rank order is not something the data can carry.
The extremes do separate: dates at 31–83% against causal at 3–24%, and only because the causal row has n=32 where the others have nine to eleven. So the descent is not refuted. Its resolution is.
What the evidence does support at n=80
If the five-row ordering is not resolved, what is?
- Source
- The same ledger, partitioned by claim structure rather than subject matter.
- Encoding
- Same interval grammar. Both samples are large enough to draw a filled point.
- Reading
- These two intervals do not overlap, so the separation is real at 95%.
Data behind this figure
| Stratum | k / n | Point | 95% interval |
|---|---|---|---|
| Predicate-bearing | 14 / 24 | 58% | 39–76% |
| Conclusion-bearing | 6 / 49 | 12% | 6–24% |
The narrower claim is the one the evidence supports, and it is harder to attack. Claims asserting a checkable predicate confirm at 58% [39–76]; claims asserting a conclusion built on one at 12% [6–24]. Those intervals do not touch. This is a two-level result at n=73, not a five-row ranking at n=9.
What it would take to fix. Halving the interval on the quantity row needs roughly n=44 quantity claims against the 9 it has; the dates row needs about n=51. That converts "we need more claims" from an aspiration into a number, and it is why the second-corpus test in the limitations is the binding constraint rather than more adjudication of these eighty.
The five-row descent is rubric-conditional. Monotonic for 3 of 3 raters using the rubric and 0 of 2 without it. Reported as a headline about the corpus, it overstates what the evidence carries.
What reproduces without the rubric is the descent inside predicate-only claims, 100 → 80 → 50 → 29 → 0, monotonic for 2 of 2 control raters as well as 3 of 3 rubric raters. Its causal cell holds only two claims; the bottom of that gradient is the least established part of it.
Two thousand random permutations of the ledger's own verdicts across the same claims produced a monotonic descent 0.7% of the time, so the observed ordering is not a chance artifact.
The control that came from outside
Two reviewers, working independently and without access to each other, arrived at the same objection: claim type and claim structure may be nearly the same variable. If so, a rubric whose decisive test is predicate or conclusion would demote the causal row by construction, and the gradient would be an artifact of the instrument rather than a property of the source.
The confound is real and measurable. 37 of 39 causal claims are conclusion-bearing (95%), against 2 of 9 quantity claims (22%). So the test runs the control: hold structure constant and recompute the descent within each stratum.
| Stratum | Descent | Ordering | Smallest cell |
|---|---|---|---|
| All claims | 60 → 45 → 36 → 22 → 9 | holds | 9 claims |
| Predicate-only | 100 → 80 → 50 → 29 → 0 | holds | 2 claims |
| Conclusion-bearing | 0 → 17 → 29 → 0 → 10 | breaks | 2 claims |
The ordering does not survive the control. Among conclusion-bearing claims alone it is not monotonic: causal scores 10%, above dates at 0% and quantities at 0%. Among predicate-only claims it holds cleanly. The reviewers' diagnosis is correct, and the honest reading is neither that the descent is real nor that it is an artifact: roughly half its gradient is structural rather than a property of the source.
One caveat neither reviewer stated, and it cuts against the strength of the refutation. The conclusion-bearing strata are tiny. Quantity holds 2 resolved claims and dates 4. A break driven by a 0% computed on 2 claims is not evidence that the ordering fails; it is evidence that this stratum cannot decide either way. The correct verdict on the control is underdetermined, not refuted, and it stays that way until the corpus is larger.
A third reviewer claim, that the retention from the pre-rubric figures is monotonic in the same direction as the descent and therefore signals a manufactured gradient, does not hold against the shipped ledger. It was computed against the author's superseded pass. Against the panel ledger the retention runs 67 / 67 / 56 / 50 / 64, which is not monotonic, and the causal row retains more than quantities or documents.
How the eighty claims resolved
What did the independent panel actually find?
- Source
- Majority consensus of three raters blind to this project.
- Encoding
- Bar length is share of the rated corpus.
- Reading
- CORE_INFERENCE, meaning predicate verified and conclusion not, is among the largest categories.
Data behind this figure
| Verdict | Claims | Share |
|---|---|---|
| CONFIRMED | 20 | 25.0% |
| CONFIRMED IN PART | 6 | 7.5% |
| CORE INFERENCE | 17 | 21.2% |
| PARTIAL | 2 | 2.5% |
| UNRESOLVED | 5 | 6.2% |
| NOT IN EVIDENCE | 3 | 3.8% |
| CONTRADICTED TO DATE | 6 | 7.5% |
| CONTRADICTED | 19 | 23.8% |
| ANTECEDENT FAILED | 2 | 2.5% |
The method
The ledger, as an instrument
All 80 claims as the independent panel rated them. Filter, sort, and select any claim to open its record, its sources, the rubric clause that decided it, and the three raters' individual votes. The headline agreement figure is 95.8%, which is a summary; the ledger shows where it actually broke. The panel split on 5 of the 80 claims it voted on, and on those the shipped verdict is a majority rather than a consensus. The provenance dimmer fades each claim in proportion to how well its evidence is documented: drag it and watch most of the corpus disappear.
Drag right and the evidence fades in proportion to how well it is documented, while every claim stays at full strength. At the far end only the 15 verdicts backed by a retrievable citation are still legible; the other 65 claims remain, with nothing visible behind them. That asymmetry is the project's central limitation, rendered as something you can operate rather than a sentence you have to take on trust.
| ID | Claim | Verdict | Type | Structure | Source | Panel |
|---|---|---|---|---|---|---|
| IRAN-WAR-START-DATE | The war began 28 February 2026 | CONFIRMED | Dates | predicate | 1 src | |
| IRAN-CEASEFIRE-COLLAPSE | The ceasefire collapsed over Hormuz | CORE INFERENCE | Dates | conclusion | 1 src | |
| VENEZUELA-PROOF-OF-CONCEPT | Venezuela was Trump's proof of concept | CORE INFERENCE | Dates | conclusion | none | |
| USCHINA-GRAND-BARGAIN | A US–China grand bargain is coming; Trump and Xi meet mid-May in Beijing | CONFIRMED IN PART | Dates | conclusion | 2+ src | |
| NDS-EXISTS-PUBLIC | A public National Defense Strategy exists and states the plan | CONFIRMED | Dates | predicate | none | |
| CHOKEPOINT-TREATIES-EXIST | US–Indonesia (13 Apr 2026) and US–Morocco (Apr 2026) agreements exist | CONFIRMED | Dates | predicate | none | |
| CLEAN-BREAK-DOCUMENT | The Clean Break memo, 1996, Perle and Feith, for Netanyahu, called for breaking Oslo and removing Saddam | CONFIRMED | Dates | predicate | none | |
| DRAFT-AUTO-REGISTRATION | Automatic draft registration begins in December | CONFIRMED | Dates | predicate | none | |
| ISRAEL-OCTOBER-ELECTION | Israeli elections in October will destabilize Netanyahu | CONFIRMED IN PART | Dates | conclusion | 1 src | |
| SHADOW-FLEET-SEIZURES | Russian-linked shadow-fleet tankers are being seized | CONFIRMED | Dates | predicate | none | |
| IRAN-HORMUZ-CLOSED | Iran closed the Strait of Hormuz | CONFIRMED | Institutions | predicate | 1 src | |
| HORMUZ-INSURANCE-MECHANISM | Hormuz was closed by insurance withdrawal, not mines | CONTRADICTED | Institutions | conclusion | 1 src | |
| DOLLAR-BW-TO-PETRODOLLAR | Bretton Woods 1944 → Nixon 1971 → petrodollar + China | CONFIRMED | Institutions | predicate | none | |
| DOLLAR-SANCTIONS-GOLD-ROTATION | Sanctioning Russia broke dollar neutrality and drove a rotation into gold | CORE INFERENCE | Institutions | conclusion | none | |
| EURODOLLAR-UNCOUNTABLE | The Eurodollar stock cannot be counted | CONFIRMED | Institutions | predicate | 1 src | |
| JAPAN-LIQUIDITY-NODE | “Japan is the world's liquidity node via the yen carry trade, exposed through Gulf energy | CORE INFERENCE | Institutions | conclusion | none | |
| YEN-US-INTERVENTION | The US has intervened to save a foreign currency, a sign of desperation | CONFIRMED IN PART | Institutions | conclusion | 1 src | |
| CEUTA-MOROCCO-PRECEDENT | Morocco has weaponized migration before | CONFIRMED | Institutions | predicate | none | |
| CEUTA-ITALY-SCHENGEN | Italy reimposed border controls, attacking free movement | CONTRADICTED | Institutions | conclusion | none | |
| CAPITAL-DEFINITIONAL-SPLIT | Capitalism and transnational capitalism are different things | CONFIRMED | Institutions | conclusion | none | |
| MILEI-CHABAD-TIES | Milei is Chabad-Lubavitch; first foreign visit to Schneerson's grave; Netanyahu backed Argentina | CONFIRMED IN PART | Institutions | predicate | 1 src | |
| NDS-PILLAR-INDUSTRIAL-BASE | “Its pillars include "supercharge the US defense industrial base" | CONFIRMED | Documents | predicate | none | |
| NSS-DONROE-COROLLARY | “A "Donroe Doctrine" Trump's corollary to Monroe | CONFIRMED | Documents | conclusion | 1 src | |
| NDS-KEY-TERRAIN | Greenland, Panama and the Arctic are targets | CORE INFERENCE | Documents | predicate | none | |
| CEUTA-AEI-BRIEF | An AEI brief laid out the method: bulldozers, unarmed entry, raise the flag | CONFIRMED | Documents | predicate | 1 src | |
| CAPITAL-BOE-1694 | Bank of England 1694 and the Glorious Revolution created modern finance | CORE INFERENCE | Documents | conclusion | none | |
| CAPITAL-WEBER-THESIS | Calvinist double predestination created capitalism | CORE INFERENCE | Documents | conclusion | none | |
| BRONZE-PERFECT-STORM | The Bronze Age collapse was a perfect storm, not one cause | CONFIRMED | Documents | conclusion | none | |
| BRONZE-INTERCONNECTION-FRAGILITY | Interconnection produces fragility, not resilience | CORE INFERENCE | Documents | conclusion | none | |
| CALHOUN-RESEARCHER-IDENTITY | A researcher ("James Poon") ran rat-utopia experiments that ended in extinction | CONFIRMED IN PART | Documents | predicate | none | |
| THIRD-TERM-VP-ROUTE-LEGAL | The vice-presidential route is legally available | CONTRADICTED | Documents | conclusion | none | |
| FIFA-RIGGED | “FIFA rigged the tournament for market appeal | CORE INFERENCE | Documents | conclusion | none | |
| HORMUZ-OIL-SHARE | The Gulf exports ~20% of the world's energy through it | CONFIRMED | Quantities | predicate | none | |
| IRAN-US-KIA-COUNT | “"Only 13 died in those six weeks, a very suspicious number" | CONFIRMED IN PART | Quantities | conclusion | 1 src | |
| HORMUZ-TOLL-PRICE | Iran charges ~$2 m per crossing | PARTIAL | Quantities | predicate | none | |
| IRAN-TROOPS-STAGED | 60,000 US soldiers are staged for invasion | CONTRADICTED | Quantities | predicate | none | |
| AI-CAPEX-GDP-GROWTH | ~50% of US GDP growth is AI data centres | PARTIAL | Quantities | predicate | none | |
| CEUTA-CROSSING-COUNT | ~50–60,000 crossed into Ceuta | CONFIRMED | Quantities | predicate | none | |
| CEUTA-ENCLAVE-POPULATION | The enclave has 80,000 residents | CONTRADICTED | Quantities | predicate | none | |
| FERTILIZER-CARRYING-CAPACITY | “Without fertilizer the world sustains only two billion | CONTRADICTED | Quantities | conclusion | none | |
| CALHOUN-TERMINAL-STATE | The colony ended in warfare | CONTRADICTED | Quantities | predicate | none | |
| IRAN-CHINA-FORCED-TRUCE | China forced the ceasefire, using Pakistan as mediator | CORE INFERENCE | Causes | conclusion | none | |
| IRAN-GROUND-INVASION-PLAN | “A full-scale ground invasion is the plan | CONTRADICTED | Causes | conclusion | none | |
| IRAN-KURD-BALOCH-PROXY | Forward operating bases will arm Kurds and Baloch as cannon fodder | CONTRADICTED TO DATE | Causes | conclusion | none | |
| IRAN-STRAIT-IS-OBJECTIVE | “Securing the strait is the real objective | CORE INFERENCE | Causes | conclusion | none | |
| IRAN-KHARG-WOULD-FAIL | Seizing Kharg Island would fail | CONFIRMED | Causes | conclusion | none | |
| RUSSIA-MUST-ENTER | Russia must enter the war on Iran's side | CONTRADICTED TO DATE | Causes | conclusion | 1 src | |
| CHINA-BRI-REINFORCE | China reinforces Iran via Belt and Road rail | CONTRADICTED | Causes | predicate | none | |
| CHINA-DEMAND-EXPOSURE | “China is existentially exposed to demand collapse | CONTRADICTED | Causes | conclusion | none | |
| CHINA-NO-TAIWAN-INVASION | China will not invade Taiwan | UNRESOLVED | Causes | conclusion | none | |
| WW3-PROBABILITY | 80–90% probability this becomes World War III | CONTRADICTED TO DATE | Causes | conclusion | none | |
| RUSSIA-TAKES-ODESSA | Russia takes Odessa and the Ukraine war ends | CONTRADICTED TO DATE | Causes | conclusion | 1 src | |
| DPRK-SECOND-FRONT | North Korea opens a second front | CONTRADICTED | Causes | conclusion | none | |
| AI-BUBBLE-ENGINEERED | The AI bubble is engineered and will be popped on a timed schedule | UNRESOLVED | Causes | conclusion | none | |
| NDS-STRANGLE-CHINA | “Pillar 2 means strangling China at Malacca | CONTRADICTED | Causes | conclusion | none | |
| NDS-FIFTH-PILLAR | A fifth pillar: manufacture three wars to offload $40 tn of debt | CONTRADICTED | Causes | conclusion | none | |
| CHOKEPOINT-TREATIES-PURPOSE | Their purpose is chokepoint control | CORE INFERENCE | Causes | conclusion | none | |
| CEUTA-SPAIN-AIRBASE-MOTIVE | “Spain refused the US its airbases and is being punished | CORE INFERENCE | Causes | conclusion | none | |
| CEUTA-POLICE-BUSSED | Moroccan police loaded migrants onto government trucks and directed them across | CONTRADICTED | Causes | predicate | none | |
| CEUTA-US-ISRAEL-AUTHOR | “The US and Israel orchestrated the surge | CORE INFERENCE | Causes | conclusion | none | |
| CEUTA-COORDINATING-AUTHOR | (by omission) The surge had a coordinating author | CONTRADICTED | Causes | conclusion | none | |
| CEUTA-IRREVERSIBLE | Ceuta is lost; there is nothing Europe can do | CONTRADICTED | Causes | conclusion | none | |
| CAPITAL-RELOCATING-ISRAEL | Capital is relocating from the US to Israel, with Argentina as fallback | NOT IN EVIDENCE | Causes | conclusion | none | |
| CAPITAL-CONVERGENCE-CHAINS | Points-of-convergence chains (Schneerson, Harvard Chabad, Adelson) | CONFIRMED | Causes | conclusion | none | |
| CLEAN-BREAK-CAUSED-IRAQ | “Clean Break causally produced the Iraq war and current Middle East policy | CORE INFERENCE | Causes | conclusion | none | |
| WAR-IS-ESCHATOLOGICAL | The Iran war is eschatological, not material | CORE INFERENCE | Causes | conclusion | none | |
| IRAN-2TN-COUNTERFACTUAL | $2 tn to Iran plus containing Israel would end the war | ANTECEDENT FAILED | Causes | conclusion | none | |
| ANTISEMITISM-ENGINEERED | Antisemitism is being deliberately engendered to force diaspora return | NOT IN EVIDENCE | Causes | conclusion | none | |
| BRONZE-SEA-PEOPLES-AGENT | The Sea Peoples overwhelmed civilizations, and refugees will do it again | CONTRADICTED | Causes | conclusion | none | |
| FERTILIZER-UNDERMODELLED | The fertilizer channel is dangerously under-modelled | CONFIRMED | Causes | conclusion | none | |
| FERTILIZER-GULF-SUBSTITUTION | Losing Gulf exports ≈ losing synthetic fertilizer | CORE INFERENCE | Causes | conclusion | none | |
| CALHOUN-HUMAN-GENERALIZATION | “Therefore young men will form groups and produce urban warfare | CONTRADICTED | Causes | conclusion | none | |
| DRAFT-OBLIGED-TO-FIGHT | “Therefore young men are obliged to go and fight | CONTRADICTED | Causes | conclusion | none | |
| THIRD-TERM-OBTAINED | Trump obtains a third term | UNRESOLVED | Causes | conclusion | none | |
| VANCE-DISCARDED | Vance will be discarded well before 2028 | CONTRADICTED TO DATE | Causes | conclusion | none | |
| MIDTERMS-DISRUPTED | The midterms will be disrupted, possibly not convened | UNRESOLVED | Causes | conclusion | none | |
| CIVIL-WAR-PROBABLE | Civil war is more probable than renaissance | NOT IN EVIDENCE | Causes | conclusion | none | |
| WORLDCUP-MESSI-CONDITIONAL | Conditional: if Argentina wins with visible rigging and Messi scores → Messi becomes president | ANTECEDENT FAILED | Causes | conclusion | none | |
| JAPAN-MASS-EUTHANASIA | Japan introduces mass euthanasia within ten years | UNRESOLVED | Causes | conclusion | none | |
| ARGENTINA-FALKLANDS | Argentina reclaims the Falklands within a year | CONTRADICTED TO DATE | Causes | conclusion | 1 src |
The method
What the evidence actually rests on
The evidence base was asserted, and almost never documented
How many claims carry a citation a reader could actually open?
- Source
- data/claims.json, comparing the original sweep's declared depth against sources verified and attached in the audit pass.
- Encoding
- Paired bars. Amber is asserted, blue is documented.
- Reading
- The original ledger recorded a depth for all 80 claims and a URL for none. 15 were documented in the audit pass.
Data behind this figure
| Verification depth | Asserted | Documented |
|---|---|---|
| Two or more sources | 16 | 1 |
| One source | 53 | 14 |
| No source | 11 | 65 |
65 of 80 claims carry no retrievable citation. 18 assert a source tier with no source recorded, which is unfalsifiable by construction.
Every rater in every arm adjudicated against these evidence notes, so a wrong note produces a wrong verdict in all arms at once and no amount of inter-rater agreement can detect it. That effect was measured, not estimated: an independent audit found 9 notes non-responsive to their own claims, and repairing them moved three verdicts with the claim and the rubric held constant.
Quotations, which nothing was checking
Rules 4 and 10 police the numbers and the verdicts. Until this edition nothing policed a quotation. A direct quotation attributed to a named living person is an assertion about the world of exactly the same kind as a percentage, and the build would ship one with nothing behind it and report no failures. Rule 14 now registers every quotation this project ships and resolves each one to its evidence.
13 of 15 registered quotations resolve to no source at all. They are named in the ledger below with an open-outline quotation mark, and they are disclosed rather than deleted, because a quotation with nothing behind it is a finding about this evidence layer and not an embarrassment to be tidied away. Filter the ledger to quoted, unsourced to read them.
The rule was written against a real case. The sibling edition of this project contains a three-word quotation of assent, attributed by name, that appears in exactly one file across the entire corpus: it is in no claim record, in neither manuscript, and in no transcript, because this project ships no transcripts. Every check that existed before rule 14 passes on that sentence.
What rule 14 cannot catch, stated here rather than discovered later: a quotation that is registered and sourced but misquoted, since the rule checks that a source exists and never that the source says it; paraphrase, which carries the same weight and no quotation marks; a fabricated source, since rule 12 checks that a source entry is well formed and not that the publication exists; and quotations under its word floor, excluded so the rule does not fire on ordinary terms of art. The validator panel lists these alongside its other blind spots.
What you keep
The single most important move
If nothing else here survives six months, let it be this. Split the claim from the register it arrives in. Then split the predicate from the inference.
Almost everything that goes wrong in reading a source like this, in either direction, comes from treating each session as one object with one truth value. It is not. It is 80 separable claims of five evidentiary kinds and they do not stand or fall together. The strategy memo is real whether or not the mystical reading is. The published defence strategy says what he says it says whether or not there is a fifth pillar. The registration statute exists whether or not conscription follows.
A reader who performs the split gets an early-warning system with a known failure mode, which is worth considerably more than a source with no failure mode you have identified. A reader who does not gets either a prophet or a crank, and both are unusable, not because the judgement is wrong but because neither tells you what to do on Tuesday.
The move costs about four seconds a claim. How was this obtained? Then, if the answer is good: does the conclusion he draws follow from it? Those two questions, in that order, are the entire apparatus of this document, and the practice path lets you run them yourself against claims the panel has already ruled on.
The apparatus is portable. The verdict is not. Nothing here establishes that confident commentators in general descend this way, because the protocol has been run on exactly one corpus. That limitation is stated at full strength in the failed-test path.
The failed test
What happened when the instrument was pointed at itself
This project specified an inter-rater test that could destroy its own headline, ran it, and failed: a rater working blind diverged on 9 of 19 claims against a declared threshold of 4, and every divergence ran the same direction. This is not an appendix. It is the most defensible thing the project owns.
The six stages
One, the test. Twenty claims drawn by a stated deterministic hash. A second rater ruled from claim text alone, with the record, verdict, tags and type withheld. Sample, counting rules, contamination and the conclusion at each outcome band were written down before any claim was looked up.
Two, the missing instrument. Seven of nine divergences sat inside the CONFIRMED / CONFIRMED_IN_PART / PARTIAL / CORE_INFERENCE cluster. Eleven verdicts existed with no written procedure for choosing among them. The diagnostic was in the data: the verdict meaning predicate verified, conclusion not appeared 7 times in one claim type and 0 times in the other four.
Three, a prediction, and its refutation.
A prediction registered in advance, and refuted
The audit predicted the descent would flatten and its ordering break. What happened?
- Source
- READJUDICATION-PREREGISTRATION.md, committed before the rubric existed and never edited, against the author's own re-adjudication (60 / 40 / 36 / 22 / 3), which the independent panel later superseded. Both series computed from data/claims.json.
- Encoding
- Dashed rule is the predicted value; dot is the value observed in the author's pass, not the shipped ledger.
- Reading
- The ordering was predicted to break. It held, and the gradient steepened from a 6.4x spread to 20x.
Data behind this figure
| Claim type | Predicted % | Observed % | Error |
|---|---|---|---|
| Dates | 80% | 60% | 20 |
| Institutions | 30% | 40% | 10 |
| Documents | 35% | 36% | 1 |
| Quantities | 40% | 22% | 18 |
| Causes | 11% | 3% | 8 |
The audit predicted the ordering would break and the descent flatten. It was refuted. The ordering held and the gradient steepened, from a 6.4x spread to 20x. Three of five predicted values were missed, worst of all quantities, predicted the most stable row and in fact losing half its confirmations.
Four, removing the author. Three raters blind to the volume, the preregistration and the author's rulings adjudicated all eighty claims against the rubric. Their pairwise agreement was 95.8%, and majority consensus resolved every claim with no ties. Author verdicts remaining in the shipped ledger: 0.
Five, the records underneath the verdicts. An independent audit of evidence-note quality, blind to all verdicts, found 9 records non-responsive to their own claims. The clearest case is the project's most-repeated boast: all three raters downgraded the claim that the war began on a stated date, taking the perfect date score with it, and external verification then showed the claim was right and the record was wrong.
Six, what survives. The narrow finding, the structure-controlled gradient, and the result below.
The failed test
Did the repair work?
The obvious way to answer is to put Test B next to the panel: a rater working blind agreed 53% of the time, three raters with the rubric are unanimous on 75 of 80 claims. Stated that way the rubric looks transformative. Stated that way it is also a confounded comparison, and this project exists to catch those.
Did the rubric raise agreement? Less than it looks
The obvious before-and-after is Test B against the panel. What does it actually compare?
- Source
- Test B result, the two-rater control arm, and the three-rater panel, all recomputed from the recorded verdicts.
- Encoding
- Same interval grammar throughout. Agreement is pairwise for two-rater conditions.
- Reading
- Only rows two and three hold the record constant. Between them the rubric is worth +4 points, and those two intervals overlap.
Data behind this figure
| Condition | What differs | k / n | Agreement | 95% interval |
|---|---|---|---|---|
| Test B | 1 rater, no record, no rubric | 10 / 19 | 53% | 32–73% |
| Control arm | 2 raters, record, no rubric | 74 / 80 | 92% | 85–97% |
| Panel | 3 raters, record, rubric | 77 / 80 | 96% | 90–99% |
Three things differ between Test B and the panel, not one: the rubric, the availability of the record, and the number of raters. The control arm is the row that isolates the rubric, because it holds the record constant and removes only the rubric. Between control and panel the rubric is worth +4 points, from 92% [85–97] to 96% [90–99], and those two intervals overlap.
Most of Test B's low agreement came from that rater not having the record, which is a finding about the evidence base rather than about the rubric.
The rubric still earned its place, on a different measurement: raters with and without it agree on only 70.6% of claims, so it changes the output substantially even where it does not measurably tighten agreement. Doing work and raising agreement are different claims, and only the first is established.
Where the panel actually split
75 of 80 claims were unanimous and 5 were not. Those 5 are the most informative rows in the ledger, because they name which parts of the verdict vocabulary are still unstable.
Where the panel actually split
Which verdict boundaries are still unstable?
- Source
- The 5 claims on which at least one of the three raters dissented.
- Encoding
- Lit cells are boundaries where a dissent occurred; blank cells were never confused.
- Reading
- 3 of 5 dissents involve NOT_IN_EVIDENCE, so the unstable boundary is whether evidence exists at all, not how good it is.
Data behind this figure
| Boundary | Dissents |
|---|---|
| Confirmed ↔ Not In Evidence | 2 |
| Contradicted To Date ↔ Not In Evidence | 1 |
| Confirmed In Part ↔ Core Inference | 1 |
| Contradicted ↔ Partial | 1 |
3 of 5 dissents involve NOT_IN_EVIDENCE. The unstable boundary is not how good is the evidence, which is what the rubric was written to fix; it is is there any evidence at all. That is a different and more tractable problem, and it is the lead for the next rubric revision. Filter the ledger to panel split to read all 5.
The failed test
The one result no adjudication dispute can touch
45% of the forecasts were built so they could never have been wrong
Before consulting any outcome, how many forecast statements can be scored at all?
- Source
- data/predictions.json. Admissibility decided on declared criteria before outcomes were consulted.
- Encoding
- Segment width is statement count. Grey segments cannot be scored.
- Reading
- Measurable today, and it involves no verdict, so nothing in the adjudication dispute touches it.
Data behind this figure
| Category | Statements |
|---|---|
| no criterion | 4 |
| undated | 4 |
| restates public record | 1 |
| statutory date | 1 |
| admissible | 12 |
| of which resolved | 3 |
| of which hits | 3 |
| void by antecedent | 1 |
| still open | 8 |
| total statements | 22 |
10 of 22 forecast statements (45%) cannot be scored at all: undated, criterion-free, or restatements of a published schedule. A fact about how the forecasts are built, available without waiting for anything to resolve.
Of the 12 admissible, 3 resolved and all 3 were hits. That is not a skill estimate. A sample with no misses cannot distinguish a calibrated forecaster from an overconfident one, so the honest verdict on the prediction record is insufficient data.
The failed test
What is validated, and what is not
| Result | Status | Established by |
|---|---|---|
| Verdicts on all 80 claims | Independent | Three raters blind to this project, the preregistration and the author's rulings. 95.8% pairwise agreement; majority resolved every claim with no ties. |
| Claim-structure labels | Independent | Two classifiers blind to every verdict, unanimous on all 80. |
| Record responsiveness | Independent | One auditor blind to every verdict. Found 9 evidence notes non-responsive to their own claims. |
| That the rubric changes outcomes | Independent | Control arm. Raters with and without the rubric agree on only 70.6% of claims. |
| 45% of forecasts inadmissible | No verdict involved | Computed from the admissibility protocol with criteria declared before outcomes. |
| The rubric itself | Provisional | Written by the audit author; the panel executes it. Panel-to-author agreement is 78%, leaving 22% divergence, above this project's own 20% withdrawal threshold. |
| The validator | Internal | 123 checks; 17 of 17 ordinary mutations caught, including three that plant an unsourced quotation; 4 of 9 adversarial mutations still uncaught and documented. |
| The claim texts | Unvalidated | Whether these are the right 80 claims, accurately extracted, is untested. Needs the raw transcripts. |
| 65 of 80 records | Unvalidated | No retrievable citation. |
| Independence beyond one model family | Unvalidated | Every rater, classifier and auditor is an instance of the same family as the author. |
The 22% panel-to-author divergence is why every rate in this edition is labelled provisional rather than settled. By this project's own declared threshold, that figure is a fail.
Two gaps still open, stated rather than skipped
- The extraction audit has not run. It needs the raw L0 transcripts, which are not in this repository. Without it nobody can say whether these are the right 80 claims or whether any claim text is a misquote. It is also the only fix for the two deepest validator blind spots, a falsified record and a falsified claim, both of which pass a green build today.
- No rater from outside this model family has been on the panel. Every rater, classifier and auditor here is an instance of the same family as the author. Same-family agreement cannot distinguish a correct adjudication from a shared prior. This is the single change that would let the rates move from provisional to settled.
The most defensible thing this project owns is not a number about one commentator. It is a demonstration that an analytical instrument can state its own failure conditions, be run against them, fail, and be rebuilt in public, and that what survives is smaller, stranger and better than what went in.
- Adjudication
- Majority of three independent raters (R1,R2,R3) applying VERDICT-RUBRIC.md, blind to the brief, to the author's Test B rulings and to the author's re-adjudication. CONTESTED = no majority; excluded from all rates.
- Structure labels
- claim_structure was assigned by two independent classifiers blind to every verdict, working from claim text alone. They agreed on all 80. It is used to control the predicate/conclusion confound in Vol I 20.6 and is never an input to any verdict.
- Provenance
- verification_depth_asserted is the original adjudicator's statement of how many sources were consulted; no URL was ever recorded for any of them. verification_depth is now computed from sources[] and means DOCUMENTED. The gap between the two is the size of the unrecorded evidentiary base.
- Map
- Systems, layers and couplings from data/systems.json and data/linkages.json. Chains and asserted non-couplings derived from VOLUME-II.md by build/derive_map.py, not hand-authored.
- Encoding
- Amber is the source's own assertion; blue is the record. Solid marks are verified; open outlines are the source's, drawn at full size and never faded. The verdict ramp is ordinal and is used for nothing else. Inherited from The Prophet and the Record, which is a design reference only: its numbers are superseded by this ledger.
- Generated
- build/html/build_consolidated.py from data/claims.json, predictions.json, systems.json, linkages.json, indicators.json and chains.json. No figure or count in this page is typed by hand.
Practice
The only path where you can be wrong
Three exercises. The first makes you a rater; the second lets you change the adjudication rule and watch the headline move; the third shows what the rubric actually did. None of them tell you anything the rest of the edition does not: they make you do it instead of read it, which is the difference between knowing that adjudication is subjective and finding out.
1 · Adjudicate it yourself
A claim and the record behind it, with both stored verdicts hidden. Rule on it, then compare against the original adjudicator and against the post-rubric ledger. Your own disagreement rate is the same measurement the project ran on itself and failed.
2 · Rule lab, the descent under a rule you choose
Accuracy is confirmed ÷ resolved, and every term in that is a choice. Change the choices and the five-row descent recomputes from the ledger. The gradient below is the ratio of the top row to the bottom, the project's headline claim is that it is steep and ordered.
3 · What the rubric moved
Every claim whose verdict changed when the rubric was applied, by claim type.
Seventeen claims became CORE_INFERENCE, the verdict for a verified predicate carrying
an unsupported conclusion. Where they came from is the whole story of the re-adjudication.