Three runs, one engine
All three runs were executed by the same engine, the run headers carry the same code hash (c5539ae8bf8ff7f6…), verified at run time, so differences in outcome are differences in the evidence, not the machinery.
| Run | Mode | Question | Claims | Verdicts | Direction |
|---|---|---|---|---|---|
| Run 0 | BACKTEST | Bank of Canada policy path, cutoff Jan 26 2023 | 16 | 10 held · 3 narrowed · 1 killed | UNDETERMINED @ 0.5 |
| Run B | LIVE | The 2026 maturity-wall coupon step-up thesis | 48 | 27 held · 13 narrowed · 4 killed | REFUTED @ 0.2 |
| Run C | LIVE | Goodwill impairment: “too little, too late” | 48 | 21 held · 19 narrowed · 7 killed | REFUTED @ 0.3 |
Why the backtest exists. Run 0 replays a question whose answer is now known, the Bank of Canada’s policy path, but locks the evidence at a cutoff of January 26, 2023, and audits only what was knowable then. It returned UNDETERMINED at 0.50: the record available on the cutoff date did not decide the question, and the engine said so instead of pretending otherwise. An engine that leaked hindsight would have scored confidently. That is the point of the run: it is the leak test the two live runs stand on.
The live runs. Run B audits the widely-circulated 2026 maturity-wall thesis for Canadian investment-grade issuers and refutes its central statistic at 0.20: every construction of the issuer-level median coupon step-up the retrievable record supports lands below the threshold the thesis needs. Run C audits the “too little, too late” critique of impairment-only goodwill accounting against TSX Composite issuers and comes back REFUTED at 0.30 on the claim as filed.
The vocabulary, once
| Term | Meaning |
|---|---|
| E0 – E3 | Evidence tiers, best to worst: E0 is a primary dated record; E3 is an uncited assertion. A conclusion inherits the tier of its weakest load-bearing claim. |
| Falsifier | Filed with every claim at entry: the observation that would kill it. A claim without one is not admitted. |
| HELD / NARROWED / KILLED | A claim’s fate. Narrowed claims survive only in a weaker form, shown beside the text as filed. Killed claims move to the graveyard and stay visible. |
| Attack vectors | Provenance (is the source real and dated), inference (does the claim follow), selection (what was left out), anachronism (was it knowable at the cutoff). |
| Tripwire | A pre-declared halt condition: if too many load-bearing claims fall, the run stops and says so rather than concluding. |
| BACKTEST / LIVE | Backtest runs lock evidence at a historical cutoff to test the engine; live runs audit a present-day thesis. |
Each run page carries the complete record: the claim matrix with every source, the graveyard, the crux that stayed contested, the attacks that scored zero, each agent’s concessions, the confidence arithmetic, and the unabridged round-by-round transcript on its own page. Nothing is summarised away.
Method kinship. The same discipline (label the evidence, file the falsifier, let the attack run) drives Flagged in Hindsight, The Brittle Network, and The Price of Divergence elsewhere on this site.