Crucible CockpitAdversarial claim auditingThree runs on the record

A forecast is a pile of claims.
Audit the pile.

The Crucible Cockpit takes one question (a rate call, a credit thesis, an accounting critique) and refuses to answer it directly. Instead it decomposes the question into individually falsifiable claims, each filed with its own falsifier and an evidence tier, and then sets agents against them: attackers working four vectors (provenance, inference, selection, anachronism), a null agent contesting thin evidence, an advocate allowed to rebuild only what the record supports. Claims are held, narrowed, or killed. The answer at the end is whatever the surviving claims can carry, at the tier of the weakest load-bearing link, with the graveyard and every rejected attack left on the page.

§ 1The runs

Three runs, one engine

All three runs were executed by the same engine, the run headers carry the same code hash (c5539ae8bf8ff7f6…), verified at run time, so differences in outcome are differences in the evidence, not the machinery.

Every run on the record. Direction is the reader’s verdict on the original question; probability is the reader’s, stated in advance of any resolution.
RunModeQuestionClaimsVerdictsDirection
Run 0BACKTESTBank of Canada policy path, cutoff Jan 26 20231610 held · 3 narrowed · 1 killedUNDETERMINED @ 0.5
Run BLIVEThe 2026 maturity-wall coupon step-up thesis4827 held · 13 narrowed · 4 killedREFUTED @ 0.2
Run CLIVEGoodwill impairment: “too little, too late”4821 held · 19 narrowed · 7 killedREFUTED @ 0.3

Why the backtest exists. Run 0 replays a question whose answer is now known, the Bank of Canada’s policy path, but locks the evidence at a cutoff of January 26, 2023, and audits only what was knowable then. It returned UNDETERMINED at 0.50: the record available on the cutoff date did not decide the question, and the engine said so instead of pretending otherwise. An engine that leaked hindsight would have scored confidently. That is the point of the run: it is the leak test the two live runs stand on.

The live runs. Run B audits the widely-circulated 2026 maturity-wall thesis for Canadian investment-grade issuers and refutes its central statistic at 0.20: every construction of the issuer-level median coupon step-up the retrievable record supports lands below the threshold the thesis needs. Run C audits the “too little, too late” critique of impairment-only goodwill accounting against TSX Composite issuers and comes back REFUTED at 0.30 on the claim as filed.

§ 2How to read a run

The vocabulary, once

TermMeaning
E0 – E3Evidence tiers, best to worst: E0 is a primary dated record; E3 is an uncited assertion. A conclusion inherits the tier of its weakest load-bearing claim.
FalsifierFiled with every claim at entry: the observation that would kill it. A claim without one is not admitted.
HELD / NARROWED / KILLEDA claim’s fate. Narrowed claims survive only in a weaker form, shown beside the text as filed. Killed claims move to the graveyard and stay visible.
Attack vectorsProvenance (is the source real and dated), inference (does the claim follow), selection (what was left out), anachronism (was it knowable at the cutoff).
TripwireA pre-declared halt condition: if too many load-bearing claims fall, the run stops and says so rather than concluding.
BACKTEST / LIVEBacktest runs lock evidence at a historical cutoff to test the engine; live runs audit a present-day thesis.

Each run page carries the complete record: the claim matrix with every source, the graveyard, the crux that stayed contested, the attacks that scored zero, each agent’s concessions, the confidence arithmetic, and the unabridged round-by-round transcript on its own page. Nothing is summarised away.

Method kinship. The same discipline (label the evidence, file the falsifier, let the attack run) drives Flagged in Hindsight, The Brittle Network, and The Price of Divergence elsewhere on this site.

← Research