The Calibrated Mind: Master Note
In briefThe Executive Edge
Markets, careers, and exams all pay for the same scarce asset: being right when it counts, for reasons you can repeat. The research is blunt about where that edge comes from, not intelligence, which mostly amplifies Motivated Reasoning, but process: base rates before narratives, severe tests before belief, written predictions before self-assessment. Everyone else is running unaudited software. This note is the audit.
I. Operating Principles
1. You run on testimony, not verification. You have never audited the earnings report, the drug trial, or the age of the Earth. Your accuracy therefore depends on judging institutions, incentives, and processes, a learnable skill, more than on raw analysis.
2. Human error is systematic, and that is good news. Random error washes out; systematic error compounds, but a predictable failure is a correctable one. The entire toolkit below exists because the failure modes repeat.
3. Intelligence is an amplifier, not a defence. Myside Bias shows almost no correlation with intelligence (Stanovich, West & Toplak, 2013), and a larger bias blind spot associates with higher cognitive ability (West, Meserve & Stanovich, 2012). More horsepower means better rationalizations. (Contested at the edges: a 2025 twin study challenges whether rationality is separable from g at all.)
4. The master bug is asymmetric scrutiny. Congenial claims face "can I believe this?", one supporting reason passes. Threatening claims face "must I believe this?", one flaw dismisses. Both feel like honest inquiry; you never experience yourself choosing to be unfair.
| Soldier mindset | Scout Mindset | |
|---|---|---|
| Reasoning is… | defence of territory | mapmaking |
| Question asked | can/must I believe? | what is actually there? |
| Failure mode | sophisticated wrongness | slower, but self-correcting |
5. Published evidence is a filtered sample. The largest replication audit to date (SCORE, Nature 2026) found 55.1% of social-science claims replicate, with effect sizes shrinking ~80%. Standard journals publish positive results 96% of the time; Registered Reports, accepted before results exist, only 44%. The gap is Publication Bias, not discovery.
6. Fix the process, not the person. Individual debiasing training yields g ≈ 0.26 on a weak literature, with poor real-world transfer. What works is structural: change the representation of the problem, add adversarial review, and keep score.
ImportantThe Severity Principle: the most portable test in this note
Evidence counts only if the process could have come out the other way. Mayo's formulation: if data agree with a claim but the method was practically guaranteed to find agreement, you have "bad evidence, no test (BENT)." Apply to backtests, testimonials, screened funds, and your own reasons for believing what you already believed. → Severe Testing
II. The Decision Protocol
graph TD
A["Claim or decision arrives"] --> B{"Base rate known?"}
B -->|"No"| B1["Find the reference class<br/>before touching the narrative"]
B1 --> C
B -->|"Yes"| C{"Severe test?<br/>Could this evidence have<br/>come out against the claim?"}
C -->|"No"| C1["Treat as ~no evidence.<br/>Check selection effects<br/>and incentives instead"]
C1 --> D
C -->|"Yes"| D{"Do I want this<br/>to be true?"}
D -->|"Yes"| D1["Raise the bar:<br/>write the strongest<br/>opposing case first"]
D1 --> E
D -->|"No"| E{"Is any branch ruinous?"}
E -->|"Yes"| E1["Ergodicity rule:<br/>cap the downside first.<br/>Expected value is the wrong tool"]
E1 --> F
E -->|"No"| F["Commit: explicit probability<br/>+ resolution date"]
F --> G["Log it → score it → review in bands.<br/>Judge the process, not the outcome"]
III. The Model Toolkit
Seeing the data
| Model | Core question | Fails when |
|---|---|---|
| Selection Effects / Survivorship Bias | Who is missing from this sample, and why? | Used as an all-purpose dismissal, name mechanism, direction, magnitude |
| Regression to the Mean | How much of this extreme is noise that won't repeat? | Real mean-reverting forces exist; reliable measures barely regress |
| Base Rates | How do cases like this usually turn out? | The reference class is genuinely ambiguous or gerrymandered |
| Map-Territory Distinction | What does this model deliberately omit? | "Just a model" used to dismiss data in favour of undocumented intuition |
Testing the claim
| Model | Core question | Fails when |
|---|---|---|
| Falsifiability (personal form) | What observation would change my mind? | Treated as a meaning test; Popper meant demarcation |
| Progressive vs. degenerating programmes | Is this belief predicting new things, or only explaining after the fact? | Degeneration is often visible only in hindsight |
| Likelihood Ratio | How much more likely is this evidence if I'm right than if I'm wrong? | Vividness substitutes for diagnosticity, most news has a ratio near 1 |
| Goodhart's Law / Campbell's Law | What happens when this proxy gets optimised? | Assuming metrics are useless rather than corruptible |
TipThe single best statistical habit
Convert every conditional probability into natural frequencies, counts out of 1,000. Physician accuracy on diagnostic inference jumped from 10% → 46% with this one change of representation (Hoffrage & Gigerenzer, 1998). Posterior odds = prior odds × likelihood ratio; the counting format does the algebra for you. → Bayesian Updating
Sizing the bet
| Model | Core question | Fails when |
|---|---|---|
| Expected Value | Probability × payoff, summed, for repeated, survivable bets | One-off decisions, nonlinear utility, unknowable probabilities |
| Ergodicity | Does the time average match the ensemble average? | +50%/−40% at even odds: EV +5%, growth 1.5 × 0.6 = 0.9, ruin |
| Opportunity Cost & marginal thinking | What does the next unit buy versus the best alternative? | Thresholds: 95% of a bridge is 0% of a bridge |
| Antifragility / optionality | Is my downside capped and my upside open? | The "stressor" is actually destruction; some things just break |
Reading the humans
| Model | Core question | Fails when |
|---|---|---|
| Skin in the Game | What happens to the adviser if this advice is wrong? | Motive treated as refuting the argument, it only discounts credibility |
| Coalitions & Preference Falsification | Is this belief a membership badge? Is public speech tracking private belief? | Never applied to yourself, then it's rhetoric, not analysis |
| Chesterton's Fence | Why does the current arrangement exist? | Burden set so high it licenses permanent obstruction, time-box the search |
| Error-correction capacity | How fast are errors found here, and what does admitting one cost? | , the single best one-question audit of any institution |
IV. The Evidence Ledger
What the replication era did to the famous findings. Cite from the left column freely, the middle column with caveats, the right column never.
| Robust | Contested | Dead |
|---|---|---|
| Anchoring (the effect, replicated larger than original) | Loss aversion's size: λ ≈ 1.96 pooled, but 1.07–1.31 under symmetric designs, not a constant | Social priming: 0 of 52 fully independent replications |
| Framing (survives at ~half size, d 1.13 → 0.62) | Ego depletion: d = 0.04 in registered replications; ~0.31 under intense protocols | Power posing on hormones/behaviour (felt power survives) |
| Sunk cost (23% vs 7% when effort was invested) | Forecasting training & teaming: d = 0.77 flipped to −0.29 under method-variance control; adversarial re-test pending | Ease-of-retrieval "availability" paradigm |
| Bias blind spot (replications larger than original) | Myside bias's independence from intelligence (one lab's paradigms; 2025 genetic challenge) | Nudge meta-effect: d = 0.43 → 0.04 bias-corrected; 8.7pp → 1.4pp at government scale |
| Conjunction fallacy (45% → 55% correct with incentives) | Motivated numeracy (preregistered replication failed) | Subliminal anchoring |
Watch outAwareness is not immunity
Kahneman named the "law of small numbers," then built Chapter 4 of his own bestseller on underpowered priming studies that collapsed. His 2017 concession: "I placed too much faith in underpowered studies… I knew all I needed to know to moderate my enthusiasm, but I did not think it through." If knowing about a bias inoculated against it, the field's founder would have been immune. Only process protects. → Replication Crisis
What forecasting research actually supports: experts beat chance only at short horizons (near-chance 3–5 years out); foxes beat hedgehogs; simple extrapolation beat both. The strongest robust predictor of accuracy is update style, frequent, small revisions (superforecasters: 5.1 forecasts per question, mean update 0.11, vs 2.0 and 0.20). Reference Class Forecasting was the highest-value tagged technique: Brier 0.17 vs 0.26.
V. Actionable Leverage
Finance & risk
- Base rate before thesis. Open every pitch with the outcome distribution for comparable firms; the vivid story about this one is a sample of one.
- Severity-test every backtest. A strategy tuned on the data that evaluates it would look good regardless: BENT, no information.
- Size for the time average. Any position with a live path to ruin is mispriced by EV logic. Survive first, optimise second.
- Decompose the thesis. List the 4–5 legs that must all hold and multiply rough probabilities. The product is always lower than the gut number, because conjunctions are less probable than they feel.
- Model second-order exposures, don't footnote them. FX translation on cross-listed holdings reprices continuously, not just at conversion; withholding tax drags every dividend. The tidy map omits these; the territory doesn't.
- Budget for invisible wins. Prevented losses never show up as outcomes, so risk management always looks like waste after it works. Say this out loud before anyone cuts the hedge.
Academics
- Test by production, not recognition. Fluency from reading feels like mastery and predicts almost nothing. Close the book and rebuild the derivation.
- Process-grade your own answers. A right answer from wrong reasoning is a Gettier pass, it will fail you on the exam variant. Log why you were right, not just that you were.
- Translate every stats problem into counts out of 1,000 until it is automatic.
Career & life
- Discount survivor advice. Every "how I got the offer" story comes from a sample with the rejections removed.
- Build intuition only where it can form. Expert intuition requires regular cues plus fast, unambiguous feedback (Kahneman & Klein, 2009). Choose co-op work with tight feedback loops; distrust confident gut-feel, including yours, everywhere else.
- Read the strongest opposing case before committing publicly. After public commitment, the identity machinery turns reading into ammunition-gathering.
- Audit institutions with one question: what does admitting an error cost here? Where the answer is "your reputation," errors are hidden until they are catastrophic.
The practice system
TipThe Prediction Log: the one non-negotiable
One text file: claim, explicit probability, resolution date, one-line reason. Review by confidence band, of your "80%" calls, did ~80% resolve true? You need ~20 resolved before the record means anything and ~50+ before drawing conclusions. This is the only defence against hindsight bias, which corrupts the memory you would otherwise audit yourself with. → Prediction Log · Calibration
Weekly drills, pick two: a Fermi estimate checked against reality · one conditional probability converted to natural frequencies · one held belief traced to its original source (expect a press release) · a Pre-Mortem on any commitment over your threshold · a Steelmanning pass an opponent would endorse as accurate.
Watch outTwo traps that eat smart people
The universal solvent: "that's selection bias / motivated reasoning / complexity" is a hypothesis, not a refutation; specify mechanism, direction, and rough magnitude or drop it. Isolated demands for rigour, if you have never demanded a preregistered replication of something you agreed with, your standards are tracking your preferences, not the truth.
VI. Vault Map
Foundations: Severe Testing · Falsifiability · Bayesian Updating · Base Rates · Likelihood Ratio · Replication Crisis
Risk: Ergodicity · Expected Value · Antifragility · Opportunity Cost · Goodhart's Law
Self: Motivated Reasoning · Myside Bias · Scout Mindset · Calibration · Prediction Log
Next actions: (1) create Prediction Log and enter five dated predictions today; (2) stub Base Rates, Selection Effects, and Ergodicity, the three highest-traffic nodes for a finance vault; (3) run one source-archaeology drill this week.