Alex Rajcoomar portfolio
Research →

Independent research · Field guide

The Calibrated Mind

You have never audited the earnings report, the drug trial, or the age of the Earth. Accuracy therefore depends on judging institutions, incentives and methods rather than facts.

Field guideCompanion to The Delayed Test

Origin / Independent1,876 words / 0 figures / 6 tablesCitation chords / 0Metadata verified by build

Not declared on file.

The Calibrated Mind: Master Note

In briefThe Executive Edge

Markets, careers, and exams all pay for the same scarce asset: being right when it counts, for reasons you can repeat. The research is blunt about where that edge comes from, not intelligence, which mostly amplifies Motivated Reasoning, but process: base rates before narratives, severe tests before belief, written predictions before self-assessment. Everyone else is running unaudited software. This note is the audit.


I. Operating Principles

1. You run on testimony, not verification. You have never audited the earnings report, the drug trial, or the age of the Earth. Your accuracy therefore depends on judging institutions, incentives, and processes, a learnable skill, more than on raw analysis.

2. Human error is systematic, and that is good news. Random error washes out; systematic error compounds, but a predictable failure is a correctable one. The entire toolkit below exists because the failure modes repeat.

3. Intelligence is an amplifier, not a defence. Myside Bias shows almost no correlation with intelligence (Stanovich, West & Toplak, 2013), and a larger bias blind spot associates with higher cognitive ability (West, Meserve & Stanovich, 2012). More horsepower means better rationalizations. (Contested at the edges: a 2025 twin study challenges whether rationality is separable from g at all.)

4. The master bug is asymmetric scrutiny. Congenial claims face "can I believe this?", one supporting reason passes. Threatening claims face "must I believe this?", one flaw dismisses. Both feel like honest inquiry; you never experience yourself choosing to be unfair.

Soldier mindset Scout Mindset
Reasoning is… defence of territory mapmaking
Question asked can/must I believe? what is actually there?
Failure mode sophisticated wrongness slower, but self-correcting

5. Published evidence is a filtered sample. The largest replication audit to date (SCORE, Nature 2026) found 55.1% of social-science claims replicate, with effect sizes shrinking ~80%. Standard journals publish positive results 96% of the time; Registered Reports, accepted before results exist, only 44%. The gap is Publication Bias, not discovery.

6. Fix the process, not the person. Individual debiasing training yields g ≈ 0.26 on a weak literature, with poor real-world transfer. What works is structural: change the representation of the problem, add adversarial review, and keep score.

ImportantThe Severity Principle: the most portable test in this note

Evidence counts only if the process could have come out the other way. Mayo's formulation: if data agree with a claim but the method was practically guaranteed to find agreement, you have "bad evidence, no test (BENT)." Apply to backtests, testimonials, screened funds, and your own reasons for believing what you already believed. → Severe Testing


II. The Decision Protocol

graph TD
    A["Claim or decision arrives"] --> B{"Base rate known?"}
    B -->|"No"| B1["Find the reference class<br/>before touching the narrative"]
    B1 --> C
    B -->|"Yes"| C{"Severe test?<br/>Could this evidence have<br/>come out against the claim?"}
    C -->|"No"| C1["Treat as ~no evidence.<br/>Check selection effects<br/>and incentives instead"]
    C1 --> D
    C -->|"Yes"| D{"Do I want this<br/>to be true?"}
    D -->|"Yes"| D1["Raise the bar:<br/>write the strongest<br/>opposing case first"]
    D1 --> E
    D -->|"No"| E{"Is any branch ruinous?"}
    E -->|"Yes"| E1["Ergodicity rule:<br/>cap the downside first.<br/>Expected value is the wrong tool"]
    E1 --> F
    E -->|"No"| F["Commit: explicit probability<br/>+ resolution date"]
    F --> G["Log it → score it → review in bands.<br/>Judge the process, not the outcome"]

III. The Model Toolkit

Seeing the data

Model Core question Fails when
Selection Effects / Survivorship Bias Who is missing from this sample, and why? Used as an all-purpose dismissal, name mechanism, direction, magnitude
Regression to the Mean How much of this extreme is noise that won't repeat? Real mean-reverting forces exist; reliable measures barely regress
Base Rates How do cases like this usually turn out? The reference class is genuinely ambiguous or gerrymandered
Map-Territory Distinction What does this model deliberately omit? "Just a model" used to dismiss data in favour of undocumented intuition

Testing the claim

Model Core question Fails when
Falsifiability (personal form) What observation would change my mind? Treated as a meaning test; Popper meant demarcation
Progressive vs. degenerating programmes Is this belief predicting new things, or only explaining after the fact? Degeneration is often visible only in hindsight
Likelihood Ratio How much more likely is this evidence if I'm right than if I'm wrong? Vividness substitutes for diagnosticity, most news has a ratio near 1
Goodhart's Law / Campbell's Law What happens when this proxy gets optimised? Assuming metrics are useless rather than corruptible

TipThe single best statistical habit

Convert every conditional probability into natural frequencies, counts out of 1,000. Physician accuracy on diagnostic inference jumped from 10% → 46% with this one change of representation (Hoffrage & Gigerenzer, 1998). Posterior odds = prior odds × likelihood ratio; the counting format does the algebra for you. → Bayesian Updating

Sizing the bet

Model Core question Fails when
Expected Value Probability × payoff, summed, for repeated, survivable bets One-off decisions, nonlinear utility, unknowable probabilities
Ergodicity Does the time average match the ensemble average? +50%/−40% at even odds: EV +5%, growth 1.5 × 0.6 = 0.9, ruin
Opportunity Cost & marginal thinking What does the next unit buy versus the best alternative? Thresholds: 95% of a bridge is 0% of a bridge
Antifragility / optionality Is my downside capped and my upside open? The "stressor" is actually destruction; some things just break

Reading the humans

Model Core question Fails when
Skin in the Game What happens to the adviser if this advice is wrong? Motive treated as refuting the argument, it only discounts credibility
Coalitions & Preference Falsification Is this belief a membership badge? Is public speech tracking private belief? Never applied to yourself, then it's rhetoric, not analysis
Chesterton's Fence Why does the current arrangement exist? Burden set so high it licenses permanent obstruction, time-box the search
Error-correction capacity How fast are errors found here, and what does admitting one cost? , the single best one-question audit of any institution

IV. The Evidence Ledger

What the replication era did to the famous findings. Cite from the left column freely, the middle column with caveats, the right column never.

Robust Contested Dead
Anchoring (the effect, replicated larger than original) Loss aversion's size: λ ≈ 1.96 pooled, but 1.07–1.31 under symmetric designs, not a constant Social priming: 0 of 52 fully independent replications
Framing (survives at ~half size, d 1.13 → 0.62) Ego depletion: d = 0.04 in registered replications; ~0.31 under intense protocols Power posing on hormones/behaviour (felt power survives)
Sunk cost (23% vs 7% when effort was invested) Forecasting training & teaming: d = 0.77 flipped to −0.29 under method-variance control; adversarial re-test pending Ease-of-retrieval "availability" paradigm
Bias blind spot (replications larger than original) Myside bias's independence from intelligence (one lab's paradigms; 2025 genetic challenge) Nudge meta-effect: d = 0.43 → 0.04 bias-corrected; 8.7pp → 1.4pp at government scale
Conjunction fallacy (45% → 55% correct with incentives) Motivated numeracy (preregistered replication failed) Subliminal anchoring

Watch outAwareness is not immunity

Kahneman named the "law of small numbers," then built Chapter 4 of his own bestseller on underpowered priming studies that collapsed. His 2017 concession: "I placed too much faith in underpowered studies… I knew all I needed to know to moderate my enthusiasm, but I did not think it through." If knowing about a bias inoculated against it, the field's founder would have been immune. Only process protects. → Replication Crisis

What forecasting research actually supports: experts beat chance only at short horizons (near-chance 3–5 years out); foxes beat hedgehogs; simple extrapolation beat both. The strongest robust predictor of accuracy is update style, frequent, small revisions (superforecasters: 5.1 forecasts per question, mean update 0.11, vs 2.0 and 0.20). Reference Class Forecasting was the highest-value tagged technique: Brier 0.17 vs 0.26.


V. Actionable Leverage

Finance & risk

  • Base rate before thesis. Open every pitch with the outcome distribution for comparable firms; the vivid story about this one is a sample of one.
  • Severity-test every backtest. A strategy tuned on the data that evaluates it would look good regardless: BENT, no information.
  • Size for the time average. Any position with a live path to ruin is mispriced by EV logic. Survive first, optimise second.
  • Decompose the thesis. List the 4–5 legs that must all hold and multiply rough probabilities. The product is always lower than the gut number, because conjunctions are less probable than they feel.
  • Model second-order exposures, don't footnote them. FX translation on cross-listed holdings reprices continuously, not just at conversion; withholding tax drags every dividend. The tidy map omits these; the territory doesn't.
  • Budget for invisible wins. Prevented losses never show up as outcomes, so risk management always looks like waste after it works. Say this out loud before anyone cuts the hedge.

Academics

  • Test by production, not recognition. Fluency from reading feels like mastery and predicts almost nothing. Close the book and rebuild the derivation.
  • Process-grade your own answers. A right answer from wrong reasoning is a Gettier pass, it will fail you on the exam variant. Log why you were right, not just that you were.
  • Translate every stats problem into counts out of 1,000 until it is automatic.

Career & life

  • Discount survivor advice. Every "how I got the offer" story comes from a sample with the rejections removed.
  • Build intuition only where it can form. Expert intuition requires regular cues plus fast, unambiguous feedback (Kahneman & Klein, 2009). Choose co-op work with tight feedback loops; distrust confident gut-feel, including yours, everywhere else.
  • Read the strongest opposing case before committing publicly. After public commitment, the identity machinery turns reading into ammunition-gathering.
  • Audit institutions with one question: what does admitting an error cost here? Where the answer is "your reputation," errors are hidden until they are catastrophic.

The practice system

TipThe Prediction Log: the one non-negotiable

One text file: claim, explicit probability, resolution date, one-line reason. Review by confidence band, of your "80%" calls, did ~80% resolve true? You need ~20 resolved before the record means anything and ~50+ before drawing conclusions. This is the only defence against hindsight bias, which corrupts the memory you would otherwise audit yourself with. → Prediction Log · Calibration

Weekly drills, pick two: a Fermi estimate checked against reality · one conditional probability converted to natural frequencies · one held belief traced to its original source (expect a press release) · a Pre-Mortem on any commitment over your threshold · a Steelmanning pass an opponent would endorse as accurate.

Watch outTwo traps that eat smart people

The universal solvent: "that's selection bias / motivated reasoning / complexity" is a hypothesis, not a refutation; specify mechanism, direction, and rough magnitude or drop it. Isolated demands for rigour, if you have never demanded a preregistered replication of something you agreed with, your standards are tracking your preferences, not the truth.


VI. Vault Map

Foundations: Severe Testing · Falsifiability · Bayesian Updating · Base Rates · Likelihood Ratio · Replication Crisis

Risk: Ergodicity · Expected Value · Antifragility · Opportunity Cost · Goodhart's Law

Self: Motivated Reasoning · Myside Bias · Scout Mindset · Calibration · Prediction Log

Next actions: (1) create Prediction Log and enter five dated predictions today; (2) stub Base Rates, Selection Effects, and Ergodicity, the three highest-traffic nodes for a finance vault; (3) run one source-archaeology drill this week.

Converted once from my own markdown note by a script outside this repository. The HTML is the record, and this page is the published form of it. How the site counts it.

↑↓ moveEnter openEsc close