Alex Rajcoomar portfolio
Library →

Personal project · Methodology

The Portable Kit

The framework turned on itself: five stages of audit, including a rescoring that drops the unresolved period to see whether the verdict survives it.

Five-stage auditSelf-audit

Origin / Personal4,241 words / 0 figures / 3 tablesCitation chords / 0Metadata verified by build

Built from a five-stage audit of one evaluation framework, including a rescoring on partial evidence.

The Portable Kit

A Five-Stage Audit and Transfer of the Leadership-Evaluation Framework

In briefAt a glance

This note compresses a 24,000-word presidential assessment (base weighted verdict 3.245 → 3.2 / 10) into a ten-tool portable kit for evaluating any leader, then stress-tests, audits, and transfers it. Key results: the verdict's sign is robust under every evaluative lens and coherent reweighting (disciplined corridor ≈ 2.5–5.0; the full favourable stack of evidence + interpretation + a mandate criterion reaches only 4.99); a self-audit found eight methodological strains (worst: a summary sentence overstating compliance evidence) that move the verdict at most to ≈ 3.4–3.5; the After-the-loss test's "rerouting" pattern replicates in a corporate case (Tornetta v. Musk); and the exercise yields three concrete reasoning adjustments plus one new anti-motivated-reasoning heuristic, the Flip-list.

Inputs: the reference document (Was Donald Trump a Good Leader, 9 Aug 2026, 756 sources), the teaching companion (How to Judge a President), the underlying research files, and the scoring arithmetic (base weighted score 3.245 → reported 3.2). All new computations in this document were run and verified against that arithmetic.

Standing rules applied throughout: events over attributions; "unestablished" is an available verdict; when two credible numbers conflict, the definitional difference is surfaced before either is used.


Stage 1: Compression and Portability

Task: the minimal toolset that still works in 2035, under different institutions and technologies, ranked by expected value (EV) for reducing systematic error, where EV ≈ (frequency of the error the tool prevents) × (size of the error) × (probability an unaided evaluator commits it).

The ranked kit (ten tools)

1. Baseline discipline

  • Operational form: no verdict on any governing statistic until you can state the inherited trend, the cycle position, and the counterfactual benchmark.
  • Why ranked first: trend-attribution is the most frequent error in all leadership evaluation, committed by both directions with near-certainty when unaided; it single-handedly generated the illusory halves of this record (unemployment records that were decelerating trend-continuation; NATO spending that inflected in 2014).
  • 2035-robust: fully, nothing about it is institutional.

Watch outFailure mode: Baseline discipline

Weaponized fatalism ("everything is trend, no one is responsible"). Guard: baselines assign the residual, they don't erase it; CBO-baseline outperformance of 0.6 pts/yr was real and credited.

2. Pre-committed rubric plus sensitivity analysis

  • Form: declare criteria, weights, and a benchmark before scoring; then publish how far any coherent reweighting moves the verdict.
  • Why second: verdict-first reasoning is the paradigm case of motivated cognition, and this is the only tool that makes it visible, the decisive finding here was that reweighting alone moves 3.2 only to 2.9–3.6 because all seven scores sit below the benchmark.
  • 2035-robust: fully.

Watch outFailure mode: Pre-committed rubric

Rubric-laundering, tuning weights after peeking at scores. Guard: Stage 5's Flip-list heuristic.

3. Enforcement tracing

  • Form: for every rule, name the enforcer; for every enforcer, name their incentive; a rule whose chain terminates in the goodwill of the constrained party is a norm, not a rule.
  • Why third: least intuitive tool with the highest institutional payoff (the GAO-finds → no-one-can-sue → Congress-abstains chain is invisible without it), and rising in importance toward 2035 as algorithmic and platform governance multiply rules-without-enforcers.

Watch outFailure mode: Enforcement tracing

Corrosive cynicism ("only enforcement matters, so norms are worthless"), norms do work until tested; the tool prices the test, it doesn't deny the norm.

4. Own-goals test

  • Form: first-pass grading against the leader's own stated objectives, quantified at announcement time.
  • Why fourth: it de-ideologizes evaluation more cheaply than any other move (the tariff verdict requires no position on protectionism).

Watch outFailure mode: Own-goals test (2035 caveat)

Leaders learn; goal statements are becoming vaguer and more revisable, and the tool fails against a leader who states no falsifiable goals, in which case that observation becomes the finding. Requires timestamping goals at announcement.

5. After-the-loss test

  • Form: the most diagnostic window on any power-holder opens in the two weeks after an adverse ruling, vote, or election; classify the response as accept / appeal / reroute / defy.
  • Evidence base here: three tariff regimes in eighteen months; a second birthright-citizenship order six weeks after the first was struck.
  • 2035-robust: fully, losses are universal.

Watch outFailure mode: After-the-loss test

Over-penalizing lawful appeal; the classification must keep "appeal" and "reroute" distinct (reroute = pursuing the rejected end through a substitute instrument).

6. Definitional-difference-first

  • Form: when credible sources publish conflicting numbers, locate the definition gap (5.9% vs 11.1% tariff rate = realized-vs-statutory; Atlanta Fed vs EPI wages = same-person panel vs fixed-percentile) before adopting either.
  • Rising EV toward 2035 as synthetic content makes number-laundering cheaper.

Watch outFailure mode: Definitional-difference-first

False equivalence, some conflicts are definitional, some are one side being wrong; the tool asks you to check which, not to split differences.

7. Durability sort (ink / pencil / compounding)

  • Form: classify every claimed achievement as statute-or-appointment (ink), decree (pencil), or self-growing liability (compounding: debt interest, verification gaps, systemic distrust); legacy = ink + compounding, with sign.

Watch outFailure mode: Durability sort

Underrating pencil that gets ratified later (executive actions sometimes get codified), treat pencil as an option on ink, not as zero.

8. Unreviewable-powers reading

  • Form: identify the discretion corners no institution can check (pardons, prerogative powers, founder super-voting shares) and weight conduct there heavily as a character measurement, because it is unconfounded by constraint.
  • 2035-robust: every system keeps such corners.

9. Convergence-and-category calibration

  • Form: confidence tracks (a) event vs attribution, (b) independence of converging methods, (c) whether a counterfactual is required; claims failing (c) get "unestablished," said aloud.
  • Worked examples: tariff incidence earns high confidence (five independent methods converge); a Trump-specific polarization increment earns "unestablished" (no counterfactual exists).

10. Shared-premise critics as the highest-probative evidence class

  • Form: seek the people who hold the leader's own premises and still object (the unitary-executive scholar against the unitary executive's use; the Federalist Society prosecutors who resigned).

Watch outFailure mode: Shared-premise critics

Survivorship, such critics are rare, so their absence is weak evidence of anything.

Discarded as non-portable

  • The specific 18/15/20/10/10/17/10 weights (context-tuned).
  • The 5.0 = "average of Reagan–Biden" benchmark (cohort-specific, see Stage 3, Finding 5).
  • Named instruments (V-Dem, Gallup), portable form is "perception thermometers vs event ledgers."
  • The mandate-as-constraint decision (portable is the obligation to decide it explicitly, not the answer).
  • All US-specific enforcement plumbing (portable form is tool 3 itself).

NoteStage 1 confidence

Moderately high on the toolset's contents; moderate on the ranking, the EV ordering rests on judgment about error base-rates, for which no measured data exist.

Watch outStage 1: strongest objections

(i) Compression strips the evidentiary discipline that made the tools credible; a user wielding the ten tools without primary-source habits gets structured vibes, which may be worse than unstructured humility. (ii) The kit was induced from one case study (n = 1) plus general knowledge, so its coverage of failure modes is untested, see Stage 5.

Watch outStage 1: thin-evidence note

Tools 4 and 5 rest on the second term's tariff sequence, whose final legal status is pending; if the Section 301 regime is upheld, "reroute" partially reclassifies as "lawful substitute authority," weakening the flagship example (not the tool).


Stage 2: Stress-Testing the Framework

Method. Hold the documented facts fixed. For each criterion, state the strongest coherent alternative definition a sophisticated defender or critic could adopt, who actually holds it, and the verdict movement it buys (Δ = weight × score change). Then separate two channels: Channel A ("the numbers changed"), pending evidence resolving one way or the other; Channel B ("the interpretation changed"), definitions or weights moving with facts fixed. All arithmetic machine-verified.

2.1 Per-criterion alternatives

Criterion (wt, base) Strongest defender alternative Held by Alt score Δ verdict Strongest critic alternative Alt score Δ
Economy (18%, 4.5) "Growth-and-opportunity" definition: aggregate growth, minority-employment records, equity returns count; median-affordability de-emphasized Supply-side economists; CEA 6.0 +0.27 "Distribution-first": bottom-80% combined tariff+OBBBA effect is the criterion 3.0 −0.27
National security (15%, 3.5) "Hard outcomes over soft power": burden-shift, centrifuges destroyed, hostages returned; discount opinion surveys Heritage, Hudson 5.0 +0.23 "Risk-adjusted": price the war, the blackout, New START at full weight 2.5 −0.15
Institutions (20%, 2.0) "Constitutional-order outcomes": the criterion measures whether checks functioned, and they did, courts bound him, he complied Shapiro; partially McConnell 5.0 +0.60 "Attempt-weighted": attempted subversion scores as subversion 1.0 −0.20
Crisis (10%, 4.0) Split the criterion: deliverable-crisis (Warp Speed) 8, containment/J6 2, weight the former Wenstrup subcommittee logic 5.0 +0.10 Weight the pre-vaccine window, when leverage was maximal 3.0 −0.10
Communication (10%, 4.0) "Effectiveness-only": agenda-setting, realignment, the comeback, trust is the electorate's problem Campaign professionals, both parties 7.0 +0.30 "Truth-content-weighted" 2.0 −0.20
Long-term (17%, 3.0) "Direction over durability": judiciary + China consensus prove the agenda sticks; EO-reversibility is ordinary politics; debt is bipartisan Cass; Krein-circa-2017 4.5 +0.26 Compounding-weighted: debt + verification dominate 1.5 −0.26
Character (10%, 2.0) The democratic answer: voters had full information twice and priced it Mandate theorists 4.5 +0.25 Fitness-threshold view (six senior appointees' testimony is dispositive) 1.0 −0.10

ImportantCoherence costs of the defender alternatives

Two of the defender alternatives carry a coherence cost that must be stated: the economy alternative survives only by excluding the incidence literature, not by disputing it (it is the framework's highest-confidence finding, five independent methods); and the institutions alternative must absorb the document's rebuttal that grading the president by the courts' resistance rewards provocation, plus Prakash's shared-premise objection. They are coherent, but they are not free.

2.2 The three lenses, isolated

  • Outcomes-only lens (econ 45 / natsec 25 / crisis 30): ≈ 4.1
  • Stewardship-only lens (institutions 50 / long-term 35 / character 15): ≈ 2.4
  • Mandate-only lens: 7–8 (two victories, a non-consecutive comeback, coalition realignment, but this lens answers "was he chosen?", a different question)

ImportantThe single most important sensitivity result

The verdict's sign is robust under both evaluative lenses and flips only under the lens that isn't an evaluation. This result did not appear explicitly in the reference document.

2.3 The two channels, quantified

Scenario Composition Result
Base document weights and scores 3.245
A+ (numbers break favourable) §301 tariffs upheld, Iran framework holds + inspections resume, sentiment recovery continues → econ 5.5, natsec 4.5, long-term 4.0 3.75
A− (numbers break adverse) second refund cycle, Iran framework collapses → econ 4.0, natsec 2.5, long-term 2.0 2.84
B+ (interpretation only) institutions→5.0, character→4.0, plus mandate added as an eighth criterion at 15% weight scoring 7.5 (others rescaled) 4.56
B− (interpretation only) institutions→1.0, character→1.0 2.95
A+ and B+ stacked all of the above simultaneously 4.99
A− and B− stacked 2.54

Reading: the evidence channel moves the verdict by at most ±0.5; the interpretation channel by at most ±1.3; and even the full favourable stack (every pending fact resolving his way, the strongest coherent redefinitions of the two lowest criteria, and the mandate converted from a constraint into a heavily-weighted criterion) lands at 4.99, a hair under the average-president line. This independently reproduces, by a different route, the reference document's 5.2 corner (which additionally reweighted).

ImportantThe disciplined corridor

≈ 2.5 to 5.0, and the interesting property is unchanged: the top of the corridor merely reaches "average."

2.4 Two additional stress tests the document did not run

First-term-only

Objection: the verdict is dragged by nineteen unresolved months. Rescoring on first-term evidence alone (fresh estimates, each ±0.5: econ 5.5, natsec 4.0, institutions 3.5, appointments ≈7 at one-third weight, 2020-election conduct ≈2 at two-thirds, crisis 4.0, comm 4.0, long-term 4.0, character 3.0) yields ≈ 4.1, corridor roughly 3.6–4.6.

ImportantFinding

The second term costs about 0.85 points; the verdict's sign does not depend on the incomplete term, but part of its margin does. Confidence: moderate, these are fresh estimates, clearly labelled.

Mandate-as-criterion

Converting the framework's most contested exclusion into a heavily-weighted inclusion (15%, scored 7.5) moves the base verdict from 3.25 to 3.88. Even the strongest structural concession to the defence is worth about +0.6.

NoteStage 2 confidence

High on the arithmetic (machine-verified); moderately high on the claim that the listed alternatives are the strongest available, a coherent redefinition may have been missed, though the space was searched from both directions.

Watch outStage 2: strongest objection

Score granularity is false precision; every input carries ±0.5–1.0 of judgment, so corridor endpoints deserve less trust than the corridor's central tendency and its sign.

Watch outStage 2: thin-evidence note

The A-channel scenarios depend on litigation and diplomacy outcomes that are pending, not merely contested.


Stage 3: Methodological Audit

The document's standards, applied to the document. Ordered by severity. For each: the standard, the strain, and whether the verdict survives the correction.

Watch outContext that disciplines this audit

An adversarial fact-check pass already found six critical errors in the draft (all corrected), which is base-rate evidence that self-review under-detects, so findings below should be read as a floor, not a ceiling.

Finding 1: "Repeatedly obeyed" overstates the document's own compliance evidence

Severity: highest. The Executive Summary states courts "repeatedly ruled against him and were repeatedly obeyed." The document's own §3.4.2 records a judicial finding of "willful and intentional noncompliance" (Abrego Garcia), a probable-cause contempt finding, and a 57-of-165 accusation count. The body holds the tension carefully; the summary resolves it in the defence's favour. This is the document committing a mild version of the error it names: treating formal compliance in the marquee cases as general fidelity. Verdict survives (the institutional score already prices this), but the summary sentence should read "and, in the highest-stakes cases, obeyed."

Finding 2: Perception-index salience contradicts the stated evidence hierarchy

Severity: high. The methodology demotes V-Dem-class indices to "thermometers of expert assessment" to be paired with hard events; yet the V-Dem reclassification appears in the Executive Summary's driver paragraph, the document's most-read sentences, with marquee placement. The companion repeats the pattern (Bright Line Watch's Canada-88/US-57 in its Canadian section). The stated hierarchy and the rhetorical hierarchy diverge. Correction: summary placement should be reserved for the event ledger; indices belong in the supporting tier. Verdict survives, the document itself demonstrates (§3.4.2) that the institutional case stands with the indices struck.

Finding 3: Deflator inconsistency in the median-household indictment

Severity: moderate, and instructive. The document attributes the 2026 headline-inflation spike to an exogenous oil shock (correctly), then deflates second-term real wages by headline CPI (+0.80% real over 17 months) in the affordability indictment. Consistency requires a core-deflated check. Running it: headline CPI over the window annualizes to ≈3.0%, core PCE to ≈3.2%, the gap nets out to roughly zero over this particular window, so the indictment survives the check empirically. But the document never ran the check, and under a slightly different window it could have flipped the sub-finding. This is the cleanest example of "the interpretation was fine, the hygiene was not."

Finding 4: Endpoint choice inflates a headline fiscal number

Severity: moderate. "Deficit rose 68 percent" uses FY2016 → FY2019. A defender choosing FY2017 (the last budget substantially Obama's) gets +48 percent. Both are defensible; the document picks the larger without flagging the choice, while elsewhere (fiscal-year attribution, −6.55% vs −8.63%) it does flag convention-dependence. The %-of-GDP framing (3.1→4.6) is endpoint-robust and should carry the claim.

Finding 5: Asymmetric circumstance adjustment

Severity: moderate, conceptually the deepest. The framework circumstance-adjusts the economy meticulously (91-month expansion, inheritance table) but applies no equivalent adjustment to the institutional criterion, where a defender would argue the environment was also unprecedented (record emergency-docket volume, nationwide injunctions, two assassination attempts, prosecutions later dismissed). The document's implicit answer, much of that environment was self-generated, is defensible but asserted rather than argued, and the cohort benchmark ("average of Reagan–Biden") embeds the same asymmetry: it adjusts economic starting conditions across the cohort but not institutional stress. A rigorous v3 would either adjust both or explicitly defend adjusting only one.

Finding 6: A latent double-count in the fiscal indictment

Severity: moderate. §2.3 lowers national security's weight partly to avoid double-counting with long-term positioning, but deficits and debt appear as negatives in both the economy criterion (§3.1.7) and long-term positioning (§3.8), with no equivalent disclosure. Rough magnitude: if fiscal deterioration were charged only once, the economy score plausibly rises ~0.5, moving the verdict ≈ +0.09. Small, but the inconsistency is real.

Finding 7: The own-goals verdict is time-sensitive and one sub-goal has since moved

Severity: low-moderate. "Tariffs failed on their own terms" bundles three goals. Manufacturing employment (failed, twice) and incidence (failed, high confidence) are stable. But the trade deficit narrowed materially in 2026 ($932B → ≈$728B twelve-month), which the document reports and then under-weights in the summary formulation. Honest form: "failed two of three stated goals; the third is mixed and recent." Verdict effect: negligible; framing effect: real.

Finding 8: The companion promotes n = 1 inductions to general signatures

Severity: moderate for the companion, nil for the document. The "five-tell cluster" for communicator-without-steward, and the "$0 investments page" as a diagnostic, were induced from a single case; the page is a website artifact, colorful, evidentially thin. Teaching artifacts compress; compression became overclaim. Fix: label induced heuristics with their n and their disconfirming test (are there high-decree, high-churn leaders who were durable stewards? Wartime cases are candidates).

Shared-premise critics who would still dissent from the corrected document

  • Deficit hawks would press Finding 6.
  • The originalists who won the tariff case would reject the "facts about courts" discount, arguing that litigating in good faith and complying is what constitutional fidelity looks like from the inside (they would score institutions nearer 3 than 2).
  • The C-SPAN-postponement logic, that contemporaneous expert judgment of a sitting president is punditry, indicts the document's own timing, a point it acknowledges (§6.4) but does not fully absorb.

NoteStage 3 confidence

Moderately high that these are the largest strains.

ImportantNet effect of the audit

The corrected verdict moves at most from 3.2 to ≈ 3.4–3.5 under every finding above applied simultaneously, the audit changes the document's precision and rhetoric more than its conclusion.

Watch outStage 3: strongest objection to the audit itself

It was conducted by the document's author; the six-critical-error base rate from the earlier external pass implies residual undetected defects, and an adversarial reader should be commissioned for anything above this stakes level.


Stage 4: Transfer and Generation

4.1 The template: a concentrated-power CEO (one page)

Same architecture, three substitutions: the benchmark cohort becomes peer large-cap founder/CEOs; the enforcer map becomes board → shareholders → courts (Delaware/securities) → regulators; the mandate analogue becomes shareholder votes (with the same rule, ratification legitimizes, it does not grade).

Criterion Wt The question Primary evidence class
Owner outcomes vs baseline 20% Returns and fundamentals against sector and index, not raw Ink: audited financials
Strategic position & durability 20% Moat, capital structure, key-person risk; ink vs pencil vs compounding Filings, debt schedule
Governance & internal controls 20% Board independence; what happens after adverse rulings; related-party discipline Court records, proxies
Execution vs stated goals 15% Announced targets, timestamped, vs delivered Own announcements vs actuals
Talent system 10% Executive churn, bench depth, whistleblower record Turnover data
Communication & capital-markets trust 10% Guidance credibility; regulator interactions Consent decrees, restatements
Conduct in unreviewable corners 5% Super-voting shares, controlled entities, personal-platform use Conduct where no one can stop them

Weights justification mirrors the original: governance = the internal controls only the CEO can maintain; durability heavy because compounding liabilities (leverage, key-person concentration) outlast tenure; outcomes capped at 20% because sector beta dominates raw returns.

4.2 Brief application: Elon Musk at Tesla

Watch outTime-stamp, stated plainly

Independently verifiable knowledge runs through roughly May 2025; nothing after that is covered, and this sketch must be refreshed against current filings before any use. (This is the framework's own vintage rule, applied to the author.)

The reasoning chain, tools in brackets:

  • Baseline: Tesla's returns and its 2024 delivery decline (~1.79M, the first annual drop, with BYD overtaking on BEV volume in late 2023) mean nothing raw; graded against the EV sector's simultaneous demand slowdown, execution reads mid-pack rather than catastrophic, a genuine moderation of the bear case, exactly as baseline discipline moderated the Trump-economy bear case.
  • Own-goals, Timestamped targets, one million robotaxis by 2020 (announced 2019, delivered zero by early 2025), and the 20-million-vehicles-by-2030 goal reportedly dropped from the 2024 impact report (moderate confidence; verify), fail the own-goals test on the announced schedule, independent of any view about the technology.
  • After-the-loss, the isomorph, The Delaware Court of Chancery rescinded the $55.8B pay package (January 2024); the response was not acceptance or mere appeal but rerouting: a shareholder re-ratification, reincorporation from Delaware to Texas, and a public "be careful where you incorporate" posture, structurally identical to the tariff sequence (lose on the instrument, rebuild the same end under a friendlier authority). The court declined to reinstate despite ratification (December 2024); appeal pending as of the knowledge window.
  • Enforcement tracing: The enforcer chain has a documented weak link: directors settled excess-compensation claims by returning ~$735M (2023), and the 2018 SEC consent decree's supervision terms produced years of contested compliance, detection robust, consequence intermittent.
  • Unreviewable corners: Conduct on a personally-owned platform, unconstrained by any board, is the pure character read the template asks for.
  • Durability sort: Ink: manufacturing scale, Supercharger network, energy business. Pencil: valuation premised on announced-not-delivered autonomy. Compounding: key-person concentration, explicitly priced by the principal himself ("uncomfortable growing Tesla in AI/robotics without ~25% voting control," January 2024).

CheckSketch verdict, provisional

Against a peer-founder benchmark of 5.0 (strong on ink assets and category creation, mid on baseline-adjusted outcomes, weak on governance/after-the-loss and own-goals punctuality) the corridor lands roughly 4–5 with wide bands, and the template's real yield is not the number but the finding that the rerouting pattern recurs across domains, which upgrades tool 5 from a political observation to a general one.

4.3 The Canadian-PM variant

Needs only substitutions, not redesign: benchmark = post-1980 PMs circumstance-adjusted; enforcer map = confidence convention, PBO, Auditor General, courts under the Charter, the Crown's reserve powers; mandate analogue = seat count vs vote share (a minority PM's "mandate" discount is the exact analogue of the plurality/majority distinction); unreviewable corners = prorogation advice and appointments. Building and scoring it is the retention exercise.

NoteStage 4 confidence

Moderately high on the template's structure; moderate on the Musk application (two facts flagged for verification; post-May-2025 record unexamined).

Watch outStage 4: strongest objections

(i) The corporate/constitutional disanalogy is real, shareholders can exit and are owed only fiduciary duties, citizens mostly cannot and are owed the whole constitutional order; the transfer is structural (the tests travel), not normative (the stakes don't). (ii) Choosing Musk risks availability bias, the template should next be run on a boring CEO to test whether it discriminates or just detects flamboyance.


Stage 5: Self-Improvement Loop

Limitation 1: Exogeneity rulings are not automatically propagated

The oil shock was ruled exogenous for inflation-blame, then a headline-CPI deflator was used for real wages three subsections later (Stage 3, Finding 3). The pattern: a judgment made in one cell of an analysis does not automatically recompute the other cells it touches.

  • Adjustment: whenever a factor is classified as exogenous/endogenous, mechanically list every other computation in the document that factor enters, and either re-run each under the classification or state why it is immune. This is a checklist behaviour, not an insight, which is exactly why it was missed.

Limitation 2: Salience allocation drifts from the stated evidence hierarchy

The methodology ranked events above perception indices, then gave an index marquee placement in the summary, and repeated the pattern in the companion (Stage 3, Findings 1–2). Summaries are where motivated reasoning hides after the body has been disciplined, because summary-writing feels like compression rather than inference.

  • Adjustment: audit executive summaries as claims, not as prose, every sentence in a summary must cite the evidentiary tier of its strongest support, and only top-tier items get driver-paragraph placement.

Limitation 3: Single-case inductions promoted to general signatures without labelling the n

The "five-tell cluster" and the corridor heuristics were built from one presidency and shipped as portable tools without their disconfirmation tests attached (Stage 3, Finding 8; Stage 4's Musk case partially, but only partially, replicates one tool).

  • Adjustment: every induced heuristic ships with three fields: n, the nearest disconfirming case searched for, and the observation that would retire it. Where no search has happened, the heuristic is labelled "candidate," not "signature."

The additional hardening heuristic

CheckThe pre-registered Flip-list

Before scoring any criterion, write down the specific, observable future evidence that would move that score one full point in each direction. If you cannot name it, you are not scoring, you are voting, and the criterion should be marked "judgment, not measurement" in the output. This closes the gap the sensitivity analysis leaves open: sensitivity analysis disciplines weights, but scores are where motivated reasoning actually lives, and flip-lists force scores to behave like forecasts, which are gradeable later.

Quick companion variant for daily use, the Actor-swap: rewrite any verdict sentence with a leader you favour substituted in; whatever makes you flinch is the part doing motivated work.


In briefStanding summary

The framework survives its own audit with a corrected corridor of roughly 2.5–5.0 and a central estimate near 3.2–3.5; the sign is robust under every evaluative lens and every coherent reweighting; the precision claimed should be one notch humbler than the original document's; and the kit's ten tools, three adjustments, and one flip-list heuristic are the durable output, the verdict was only the training exercise.


  • Was Donald Trump a Good Leader, the 24,000-word reference document this kit was induced from
  • How to Judge a President, the teaching companion (nine lessons, Canadian lens)
  • Baseline discipline · Pre-committed rubric · Enforcement tracing · Own-goals test · After-the-loss test
  • Definitional-difference-first · Durability sort · Unreviewable-powers reading · Convergence-and-category calibration · Shared-premise critics
  • Flip-list · Actor-swap
  • Tornetta v Musk, the corporate isomorph of the rerouting pattern
  • Presidential greatness surveys, methodology, partisan-composition problem and the Republican-respondent breakout
  • Motivated reasoning, countermeasures, where the flip-list and actor-swap belong in a broader toolkit

Converted once from my own markdown note by a script outside this repository. The HTML is the record, and this page is the published form of it. How the site counts it.

↑↓ moveEnter openEsc close