The Delayed Test
A synthesis of six research files, for a reader who intends to get better at learning, judging, and seeing the world clearly.
Source key. [Doc 1] Durable Learning: Adversarial Panel Review. [Doc 2] Human Judgment Under Uncertainty. [Doc 3] Financial Manias and Capital Cycles. [Doc 4] Judgment and Decision Making Under Uncertainty: A Four-Voice Adversarial Review. [Doc 5] What Actually Makes Complex Material Stick. [Doc 6] The Anatomy of Manias and Capital Cycles.
A note on evidence. A Verified tag below means the source file verified that claim against retrieved text and I am carrying its grade forward unaltered; I have not re-verified anything independently. Where a file marks a figure second-hand, contested, or inferred, that flag travels with the figure in the sentence where it appears, because the files' authors asked for exactly that discipline and they were right to.
Phase 1: Thematic Mapping
Core truths
-
Verified The feeling of learning and the fact of learning come apart, and the gap shows only at a delay. Rereading beat self-testing at five minutes, 81 percent to 75, and lost at one week, 42 to 56; students who reread a passage over fourteen times were beaten by students who read it about three times and spent the difference retrieving, and the rereaders were the most confident group in the study [Doc 1]. Every file returns to some version of this crossover.
-
Verified Confidence has one durable defect: ranges stated too narrowly. Of the three overconfidences, overestimation and overplacement flip sign with task difficulty; only overprecision persists [Doc 2, Doc 4]. And calibration is a property of a person paired with an environment, not a trait: weather forecasters are described in this literature as superbly calibrated because they make thousands of forecasts that resolve within a day and are scored [Doc 4].
-
Verified Structure beats intention. Mechanical combination of information outperforms holistic expert judgment by about 10 percent across tasks, judges, and experience levels [Doc 2, Doc 4]. Restating probabilities as counts out of 1,000 raised Bayesian performance by an average of 37 percentage points, a larger effect than any training in the files [Doc 2]. Raising the stakes fixes nothing: no replicated study has removed a reasoning error by paying people more [Doc 2].
-
Verified Findings shrink as they travel toward you. The same spacing intervention fell from g = 0.43 taught in isolation to g = 0.24 embedded in a real course [Doc 5]; published nudge effects fell from 8.7 percentage points in journals to 1.4 in government trials, with publication bias accounting for the entire gap [Doc 2].
-
Verified In markets, price alone says almost nothing, credit says a great deal, and nothing says when. After a real doubling, a national market was more likely to double again over five years, about 26 percent, than to halve, about 15 [Doc 3]. When credit growth and asset-price growth are jointly extreme, one-year crisis probability moves from 4 percent to 13 or 14, and three-year probability reaches 45 percent for business credit and 37 for household, with most alarms still false [Doc 3]. Where the leverage sits governs how much a bust destroys [Doc 3, Doc 6]. No feature in either mania file produces a date.
-
Verified Much of what everyone knows did not survive checking. Learning styles, the learning pyramid, ego depletion, the universal 2:1 loss-aversion constant, and the hot hand fallacy, which reversed when the researchers' own biased estimator was corrected [Doc 1, Doc 2, Doc 4, Doc 5].
The tensions
-
[TENSION] Difficulty helps only when success is possible. The files celebrate effortful retrieval and simultaneously record that tests work best when they permit high levels of performance [Doc 1]. Struggle is a dose, not a virtue, and the right dose depends on what you already know.
-
[TENSION] The correct method expires as you improve. Worked examples rescue the novice, then stop helping, then hurt [Doc 1, Doc 5]. Growth demands firing the scaffolding that produced it, on schedule, against comfort.
-
[TENSION] Is the flaw in the mind or in the interface? A 37-point improvement from reformatting suggests the interface; the reply is that spectacles do not prove myopia is a myth [Doc 2]. Unresolved, and the practical conclusion survives either way: redesign the interface.
-
[TENSION] The best-loved technique is unproven exactly where a quantitative student needs it most. Retrieval practice in mathematics sits at g = 0.18 with a confidence interval crossing zero, and the reviewing panel ended its session divided over whether that is a real boundary or an artifact of studies that never demanded retrieval [Doc 5].
-
[TENSION] Techniques reverse; principles hold. Interleaving is worth d = 0.83 in one randomized classroom trial and g = -0.39 on vocabulary [Doc 5]. The principle underneath the technique: practise the discrimination you will be tested on, and block what contains no discrimination.
-
[TENSION] The market is efficient and fragile at once, depending on the unit of analysis. Whole indices hand the skeptic the win; sectors net of market hand it to the historian; most public arguments about bubbles are two people using different units without noticing, a synthesis one file offers while flagging it as its own inference [Doc 6].
-
[TENSION] Stories predict nothing and remain indispensable. The new-era narrative failed every discrimination test and was retained anyway, demoted from cause to pointer: it names the one number to audit [Doc 3, Doc 6].
The hidden lessons
-
[INSPIRATION] You can build laboratory instruments for a sample size of one. A prediction with a probability and a date, a score written down before marking, one decision made twice a week apart: apparatus that measures what introspection cannot reach.
-
[INSPIRATION] Being wrong on schedule is a skill. A 90 percent claim resolving true at 60 percent is a successful detection of overprecision, not a failure [Doc 2]; a warning that is wrong most of the time can still be the best instrument in existence [Doc 3]. The educated response to error is bookkeeping, not shame.
-
[INSPIRATION] The files' deepest teaching is their conduct. Each ranks its own weakest claims, keeps a revision history, and confesses in public when a load-bearing number turns out to live in the wrong paper. Doubt practised as craft rather than mood is a learnable discipline, and possibly the learnable discipline.
-
[INSPIRATION] Audit the instrument before diagnosing the world. In the hot hand, the bias sat in the researchers' measure, not in the players [Doc 2, Doc 4].
-
[INSPIRATION] Everything durable is scored late. Retention at a week, calibration at resolution, solvency at the turn of the credit cycle. Patience here is not a temperament. It is a method.
Phase 2: Synthesis Panel
THE STORYTELLER: One spine or nothing, and the spine is the delayed test. The same event recurs at three scales: a student rereads and feels ready; a manager states a range and feels sure; a market watches prices climb and feels vindicated. Fluency, confidence, euphoria: one illusion in three costumes, each exposed only when the late exam arrives. Open with the 2006 experiment, because it holds the whole argument in miniature, including the cruelty that the losing group was the most confident. Then turn, from the illusion to the instruments people built against it, and end with the people who keep the instruments.
THE MENTOR: Agreed, on one condition. If the middle movement reads as a list of study hacks, we will have manufactured the exact object these files warn about: a technique stripped of its moderators and sold on a poster. Every instrument must carry its principle and its boundary in the same breath. Retrieval without prior knowledge is guessing. Interleaving reverses on vocabulary. A checklist only works if you keep it after it embarrasses you. And the reader must leave with agency, because the files' most hopeful finding is that calibration is built, not inherited.
THE SCHOLAR: Three guardrails, none negotiable. First, numbers keep their grades: nothing the files mark second-hand gets promoted by handsome prose, and the flag goes in the sentence, not in a footnote. Second, live disputes stay live: mathematics retrieval, growth mindset, the trainability of forecasting. Third, the files' self-corrections are the strongest material we hold; one of them discovered that its most-cited figure did not appear in the paper it had been attributed to, and printed the failure. Lose that and the piece loses its reason to exist.
THE STORYTELLER: Then the corrections are not a footnote. They are the final movement, and the ending is the ethic rather than the toolkit.
THE MENTOR: Structure settled: an overture on the crossover; Movement I, the illusion at three scales; Movement II, the instruments with the boundaries fused on; Movement III, verification as character; a coda that returns to the delayed test.
Resolution. Four movements in continuous prose. [Doc n] citations inline. Disputes flagged where they occur. No technique without its failure condition.
Phase 3: Master Narrative
The Delayed Test
Overture: two exams
Begin with an experiment so plain it could pass for homework. In 2006, the psychologists Roediger and Karpicke had students study a short prose passage. Some read it again and again. Others closed the text and wrote down what they could remember. Five minutes after studying, the rereaders led, 81 percent to 75. One week later the order had inverted: the tested group recalled 56 percent of the material against the rereaders' 42, and in a second experiment the gap widened to 61 against 40, with a detail worth reading twice: the losing group had read the passage about fourteen times, the winning group about three [Doc 1]. Then comes the finding that turns a memory study into a warning. Asked to predict their own future recall, the rereaders were the most confident students in the room, and they finished last [Doc 1].
There are two exams hiding in that experiment. The five-minute exam measures how learning feels. The one-week exam measures what learning is. The two disagree, in a known direction, and the disagreement is silent: nothing in the experience of rereading tells you that the ease you feel is recognition rather than possession.
The six files behind this essay study three subjects that look unrelated: how memory holds complex material, how people judge under uncertainty, and how financial manias inflate and break. Read together, they are one subject. Each examines the gap between the immediate exam and the delayed one, at a different scale: a student with a highlighter, a professional with a forecast, a market with a story it loves. And all six arrive, by different roads, at the same conclusion. The gap cannot be closed by feeling more carefully. It closes only from outside, with instruments that score you later, in writing.
I. One illusion, three costumes
Start with the student, because the student's version has been measured best. Surveyed on their habits, 84 percent of university students named rereading as a primary study strategy; 11 percent reported testing themselves [Doc 5]. This is not laziness. It is trust in an instrument that happens to be broken: the feeling of fluency. When Kornell and Bjork taught people to recognize painters' styles, spaced study beat massed study for 85 percent of participants, yet 83 percent insisted afterward that massing had served them at least as well [Doc 5]. The judges were shown the verdict and voted against it. Fluency is data about the present misread as a promise about the future, and the files give the distinction a name worth keeping for life: performance is what you can measure now; learning can only be measured at a delay [Doc 5].
The professional wears the second costume. Asked for ranges they were 90 percent sure of, a sample of roughly two thousand managers captured the truth somewhere between 40 and 60 percent of the time, a figure the file grades as second-hand and asks you to check before repeating [Doc 2]. Overprecision is the one form of overconfidence that survived scrutiny; the two more famous forms, overrating your score and overrating your rank, flip sign depending on how hard the task is, which means "people are overconfident" was never a finding so much as an average over a sign-flipping effect, as one panel put it [Doc 2, Doc 4]. What makes overprecision interesting rather than merely damning is its exception. Weather forecasters appear in this literature under the word "superb": calibrated almost perfectly, year after year [Doc 4]. Not because meteorology attracts the humble, but because the job is itself an instrument: thousands of forecasts, each resolving within a day, each scored [Doc 4]. Calibration turns out to be a property of a person joined to an environment. Change the environment and the property moves. That reframing may be the most hopeful sentence in all six files, because environments can be built.
The third costume is collective, and here the files spring a trap. You expect the mania literature to complete the pattern: crowds gone mad, tulips traded for houses, folly obvious in hindsight. The files refuse. Checked against price data, the canonical story collapses first: the rare tulip bulbs everyone cites depreciated at 32 percent a year after the peak, which one economist found close to the normal schedule for a horticultural product [Doc 6]. Checked against base rates, the folk rule collapses next: following a real doubling, national markets went on to double again over five years more often, about 26 percent of the time, than they halved, about 15 [Doc 3]. Tulipmania survives not as evidence about markets but as evidence about readers: a story so retellable that four centuries passed it along without an audit, which makes it fluency's largest monument. What actually separates the booms that break from the booms that build is duller and more precise, and it belongs to the next movement. But the market's version of the rereader's confidence deserves its name now: the new-era story, present in every boom, including the booms that turned out to be right [Doc 3, Doc 6].
Before the instruments, close the loop on why they are necessary. Intelligence does not protect: among teachers, more general knowledge predicted more belief in neuromyths, not less [Doc 5]. Motivation does not protect: across 74 experiments, no replicated study made a reasoning error disappear by raising the payment [Doc 2]. Sincerity routes through the same broken gauge as everything else. What protects is structure.
II. The instruments
Everything durable in these files converts into apparatus: an arrangement that forces the delayed test to happen early, and on the record. Each instrument arrives fused to a boundary, and the boundary is not a caveat to skim. It is the part that keeps the instrument from becoming the next poster.
For memory. The first instrument is a blank page. After studying a topic, close everything and reconstruct it, the criteria, the rule, the consequence, in your own hand; then check the reconstruction against the source and mark it. The checking is not garnish: feedback nearly doubles the measured effect, g = 0.73 with it against 0.39 without [Doc 5]. The principle underneath is strange enough to state plainly: a test is not a measurement that leaves the thing measured untouched. A test causes learning [Doc 1, Doc 5]. Two boundaries come fused on. Retrieval works when success is possible; a blank page attempted on material you never encoded is guessing, so a new topic starts with worked examples, not heroics [Doc 1]. And in mathematics the effect is simply not established: the pooled estimate is g = 0.18 with a confidence interval crossing zero, and the reviewing panel closed still divided over whether that marks a true boundary or an artifact of studies that never really demanded retrieval [Doc 5]. Expect the blank page to earn its keep on concepts, criteria, and structure, and to earn less on computation.
The second instrument is a calendar. The spacing literature's central finding is not an interval; it is a function: the best gap between sessions scales with how long you need to remember, and the literature declines, explicitly, to hand you your number [Doc 1, Doc 5]. So the calendar calibrates itself. Touch a topic at day zero, day three, day ten, and let the day-ten failure rate tell you whether your gap has outrun your retention; failure there is not a setback but the reading on the dial [Doc 1]. One consumer note the files deliver with visible satisfaction: expanding-interval schedules performed no better than a uniform rota, a difference of g = 0.034, so a dated sheet of paper is competitive with the subscription [Doc 5].
The third is a shuffle. Build practice sets in which no two consecutive problems take the same treatment, so the first thing each problem demands is a decision about what kind of problem it is. In the strongest classroom trial in the files, a pre-registered randomized experiment across 54 classes, the shuffled group scored 61 percent to the blocked group's 38 on a surprise test a month later, d = 0.83 [Doc 5]. The shuffle feels worse while you do it, which is the point: it moves the discrimination, the choice of tool, off the worksheet's heading and into you. The boundary is sharp. For word-like material, terminology, defined terms, vocabulary, the effect reverses, g = -0.39, so block your terms and shuffle your problems [Doc 5]. Notice that the technique just became a principle: practise the discrimination you will be tested on.
The fourth is a switch. Worked examples are the right instrument for a novice and become dead weight, then drag, as competence arrives; the literature calls this the expertise reversal effect [Doc 1, Doc 5]. The switch has a readable signal: when studying a solution stops producing "I would have done that differently" and produces only "yes, that is what I did," stop reading solutions on that topic [Doc 5]. There is a quiet dignity in this finding. The evidence itself instructs you to outgrow your instruction.
Beneath all four sits the substrate. A single night of lost sleep degraded next-day encoding on the order of a fifth in the imaging study the files cite [Doc 5], so the all-nighter charges you twice: once on the exam it was meant to serve, and again on everything you try to learn the following day. The files' careful version of the sleep advice is deliberately thin, and thin is what survives scrutiny: do not trade sleep for study hours [Doc 1].
For judgment. The first instrument here is a log, and the files are strict about its anatomy: a forecast without a probability, a date, and a score is not a forecast; it is a mood [Doc 2]. Write predictions down with all three, resolve them on schedule, and score them. The observable that tells you the log is real: within a month or two you can name, from your own records, a specific claim you held at 90 percent and got wrong [Doc 4]. If you cannot, your questions are too easy and the log is flattering you. One caution the files insist on: the celebrated claim that an hour of training improves forecasting accuracy is now contested by a reanalysis of the same tournament that produced it, so keep the behaviour that survived, frequent small updates against a score, and hold the advertised effect loosely [Doc 4].
The second is a count. Take any conditional probability, a test's false-positive rate, a fund's odds of beating its index, and rewrite it as people out of 1,000 before you reason about it. Reformatting alone moved Bayesian accuracy by an average of 37 percentage points across four problem structures [Doc 2], and training in the format held for months while formula training decayed within weeks [Doc 4]. Theorists still argue about why. The practice is free either way.
The third is a checklist written before the case. Choose four or five criteria in advance, score them independently, sum, and only then consult your intuition. Mechanical combination beats holistic judgment by about 10 percent, an edge that held across judgment tasks, types of judges, and levels of experience [Doc 2, Doc 4]. Ten percent is not annihilation, and the files preserve the concession in full: the clinicians often matched the model [Doc 2]. But the edge is consistent, and it comes with two boundaries. The rule needs a scorable outcome and a stable environment to exist at all [Doc 2]. And the human problem is not building the model but keeping it: people abandon an algorithm after watching it err, even when they have also watched it outperform them [Doc 4]. The discipline is not the checklist. The discipline is keeping the checklist after it embarrasses you. The files add a diagnostic: if your checklist never disagrees with your gut, it is decorative [Doc 2].
The fourth is a second look. Judge the same case twice, days apart, without consulting your first answer, and measure the spread. Variability between two judgments of identical material is noise by definition, and professionals show far more of it than they or their institutions expect [Doc 2, Doc 4]. You cannot introspect your own variance. You can measure it, and once measured, the remedy is mechanical: average with yourself.
For the world. The instruments scale up. The mania files' central device is a watch file: a handful of credit series maintained by hand, because the one signal that survived their adversarial review is credit growth interacted with asset prices, both extreme together [Doc 3]. What distinguishes this literature from the folklore is that the numbers arrive with their error rates printed on the casing. Joint overheating moves a country's one-year crisis probability from 4 percent to 13 or 14, and its three-year probability to 45 percent for business credit, 37 for household [Doc 3]. Read the failure rates in the same breath: a majority of alarms are followed by no crisis, and roughly a third of crises arrive unannounced, error rates the file computes from the paper's own published figures [Doc 3]. The indicator survives the hindsight objection, because it holds when estimated only on data available at the time, ending before 2008; the file calls that robustness check the single most important sentence it contains [Doc 3]. And the files refuse the easy transfer: a regulator, whose errors are priced asymmetrically, should act on a 45 percent signal; an investor who de-risks on it is wrong more often than right and pays in forgone returns, so for an individual it is a dial for attention and position size, never a date [Doc 3]. A second dial sits deeper: where the leverage sits. Busts financed by credit, housing credit above all, destroy multiples of what equity manias destroy, so severity is legible in advance even when timing never is [Doc 3, Doc 6].
Within a single sector, the files offer a pre-commitment. Crash probability, defined precisely as a 40 percent drawdown within two years, rises from a base rate of 14 percent to 53 after a run-up of 100 percent net of the market, and to 80 percent among the fifteen episodes above 150, a sample the file marks small enough to treat as directional only [Doc 3, Doc 6]. Yet the same study concedes, in its authors' words, that Fama is correct: average forward returns after run-ups are not unusually low [Doc 6]. Both statements are true of one distribution. The mean stays put while the left tail fattens, which is why the instrument is sizing, never shorting. So write a dated note, now, stating what you will do at each threshold, and hear how severely the files score the practice: if you amend the note as the number approaches, the amendment is not a footnote to the experiment. It is the result [Doc 6].
For the stories the world tells, one more instrument: the falsifier, written beside the narrative. New-era stories predict nothing; they attend every boom, including the ones that were right [Doc 3]. But each story names, if you listen closely, the single number on which it depends. In the 1840s it was the traffic takers' estimates, checkable and almost never checked [Doc 6]. In 2000 it was the claim that internet traffic doubled every hundred days, repeated by an executive while measured traffic doubled roughly yearly [Doc 6]. The story is not the signal. The story is the map to the audit.
One instrument remains, and it supervises the others. Before you mark any self-test, score any checklist, resolve any forecast: write down what you expect. Then mark, and keep both columns. This five-second habit converts the fluency illusion from an ambient hazard into a number, your own, per topic, and the files rank it as the highest-leverage practice on their lists for one reason: it is the instrument that tells you whether the rest are working [Doc 1, Doc 5].
III. The keeper of the score
The strangest thing about these six files is not any finding in them. It is their conduct. Every claim carries a grade: verified against retrieved text, reported second-hand, inferred, or remembered and therefore suspect. Every file ends by ranking its own weakest claims, weakest first. One keeps a revision history in which the author records, checkpoint by checkpoint, what verification did to his own confidence, including the entry nobody enjoys writing: the figure he had cited three times, the one carrying more weight than any other in the document, turned out not to appear in the paper he had attributed it to. His note names the condition, secondary-sourced and load-bearing, "the worst combination in this document," and lets the wound stay visible [Doc 2]. Then he permits himself one generalization, flagged as an inference from two instances: the most quotable numbers are the least verifiable ones [Doc 2].
That sentence scales. In the historian's summary that closes one file, whole fields follow the same arc: the research corrects itself in the journals while the paperback keeps the wide claim. Base-rate neglect narrowed into representation dependence; overconfidence narrowed into overprecision; nudging kept its logic and lost an order of magnitude; and in each case the researchers did the honest thing where few readers look [Doc 2]. Meanwhile the learning pyramid, a set of retention percentages neither learning file could trace to any study, has survived roughly sixty years inside teacher training [Doc 1], and belief in learning styles still runs near 90 percent among educators, years after the definitive review found the required evidence absent [Doc 5]. Fabrications persist because they are useful to repeat, and the defence is not brilliance but a habit: ask where a number lives before you let it move you. One panelist compresses this into a prior you can carry anywhere: assume the effect is real, assume it is roughly half the advertised size in your setting, and assume the branded version has already drifted from the tested one [Doc 5].
The literature even supplies a parable for its own practice. For decades the hot hand fallacy, fans seeing streaks where none existed, stood among the field's favourite proofs of human bias. Then two economists showed that the canonical measure itself carried a subtle statistical bias, and once they corrected it, the longstanding conclusions reversed [Doc 2]. The error had lived in the instrument, and the instrument belonged to the researchers; the replications missed it because they reused the measure. The right lesson is not contempt for the field, which caught the error with its own tools and published the reversal. The lesson is the direction of the audit: before diagnosing the world's irrationality, check the measure, especially when the measure is you [Doc 2, Doc 4].
What all of this asks of you is a changed relationship with being wrong. In these files, error is not an event to survive; it is data arriving on schedule. The 90 percent claim that resolves false is your calibration log functioning [Doc 2]. The alarm followed by no crisis is the watch file functioning [Doc 3]. The failed day-ten retrieval is the calendar reporting your true retention [Doc 1]. A person who keeps score on themselves feels worse at five minutes and knows more at a week, which is the crossover from the overture, applied now to character. The files model the endpoint: authors who downgrade their own best numbers in public, and become more credible for it, not less.
Coda: what the late exam selects for
Everything worth having in these pages is scored late. Retention is scored at a week, not at the end of the study session. Calibration is scored at resolution, not at the moment of the confident forecast. Solvency is scored across the credit cycle, not at the top of it. The five-minute exam rewards the reread, the tight interval, the story at the peak. The delayed exam rewards the blank page, the dated probability, the written falsifier. You do not get to choose which exam matters. You choose which one you practise for.
Here is the encouragement, and the files earn it rather than assert it. Not one instrument in this essay requires brilliance. A blank sheet. A dated rota. A log with probabilities and resolution dates. A count out of 1,000. Criteria written before the case. A sentence naming what would prove you wrong. The advantage these confer is not intelligence; it is an earlier acquaintance with the truth. The students who read the passage three times beat the students who read it fourteen because reality was allowed to mark their work while the marking was still cheap [Doc 1]. That is the whole method, at every scale these files reach: invite the delayed test in early, keep its scores in your own handwriting, and let the person who emerges from a decade of that be the one the tests built.
Phase 4: Self-Audit
The densest remaining paragraph. The watch-file paragraph in Movement II asks the reader to hold five probabilities and two error types in one breath, and it reads like a briefing where the rest of the essay reads like prose. The evocative repair is sitting in the source file: Germany in 2007 was nowhere near its own danger zone and was pulled under anyway, through neighbours who were deep inside theirs [Doc 3]. One country's story, told in three sentences before the rates, would let the numbers land as fate rather than arithmetic.
Wisdom lost, with justifications.
- Prospect theory's replicated core, reference dependence and the death of the fixed 2:1 loss-aversion constant [Doc 2, Doc 4], was cut entirely because it describes how choices feel rather than what a practice should be, and the narrative's spine is practice.
- The full moderator structure of retrieval practice, recall format roughly halving the effect, the one-day delay threshold, the thin non-WEIRD evidence base [Doc 5], was compressed to feedback alone, because a narrative can carry one moderator before it becomes a manual, and the manual already exists in the source file.
- The severity half of the mania literature, Canada passing through the same household boom as the United States without a banking crisis, and Lehman's quarter-end dressing of its own reported leverage [Doc 3], was reduced to two clauses, because the file marks the Canadian explanation contested, and a confident compression of a contested case is the exact move this essay argues against.
A reflexive note the audit owes the reader. This essay is itself the operation the files warn about: research packaged into memorable form, which is how moderators get lost. The defence attempted here was fusing every instrument to its boundary inside the same sentence. It is a partial defence. A reader who retains only the metaphors is holding the branded version, and the cure is the one the files prescribe: before acting on any number in this essay, go back to the file, and before citing the file, go back to the paper.