Chapter 49 of 60

How Do You Know the Diagnosis Is Right?

Concepts

CHAPTER 49 โ€” How Do You Know the Diagnosis Is Right?

PART VIII โ€” Building the AI Debugger

PURPOSE

Sets the four-part verification bar (prediction-match ledger, independent evidence, reversal-or-cap, human sign-off) that promotes a last-branch-standing into a SIGNED diagnosis โ€” or honestly caps it at PROVISIONAL.

CENTRAL QUESTION

What verification bar turns a surviving hypothesis into a signed diagnosis โ€” and what evidence counts as independent enough to clear it?

UNIQUE CLAIM

Only this chapter defines verification as trial-with-transcript: winner matches all verification predictions while every loser mismatches โ‰ฅ1 with cited deciding runs, plus a second-stream confirmation and a ร—3 reversal (or explicit INFEASIBLE cap) โ€” with the dated ledger as preregistration against HARKing.

DEBUGGING OBJECT

Stale-snapshot H1 as survivor for refund-quote-041: constructed SIGNED ledger P1 absent-list + P2 k-expansion 0/3 + P3 wording 0/3 + P4 fresh re-index 3/3 by independent operator + P5 stale-reintroduction 3/3 return; contrasted with reversal-infeasible production-permission case capped PROVISIONAL (“treat as cause for prevention; don’t cite externally”).

CONCEPTS INTRODUCED (only genuinely new here)

  • Four-part bar: ledger (P1โ€“Pn dated match/mismatch) + independence (different log/probe/operator from discovery) + reversal-or-cap + impact-graded human sign-off on raw records
  • Multi-cause H3 as live hypothesis (staleness + chunking/filter jointly) with partial-return signature; relief-is-not-proof rule

CONCEPTS DEVELOPED / REUSED (with source chapter)

  • Ch41 two-direction intervention vs single-direction relief; Ch47โ€“48 predictions/trials ledgerized; Ch46 I-4 post-dating policed; Cook/Allspaw hindsight critique operationalized as PROVISIONAL

PREREQUISITES

Hypothesis space with dated predictions + execution/trial logs + independence-stream pointer + impact class.

LOCAL INVARIANTS

  • Ledger every prediction to its outcome; exonerate every loser with its deciding run; confirm via a stream unused in discovery; reverse ร—3 or justify infeasibility aloud; humans sign high-impact on hashes/verbatim/trial counts โ€” assistants never self-sign.

FAILURE MODES (this chapter’s specific ones)

  • Relief-as-proof (“fixed, therefore diagnosed”); self-confirming same-stream verification; undated matches; skipped reversal kept silent; assistant-signed verdicts; generality creep (“verified here” โ†’ everywhere).

DIAGNOSTIC METHOD (3-6 steps)

  1. Ledger dated predictions; match each to outcome; cite loser exonerations.
  2. Confirm from an independent stream/operator re-running the bundle.
  3. Reverse the convicted condition ร—3 (or record INFEASIBLE + cap at PROVISIONAL).
  4. Human sign-off by impact; write license scope (this bundle, these versions).

RESEARCH-DERIVED IDEAS (papers/findings with bounds)

  • Nosek et al., Preregistration Revolution, PNAS 2018 โ€” discovery evidence is “used up”; confirming on it tells nothing new; dated ledger = preregistration against postdiction-as-prediction.
  • Kerr, HARKing, PSPR 1998 โ€” names prediction-after-outcome; sub-types CHARKing (construct a story to fit the result โ€” the incident-doc common one) / RHARKing (retrieve) / SHARKing (suppress).
  • Gelman & Loken, Garden of Forking Paths, 2013 โ€” analysis choices contingent on the data (which trials, which metric, which threshold) inflate false positives even with an honest pre-stated hypothesis and no p-hacking โ€” so the ledger dates the whole plan (observable + trial count + match criterion), not just the hypothesis.
  • Cook, How Complex Systems Fail, 2000 โ€” catastrophe needs multiple contributors; rarely one isolated root cause; backs H3 + PROVISIONAL honesty.
  • Allspaw, Blameless PostMortems, 2012 โ€” single “root cause” usually hindsight bias; counterfactual “if only X” describes a non-world.
  • Leveson, Engineering a Safer World (STAMP), 2011 โ€” accidents as inadequate constraint enforcement by the control structure, not component-failure chains โ€” the formal version of “find the root cause is the wrong frame”; CAST analysis method โ†’ Ch59.

EXPERIMENT / LAB (actual lab, H-structure)

Lab 49 (PROPOSED): ledger one past “root cause”. H1: all four parts present (SIGNED stands); H2: independence/reversal missing (downgrade PROVISIONAL); H3: loser exoneration uncited (UNSUPPORTED). Reconstruct P1โ€“Pn with dates, name independence stream, attempt/justify reversal, sign. Re-asserted cause without ledger is not completion.

COMPANION TOOL (name + accepts/can-establish/cannot-establish)

Diagnosis Validation Protocol Runner โ€” accepts: space + dated predictions, trial logs, independence pointer, impact class. Can establish: whether this case meets the bar (this case only). Cannot establish: cross-case generality, future stability, absence of unlisted causes (Ch47’s job); never uses explanations, scores, correlations, single runs, agreement, relief.

PREVENTION ARTIFACT

Ledger verdict SIGNED / PROVISIONAL (missing part named) / UNSUPPORTED with sign-off (name + date); SIGNED โ†’ Part IX artifact; PROVISIONAL ships with cap written on it.

READER OUTCOME (testable phrasing)

Given one past diagnosis, reader produces a dated P1โ€“Pn match table, cites each loser’s deciding run, names the independence stream, attempts or justifies-no reversal, and records a signed SIGNED/PROVISIONAL/UNSUPPORTED verdict with scope.

DEPENDENCIES

Ch47โ€“48 (space + runs); Ch41 (two-direction intervention); Ch46 (I-2/I-6 independence); Ch1 ladder (ladder rung clarified by reversal); Ch45 (the prediction-match ledger IS the Ch45 one-diagnostic-case record with hypothesis-and-verification fields filled โ€” same object, later stage; audit ยง9/ยง269); Ch59 (CAST โ€” the incident-analysis method STAMP feeds).

FORWARD BRIDGE

Case skill โ‰  benchmark skill; Ch50 designs AIDebugBench โ€” tasks, ledger rubrics, contamination controls โ€” reporting zero scores because none were run.

EVIDENCE / RESEARCH REQUIREMENTS

Reader’s own ledger with second-reader independence check; constructed SIGNED + PROVISIONAL pair only, no measured runs.

ANTI-CLAIMS / LIMITS

One verdict covers one case under one bundle/version set; later rolls/re-indexes/config changes bound the license until re-verified. UNKNOWN wherever predictions post-date outcomes or independence absent. No generality, no exemption from re-verification.

Explain this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Apply this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Part VIII โ€” Building the AI Debugger

The diagnosis that survived everything except verification

Chapter 48 split the space to one survivor โ€” but survival is not proof. A practitioner holds H1 (stale snapshot) as the last branch standing for the fabricated citation: the chunk was absent, expansion changed nothing, wording swaps changed nothing. She ships the re-index, the symptom disappears for a week, and the postmortem records “root cause: stale index.” A month later the fabrication returns with a fresh index. The original evidence never distinguished “stale snapshot caused it” from “stale snapshot accompanied it” โ€” no prediction-match ledger, no independent verification evidence, no human sign-off on the causal claim. The branch survived; the diagnosis was never verified.

OBSERVATION: H1 survived elimination; the re-index coincided with symptom relief once; no pre-registered verification predictions, no independent evidence stream, no reversal test recorded. HYPOTHESIS H1 (causal): the stale snapshot caused the fabrication. H2 (accompaniment): staleness correlated while another factor (e.g., chunking rule, ingestion filter) caused it. H3 (multi-cause): staleness plus a second factor jointly required. INFERENCE: none yet โ€” H1/H2/H3 predict different reversal and independence signatures: H1 predicts reintroducing staleness reintroduces the symptom; H2 predicts it does not; H3 predicts partial return.

This chapter’s question: what verification bar turns a surviving hypothesis into a signed diagnosis โ€” and what evidence counts as independent enough to clear it?

Why “it worked after the fix” fails first

The obvious move โ€” citing post-fix relief as proof โ€” fails because relief is a downstream symptom, and downstream symptoms never diagnose (book-wide hard rule). Five insufficiencies hide behind relief:

  1. Coincident relief. The re-index shipped with a config refresh, a restart, and a prompt tweak. Which intervention caused relief is unassigned โ€” multi-variable disposal proves nothing.
  2. Single-run confirmation. One good week cited as verification. Nondeterministic systems need repeated confirmation runs plus a reversal (reintroduce โ†’ symptom returns โ†’ remove โ†’ relief returns) before causal language.
  3. Dependent re-evidence. “Verifying” with the same log, same assistant, same metric that produced the hypothesis. Verification evidence must be independent of discovery evidence โ€” different stream, different tool, or different operator.
  4. No ledger. Predictions scattered across chat, runs, and memory. Without a prediction-match ledger (prediction โ†’ outcome โ†’ match/mismatch, dated), nobody can audit what was actually confirmed. This is the postdiction-as-prediction error: Nosek and colleagues frame it as the core threat to credibility โ€” generating a hypothesis from existing observations and then presenting it as if it had been tested against new ones (Nosek et al., 2018). Kerr named the specific move HARKing โ€” hypothesizing after the results are known โ€” and its sub-types; the one that dominates incident docs is constructing a “root cause” story to fit the relief (Kerr, 1998). But HARKing is about the hypothesis; Gelman and Loken showed the same false “verified” can come from the analysis choices made after seeing the runs โ€” which trials to count, which metric, which match threshold โ€” even with an honest pre-stated hypothesis and no p-hacking (Gelman & Loken, 2013). So the dated ledger is a preregistration of the whole plan: not just the hypothesis, but the observable, the trial count, and the match criterion. It makes it impossible to pretend a post-hoc match was foreseen โ€” or a post-hoc analysis path.
  5. Unsigned high-impact claims. Causal verdicts affecting production, refunds, or access shipped on assistant authority. High-impact diagnoses require human verification against raw records โ€” always, including here.

OPINION: most “root causes” in incident docs are last-hypotheses-standing wearing verification costumes. A diagnosis without a ledger is a rumor with a timestamp.

The incident-analysis field has been saying a version of this for years. Cook’s How Complex Systems Fail argues that catastrophe in a complex system requires multiple contributing failures, so there is rarely an isolated root cause to find (Cook, 2000). Allspaw’s Blameless PostMortems and Dekker’s Field Guide to Understanding ‘Human Error’ both warn that naming a single “root cause” after the fact is usually hindsight bias โ€” the counterfactual “if only X” describes a world that did not exist (Allspaw, 2012; Dekker, 2014). Leveson’s STAMP formalizes the point: a complex failure comes from inadequate enforcement of constraints across the whole control structure, not one broken component, so “find the root cause” is often the wrong frame (Leveson, 2011); Chapter 59 develops its analysis method, CAST. That is precisely why this chapter carries H3 (multi-cause) as a live hypothesis and demands a reversal test: to distinguish “stale snapshot was one condition among several” from “stale snapshot was the cause.”

The mental model: verification is a trial with a transcript. The diagnosis is the charge; the ledger is the transcript; independent evidence is the second witness; the human sign-off is the verdict. Surviving elimination gets the case to court. Only prediction-matches on independent evidence, including a reversal where feasible, convict. That ledger is not a new artifact โ€” it is the Chapter 45 diagnostic-case record with the hypothesis-and-verification fields filled: same object, later stage.

The method: the verification bar in four parts

A signed diagnosis requires all four, recorded in the ledger:

  1. Prediction-match ledger. Every live hypothesis’s pre-dated predictions listed with outcomes and match/mismatch verdicts. Minimum to convict: the winner matches on all its verification predictions; every loser mismatches on at least one (exoneration cited to its deciding run).
  2. Independence of verification evidence. At least one confirming observation from a stream not used to generate the hypothesis: different log, different probe, different operator re-running from the bundle, or the assistant’s proposal confirmed by human-executed runs. Same-stream confirmation is corroboration theater.
  3. Reversal test (where feasible). Reintroduce the convicted condition and observe symptom return (ร—3 trials where nondeterministic), then remove and observe relief. Where reversal is unsafe or impossible (production harm, destructive state), record the infeasibility explicitly and cap the verdict at PROVISIONAL โ€” never upgrade by enthusiasm.
  4. Human sign-off by impact. Low-impact (notebook note, draft config): practitioner sign-off on the ledger suffices. High-impact (production change, customer correction, access/refund action): a human verifies the deciding records raw (hashes, verbatim lines, trial counts) and signs name + date. No assistant self-sign-off, ever.
    flowchart TD
    W["surviving hypothesis from Ch48 elimination"] --> P{"every live hypothesis's PRE-DATED predictions matched to outcomes?"}
    P -->|"no / post-dated"| UNS["UNSUPPORTED โ€” reopen the space, not the production change"]
    P -->|yes| L{"every loser exonerated with its deciding run cited?"}
    L -->|no| UNS
    L -->|yes| I{">=1 confirming observation from an INDEPENDENT stream / operator / tool?"}
    I -->|no| PROV["PROVISIONAL โ€” same-stream confirmation is theater; name the missing part"]
    I -->|yes| RV{"reversal: reintroduce the condition -> symptom returns x3, then remove -> relief?"}
    RV -->|"infeasible / unsafe"| PROV
    RV -->|"done, matches x3"| SO{"impact class?"}
    SO -->|low| SIGN["SIGNED โ€” practitioner sign-off on the ledger"]
    SO -->|high| HV["human verifies deciding records raw (hashes, verbatim, trial counts), signs name + date"]
    HV --> SIGN
  
PREDICTION-MATCH LEDGER (worked sketch; constructed, not measured):
Case refund-quote-041. Winner H1 stale-snapshot; losers H2/H3.
P1 (H1): gold absent from ranked list ............ OBSERVED absent โ€” MATCH
P2 (H1): k-expansion changes nothing (ร—3) ........ 0/3 changed โ€” MATCH
P3 (H1): wording-swap changes nothing (ร—3) ....... 0/3 changed โ€” MATCH
P4 (H1, verification): fresh re-index restores
    citation ร—3, INDEPENDENT operator reruns ..... 3/3 โ€” MATCH (indep. stream)
P5 (H1, reversal): reintroduce stale snapshot โ†’
    fabrication returns (ร—3) ..................... 3/3 โ€” MATCH
H2 exonerated at P1/P2. H3 exonerated at P3.
INDEPENDENCE: P4 by second operator from bundle alone. SIGN-OFF: ___ (human).
RULE: all P1โ€“P5 match + independence + reversal = SIGNED; any missing
part caps at PROVISIONAL (record which).

OBSERVATION (constructed illustration, not a measured run): the ledger above matches on all five with independent confirmation and reversal. UPDATED BELIEF: H1 verified for this instance under this bundle and version set; H2/H3 exonerated here โ€” licensed claim only, no generality; without P4/P5 the same table would read PROVISIONAL, not SIGNED.

No model explanation of why the fix worked, no confidence score on the diagnosis, no score deltas, no correlation statistics alone, no single confirming run, no second assistant’s agreement, and no downstream quiet (“tickets dropped”) substitutes for ledger matches on independent evidence with reversal. Matches, witnesses, and signatures decide.

Example: capping an honest PROVISIONAL

Contrast the constructed SIGNED case with a reversal-infeasible one: the convicted condition is a production permission escalation that cannot be reintroduced safely. The practitioner records P1โ€“P4 matches with independent confirmation, marks P5 INFEASIBLE with the reason (“reintroducing grants live access”), and signs PROVISIONAL with the cap stated: “treat as cause for prevention purposes; do not cite as proven cause in external reports.” The cap is not failure โ€” it is the method working. PROVISIONAL with a named missing part beats SIGNED on theater.

Research lineage: preregister the prediction, reverse the condition

Independence of verification evidence is the prediction/postdiction line. Nosek’s distinction is not just about timing โ€” it is that a hypothesis built from a body of evidence has already “used up” that evidence, and confirming it against the same evidence tells you nothing new (Nosek et al., 2018). The requirement for a second stream, a second operator, or human-executed runs confirming an assistant’s proposal is that principle operationalized.

The reversal test is an intervention, not a re-observation. Reintroducing the convicted condition and watching the symptom return is Causal Testing (Chapters 1 and 41) and Zeller’s cause-effect-chain logic: a difference is promoted to a cause only when changing it changes the outcome, in both directions where feasible. Post-fix relief is a single-direction, single-trial observation; the reversal makes it a two-direction, repeated intervention.

PROVISIONAL is the honest verdict when the complex system won’t let you reverse. Cook and Allspaw’s point cuts both ways: sometimes you genuinely cannot isolate one cause because the failure had several, and sometimes you cannot safely reverse the one you found. Both cases end at PROVISIONAL with the reason named โ€” which is the method working, not failing.

Lab 49: ledger one of your own past diagnoses (proposed)

PROPOSED, not executed: no author-measured results are reported. The evidence this chapter requires is the reader’s own ledger.

Setup. Take one past diagnosis you recorded as “root cause” (any domain). The review regime (memory/impression vs. this chapter’s four-part ledger) is the independent variable; the case materials are controlled. A second reader checks independence.

Task.

  1. Before opening materials, write H1/H2/H3 about your diagnosis with distinct predicted ledger outcomes: H1: “ledger shows all four parts present (SIGNED stands)”; H2: “ledger shows missing independence or reversal (downgrade to PROVISIONAL)”; H3: “ledger shows a loser’s exoneration uncited (verdict UNSUPPORTED).”
  2. Reconstruct P1โ€“Pn with dates (or mark post-dated), identify the independence stream (or its absence), attempt/justify-no reversal, and obtain sign-off.
  3. Record the final verdict: SIGNED / PROVISIONAL (missing part named) / UNSUPPORTED.
Hypothesis Predicted ledger signature FORECAST OBSERVATION UPDATED BELIEF
H1 stands signed all four parts present ___ ___ live/exonerated
H2 downgraded independence/reversal missing ___ ___ live/exonerated
H3 unsupported exoneration uncited ___ ___ live/exonerated

Success criterion. A filled ledger with dated predictions, independence stream named, reversal attempted or infeasibility justified, and a signed verdict with caps stated. A re-asserted “root cause” without the ledger is explicitly not completion.

Companion tool: Diagnosis Validation Protocol Runner

What it accepts: the hypothesis space with dated predictions, the execution/trial logs, the independence-stream pointer, and the impact class. What it performs: it matches predictions to outcomes row by row, checks that every loser’s exoneration cites its deciding run, verifies the independence stream differs from discovery, evaluates reversal (done / infeasible-with-reason), and emits SIGNED / PROVISIONAL / UNSUPPORTED with missing parts named. What it can establish: whether a diagnosis meets the verification bar โ€” for the examined case only. What it cannot establish: cross-case generality, future stability under version change, or the absence of unlisted causes (coverage belonged to Ch47). It never treats explanations, scores, correlations, single runs, agreement, or symptom relief as verification. How its output changes your next action: SIGNED โ†’ convert to prevention artifact (Part IX); PROVISIONAL โ†’ ship prevention with the cap written on it; UNSUPPORTED โ†’ reopen the space, never the production change.

Paper form, sufficient for this chapter:

Case ___ Bundle ___ Impact: low / high
LEDGER (prediction โ†’ outcome โ†’ match?): P1 ___ P2 ___ P3 ___ P4 ___ P5 ___
LOSERS exonerated (run cited): ___  INDEPENDENCE stream: ___
REVERSAL: done (ร—3 ___) / INFEASIBLE (reason ___)  SIGN-OFF (human, date): ___
VERDICT: SIGNED / PROVISIONAL (missing ___) / UNSUPPORTED

Where a software implementation does not yet exist in the reader’s stack, this record is the tool. No ledger, no license to claim cause.

Reusable procedure: verify every diagnosis the same way

  1. Ledger predictions โ€” dated before runs; match each to its outcome.
  2. Exonerate losers โ€” every non-winner cites its deciding run.
  3. Confirm independently โ€” a second stream, tool, or operator re-runs from the bundle.
  4. Reverse where feasible โ€” reintroduce ร—3 or justify infeasibility and cap.
  5. Sign by impact โ€” human on raw records for high-impact; name and date recorded.

Failure modes

  • Relief-as-proof. “Fixed, therefore diagnosed.” Relief is a symptom; ledgers are proof.
  • Self-confirming streams. Verifying with the discovery evidence. Independence or theater โ€” no third option.
  • Undated predictions. Matches claimed without timestamps. Post-dated matches are inadmissible.
  • Skipped reversal, silent. No reversal, no infeasibility note, full causal language anyway. Missing parts cap verdicts aloud.
  • Assistant-signed verdicts. The proposer signing its own case. Humans sign; assistants propose.
  • Generality creep. “Verified here” slides into “verified everywhere.” Licenses cover the examined bundle and version set โ€” nothing wider.

Limits, per contract: one verdict covers one case under one bundle/version set; it warrants the causal claim stated, not neighboring claims; any later version roll, re-index, or config change bounds the license until re-verified. UNKNOWN wherever predictions post-date outcomes or independence is absent.

References

Debugging Checklist

  • Every prediction dated before its run, matched to an outcome?
  • Winner matches all verification predictions; losers exonerated with cited runs?
  • At least one confirming observation from an independent stream/operator?
  • Reversal done ร—3, or infeasibility justified with a PROVISIONAL cap?
  • No forbidden sources (explanation/scores/correlation/single-run/agreement/symptom) as verification?
  • Verdict recorded SIGNED / PROVISIONAL / UNSUPPORTED with missing parts named?
  • Human sign-off on raw records for high-impact (name + date)?
  • License scope written (this bundle, these versions, no wider)?

What This Chapter Established

  • The verification bar: prediction-match ledgers, independence of verification evidence, reversal-or-cap, and human sign-off by impact โ€” demonstrated on the constructed SIGNED ledger plus the honest-PROVISIONAL contrast, no measured runs claimed.
  • The trial-with-transcript mental model and the relief-is-not-proof rule.
  • Lab 49 as a proposed past-diagnosis re-ledger the reader executes; the Diagnosis Validation Protocol Runner contract (accepts/performs/can-establish/cannot-establish/next-action).
  • What was NOT proved: any cause beyond the examined instance, any generality across versions, or any exemption from re-verification after change. One bar set; nothing universal.
  • Research grounding: “prediction-after-outcome” is postdiction-as-prediction / HARKing (Nosek et al.; Kerr โ€” the incident-doc common sub-type is constructing the root-cause story), and analysis choices made after seeing the data are a second route to a false “verified” even with an honest hypothesis (garden of forking paths โ€” Gelman & Loken), so the dated ledger preregisters the whole plan โ€” observable, trial count, match criterion; independence of verification evidence is the prediction/postdiction line โ€” confirming against the discovery evidence tells you nothing new; the reversal test is a two-direction intervention (Causal Testing / Zeller); and PROVISIONAL is the honest verdict when a complex failure has several contributors (Cook; Allspaw; formally, STAMP โ€” Leveson) or the condition cannot be safely reversed.

Next

Individual diagnoses can now be verified โ€” but a debugger that verifies one case well may still fail the next fifty. Case skill and benchmark skill are different claims, and only the second can be stated quantitatively. Chapter 50, “AIDebugBench,” designs the benchmark that would measure debugging skill across task families โ€” its tasks, rubrics, and contamination controls โ€” while reporting no scores, because none have been run.