Chapter 59 of 60

The Full AI Incident Investigation

Concepts

CHAPTER 59 β€” The Full AI Incident Investigation

PART X β€” The Debugging AI Playbook

PURPOSE

Staffs the full-budget procedure β€” four separated roles, five scheduled phases, one published packet with probable cause plus a separate contributing-factors list β€” converting a regulator-owed weekend incident into certified prevention with bounded generality.

CENTRAL QUESTION

What is the complete procedure β€” roles, artifacts, schedule, and publication β€” that converts a full incident into certified prevention?

UNIQUE CLAIM

Only this chapter imports the NTSB template wholesale (party-system roles; factual manifest before any analysis; probable cause + separately enumerated contributing factors; blame-independence by cultural choice) fused with STAMP/CAST systemic causation where a missing coordination link (H3) ships its own clause as a control-structure defect.

DEBUGGING OBJECT

Weekend duplicate refunds across hundreds of accounts, two services disagreeing, no shared request ID, guardrail verdicts absent on the disputed edge: H1 contract failure (A sends violate v2 β€” constructed: scope field missing, H1-shaped) vs H2 consumption failure (B deviant on valid input β€” exonerated: protocol-following on violating input) vs H3 coordination/evidence failure (join missing β†’ ownership undecidable β†’ instrumentation repair ships alongside schema v3 + join-key repair).

CONCEPTS INTRODUCED (only genuinely new here)

  • Roles day 0 (commander: schedule/order/publication; custodian: freezes/hashes/fidelity, diagnoses nothing; investigator: hypotheses/trials; prevention owner: clauses/tests/rollout named before causation known; no head-sharing on SEV-1)
  • Phases (1 Assemble days 0–1 manifest-first; 2 Isolate days 1–3 boundary + bisection + minimization Γ—3 verbatim-cited; 3 Repair+calibrate days 3–5 single-variable clauses/tests + joint ledgers; 4 Remediate+publish days 5–7 scoped remediation + packet with labeled INFERRED, named UNKNOWNs/residuals, generality boundary)
  • Packet rule: unlabeled inference anywhere reopens the investigation; “never again” forbidden

CONCEPTS DEVELOPED / REUSED (with source chapter)

  • Full-depth composition of Ch52 records + joins, Ch53 bidirectional tests, Ch54 calibrated clauses, Ch55 joint ledgers, Ch56 live record, Ch57–58 opening artifacts, Ch43 boundary routing, Ch49 hindsight discipline; deepest rung of the Ch57β†’58β†’59 budget ladder

PREREQUISITES

Role roster + evidence manifest with fidelity labels + H1/H2/H3 boundary/join signatures + trial/bisection logs + calibration sets + ledger inputs + draft packet.

LOCAL INVARIANTS

  • Manifest before theory; single-variable interventions with declared trials; every verdict verbatim-cited; every repair bidirectional; every edge calibrated; every repair joint-ledgered; remediation scoped from frozen records; generality bounded to this class/edges/clauses.

FAILURE MODES (this chapter’s specific ones)

  • Role collapse (self-certifying timelines); manifest skipping; blame-first routing (team investigates its own innocence); silent remediation; promise prevention (unversioned “we will add checks”); unbounded schedule; generality smuggling; multi-variable incident repair.

DIAGNOSTIC METHOD (3-6 steps)

  1. Staff four separated roles day zero.
  2. Custodian freezes manifest with fidelity labels before theory.
  3. Investigator isolates H1/H2/H3 by record signature, β‰₯3 trials, verbatim citation.
  4. Prevention owner repairs + calibrates single-variable with bidirectional tests + joint ledgers.
  5. Commander remediates from frozen scope and publishes the packet with bounded generality.

RESEARCH-DERIVED IDEAS (papers/findings with bounds)

  • NTSB Investigative Process (party system; factual report before analysis; probable cause + separately-listed contributing factors) β€” investigations “not conducted for the purpose of determining the rights, liabilities, or blame of any person or entity” (49 CFR Β§ 831.4(c)); reports inadmissible in damages suits (49 U.S.C. Β§ 1154(b)); purpose is recurrence prevention; internal teams lack subpoena/statutory immunity, so no-blame is deliberate culture. Bounds: physical-transport model with regulated recorders. [Statute attribution corrected 2026-09-07 β€” the “no rights/liabilities” rule is 49 CFR 831.4(c), not Β§1154(b).]
  • Leveson, Engineering a Safer World, MIT 2011 (STAMP) β€” accidents = inadequate safety-constraint enforcement across a control structure, not one broken part; H3 is that defect.
  • Leveson, CAST Handbook, MIT PSAS 2019 β€” investigate why constraints/feedback loops were inadequate; CAST/STAMP analyses exist for 737 MAX (Rose 2024), Fukushima (Uesako 2016), 2008 financial crisis (Spencer 2012) β€” all single-cause-framing failures; prevention plural with residuals. Bounds: one of several systemic models (FRAM/Hollnagel, AcciMap/Rasmussen are alternatives); STAMP/CAST-on-AI is early (STPA-for-frontier-AI, arXiv:2506.01782, 2025, still exploratory).
  • Allspaw 2012 + Cook 2000 (via Ch56) β€” no team narrative as evidence; contributing factors stay a list, never collapsed.
  • The published packet IS the book’s single diagnostic-case record (Ch45) at its last stage β€” seven-slot crash dump at capture β†’ Ch1 hypothesis record β†’ Ch24 case file β†’ Ch34 evidence ledger β†’ Ch36 trajectory β†’ Ch49 preregistered ledger β†’ Ch59 published account. Resolves ledger Ch45β†’Ch59.

EXPERIMENT / LAB (actual lab, H-structure)

Lab 59 (PROPOSED): packet drill on one significant past/multi-service staged incident, all four roles staffed (solo readers rotate with role-labeled notes), compressed one-day-per-phase max, manifest before theorizing, β‰₯3 trials where records permit. H1: send-violation at edge ___; H2: valid-input deviant use; H3: unjoinable/ordering-sensitive. Success = auditable packet (manifest, verbatim verdicts, bidirectional repairs, calibrated clauses, ledger deltas, remediation scope, residuals, generality boundary). Thorough narrative alone is not completion.

COMPANION TOOL (name + accepts/can-establish/cannot-establish)

Full Incident Investigation Checklist β€” accepts: roster, manifest + fidelity, H1/H2/H3 + signatures, trial/boundary/bisection logs, repairs + joint predictions, calibration sets, ledger inputs, draft packet. Can establish: whether this incident is fully accounted with certified prevention (incident + edges + clauses only). Cannot establish: cross-incident immunity, threshold permanence, unknown-class completeness; never uses narratives, confidence, agreement, single runs, calm.

PREVENTION ARTIFACT

Published packet (roles, manifest bundles frozen/reconstructed-UNKNOWN + joins, H1/H2/H3 with verbatim records, repairs bidir / Γ—3, clauses v___ + cal sets, ledger Q/L/S deltas, remediated accounts, residuals, generality boundary) + scheduled threshold review.

READER OUTCOME (testable phrasing)

Given one significant incident, reader staffs four roles, freezes a fidelity-labeled manifest before theorizing, isolates H1/H2/H3 with β‰₯3 trials and verbatim cites, verifies each repair bidirectional with calibrated clauses and joint ledgers, and publishes a packet a fresh reader can audit end-to-end.

DEPENDENCIES

Ch52/53/54/55/56/57/58 procedures; Ch43 boundary routing; Ch49 verification bar.

FORWARD BRIDGE

Three budgets, one method (triage ten, isolate sixty, publish in days); Ch60 files every instrument in one workbench with a selection guide and bounds on the label.

EVIDENCE / RESEARCH REQUIREMENTS

Reader’s own retrospective packet; constructed weekend-duplicates only, no measured runs.

ANTI-CLAIMS / LIMITS

One packet covers one incident under one manifest; certifies this class on these edges under these clauses β€” never immunity, permanence, or unknown classes. UNKNOWN wherever fidelity was reconstructed.

Explain this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Apply this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Part X β€” The Debugging AI Playbook

The incident that owes a published record

Chapter 58 ended with isolation and queued artifacts: one defect, one repair candidate, deeper questions deferred. Now the incident that defeats the hour: a weekend of duplicate refunds across hundreds of accounts, two services disagreeing on whose handoff dropped scope, a regulator asking what happened, and leadership asking what prevents recurrence. Triage routed it; the hour isolated one edge; neither suffices. This incident owes customers remediation, the team prevention, and the record an account β€” with named roles producing named artifacts on a visible schedule.

OBSERVATION: duplicate authorizations span many accounts over two days; service A’s logs show schema-valid sends, service B’s logs show scope-less receipts; no shared request ID joins the two; guardrail verdict events are absent on the disputed edge. HYPOTHESIS H1 (contract failure): sends violate the edge schema (A owns the repair). H2 (consumption failure): sends are valid; B’s consumption diverges from protocol (B owns the repair). H3 (coordination/evidence failure): the join is missing so ownership is undecidable from current records β€” the first repair is instrumentation, not assignment. INFERENCE: none yet β€” H1/H2/H3 predict different boundary-and-join signatures and separate only under the full procedure with roles and frozen artifacts.

This chapter’s question: what is the complete procedure β€” roles, artifacts, schedule, and publication β€” that converts a full incident into certified prevention?

Why “thorough heroics” fails first

The obvious move β€” the best engineer investigating exhaustively alone β€” fails because completeness without roles produces evidence nobody trusts. Six defects hide behind solo thoroughness:

  1. Role collapse. Investigator, evidence custodian, communicator, and repair owner in one head. Each duty corrupts the others; the timeline becomes self-certifying.
  2. Artifact improvisation. Each incident invents its own record shapes. Cross-incident comparison dies; the toolkit (Ch60) has nothing uniform to index.
  3. Blame-first routing. Ownership assigned before H1/H2/H3 separate. The assigned team then investigates its own innocence β€” findings discounted in advance.
  4. Silent remediation. Customers remediated without a published account. Trust billed without receipt; the regulator returns.
  5. Prevention by promise. “We will add checks” with no clause versions, calibration sets, or CI links. Promises are not guardrails (Ch54’s rule, enforced here at incident scale).
  6. Unbounded investigation. No schedule, no definition of done. The incident stays open for a quarter, blocking the conveyor with perfectionism.

OPINION: a full investigation without named roles and a publication date is a private education billed as organizational learning. Staff it, schedule it, publish it.

The mental model: the full investigation as the book’s method at full depth β€” every chapter’s instrument played once, by its assigned role, into a published packet. KEPT at this budget (nothing skipped except generality claims): Ch52 records assembled across services with joins repaired; Ch53 conveyor run to bidirectional tests; Ch54 clauses registered and calibrated per committing edge; Ch55 joint ledger with the incident’s cost booked; Ch56 live record audited; Ch57–58 records consumed as opening artifacts. SKIPPED honestly: cross-incident generality (“this can never happen again” is forbidden β€” the packet claims this defect class on these edges under these clauses, nothing more).

The published packet is the book’s single diagnostic-case record (Chapter 45) at its last stage β€” the same object that was a seven-slot crash dump at capture, a hypothesis record in Chapter 1, a case file in Chapter 24, an evidence ledger in Chapter 34, a trajectory in Chapter 36, a preregistered ledger in Chapter 49 β€” now an account with named roles, cited records, and a publication date.

This is the structure of a formal accident investigation. The NTSB separates the factual report (what the recorders show, published first, no analysis) from the analysis and the probable cause finding, states contributing factors as a separate list rather than collapsing everything into one cause, and is barred by statute from apportioning blame or liability β€” its sole purpose is preventing recurrence. Every one of those is a rule this chapter imports: manifest before theory, H1/H2/H3 held plural, blame routed by record signature not advocacy, prevention as the deliverable.

The method: roles, phases, and the published packet

Staff four roles, run five phases on schedule, publish one packet:

  1. Roles (day 0). Incident commander (schedule, order, publication), evidence custodian (freezes, hashes, fidelity labels β€” diagnoses nothing), investigator (hypotheses, interventions, trials), prevention owner (clauses, tests, rollout β€” named before causation is known, so prevention has an author regardless of blame routing). No role shares a head with another on SEV-1.
  2. Phase 1 β€” Assemble (days 0–1). Custodian freezes all bundles, repairs joins (shared IDs backfilled only where logs permit; the rest marked UNKNOWN-fidelity), and publishes the evidence manifest. Nothing theorizes before the manifest.
  3. Phase 2 β€” Isolate (days 1–3). Investigator runs H1/H2/H3 with pre-written signatures across boundary records (Ch43’s routing), environment bisection, and minimization β€” β‰₯3 trials per cell, distributions recorded. Blame routes by record signature, never by terminal position or team advocacy.
  4. Phase 3 β€” Repair and calibrate (days 3–5). Prevention owner registers guardrail clauses per committing edge with calibration sets and thresholds (Ch54), promotes bidirectional regression tests (Ch53), and books the joint cost ledger for each repair (Ch55). Each repair ships single-variable with pre-written joint predictions.
  5. Phase 4 β€” Remediate and publish (days 5–7). Commander oversees customer remediation from the frozen scope list, then publishes the packet: timeline from hashes, H1/H2/H3 verdicts with deciding records cited verbatim, repairs with bidirectional verdicts, clauses with calibration sets, ledger deltas, residual risks named, and the generality boundary stated. INFERRED labeled throughout; UNKNOWNs named, never backfilled.
    flowchart TD
    R["day 0: staff four separated roles β€” commander, evidence custodian (diagnoses nothing), investigator, prevention owner (named before causation is known)"] --> P1["Phase 1 (d0-1): custodian freezes all bundles, repairs joins, publishes the evidence manifest β€” nothing theorizes first"]
    P1 --> P2["Phase 2 (d1-3): investigator runs H1/H2/H3 by RECORD SIGNATURE across boundary records + bisection + minimization, >=3 trials per cell"]
    P2 --> RT{"blame routes by signature, never by terminal position or team advocacy"}
    RT -->|"send violates the edge schema"| H1["H1 contract failure β€” sender-side repair"]
    RT -->|"valid input, deviant consumption"| H2["H2 consumption failure β€” receiver-side repair"]
    RT -->|"join missing / ordering-sensitive"| H3["H3 coordination + evidence failure β€” instrumentation ships as its own clause"]
    H1 --> P3["Phase 3 (d3-5): prevention owner registers calibrated guardrail clauses per committing edge, promotes bidirectional tests, books the joint cost ledger β€” one variable per repair"]
    H2 --> P3
    H3 --> P3
    P3 --> P4["Phase 4 (d5-7): commander oversees remediation from the frozen scope list, then publishes the packet β€” INFERRED labeled, UNKNOWNs named, generality boundary stated"]
  
FULL INVESTIGATION PACKET (published per SEV-1/selected SEV-2):
incident ___ | roles: commander ___ custodian ___ investigator ___ prevention ___
evidence manifest: bundles ___ (frozen ___ / reconstructed-UNKNOWN ___), joins ok/broken ___
H1/H2/H3 verdicts: ___ / ___ / ___ (deciding records cited verbatim ___)
repairs: ___ (pre ___/___ red, post ___/___ green, x3) | clauses: ___ v___ (cal set ___)
ledger: Q___ L___ S___ deltas ___ | remediated accounts ___ | residual risks ___
GENERALITY BOUNDARY: this class, these edges, these clauses. Nothing universal.
RULE: unlabeled inference anywhere in the packet reopens the investigation.

OBSERVATION (constructed illustration, not a measured run): boundary checks show service A sends missing the scope field against schema v2 β€” H1-shaped; service B’s consumption follows its protocol given violating input (H2 exonerated here); the missing join delayed routing by two days (H3 as structural finding with its own instrumentation repair). UPDATED BELIEF: H1 supported for this instance (edge-contract failure, A-side repair); H2 exonerated here; H3 supported as process cause (join repair ships as its own clause with its own calibration). Blame routed, evidence repaired, prevention plural β€” the packet holds all three without merging them.

No team narrative (“service B should have coped”), no executive confidence, no inter-team agreement as verdict, no single confirming run, and no downstream calm (“refunds stopped”) enters the packet as evidence. Records route; prose accounts.

Example: the weekend duplicates, fully worked

The staffed team executes phases without role collapse:

# full procedure: manifest, isolate, repair, publish (roles separated throughout)
manifest = custodian.freeze(services=["A", "B"], window="48h")  # OBSERVATION only
print(manifest.joins)  # MEASUREMENT: A->B join broken, fidelity labeled per bundle
# Investigator: H1/H2/H3 with pre-written boundary signatures, >=3 trials per cell
verdicts = isolate(manifest, hypotheses=[H1_contract, H2_consumption, H3_join],
                   trials=3, cite_verbatim=True)  # blame by signature, never by position
# Prevention owner: one clause + one test per repair, each single-variable, joint ledger
for repair in verdicts.repairs:
    register_clause(repair.edge, calibrate_on=frozen_set, trials=3)  # Ch54 discipline
    promote_test(repair.repro, bidirectional=True, trials=3)  # Ch53 discipline
    book_ledger(repair, quantities=["Q", "L", "S"])  # Ch55 discipline
publish(packet(manifest, verdicts, repairs, ledger, residuals))  # commander; INFERRED labeled

In the constructed case the packet ships on day seven: schema v3 with reject-and-ask on the Aβ†’B edge, the join key repaired across services with backfill-fidelity labels, regression tests bidirectional-green, calibration sets hashed into the clauses, remediation scoped from frozen records, and two residual risks named (unlisted-agent edges, traffic-regime threshold drift). The licensed claim covers this incident’s class on these edges β€” stated on the packet’s final page, where generality goes to be bounded.

Lab 59: packet drill with pre-written routing predictions (proposed)

PROPOSED, not executed: no author-measured results are reported. The evidence this chapter requires is the reader’s own retrospective packet.

Setup. Take one significant past incident (or a multi-service staged failure). Staff all four roles β€” solo readers rotate roles explicitly with role-labeled notes, never merged. The procedure (habitual postmortem vs. this chapter’s phases) is the independent variable; incident and records are controlled.

Task.

  1. Before assembling, write H1/H2/H3 with distinct predicted boundary-or-join signatures and name the deciding record each requires.
  2. Run the phases on a compressed schedule (one day per phase maximum); the custodian’s manifest precedes all theorizing; every cell runs β‰₯3 trials where records permit.
  3. Publish a packet a fresh reader can audit: every verdict traceable to a cited record, every repair bidirectional, every UNKNOWN named.
Hypothesis Predicted packet signature FORECAST OBSERVATION (Γ—3 trials) UPDATED BELIEF
H1 contract send-violation at edge ___ ___ ___ ___ ___ live/exonerated
H2 consumption valid input, deviant use ___ ___ live/exonerated
H3 join/coordination unjoinable or ordering-sensitive ___ ___ live/exonerated

Success criterion. A published packet with manifest, verbatim-cited verdicts, bidirectional repairs, calibrated clauses, ledger deltas, remediation scope, residuals, and a stated generality boundary. A thorough narrative without these artifacts is explicitly not completion.

Companion tool: Full Incident Investigation Checklist

What it accepts: the role roster, the evidence manifest with fidelity labels, H1/H2/H3 with pre-written signatures, trial logs, boundary and bisection records, repair candidates with joint predictions, calibration sets, ledger inputs, and the draft packet. What it performs: it verifies role separation, manifest-before-theory ordering, single-variable interventions with declared trials, verbatim record citation per verdict, bidirectional test proof per repair, calibrated clauses per committing edge, joint ledger per repair, remediation scoping from frozen records, and generality bounding β€” reopening the investigation on any unlabeled inference. What it can establish: whether the incident is fully accounted and its prevention certified β€” for the examined incident, edges, and clauses only. What it cannot establish: cross-incident immunity, threshold permanence across traffic regimes, or completeness over unknown classes. It never treats narratives, confidence, agreement, single runs, or calm as packet evidence. How its output changes your next action: packet-complete routes to publish and schedule threshold review; gaps route to the owning phase with one queued intervention; residuals route to the backlog with named monitors β€” each tracked by ID.

Paper form, sufficient for this chapter:

Incident ___ | roles staffed separately y/n ___ | manifest ___ bundles (UNKNOWN: ___)
H1/H2/H3: ___ / ___ / ___ (records cited ___)
Repairs: ___ (bidir ___/___) | Clauses: ___ (cal ___) | Ledger Q/L/S: ___/___/___
Remediated ___ | Residuals ___ | Generality: this class, these edges. PUBLISHED ___

Where a software implementation does not yet exist in the reader’s stack, this record is the tool. Staff it, schedule it, publish it.

Research lineage: formal accident investigation and systemic causation

The NTSB template. Formal transportation-accident investigation supplies this chapter’s skeleton. The party system assigns distinct organizations and individuals to distinct working groups (structures, powerplants, operations, human factors) rather than one investigator covering everything β€” the four-role separation. The factual report is published before any analysis and contains no probable-cause finding β€” manifest before theory. The final report states a probable cause plus a separately enumerated list of contributing factors β€” H1/H2/H3 held plural, not merged. And the Board’s rules bar its investigations from being “conducted for the purpose of determining the rights, liabilities, or blame of any person or entity” (49 CFR Β§ 831.4(c)), with a companion statute keeping its reports out of damages suits (49 U.S.C. Β§ 1154(b)): the investigation exists to prevent recurrence, which is exactly why this chapter routes blame by record signature and names a prevention owner before causation is known.

Systemic causation (STAMP/CAST). Leveson’s Engineering a Safer World (MIT Press, 2011) argues that in software-intensive systems, accidents rarely reduce to a single broken component β€” they result from inadequate enforcement of safety constraints across a control structure. CAST (Causal Analysis based on System Theory) investigates why the constraints and feedback loops were inadequate, not just which part failed. This is the research basis for H3 as a first-class outcome: “the join was missing so ownership was undecidable” is a control-structure defect, and it ships its own clause with its own calibration rather than being folded into the A-side or B-side repair. CAST has been applied to the 737 MAX, Fukushima, and the 2008 financial crisis β€” all cases where single-cause framing failed.

Blameless analysis, again. The Chapter 56 lineage (Allspaw’s blameless postmortems on Dekker’s Just Culture; Cook on hindsight and root-cause-as-artifact) carries directly into the packet: “no team narrative enters as evidence,” “INFERRED labeled throughout,” and “residuals exist, name them” are the written-artifact form of blamelessness at incident scale.

Bounds: the NTSB model is for physical-transportation accidents with regulated recorders and legal independence; a company’s internal investigation has neither subpoena power nor statutory blame immunity, so the “no blame” property has to be a deliberate cultural choice. STAMP/CAST is a mature framework in safety engineering β€” and one of several systemic accident models (FRAM and AcciMap are alternatives) β€” but its application to AI systems is early, with the first structured attempts (e.g. STPA applied to frontier-AI development, arXiv:2506.01782, 2025) still exploratory. The transferable core: separate factual from analytical, publish the facts first, keep causes plural, and treat missing coordination as a real cause β€” not any specific regulatory apparatus.

Reusable procedure: the full AI incident investigation

  1. Staff β€” four roles, separated, named on day zero.
  2. Assemble β€” custodian’s manifest with fidelity labels before theory.
  3. Isolate β€” H1/H2/H3 by record signature, β‰₯3 trials, verbatim citation.
  4. Repair + calibrate β€” one variable per repair, bidirectional tests, calibrated clauses, joint ledgers.
  5. Remediate + publish β€” scope from frozen records, packet with labeled inference and bounded generality.

Failure modes

  • Role collapse. One head, all duties. Self-certifying timelines.
  • Manifest skipping. Theorizing before freezing. Evidence shaped to fit the theory.
  • Blame-first routing. Assignment before separation. Investigations of innocence, findings discounted.
  • Promise prevention. “We will add checks” without versions, sets, or CI links. Prose as guardrail.
  • Unbounded schedule. No definition of done. Perfectionism starving the conveyor.
  • Generality smuggling. “Never again” in the packet. The forbidden claim; residuals exist, name them.
  • Silent remediation. Accounts fixed, nothing published. Trust unbilled, regulators unmoved.
  • Multi-variable incident repair. Edges, prompts, and thresholds in one deploy. Attribution destroyed at scale.

Limits, per contract: one packet covers one incident under one evidence manifest; it certifies this class on these edges under these clauses β€” never immunity, never permanence across regimes, never unknown classes. UNKNOWN wherever fidelity was reconstructed rather than frozen.

References

  • National Transportation Safety Board. The Investigative Process. https://www.ntsb.gov/investigations/process/Pages/default.aspx (party system; factual report vs. analysis vs. probable cause). See 49 CFR Β§ 831.4(c) β€” investigations not conducted to determine rights, liabilities, or blame β€” and 49 U.S.C. Β§ 1154(b) β€” Board reports inadmissible in damages suits.
  • Nancy Leveson. Engineering a Safer World: Systems Thinking Applied to Safety. MIT Press, 2011. https://doi.org/10.7551/mitpress/8179.001.0001 (STAMP accident model; safety constraints and control structure).
  • Nancy Leveson. CAST Handbook: How to Learn More from Incidents and Accidents. MIT PSAS, 2019. http://sunnyday.mit.edu/CAST-Handbook.pdf
  • John Allspaw. Blameless PostMortems and a Just Culture. Etsy Code as Craft, 2012. (Cross-ref Ch 56 β€” carried into the packet’s evidence rules.)
  • Richard Cook. How Complex Systems Fail. University of Chicago, 2000. (Cross-ref Ch 56 β€” root-cause-as-artifact; why contributing factors stay a list.)

Debugging Checklist

  • Four roles staffed and separated (roster published)?
  • Evidence manifest frozen before any theorizing (fidelity labeled)?
  • H1/H2/H3 pre-written with distinct boundary/join signatures?
  • All cells run β‰₯3 trials with distributions recorded?
  • Every verdict cites its deciding record verbatim?
  • Each repair verified bidirectional (red pre, green post)?
  • Clauses registered with calibration sets + thresholds as setup choices?
  • Joint ledger (Q/L/S) booked per repair?
  • Remediation scoped from frozen records (not estimates)?
  • Packet published with labeled inference, named UNKNOWNs, bounded generality?
  • No narrative, confidence, agreement, single runs, or calm cited as evidence?

What This Chapter Established

  • The full AI incident investigation: four roles, five phases, one published packet β€” demonstrated on the constructed weekend duplicate-refund incident, no measured runs claimed.
  • The contract/consumption/coordination separation (H1/H2/H3) with verbatim citation and the manifest-before-theory ordering rule.
  • Lab 59 as a proposed packet drill the reader executes; the Full Incident Investigation Checklist contract (accepts/performs/can-establish/cannot-establish/next-action).
  • What was NOT proved: any cross-incident immunity, any threshold permanence, or any unknown-class completeness. One incident fully accounted; nothing universal.
  • Position in the arc: Chapters 57–58 compress the method to ten and sixty minutes with honest skips; this chapter spends the full budget with nothing skipped but generality. The playbook’s deepest procedure, bounded on its final page.
  • Research grounding: the NTSB template (party system, factual report before analysis, probable cause + separate contributing factors, investigations not conducted to determine rights/liabilities/blame β€” 49 CFR Β§ 831.4(c)) and systemic causation (Leveson STAMP/CAST, one of several systemic models) as the basis for H3 as a first-class control-structure cause; the published packet is the Chapter 45 diagnostic-case record at its last stage.

Next

Three budgets, one method β€” triage in ten, isolate in sixty, account fully in days. Each procedure consumed the instruments of Parts I–IX without restating them, which raises the book’s final question: what exactly has the practitioner accumulated across sixty chapters, how is it selected under pressure, and what did it actually prove? Chapter 60, “The Debugging AI Toolkit,” inventories the whole kit with a selection guide; what the book established β€” and what it left provisional β€” is its chapter’s to establish, not this one’s.