← Jev From First Principles

Decision Graphs

What must be recorded so a graph of decisions can be inspected, replayed and audited?

Decision Graphs

Design draft: this chapter builds the graph machinery and runs a fake-provider replay study on CPU. No model runs; every provider is a deterministic or seeded fake. The numbers measure the machinery, not any model.

The problem

Chapter 25 showed a pipeline whose every error is attributable to one stage. But a list of stages is a weak container for that attribution: it cannot say why a verdict depended on some evidence and not other evidence, and it cannot be replayed after the fact to check whether the world has moved.

This chapter asks: what must be recorded so a graph of decisions can be inspected, replayed and audited?

The answer shapes what counts as evidence everywhere else in the book: Chapter 35’s verdicts cite ledger rows; a ledger row you cannot replay is a claim you cannot check.

What we expect and why

Our starting hypothesis is that replays disagree for sampled providers and agree for deterministic ones, and drift is detectable by state hash.

Three papers frame it:

  1. Besta et al. (2023) — GoT treats LLM thoughts as vertices of a graph with dependency edges, enabling synergies and feedback loops. The graph is the structure, over model calls rather than over typed decisions. (Read at abstract level.)

  2. Opsahl-Ong et al. (2024) — MIPRO optimises multi-stage LM programs by treating the graph of modules as the unit of optimisation. A graph of decisions is therefore something you optimise, not just record. (Read at abstract level.)

  3. Buneman, Khanna and Tan (2001) — the classic provenance vocabulary: why-provenance is the set of source data that contributed to an output; where-provenance is the set of locations the output was copied from. (Verified via Springer DOI 10.1007/3-540-44503-X_20, ICDT 2001, pp. 316-330; a classic, quoted only at the abstract/summary level.)

The build

src/arbiter/graph.py defines a GraphRecord — question, answer set, decision, provider and seed, calibration record id, upstream dependencies, consumed state fields, a state hash, a timestamp, and whether the provider is deterministic — and a DecisionGraph that records, persists, replays and explains them.

from arbiter.graph import (
    DecisionGraph,
    categorise,
    deterministic_fake,
    replay,
    run_and_record,
)
from arbiter.retrieve import mrr, rank
    # 1. Record a three-node chain with fake providers.
    print("1. record a decision graph")
    g = DecisionGraph()
    state: dict = {}
    prev = "root"
    for i in range(3):
        rec = run_and_record(
            node_id=f"n{i}", kind="decision", question=f"q{i}",
            answer_set=("a", "b", "c"), provider=deterministic_fake,
            provider_name="deterministic", seed=None, state=state,
            evidence=(prev,), consumed_fields=tuple(sorted(state.keys())),
            deterministic=True,
        )
        g.add_record(rec)
        prev = f"n{i}"
    for nid in g.topo_order():
        r = g.records[nid]
        print(f"   {nid}: {r.question} -> {r.decision} hash={r.state_hash}")
    # 2. Why and where provenance.
    print("2. provenance")
    why = g.explain("n2", mode="why")
    where = g.explain("n2", mode="where")
    print(f"   why  of n2 -> {why['why']}")
    print(f"   where of n2 -> {where['where']}")
    # 3. Replay with the stored seeds: full agreement.
    print("3. replay")
    out = replay(g, lambda rec: deterministic_fake)
    print(f"   decision agreement: {sum(o.decision_agreed for o in out)}/{len(out)}")
    print(f"   categories: {categorise(out)}")

The walkthrough prints:

1. record a decision graph
   n0: q0 -> c hash=fd2079a3096d5abb
   n1: q1 -> b hash=f2e766f5a819b5ee
   n2: q2 -> a hash=a19ec48d047e48c1
2. provenance
   why  of n2 -> ['n0', 'n1']
   where of n2 -> ['n0', 'n1']
3. replay
   decision agreement: 3/3
   categories: {'no_change': 3}

The replay study

python examples/ch26-decision-graphs/run_ch26.py records 50 graphs of five nodes each (deterministic and seeded-sampled fake providers) and replays them two ways: with the stored seeds, and with the sampled nodes’ seeds changed by one. results/ch26.jsonl:

Quantity Value
replay agreement, stored seeds (250 nodes) 1.000
replay agreement, changed seeds (250 nodes) 0.684
agreement loss, changed seeds (paired bootstrap) −0.316 [−0.376, −0.264]
nodes whose state hash drifted 0.696
graphs with at least one drifted node 48 / 50
mean explain time 0.002 ms

What surprised us

  1. Deterministic nodes survive replay; sampled nodes do not, by the design of the fakes. With the stored seed, every node reproduces (1.000). Change the seed on a sampled node and its decision flips 79% of the time, dragging the agreement down to 0.684. The riddle “is this decision replayable?” is answered by one field: deterministic.

  2. Drift detection is broader than decision disagreement. 0.696 of nodes show a changed state hash, more than the 0.316 whose own decision changed. A changed decision upstream changes the hash of every downstream node, even one whose own decision is unchanged. That is exactly what an auditor wants: a trace of nodes touched by a drift, not just the one node where it started.

  3. The hash is over the decision sequence, not over provider internals. The first version of the hash included the provider name and seed in the node entry, which made replay structurally unequal for no semantic reason; the current hash is over the ordered decisions, so it detects decision drift and nothing else. This is the “what must be recorded” question answered precisely: enough to reproduce the decisions, and a checksum over those decisions.

  4. Explain is the cheap direction. Recovering the ancestors (why) and consumed fields (where) of any output is microseconds on a stored graph. The expensive direction is replay, which re-runs providers; explain reads the record.

  5. Buneman’s split applies cleanly. why of node n2 is its transitive dependencies; where is the state fields it consumed. For a chain both are the same list, but in a diamond — two routes into one node — why would carry both branches while where would carry only what the node actually read. The vocabulary is exactly the one needed to audit a decision.

Wrong / Correct. Wrong: “A decision log is a record of what was returned.” Correct: “A decision record is a replayable unit: the question, the provider and its seed, the dependencies, the consumed state, a checksum and a timestamp. A log you cannot replay is a log you cannot audit.”

The auditable-artifact contract

schemas/decision-record.schema.json is the JSON Schema a record must satisfy: question, answer_set, decision, provider, provider_seed, calibration, evidence, consumed_fields, state_hash (a 16-hex sha256), timestamp, deterministic. tools/replay.py replays a stored graph and prints the disagreement categories. Together they make a decision an artifact you can check.

What to carry forward

What is the minimum a decision must record to count as evidence? On this build: the question and answer set (so the decision is meaningful), the provider and seed (so it can be reproduced), the dependencies and consumed fields (so why/where can be recovered), and a checksum over the decision sequence (so drift is detectable). Chapter 34 composes these graphs into a program whose side effects are gated by decisions, and replays the whole graph from Chapter 26’s records.

Close by

What is the minimum a decision must record to count as evidence? The record above. The next chapter pauses the composition thread to ask what the per-decision evidence costs when decisions are made over time.

Limitations

  • The replay study uses fake providers only; the agreement rates are properties of those fakes, not of real providers. A real, sampled provider’s replay behaviour is PENDING_RUN.
  • GoT and MIPRO were read in full text after drafting (Session R), which found no contradiction; Buneman et al. 2001 is verified via DOI and cited at the summary level (a classic, not on arXiv).
  • The state hash is over the decision sequence; it detects decision drift, not internal score drift within a provider.
  • The graph is an in-memory/JSONL structure, not a database; Chapter 27’s accounting compares storage costs.