Decisions About Decisions
How does uncertainty propagate when one decision consumes another?
Decisions About Decisions
Design draft: this chapter argues from the literature and from arithmetic. It runs no provider. Every number below is a consequence of a constructed contingency table or an assumed conditional table, and the assumptions are swept and stated next to the results. None of it is evidence about any model. The provider experiment is deferred.
The problem
Chapter 16 kept decide a pure function: it takes a state, asks a provider, returns a typed outcome. It declined to say how one decision feeds another. Chapter 22 showed why that matters: a filter decides which candidates a later decision reads, and its errors become the later decision’s errors.
This chapter asks: can the confidence of a chain of decisions be computed from its parts?
The naive answer is the product rule. Two decisions each right 90% of the time give a chain right 0.9 × 0.9 = 0.81 of the time. The question is when that arithmetic is the right model and when it is a fiction.
What we expect and why
Our starting hypothesis is that errors are positively correlated, so the product rule is wrong, and passing evidence beats passing the answer.
Three papers frame it:
-
Ross et al. (2010) — DAgger studies sequential prediction where future observations depend on previous predictions, which “violate the common i.i.d. assumptions made in statistical learning”. Compounding error is the expected failure. (Read at abstract level.)
-
Wang et al. (2022) — self-consistency samples many reasoning paths and picks the most consistent answer, using agreement as a confidence signal. Agreement is one way to estimate a stage’s confidence without a calibrated head. (Read at abstract level.)
-
Huang et al. (2023) — LLMs struggle to self-correct reasoning without external feedback, and at times performance degrades after self-correction. A downstream stage cannot rescue an upstream error by itself. (Read at abstract level.)
The first two motivate the chain; the third warns that the useful direction of information is forward.
The assumption, stated first
The product rule is not a theorem. It is the assumption of independence: it says the events “A is correct” and “B is correct” are statistically independent. Under any dependence it is simply wrong, and the size of the error is a function of the dependence.
The arithmetic lives in src/arbiter/chain.py. It works on a constructed 2×2 table: of n items, a that A got right, b that B got right, and both that both got right. The marginals and the joint determine everything; the product rule asserts both = a·b/n.
from arbiter.chain import (
ChainTable,
ConditionalTable,
chain_table_from_conditional,
ece,
fully_correlated_chain_accuracy,
independent_chain_accuracy,
joint_range,
phi_coefficient,
)
N, A, B = 100, 90, 90
# 1. The product rule is exact only at independence.
print("1. the product rule vs the exact joint (n=100, A correct 90, B correct 90)")
lo, hi = joint_range(N, A, B)
print(f" valid both-correct range: [{lo}, {hi}]; the product assumes {A * B / N:.0f}")
for bc in (lo, 81, hi):
t = ChainTable(N, A, B, bc)
print(f" both_correct={bc}: joint={t.p_joint:.2f} product={t.product_prediction:.2f} "
f"error={t.product_error:+.3f} phi={phi_coefficient(t.p_a, t.p_b, t.p_joint):+.3f}")
# 2. Longer chains: independent product vs fully correlated joint.
print("2. chain length")
for accs in ([0.9, 0.9], [0.9, 0.9, 0.9]):
ind = independent_chain_accuracy(accs)
corr = fully_correlated_chain_accuracy(accs)
print(f" k={len(accs)}: independent={ind:.3f} fully_correlated={corr:.3f} gap={corr - ind:+.3f}")
# 3. What B receives (assumed conditional table).
print("3. passing the argmax vs the evidence (ASSUMED q, q')")
modes = (
("argmax", ConditionalTable(q=0.92, q_prime=0.2)),
("evidence", ConditionalTable(q=0.86, q_prime=0.7)),
)
for name, cond in modes:
t = chain_table_from_conditional(N, A, cond)
print(f" {name:<9} q={cond.q} q'={cond.q_prime} dep={cond.dependency:.2f} "
f"joint={t.p_joint:.2f} product={t.product_prediction:.2f} error={t.product_error:+.3f}")
# 4. The chain score is not calibrated on constructed data.
print("4. the product score is not calibrated (LABELLED hand example)")
scores = [0.81, 0.81, 0.36, 0.36, 0.81, 0.36]
correct = [True, True, False, False, True, False]
print(f" ECE(product score)={ece(scores, correct, n_bins=2):.3f}")
The walkthrough prints:
1. the product rule vs the exact joint (n=100, A correct 90, B correct 90)
valid both-correct range: [80, 90]; the product assumes 81
both_correct=80: joint=0.80 product=0.81 error=+0.010 phi=-0.111
both_correct=81: joint=0.81 product=0.81 error=+0.000 phi=+0.000
both_correct=90: joint=0.90 product=0.81 error=-0.090 phi=+1.000
2. chain length
k=2: independent=0.810 fully_correlated=0.900 gap=+0.090
k=3: independent=0.729 fully_correlated=0.900 gap=+0.171
3. passing the argmax vs the evidence (ASSUMED q, q')
argmax q=0.92 q'=0.2 dep=0.72 joint=0.83 product=0.77 error=-0.065
evidence q=0.86 q'=0.7 dep=0.16 joint=0.77 product=0.76 error=-0.014
4. the product score is not calibrated (LABELLED hand example)
ECE(product score)=0.275
The result, with the assumption next to it
run_ch24.py sweeps the constructed parameters and writes results/ch24.jsonl. Every row carries a mode: constructed_arithmetic label. Read the numbers as statements about the table, not about any provider.
Sweep over the joint (marginals 0.90, 0.90; only both_correct varies):
| both_correct | joint | product | product error | phi |
|---|---|---|---|---|
| 80 | 0.80 | 0.81 | +0.010 | −0.111 |
| 81 | 0.81 | 0.81 | +0.000 | +0.000 |
| 90 | 0.90 | 0.81 | −0.090 | +1.000 |
The product is exact at phi = 0 and wrong everywhere else. At perfect correlation (phi = +1) it understates the joint by 0.09; at the antithetic end it overstates it by 0.01. The full sweep is in the results file; the error is monotone in the correlation.
Chain length (each stage assumed 0.90): the independent product is 0.810 for two stages and 0.729 for three; the fully-correlated joint is 0.900 for both. The gap grows with the number of stages (0.090, then 0.171). This is the one place a “compounding” story shows up — but only because the fully-correlated model is the alternative, and both are assumptions.
Passing modes (assumed conditional accuracies q = P(B right | A right), q' = P(B right | A wrong)):
| B receives | q | q' | dependency | joint | product | error |
|---|---|---|---|---|---|---|
| argmax | 0.92 | 0.20 | 0.72 | 0.828 | 0.763 | −0.065 |
| distribution | 0.90 | 0.45 | 0.45 | 0.810 | 0.770 | −0.040 |
| evidence | 0.86 | 0.70 | 0.16 | 0.774 | 0.760 | −0.014 |
What surprised us
-
The preregistered direction was wrong. We predicted that under positive correlation the product underestimates the failure rate (i.e. overestimates accuracy). The arithmetic says the opposite: positive correlation of the outcomes pushes the joint up (toward
min(p_a, p_b)), so the product understates the joint and overstates the failure rate. The product is conservative under positive correlation, not optimistic. P2 is refuted, and the refutation is the finding. -
Passing the evidence does not give the higher joint — it gives the lower dependency. P4 is partly refuted. The evidence mode has the smallest dependency (0.16) but the lowest joint (0.774), because B re-decides on its own and no longer benefits from A being right. Passing the argmax gives the highest joint (0.828) and the worst dependency (0.72). That is a trade-off, not a ranking: an argmax chain is better when A is usually right and worse when A is not, and the chapter’s assumed numbers put A at 0.90. This “refutation” is of a claim about our own assumed tables, not about any provider. With other values of
qandq'the ranking of the three modes changes, so what the result really shows is that the prediction was not stated precisely enough to bear on real systems. Whether passing the evidence helps is an empirical question for a provider run, and it stays open. -
The dependency sweep isolates the mechanism. Fixing
q = 0.92and raisingq'from 0.20 to 0.92 moves the product error from −0.065 to +0.000: as B becomes robust to A’s errors, the product rule becomes exact. The whole effect is the dependency between the stages, nothing else. -
The product score is not calibrated. On a constructed stream where the naive product is the chain score, the out-of-sample ECE is 0.132 (the hand example in the walkthrough gives 0.275 on six points). Fitting a monotone map on one half drops it to 0.005 on the other. The product is a number, not a probability.
The controller interface Chapter 16 declined
A chain needs something Chapter 16 did not specify: a controller that decides what a stage passes forward. This chapter’s three modes are three controller policies — pass the argmax, pass the distribution, pass the evidence. The interface is small and typed:
- a stage returns
(value, score, evidence); - the controller chooses which of the three the next stage receives;
- the choice is where the dependency comes from, and it is a design decision, not a property of
decide.
Naming it makes the earlier omission concrete: decide is a function; a chain is a function plus a controller. Chapter 26 (decision graphs) generalises the controller to a graph of stages.
Wrong / Correct. Wrong: “The chain’s confidence is the product of the stages’ confidences.” Correct: “The product is the independence assumption. The chain’s joint is
p_A · P(B right | A right), and the controlling quantity is the dependency between the stages. The product is exact only when that dependency is zero.”
The distinction this chapter keeps
A score is not a probability. The product of two calibrated confidences is not a calibrated confidence for the chain; it needs its own fitted map and its own calibration split (as Chapter 10 argued for single decisions).
Arithmetic is not evidence. The sweep shows what follows if a given correlation holds. Measuring the correlation between two real stages needs a provider run, which is deferred. The arithmetic constrains the design; it does not report on a model.
What to carry forward
What is the right type for a chained decision’s confidence? On this evidence it is not a product of the parts: it is a joint that depends on a conditional structure the language must make explicit. Chapter 25 builds a pipeline where the stages are typed and the controller is visible; Chapter 26 replays a graph of such stages.
Close by
What is the right type for a chained decision’s confidence? A joint over the chain conditioned on how each stage consumes the last — a type the product rule cannot express.
Limitations
- No provider was run. This chapter is arithmetic on constructed tables.
results/ch24.jsonlis labelledmode: constructed_arithmetic; the recalibration arm is labelledsimulated: true. Nothing here is OBSERVED evidence about a model. - The correlation is assumed, not measured. The sweep shows the consequence of each value; the value itself for any real pair of decisions requires a provider run (deferred, PENDING_RUN in the metadata).
- The passing-mode conditional tables are assumed.
(q, q')for argmax, distribution and evidence are modelling choices, not measurements; the chapter’s mode ranking would change with different values. - The recalibration arm uses constructed data. Its ECE figures are properties of the simulation.
- The three required papers were read in full text after drafting (Session R), which found no contradiction with the chapter’s claims.
- The controller interface is introduced by name but not implemented as a language feature here; Chapter 26 generalises it.