← Jev From First Principles

Programming with Decisions

On the evidence, what is a decision model, and is decision a programming primitive?

Programming with Decisions

This chapter is a verdict over the ledger, not an experiment: one mechanical extractor (run_ch35.py) reads the metadata preregistrations and writes results/ch35.jsonl; every number below is from that file or from a cited ledger row. No model ran. The hypothesis classification follows the author’s floor for Session E: no hypothesis is decided above INSUFFICIENT_EVIDENCE without a real measured provider.

The question

On the evidence, what is a decision model, and is decision a programming primitive? The book must answer from what it actually measured, and the answer is allowed to be narrower than the investigation set out to build.

What the evidence is

run_ch35.py reads all thirty-five chapters’ metadata. The mechanical snapshot (results/ch35.jsonl, summary row):

1. mechanical snapshot (metadata only)
   chapters 35; with result files 25
   verdict counts {'INSUFFICIENT_EVIDENCE': 12, 'PARTIALLY_SUPPORTED': 9, 'PARTIALLY_REFUTED': 3, 'REFUTED': 1, 'SUPPORTED': 3, 'FULLY_SUPPORTED': 1}
2. hypothesis verdicts (all INSUFFICIENT_EVIDENCE)
   H1 INSUFFICIENT_EVIDENCE: no trained decision model was measured (Ch7/15/28 deferred)
   H2 INSUFFICIENT_EVIDENCE: bug-class value shown by construction/deterministic tests only (Ch32)
   H3 INSUFFICIENT_EVIDENCE: transfer experiment that would decide it never ran (Ch15)
   H4 INSUFFICIENT_EVIDENCE: no decision family on a measured Pareto front (Ch28 compile-only)
   L1 INSUFFICIENT_EVIDENCE: for decide reduced to library forms (Ch23); survivors' discipline unmeasured
   L2 INSUFFICIENT_EVIDENCE: fail-silently-under-shift behavioural claim unmeasured (Ch17)
   L3 INSUFFICIENT_EVIDENCE: compiler behaviour within supplied profiles only (Ch33)
3. prediction hit-rate framing
   chapter-level experiments with a recorded verdict: 17
   hypothesis-level verdicts above INSUFFICIENT_EVIDENCE: 7 -> 0 by the author's floor

The import and the three step blocks that produce it are in examples/ch35-programming-with-decisions/walkthrough_ch35.py, shown verbatim below; the output above is byte-identical to its live run.

import yaml
from pathlib import Path
import sys
    # 1. the mechanical snapshot
    print("1. mechanical snapshot (metadata only)")
    print(f"   chapters {len(rows)}; with result files "
          f"{sum(1 for r in rows if r['has_results'])}")
    print(f"   verdict counts {dict(counts)}")
    # 2. the hypothesis verdicts, under the author's floor
    print("2. hypothesis verdicts (all INSUFFICIENT_EVIDENCE)")
    for name, note in HYPOTHESES:
        print(f"   {name} INSUFFICIENT_EVIDENCE: {note}")
    # 3. the hit-rate framing: verdict counts, not one number
    decided = sum(v for k, v in counts.items()
                  if k in ("SUPPORTED", "PARTIALLY_SUPPORTED",
                           "REFUTED", "PARTIALLY_REFUTED", "FULLY_SUPPORTED"))
    print("3. prediction hit-rate framing")
    print(f"   chapter-level experiments with a recorded verdict: {decided}")
    print(f"   hypothesis-level verdicts above INSUFFICIENT_EVIDENCE: "
          f"{sum(1 for _ in HYPOTHESES)} -> 0 by the author's floor")

What the book actually measured breaks into three kinds:

  1. Real providers, measured — the tuned TF-IDF+LR on 77-way intent (0.8779, Ch4), the zero-shot embedding provider (Ch5), NLI (Ch6), the calibrated and shifted provider rows (Ch10), and the Part VI CPU encoders and cross-encoder on SciFact (Ch21-25). These are OBSERVED.
  2. Deterministic model-free logic — Ch16-19, 22, 23, 31-34: construct semantics, validators, routers, the compiler, the semantic program, all over supplied scores or hand-built stubs. These are ARITHMETIC/mode: simulation/mode: analytic and are never evidence about a model.
  3. Deferred or compile-only — Ch7’s model sweep, Ch13-15’s GPU runs, Ch28’s winner table (17 cells from earlier rows; decision-model cells NOT_OBSERVED). These are PENDING_RUN or inventory.

The hypothesis verdicts

Every hypothesis is classified INSUFFICIENT_EVIDENCE. The reason is the same across all seven: the hypothesis needs a real measured provider (a decision model, an LLM judge, a cascade on real traffic), and the book never measured one — the GPU runs were deferred, and the router/compiler/program chapters ran on supplied or stubbed numbers. The verdict table (full rows in results/ch35.jsonl):

H Verdict Deciding chapters What would have decided it
H1 nothing new INSUFFICIENT_EVIDENCE 7, 15, 28 a trained decision model beating NLI/first-token/classifier on held-out decisions, 3 seeds
H2 useful abstraction INSUFFICIENT_EVIDENCE 8, 16, 17, 32 measured bug class + measured provider swap without caller changes
H3 general decision capability INSUFFICIENT_EVIDENCE 5, 15 transfer: k diverse tasks > k relabelings of one task
H4 compromise region INSUFFICIENT_EVIDENCE 28 a decision family on a measured Pareto front
L1 syntax adds nothing INSUFFICIENT_EVIDENCE 16, 17, 18, 19, 22, 23 any surviving construct’s discipline measurably preventing failure
L2 uncertainty must be enforced INSUFFICIENT_EVIDENCE 16, 17, 32 enforced vs optional handling differing behaviourally under shift
L3 provider selection is compiler’s job INSUFFICIENT_EVIDENCE 33 compiler plans matching real hand-built cascades on real profiles

The chapter-level rows that did decide something are recorded as chapter verdicts: Ch23 REFUTED for decide as a construct (its library forms reproduce it exactly — ledger 23.6), Ch21/22/24 PARTIALLY_REFUTED their own pipeline predictions, and Ch26, 33, 34 SUPPORTED theirs (graph replay; the compiler within its profiles; the semantic program’s boundary). Those verdicts are about constructed or simulated systems and are reported as such; they are not hypothesis verdicts.

Prediction-versus-outcome

The preregistrations predicted results chapter by chapter; the hit rate is reported as the verdict distribution, not a single number, because the distribution is the honest summary: of the 34 investigated chapters, 12 stand at INSUFFICIENT_EVIDENCE (design drafts whose runs never happened), 17 recorded a chapter-level verdict (9 PARTIALLY_SUPPORTED, 3 SUPPORTED/FULLY_SUPPORTED, 3 PARTIALLY_REFUTED, 1 REFUTED), and every row is mechanical from the metadata (results/ch35.jsonl; Chapter 35 is the summary itself and is not counted). The book was surprised where it tested: Ch22’s risk-controlled decision filter covered 2.6% of threshold pairs and reached end-to-end 0.202 (a P3 refutation), Ch23’s construct was rejected, Ch24’s product rule was refuted by its own arithmetic, Ch30’s cascade guarantee did not compose, Ch33’s cascade was Pareto-dominated. Where the book only designed, it was not surprised — because it did not test.

Protocol deviations are part of the record: Ch21’s run predated its preregistration edit (disclosed in that chapter and its metadata); Ch28 selected best-of-config on test (disclosed — an inventory, not a leaderboard); Ch34’s sampling seed deviated from its prereg placeholder (disclosed). Each was reported, none was hidden; that is the extent of what the ledger would have been accused of and what the audit below checks.

The audit (leakage, underspecification, weak baselines)

Three papers frame the self-audit: Kapoor & Narayanan (2207.07048) on data leakage; D’Amour et al. (2011.03395) on underspecification; Lipton & Steinhardt (1807.03341, reused from Ch4) on scholarship trends. The audit’s conclusions:

  1. Leakage — the strong item is the SciFact dev-set reuse. The same 188 labelled SciFact dev claims served as the test split in Chapters 21, 22, 23 and 25. Each chapter touched it once and counted the touch, but across Part VI it behaved as a development set: later designs (which filter, which threshold, the Part VI pivot) were informed by earlier results on it. What that does to the Part VI claims: they are pipeline results on one corpus, positively selected by a process that saw the target; they are NOT evidence about Jev or H1-H4 (the chapters already say this), and their internal comparisons (filter vs filter, cascade vs single) must be read as within-corpus, not cross-corpus. A second, weaker item: most label and wording variants were authored by one writer with the chapter’s claims in view (Ch5 reports the 13-29 point wording swings, which is the mitigation, not the cure).
  2. Underspecification — present and measured. D’Amour’s point is that predictors equal on held-out behave differently deployed; the book’s strongest analogue is Ch10: calibration that held in-distribution degraded under shift for every measured provider, and conformal coverage fell under shift (0.9088/0.9107/0.9166 → 0.5687/0.5611/0.8550, ledger 11.9/11.10). Equal-in-distribution does not mean equal-deployed; the book says so where it measured it.
  3. Weak baselines — engaged, not always fixed. Ch4’s majority bar is the boring baseline; Ch28’s table is honest about NOT_OBSERVED cells; but the fact that 15 of 17 filled Ch28 cells reprint Ch4/Ch5/Ch6 rows means the “where the tiny classifier wins” section is thin, and the LLM-judge and decision-model sections rest on NOT_OBSERVED — the audit says plainly the headline comparison the table invites is not supported by it.

What would not survive the audit: every Part VI number read as evidence about decision models (they are pipeline results); every Ch29-33 number read as a finding about real routers, cascades or compilers (they are consequences of supplied, mostly assumed profiles); and any sentence that turned a probability into a certainty — the ledger rows the audit script checks are the guard against that. The mechanical check is python examples/audit_claims.py, which verifies that every ledger row citing a results/chNN.jsonl selector points at an existing row (Section 7 of the finalisation).

The ten questions

  1. Is a “decision model” distinct from known categories? INSUFFICIENT_EVIDENCE — no decision model was measured; the categories’ competition was only partially probed (Ch6: NLI lost to embeddings).
  2. Is “decision” a useful abstraction anyway? Argued and built (typed answers, mandatory uncertainty, provider independence; Ch8, 16-19, 32); as a measured claim, INSUFFICIENT_EVIDENCE — the bug-class value is construction/deterministic evidence (Ch32’s 5/5).
  3. First-class constructs or library calls? Mixed, and decided construct by construct: Ch23 rejected for decide (library form); if decide, match decide, where decide survived as library compositions with enforced semantics.
  4. Which syntax survived? decide (Ch16), if decide (Ch17), match decide (Ch18), where decide (Ch22, typed and abstaining — LOTUS sem_filter, narrower). Survival required tested code, mandatory arms/ties, and honest prior-art placement.
  5. Which construct was rejected, by which rule? for decide (Ch23): its three candidate differences (lazy order, early stop, ranked iteration) each had an exact library equivalent (filter, takewhile/any, sorted); equivalence was measured over supplied scores; rejection is a result, not a failure.
  6. Which provider per decision family? As measured: tuned LR for closed intent sets (0.8779, Ch4); embeddings for unseen labels (Ch5); the cross-encoder for relevance over the Part VI corpus — and nothing beyond that. Ch28’s table is an inventory with honest NOT_OBSERVED cells.
  7. What role do embeddings play? Candidate generation and cheap filtering (Part VI): retrieval over-recalls, decisions filter (Ch21-22); retrieval error is separate from decision error and each is reported separately.
  8. What role do generative LLMs play? Never measured here. The router/cascade/chapters (Ch28-33) treat them as supplied-profile entries only; the 2.6%-coverage risk-filter result (Ch22) is the closest the book came to an LLM-free operating-point lesson.
  9. How should uncertainty propagate? By construction: abstention is a first-class outcome (Ch11), match ties resolve to CONFLICTING_EVIDENCE (Ch18), the product rule does not compose (Ch24), cascades’ risk control does not compose either (Ch30) — propagation must carry distributions, not point wins; unmeasured as a behavioural claim (L2).
  10. What exists now that did not exist before? See the narrowest description below.

Four lists

What Jev gave us — the provocation and the terminology (decision models, typed decisions); the reason the investigation existed. Nothing measured here confirms or refutes its claims.

What existing ML already provided — classification (Ch4 LR bar), zero-shot embeddings (Ch5), NLI (Ch6), calibration maps (Ch10), conformal prediction (Ch10), BM25 and cross-encoders (Part VI), contextual bandits (Li et al. §4, read in full: the audit sample’s unbiasedness needs random logged traffic), cost-based optimisation (Selinger 1979; Palimpzest), declarative pipelines (Ludwig, Cedar, DSPy, LOTUS, DocETL, Palimpzest), agent loops (ReAct, PAL, Toolformer).

What this book discovered — measured: calibration degrades under shift for every measured provider; conformal coverage falls under shift; zero-shot wording swings results 13-29 points; the risk-controlled decision filter covered 2.6% of threshold pairs; the escalator cascade’s guarantee does not compose (naive 0.577 vs measured 0.830); the Part VI filter choice matters more than the syntax. Constructed: the Arbiter library (contract → types → control → match → abstain → calibration → retrieval → pipeline → graph → router → cascade → spec → compiler → semantic program), the Chapter 26 replay graph (provenance, why/where, drift by state hash), the decision-spec validator (5/5 seeded bug classes), the semantic program with its permit/deny/escalate boundary and 0 free-text decisions against an analytical ReAct count. Mechanical: a prediction-versus-outcome table every chapter contributed to.

What remains speculation — every H1-H4 and L1-L3 verdict; every Ch29-33 number as a statement about real systems; the shared-hidden-state hypothesis (Ch13, 27); the compression thesis (Ch14-15); hosted Jev’s behaviour; the judge/cascade/decision-model columns of the Ch28 table.

The SciFact disclosure, stated plainly

The same 188 labelled SciFact dev claims served as the test split in Chapters 21, 22, 23 and 25. Each chapter touched it once and counted the touch, but across Part VI it behaved as a development set: later designs were informed by earlier results on it, so the Part VI comparisons are within-corpus and positively selected on the target. They are pipeline-level results about retrieval and filtering on one corpus; they are not evidence about Jev or about H1-H4, and the strength of the Part VI claims is correspondingly the strength of a single-corpus development-set experiment — stated here so Chapter 35 cannot be read otherwise.

Constructs: survived and rejected

Construct Verdict Ledger
decide survived 16.x
if decide survived (mandatory arms) 17.7
match decide survived (mandatory uncertain arm, tie rule) 18.x
where decide survived as the typed, abstaining filter 22.8 (LOTUS sem_filter, no novelty)
for decide rejected — library filter/takewhile/sorted 23.6

The narrowest defensible description

On the evidence of this book, the thing that was built is:

A tested library of typed-decision primitives — an answer contract with mandatory uncertainty handling, abstention, calibration against shift, decision graphs with replayable provenance, a declarative decision spec with a validator, a cost-based compiler over supplied profiles, and a semantic program with a permit/deny/escalate runtime boundary — together with measured CPU results showing that calibration and coverage degrade under shift, that zero-shot wording swings accuracy, and that on one small corpus a typed filtering decision beats no filtering on cost at a small accuracy price, while everything claimed about decision models remains undecided for lack of a measured one.

That is narrower than the book set out to build, and it is what the ledger supports. The book’s one rule-protecting result stands on its own: for decide reduced to library forms, and the surviving constructs earned their place with tested code, not with syntax.

Wrong / Correct. Wrong: “The book showed how to program with decisions.” Correct: “The book showed how to program with typed decision objects, and was honest that the question the title asks — whether decision deserves to be a programming primitive in the Jev sense — is left INSUFFICIENT_EVIDENCE: the syntax that survived is the part the tests paid for, and the constructs that need a real decision model to justify them are the ones the book could not measure.”

Close by

What did we build that did not exist when the investigation began? A replayable decision graph with why/where provenance, a spec validator that catches bugs a library call cannot, and a semantic program with a permit/deny/escalate boundary running on the surviving construct set — all tested, all model-free, and all narrower than the question they were built to answer.

Limitations

  • Everything above is a summary of prior chapters; no new measurement was made this chapter beyond the mechanical extractor (results/ch35.jsonl).
  • The hypothesis verdicts follow the author’s Session E floor; a reader who counts construction and deterministic tests as evidence would move H2 and L1 partially up, and the chapter says so rather than hiding the judgement.
  • D’Amour (2011.03395) and Kapoor (2207.07048) are read at abstract level; Lipton (1807.03341) was read in Ch4 (§1, §5.1). The audit uses their framing, not their full taxonomies.
  • The prediction-versus-outcome table is a verdict distribution, not a single hit-rate number; the distribution is the honest summary and is stated as such.
  • Assumptions inherited from earlier chapters were re-checked during the finalisation audit (evidence/book-audit.md); any check that failed is listed there.