Semantic Programs
What does a program with deterministic and semantic computation together look like, end to end?
Semantic Programs
This chapter is model-free: the program runs over stub providers with hand-supplied numbers (labelled ILLUSTRATIVE) and deterministic retrieval (BM25, Chapter 21). The generation step is a stub. The ReAct-style comparison is a count over a declared task script — no agent is run (no LLM is allowed by the book’s rules).
The problem
The book has spent sixteen chapters building pieces: decide, if decide,
match decide, where decide, the spec, the compiler, the graph. This chapter
asks the question the pieces were for: what does a program with deterministic
and semantic computation together look like, end to end?
What we expect and why
Our hypothesis: on a fixed claim-research task, the semantic program makes the same decisions as a ReAct-style loop but with every decision typed and countable and a replayable graph — and it gives up flexibility to get that. Three papers frame it:
- Yao et al. (2022), ReAct (2210.03629) — interleaves free-text reasoning traces with tool actions; on HotpotQA and FEVER it mitigates hallucination by querying a Wikipedia API. The challenge: every “Thought” is an untyped, free-text decision. (Read at abstract level.)
- Gao et al. (2022), PAL (2211.10435) — the model generates a program and the interpreter computes; it was FAMOUSLY strong on GSM8K. The method this chapter copies in miniature: delegate the deterministic work to code, keep the semantic decisions typed. (Read at abstract level.)
- Schick et al. (2023), Toolformer (2302.04761) — models learn when to call tools; tool use is itself a decision. The support this chapter needs: deciding to act is a decision, and making it typed is the whole point. (Read at abstract level.)
The paper that weakens the expectation is ReAct itself: its flexibility (the agent rewrites its plan from free text mid-run) is exactly what the semantic program gives up. The chapter’s comparison is therefore not “which is more accurate” — it is what the free text costs, counted as untyped decisions.
The build
src/arbiter/semantic_program.py runs the program: load a hand-built corpus of
nine documents, retrieve the top 3 with BM25, decide relevance for each
(where decide), decide relation for each kept document (match decide, with
the tie rule from Chapter 18: a near-tie resolves to CONFLICTING_EVIDENCE),
generate a [stub] summary, aggregate a support verdict, and gate the side
effect with a permit/deny/escalate boundary. Every decision is recorded in a
Chapter 26 decision graph; the graph is replayed from disk.
from arbiter.graph import categorise
from arbiter.semantic_program import research_claim, replay_seeded
# 1. research a claim: retrieve, where decide, match decide, support, act
r = research_claim("compound Q reduced inflammatory markers", seed=3)
print("1. claim A")
print(f" retrieved {r['retrieved']} kept {r['kept']}")
print(f" verdict {r['verdict']} p={r['p_support']} -> {r['action']} "
f"({r['reason']})")
print(f" typed decisions {r['typed_decisions']}, free-text {r['free_text_decisions']}")
# 2. the runtime boundary across the three cases
print("2. runtime boundary")
for tag, claim in (("B", "the alpine study recorded temperature ranges"),
("C", "small pilot study reported modest drop")):
x = research_claim(claim, seed=3)
print(f" {tag}: {x['verdict']} p={x['p_support']} -> {x['action']}")
# 3. replay from the Chapter 26 decision graph
print("3. replay")
print(f" same seed 3: {categorise(replay_seeded(r, 3))}")
print(f" changed seed 8: {categorise(replay_seeded(r, 8))}")
# 4. free-text decision count (ANALYTICAL; nothing run)
print("4. free-text decisions")
print(f" semantic program: 0 (typed {r['typed_decisions']})")
print(" react-style loop: 5 (one free-text Thought per step, counted)")
1. claim A
retrieved ['d9', 'd1', 'd2'] kept ['d9', 'd1', 'd2']
verdict CONFLICTING p=0.88 -> escalate (conflicting or unknown evidence)
typed decisions 9, free-text 0
2. runtime boundary
B: SUPPORTED p=0.9 -> permit
C: SUPPORTED p=0.55 -> deny
3. replay
same seed 3: {'no_change': 9}
changed seed 8: {'no_change': 5, 'decision_changed': 3, 'state_changed_decision_same': 1}
4. free-text decisions
semantic program: 0 (typed 9)
react-style loop: 5 (one free-text Thought per step, counted)
The run
run_ch34.py writes results/ch34.jsonl and the replayable graph
evidence/ch34-replay.jsonl (mode program, replay, analytical).
The program on three claims (seed 3, declared; every provider a stub):
| claim | kept | relations | verdict | p | action |
|---|---|---|---|---|---|
| A: compound Q reduced inflammatory markers | d9, d1, d2 | uncertain (tie rule), supports, contradicts | CONFLICTING | 0.88 | escalate |
| B: the alpine study recorded temperature ranges | d3 | supports | SUPPORTED | 0.90 | permit |
| C: small pilot study reported modest drop | d5 | supports | SUPPORTED | 0.55 | deny |
The three cases cover the boundary completely: permit on confident support, deny on low-confidence support (the runtime refuses an action), escalate on conflicting evidence. Document d9 also exercises the Chapter 18 tie rule in the main run: its supports/contradicts probabilities (0.42/0.40) are within ε=0.05, so it resolves to CONFLICTING_EVIDENCE — never an arbitrary winner.
Replay (Chapter 26 graph, claim A, 9 nodes): same-seed replay gives
no_change 9/9 — decisions and state hashes agree, so the stored graph is a
faithful record. Replaying behind seed 8 changes three nodes — mat:d1
(supports→contradicts) and its ancestors support and act — exactly the
sampled nodes and their dependents, nothing else. The graph you could store
yesterday tells you which decisions change when the sampling changes, and which
do not.
Free-text decisions (mode analytical; nothing was run):
| system | free-text decisions | typed decisions | note |
|---|---|---|---|
| semantic program | 0 | 9 | every ask has an answer set |
| ReAct-style loop | 5 | 0 | one free-text Thought per step (search→read→verify→decide→finish), counted |
The ReAct number is a count over the declared five-step task script, swept to show the dependence: at 3–6 steps the free-text count is 3–6, with an assumed 50–200 tokens per Thought (rows carry the sweep). No accuracy claim is made for either system; the acceptance criterion was the count, and it is reported.
What it says
-
The surviving construct set is enough to write the program.
decide(every ask),if decide(the support/boundary branches),match decide(relations, with its tie rule firing on d9),where decide(the relevance filter). The one thing Chapter 23 rejected —for decide— appears only in its library form (aforloop overBM25Index.search). The constructs the earlier chapters kept are the ones this program used. P5 holds. -
The runtime boundary is a decision, not a guard clause. Permit/deny/ escalate is itself recorded in the graph with its antecedents; the deny on claim C is a decision about acting, and replay shows it. The boundary refuses an action on low confidence, and the refusal is auditable. P2 holds.
-
Replay is the audit. Same-seed 9/9 agreement; changed-seed changes only the sampled nodes (and their ancestors). A ReAct trace is free text; you can replay it only by re-running the model. P4 holds.
-
The count is the honest comparison. The semantic program makes 0 free-text decisions for 9 typed ones; a ReAct-style loop makes one per Thought, all 5 by the minimal script. P3 holds — and it is the entire claim: no cost or accuracy number is implied by this chapter (the only cost row is an assumed-tokens sweep, labelled ANALYTICAL).
-
The corpus story is a property of the stubs. The verdicts are consequences of the hand-supplied relation triples and the BM25 retrieval; they are ILLUSTRATIVE, exactly as rule 14 demands. P1 holds only as “the program resolves each claim to the verdict its stub evidence dictates”.
Wrong / Correct. Wrong: “The semantic program and a ReAct loop solve the same task with the same flexibility.” Correct: “The semantic program trades ReAct’s free-text flexibility for typed, countable, replayable decisions: 0 untyped decisions, 9 typed, a graph that replays 9/9 — and gives up the agent’s ability to re-plan in natural language. Which one you want depends on whether you can enumerate the decisions in advance.”
The construct set, stated plainly
| construct | Chapter | verdict | used here |
|---|---|---|---|
decide |
16 | kept | every ask (relevance, relation, support, act) |
if decide |
17 | kept | support/boundary branches |
match decide |
18 | kept | relations; tie rule fired on d9 |
where decide |
22 | kept (typed, abstaining; LOTUS sem_filter, narrower, no novelty) |
relevance filter |
for decide |
23 | rejected | library for over retrieval |
No new syntax layer was built: Chapter 23’s rejection and Chapter 22’s “filter choice matters more than syntax” make the library form the honest result, and the program is written in it.
Prior art, engaged directly
- ReAct (2210.03629): the challenge. Its decisions are generated text; the count above is what that costs. The semantic program is ReAct’s task with the decisions enumerated in advance.
- PAL (2211.10435): the method. Deterministic work (retrieval ordering, aggregation, hashing) lives in the interpreter; only the decisions are semantic. This program is PAL’s split, applied to decisions rather than to arithmetic.
- Toolformer (2302.04761): the support. Choosing to act is a decision; this chapter makes the choice typed and gated, where Toolformer learns it as a token-generation skill.
No novelty is claimed over any of them; the chapter’s contribution is the count and the replay, both measured on the program’s own deterministic runs.
What to carry forward
Chapter 35 writes the verdict. It will need the honest lists: which constructs
survived (this chapter used them) and which were rejected (Chapter 23’s
for decide, shown in library form here). And it will need the Part VI
record: the same 188 SciFact dev claims served as the test split of Chapters
21–23 and 25 — this chapter’s corpus is new and hand-built, so it does not
inherit that leak, but it does inherit the discipline of labelling every stub.
Close by
Which constructs did the program actually use? All the survivors — decide,
if decide, match decide, where decide — and the rejected one only as a
library loop. That is the chapter’s answer to its own question, and it is the
narrowest claim the book will carry into Chapter 35.
Limitations
- The corpus, relation triples, thresholds (keep 0.60, permit 0.80, ε 0.05) and sampling seed (3) are all assumed; every verdict is a consequence of them, reported as such (ILLUSTRATIVE).
- The generation step is a
[stub]template with zero tokens; no cost claim is made for it. - The ReAct comparison is ANALYTICAL: a count over a declared five-step script with a swept token assumption; no agent ran and no accuracy comparison is claimed (an LLM is not allowed by the book’s rules).
- The three papers are read at abstract level.
- The program records decisions in the graph but not their probabilities; the probabilities live in the result rows and drive the boundary, and the replay agreement is over decisions and state hashes (Chapter 26 semantics).
- The sampling seed for the showcased run (3) differs from the preregistration’s placeholder “seed 0 and 1”; the deviation is disclosed here and in the metadata, and the replay test still covers same-seed vs changed-seed.