The Decision Expression
Give the decision a calling convention: one total expression that returns a value with provenance or a typed refusal, and say what == means on it.
Design draft: this chapter argues from the literature and from earlier measured results; its own experiment has not been run.
You have four chapters of machinery and no way to call it. The contract validates answers. The abstain gate prices refusal. The unknown resolver names three ways of not deciding. Each is a function you could call, but nothing says in which order, what happens when one step fails, or what the caller is allowed to ignore. So here is the failure this chapter exists to prevent, in the smallest form that still runs:
probs = model(state) # {"refund": 0.40, "exchange": 0.35, "other": 0.25}
choice = max(probs, key=probs.get)
act_on(choice) # issues a refund on a 40% plurality
That code type-checks and runs. It acts on a weak winner because nothing in its shape asks whether the winner is strong enough. It keeps no record of which model answered. And it is tied to the shape one provider happens to return: give it a provider that declines to answer, and it crashes on a structure it never expected. Each of those is a bug an earlier chapter has already named. This chapter collects them into one calling convention, the decision expression, and writes down exactly what it guarantees.
The expression is decide(question, state, choices, provider, ...). It returns an Outcome: a Decision with value, score, alternatives, provider, evidence and trace, or a typed non-answer (Refusal, Abstain, NoneOfTheAbove, Unknown). There is no bare-value path. The claim under test, preregistered as JEV-16-01, is that this adds three things over a plain function returning a distribution: mandatory handling of abstention, provenance on every value, and provider independence. None of the three is new to computer science. The combination may be the contribution, and Chapters 17 to 19 will decide whether even that needs syntax, or whether the library form is enough.
What the literature already built
The program-over-distributions view is prior art. Four sources bear on this chapter: three that shape the design and one that argues against the whole enterprise.
Language model cascades: the framing, without the bill
DOCUMENTED: Dohan and colleagues argue that compositions of prompted models (chain-of-thought, verifiers, selection-inference, tool use) are probabilistic programs over string-valued random variables. They implement this as a trace-based language embedded in Python, with sample and observe operations. [Dohan et al., arXiv:2207.10342, Abstract, §§1–3, 4]
Their discussion is candid about the price. Beyond rejection sampling nothing is evaluated, and efficient inference over string-valued variables is the core technical challenge. Their STaR interpretation shows what cascades buy: rationale generation becomes stochastic EM, an E-step that imputes thoughts by rejection sampling plus an M-step that updates parameters, which opens tuning methods to every cascade and not only chain-of-thought.
We take the framing and leave the bill. Our expression performs no inference at all. It reads one distribution per call. A cascade conditions and infers; decide() asks and records. The day a decision program needs multi-step inference over latent thoughts, it will have outgrown this chapter and moved into Dohan’s.
LMQL: the closest existing construct
DOCUMENTED: LMQL is the nearest language-level construct, so the comparison must be direct, because this chapter may not claim novelty over it. LMQL separates a front end from a back end. Users write queries with decoders (argmax, sample, beam), where-clause constraints over hole variables, and control flow. The runtime generates token masks from eager partial-evaluation semantics and prunes the search space, cutting inference cost and latency by 26 to 80% while retaining or improving accuracy. [Beurer-Kellner et al., arXiv:2212.06094, Abstract/§1, contributions, §§2–4, §6]
Two of its features sit close to ours. Its distribute clause reads a probability distribution over a support set for classification, which is the same shape as our option readout. Its scripted beam search jointly optimises holes and control flow, much as later chapters will have to handle composed decisions.
Three differences matter.
- LMQL constrains generation;
decide()constrains calling. No text is produced, parsed or masked. - LMQL’s own semantics admits that some constraints cannot be enforced eagerly and fall back to backtracking. A construct that sometimes degrades silently is precisely what our total outcome type refuses to be.
- LMQL programs stay portable across models by abstracting tokenisation. Our expression stays portable across providers by abstracting the answer shape. That is why provider independence, not token masking, is one of the three tested properties.
Probabilistic programming: borrow the vocabulary, reject the machinery
DOCUMENTED: van de Meent and colleagues give the vocabulary to borrow and the machinery to reject. Their language has sample for unobserved random variables and observe for observed ones, with evaluation rules such as likelihood weighting that draw from the prior and accumulate weights. It has an explicit interface between program executions and an inference controller, so many inference algorithms can run against one program. Conditioning, they argue, is the foundational computation. [van de Meent et al., arXiv:1809.10756, Abstract, §§2.1, 4.1]
We borrow exactly that: a decision expression conditions a caller on a provider’s answer the way observe conditions a model on data. We reject the rest. There is no sampling, no weighting, no inference controller and no posterior. A decision is not a posterior, and treating it like one would license arithmetic (averaging decisions, conditioning decisions on decisions) that this book has not earned. Chapter 24 will ask whether that arithmetic can ever be earned, and to answer it will have to reintroduce the controller interface this chapter declines.
Constrained decoding: the paper that keeps us honest
DOCUMENTED, and it challenges the whole enterprise: Xiong and colleagues benchmark constrained decoding on small on-device models and find that syntactic guarantees do not buy semantic correctness, and can actively destroy it. Their Constraint-Induced Regression names the failure: the mask preserves validity while the meaning degenerates, concentrated in code and DSL tasks. Constraints, they conclude, are a reliability layer and not a universal enhancement. [Xiong et al., IJCAI 2026, Abstract, §1, §§4.3–4.4]
This is the paper that keeps Chapter 16 honest about what it is not claiming. Our three added things deliberately exclude shape enforcement. Nothing here is about making outputs parse. A construct that guaranteed parsing but not meaning would be LMQL’s territory, or StructureBench’s warning.
If you have read Language, you know the executable-contract rung this chapter stands on. If you have read DSPy From First Principles, you know the declare-then-compile split: signatures state behaviour, compilation chooses how. decide() is the signature; Chapters 29 to 33 are the compiler. Neither book is re-taught here.
Wrong: “A typed wrapper around a model call is ceremony, because the floats are the same either way.”
Correct: “The floats are the same; the bugs you can still write are not. Ceremony that changes which failures compile is semantics, not decoration.”
What the earlier chapters already settled
Chapter 16 adds no new mathematics. It fixes the order in which earlier pieces run and the type that comes out.
- The Decision Contract gave request shapes, result validation, winner-is-argmax, and
Refusalas a first-class provider answer. - Abstain gave the total
Resolvedtype and threshold arithmetic over supplied scores, tested without models. - Unknown gave reserved-label normalisation with missing-state precedence.
- Calibration gave the provenance view: evidence records, never certificates.
- Chapter 7 set the testing precedent this chapter follows: its provider was built and unit-tested against a fake model with known logits while the real model sweep waited.
The evaluation order
decide() runs the same seven steps on every call. The order is written down in docs/semantics.md precisely enough that the syntax of Chapters 17 to 19 can compile to it.
flowchart TD
A[decide question/state/choices] --> B{validate request}
B -- invalid --> X[ValueError]
B -- valid --> C[provider.decide]
C --> D{validate result}
D -- invalid --> X
D -- valid --> E{Refusal?}
E -- yes --> F[return Refusal]
E -- no --> G{missing state?}
G -- yes --> H[return Unknown]
G -- no --> I{reserved winner?}
I -- yes --> J[return Abstain/NOTA/Unknown]
I -- no --> K{below threshold?}
K -- yes --> L[return Abstain]
K -- no --> M[return Decision + provenance]
In words: validate the request; call the provider; validate its result against this request; pass a provider Refusal through unchanged, because it is the provider’s answer and not the caller’s; resolve missing mandatory state to Unknown before any score is read; resolve reserved labels to their typed outcomes; send a winner below the threshold to Abstain; and otherwise return the Decision with the caller’s evidence appended, never substituted.
Precondition violations are ValueError, never outcomes. An empty state or a one-option question is a programming error, not a decision.
A complete example
The implementation is src/arbiter/expr.py, tested by 13 tests in tests/semantics/. Everything below is examples/ch16-the-decision-expression/walkthrough_ch16.py, which you can run as it stands. The providers are fakes that return distributions you hand them, so what the example shows is the type system, not any real provider.
First the two providers. FakeProvider returns whatever distribution it was given. RefusingProvider cannot answer and says so in the contract’s way.
class FakeProvider:
"""Returns a distribution you supply, for any question about those labels."""
name = "fake"
def __init__(self, labels, probs, model):
self.labels, self.probs, self.model = list(labels), list(probs), model
def decide(self, request: DecisionRequest) -> DecisionResult:
dist = dict(zip(self.labels, self.probs))
answers = {}
for qid, question in request.questions.items():
labels = sorted(question.criteria)
answers[qid] = choice_result_for(self.name, labels, [dist[l] for l in labels])
return DecisionResult(model=self.model, answers=answers)
class RefusingProvider:
"""A provider that cannot answer: it says so, in the contract's way."""
name = "refuser"
def decide(self, request: DecisionRequest) -> DecisionResult:
refusal = Refusal(reason=RefusalReason.LABELS_UNSEEN,
message="never fitted on these labels")
return DecisionResult(model="refuser-1", answers={q: refusal for q in request.questions})
Then the five situations the chapter cares about, in one main():
QUESTION = "What does the customer want?"
STATE = "my card was charged twice"
CHOICES = ["refund", "exchange", "other"]
def main() -> None:
# 1. A confident provider: the expression returns a Decision with provenance.
confident = FakeProvider(CHOICES, [0.70, 0.20, 0.10], model="confident-1")
d = decide(QUESTION, STATE, CHOICES, confident, threshold=0.5, trace="walkthrough")
print("1. confident provider")
print(f" {type(d).__name__}: value={d.value!r} score={d.score:.2f}")
print(f" provenance: provider={d.provider!r} trace={d.trace!r}")
print(f" evidence: {d.evidence[0]!r}")
# 2. A weak winner: the same call returns Abstain; the plain function acts.
weak = FakeProvider(CHOICES, [0.40, 0.35, 0.25], model="weak-1")
a = decide(QUESTION, STATE, CHOICES, weak, threshold=0.5)
probs = plain_decide(STATE, CHOICES, weak)
acted_on = max(probs, key=probs.get)
print("2. weak winner (0.40 / 0.35 / 0.25), threshold 0.5")
print(f" decide() -> {type(a).__name__}(reason={a.reason.value!r})")
print(f" plain function -> acts on {acted_on!r} at {probs[acted_on]:.2f}, no abstain arm")
# 3. A provider that cannot answer: typed Refusal versus a crash.
refuser = RefusingProvider()
r = decide(QUESTION, STATE, CHOICES, refuser)
print("3. provider refuses")
print(f" decide() -> {type(r).__name__}(reason={r.reason.name}, message={r.message!r})")
try:
plain_decide(STATE, CHOICES, refuser)
except Exception as exc: # the plain caller assumed a distribution
print(f" plain function -> {type(exc).__name__}: caller assumed a distribution")
# 4. Two providers, same answer: == compares records, same_choice compares answers.
other = FakeProvider(CHOICES, [0.70, 0.20, 0.10], model="confident-2")
d2 = decide(QUESTION, STATE, CHOICES, other, threshold=0.5, trace="walkthrough")
print("4. two providers choose the same label")
print(f" d == d2 -> {d == d2} (different provider and model)")
print(f" same_choice(d,d2) -> {same_choice(d, d2)}")
print(f" same_choice(d, a) -> {same_choice(d, a)} (a Decision never agrees with an Abstain)")
# 5. Swapping the provider changes no caller code.
def caller(provider):
return decide(QUESTION, STATE, CHOICES, provider, threshold=0.5)
print("5. the caller never changes")
for p in (confident, weak, refuser):
print(f" caller({getattr(p, 'model', 'refuser-1')}) -> {type(caller(p)).__name__}")
And what it prints, verbatim:
1. confident provider
Decision: value='refund' score=0.70
provenance: provider='confident-1' trace='walkthrough'
evidence: 'provider-returned distribution; not a truth guarantee'
2. weak winner (0.40 / 0.35 / 0.25), threshold 0.5
decide() -> Abstain(reason='low_confidence')
plain function -> acts on 'refund' at 0.40, no abstain arm
3. provider refuses
decide() -> Refusal(reason=LABELS_UNSEEN, message='never fitted on these labels')
plain function -> AttributeError: caller assumed a distribution
4. two providers choose the same label
d == d2 -> False (different provider and model)
same_choice(d,d2) -> True
same_choice(d, a) -> False (a Decision never agrees with an Abstain)
5. the caller never changes
caller(confident-1) -> Decision
caller(weak-1) -> Abstain
caller(refuser-1) -> Refusal
Read the output top to bottom.
- A confident provider returns a
Decision. The value arrives with its score, its provider, its trace and the evidence note the contract attaches, so a later reader can tell where the answer came from. - A weak winner, 0.40 against 0.35 and 0.25, becomes an
Abstainunder a 0.5 threshold. The plain function, given the same provider, acts onrefundat 0.40 because nothing in its shape lets it do otherwise. This is the opening bug, produced on demand. - A refusing provider returns a typed
Refusalfromdecide(). The plain function crashes with anAttributeError, because it assumed the provider would always hand back a distribution. That is a quieter version of coupling to a provider’s shape: the caller broke on something the provider was entitled to do. ==andsame_choiceanswer different questions, and the next section explains why both exist.- The caller never changes. One function, three providers, three different outcome types, no edits.
What == means
The prompt for this chapter asks what == means on a decision, and the answer is two functions, because callers ask two different questions.
== on a Decision compares the full record: value, score, alternatives, provider, evidence and trace. Two providers that both chose refund are therefore not ==. same_choice(a, b) compares answers. Two Decisions agree when their values agree. Anything else agrees only under ==, so a Decision never agrees with an Abstain.
Line 4 of the output shows both. The two providers chose the same label, so same_choice(d, d2) is True while d == d2 is False. The rule of thumb fits in a sentence: audit trails compare with ==, and business logic branches with same_choice.
Mixing them up produces exactly the provenance bugs the chapter is about. Deduplicate decisions with == and you keep two copies that agree. Deduplicate with same_choice and you silently drop the second provider’s evidence. Branch on == to ask “did both pick refund?” and you miss a legitimate agreement.
The three bug classes, as tests
The preregistered checklist asks what each version makes impossible, with a concrete failing example on the plain-function side. These are the three tests, exactly as they stand in tests/semantics/test_expr.py.
P1, a dropped abstention:
def test_plain_function_permits_dropped_abstain():
"""P1 failing example: nothing in the plain shape forces abstain handling."""
labels = ["refund", "exchange", "other"]
p = FakeProvider(labels, [0.40, 0.35, 0.25])
probs = plain_decide("state text", labels, p)
# The natural caller code below runs fine and silently acts on a 40% plurality:
acted = max(probs, key=probs.get)
assert acted == "refund"
assert "abstain" not in probs and max(probs.values()) < 0.5
# The expression, given the same provider and a threshold, abstains instead.
assert isinstance(decide("What?", "state text", labels, p, threshold=0.5), Abstain)
P2, no provenance:
def test_plain_function_carries_no_provenance():
"""P2 failing example: the distribution knows nothing about its source."""
p = FakeProvider(["refund", "exchange"], [0.7, 0.3])
probs = plain_decide("state text", ["refund", "exchange"], p)
assert isinstance(probs, dict)
assert not hasattr(probs, "provider") and not hasattr(probs, "evidence")
P3, coupling to the provider’s shape:
def test_plain_function_couples_caller_to_provider_shape():
"""P3 failing example: a provider returning a different shape breaks it."""
p = FakeProvider(["refund", "exchange"], [0.7, 0.3])
probs = plain_decide("state text", ["refund", "exchange"], p)
with pytest.raises((KeyError, TypeError)):
probs["refund_nonexistent_key"] # caller assumed labels the provider never promised
Each is then answered by the expression form, where the same bug is unrepresentable: totality, provenance by construction, and shape-indifferent calling.
Notice what these tests do not show, because a checklist chapter has to say so. They do not show that real providers honour the contract; Chapter 8’s property suite does that per provider. They do not show that thresholds survive a shift in the data; Chapter 11’s run is pending. They do not show that unknown labels are detected; Chapter 12’s run is pending. The expression inherits every upstream guarantee and every upstream gap. Composition is not laundering.
The experiment, designed but not run
JEV-16-01 is preregistered with status: NOT_RUN; the full block is in metadata/16-chapter.yaml.
The preregistered comparison is the bug-class checklist: what each version makes impossible, with concrete failing examples on the plain-function side. The software half of that checklist is DEMONSTRATED_IN_TESTS: 13 tests, each naming the bug it exhibits. What remains NOT_RUN is the provider-level half, the same checklist executed against real providers, which is what could move L1.
Predictions (HYPOTHESIS):
- P1: totality removes silent abstain-drops.
- P2: provenance rides every decided value.
- P3: provider swaps change no caller code.
Refutation: P1 fails on any bare-value path for a refused decision; P2 on any provenance-free decided value; P3 on any swap that requires caller edits.
The L1 verdict, whether syntax adds anything over this library form, is explicitly not decided here. If Chapters 17 to 19 find no behavioural difference between enforced and optional handling, the library form wins and L1 stands. That is the honest scope of a semantics chapter: it fixes meaning so later chapters can test force.
PENDING_RUN: result for JEV-16-01, P1–P3 (provider-level) Command:
python examples/ch16-the-decision-expression/run_ch16.py --providers tfidf-lr,embed-sim --checklist abstain,provenance,swapFills:results/ch16.jsonlStage plan (CPU-only, reusing frozen Ch4/Ch5 providers; each ≤15 min): S1 run checklist against real providers; S2 record failing-example rows; S3 write rows.
What would change your mind
A bare-value path anywhere in decide() would refute P1 on the spot. Totality is a property of the code, and the tests assert it on every path they cover.
A real provider whose result cannot be validated without caller-side surgery would refute P3’s independence claim in practice, though not in principle. The contract would be intact and the provider nonconformant, which is exactly the distinction Chapter 8’s validate_result exists to draw.
A subtler defeat is convention collapse: callers unwrapping every Outcome with a helper that defaults abstains to the argmax. The type system cannot stop a determined caller from reintroducing the bug it was given tools to avoid. Chapter 17’s enforced-versus-optional experiment measures exactly that failure mode.
The largest mind-change available belongs to Chapter 17. If enforced abstain-handling shows no behavioural difference from optional handling, the expression’s first added thing evaporates, and with it much of the case for syntax.
No pivot is proposed. The software is demonstrated, the providers are pending, and there is no new measured evidence.
Limitations
JEV-16-01 is NOT_RUN at provider level, and results/ch16.jsonl does not exist.
Required papers are PARTIAL by section. Dohan’s inference machinery, LMQL’s appendices and the probabilistic-programming book’s later chapters were not reproduced. StructureBench concerns small on-device generative models and not decision providers; the analogy is explicit, not an equivalence.
Fake providers in tests prove shape, not behaviour. A real provider can still be miscalibrated, slow or wrong, and Chapters 4 to 10 own those properties.
decide() performs no inference, no conditioning across calls and no cross-decision combination; Chapter 24 owns that question. No parser was written: the expression is a library calling convention, and whether syntax adds anything is the experiment of Chapters 17 to 19, not an assumption of this one. Compile-time enforcement of the match arms is future work and is not a property of this module. The distribute-style readout comparison with LMQL is structural, not benchmarked.
What the next chapters inherit
Chapter 17 inherits decide() as the compile target for if decide: branch semantics, threshold guarantees, and the enforced-versus-optional experiment that decides L1’s fate. Concretely, if decide must desugar to a match on the Outcome type with no wildcard arm that silently swallows Abstain. Otherwise it is compiling to the plain function’s semantics and not to this chapter’s.
Chapter 18 inherits the Outcome type for match decide: exhaustiveness over five members, not two, with the compiler rejecting a match that forgets Unknown or Refusal.
Chapter 19 inherits the question the prompt leaves open: what == means per decision type, and which types earn primitives.
The closing question nearly answers itself. What the expression guarantees that a plain function does not is totality, provenance and independence, demonstrated in tests and pending in providers. Whether that guarantee needs syntax is the next three chapters’ problem, and they inherit a semantics precise enough to test it against.