The Decision Contract
A provider can answer, refuse, or violate the contract. Fake-model swaps test which distinctions the caller actually needs.
You replace the refund predicate with a classifier. The next ticket arrives, but your question names a label the classifier has never seen. Should the application guess, catch an exception, or pass the ticket to a human? A distribution over the wrong labels would look reassuring while answering a different question.
Stop Generating kept that application’s action separate from its provider. The Jev Provocation supplied a first request and response shape. The Smallest Decision and Zero-Shot Decisions then supplied unlike providers. The First Token exposed a further problem: some ways of presenting an option set cannot be read by a particular provider. You need to know whether it answered your request at all before asking how good its answer was.
This chapter tests the interface with deterministic fakes and recorded scores. OBSERVED means code executed here on CPU; it never means an NLI, embedding or language model was evaluated in this session. Hosted Jev is a wire-shape stub. Accuracy, calibration, model latency and live hosted behavior are NOT_OBSERVED.
The expectation and its limits
HYPOTHESIS — H-contract: a neutral contract needs an explicit way for a provider to decline a valid request, and presentation belongs to the provider. The committed prediction expected two or three providers to break the first draft. It also predicted that provider-chosen presentation would need fewer caller changes than caller-required presentation.
The separation has prior art. DOCUMENTED: DSPy signatures declare a task while modules and compilation determine how a model implements it. The book imports that distinction from DSPy From First Principles, rather than claiming it as a new mechanism. Khattab et al., §3.
Meyer supplies the vocabulary of caller obligations, supplier guarantees and invariants. His discussion also permits a broader interface that handles expected special cases. DOCUMENTED: contract theory does not uniquely require refusal to be a returned value; responsibility boundaries remain a design choice. The primary article was accessible in this resumption, correcting the earlier reading note. Meyer, pp. 42–44 and 49–51.
DOCUMENTED: JSONSchemaBench tests whether engines admit valid instances and reject invalid ones; both excessive restriction and insufficient restriction are failures. That is why our suite includes supported requests alongside impossible ones. Geng et al., §5.3.
The recent structured-output study by Chavan provides the strongest warning: schema-valid output can omit part of a requested task, and field-presence metrics can miss empty substance. Our assertions check consistency, not truth. Chavan, §§4.3–5.1.
All required papers were read from primary full-text copies, with the inspected sections recorded as PARTIAL for this resumption. The modern paper was read FULL_TEXT. The chapter’s source-verification companion records the sections, fetch failures and correction of an inherited misattribution to JSONSchemaBench.
What the contract promises
PROPOSED: retain the earlier question and answer names, and put a checked
boundary around them. ContractProvider.decide returns an answer or Refusal
under every requested question id. validate_result binds the answer to the
request: its probability keys must cover exactly the declared options. Checking
only that the returned distribution sums to one cannot detect omitted options.
Decision[State, Value] projects a choice into value, score, alternatives,
provider, evidence and trace, alongside its state and declared answer set.
Score is the selected probability. Alternatives retain the request’s label order.
Evidence records provenance; trace identifies the checked request. Neither is a
proof that the label means what you intended.
The outcomes stay distinct:
| Situation | Representation | What the caller learns |
|---|---|---|
| Invalid question or blank state | Precondition exception | Repair the request |
| Supported choice | Choice answer and distribution | A declared option was selected |
| Provider cannot satisfy a valid request | Typed refusal, without a distribution | Route or record the inability |
| Explicit abstention option selected | Abstention.ABSTAIN as a declared value |
An answer inside the declared space |
| Broken code or transport failure | Exception | An expected refusal must not conceal a defect |
Refusal describes inability to supply the requested answer; abstention reserves an answer value. This chapter chooses their representation. Later chapters determine their policy and semantics. No calibration or open-set claim follows from the type.
Presentation is optional provider evidence, not a mandatory prompt field for every classifier. The option scorer reports its printed markers and explicitly refuses an impossible letter readout. A caller may impose a presentation through the experimental wrapper; the provider must honour it or refuse.
This maps directly to Language’s Representation Ladder:
| Element | Rung |
|---|---|
| State, instructions and option descriptions | natural prose |
| Question ids and declared options | structured requirements |
| Value, score, alternatives and refusal fields | schemas and constraints |
| Running checks of distributions and request binding | executable contracts |
| Provider, evidence and trace records | machine-native operational state |
| Application action after the outcome | state-transition representations |
The contract is not a formal specification or proof of semantic correctness.
planning/decision-contract.md explains the mapping and extension points.
The contract, running
The table above is easier to believe when each row is a call. Everything below is examples/ch08-the-decision-contract/walkthrough_ch08.py, which you can run as it stands. The providers are deterministic fakes that return whatever they are told to, so the example shows which distinctions the contract draws and says nothing about any real provider.
from arbiter.contract import (
Abstention,
ChoiceAnswer,
ChoiceQuestion,
DecisionRequest,
DecisionResult,
Refusal,
RefusalReason,
choice_confidence,
decision_for,
validate_result,
)
CRITERIA = {"billing": "Charges and payments", "returns": "Wrong or damaged items"}
def request_for(criteria, state="my card was charged twice"):
return DecisionRequest(
state=state,
questions={"route": ChoiceQuestion(id="route", instructions="Which team handles this?", criteria=criteria)},
)
def answer_with(probabilities, choice=None):
"""A provider's answer. `choice` defaults to the argmax, as the contract requires."""
choice = choice or max(probabilities, key=probabilities.get)
return DecisionResult(
model="fake-1",
answers={"route": ChoiceAnswer(choice, probabilities, choice_confidence(list(probabilities.values())))},
)
def attempt(label, build):
try:
build()
print(f" {label:<34} -> accepted")
except ValueError as exc:
print(f" {label:<34} -> ValueError: {exc}")
def main() -> None:
request = request_for(CRITERIA)
# 1. A bad request is the caller's bug, so it raises before any provider is asked.
print("1. an invalid request is a precondition error")
attempt("blank state", lambda: request_for(CRITERIA, state=" "))
attempt("only one option", lambda: request_for({"billing": "Charges"}))
# 2. A supported choice: the answer is bound to THIS request and projected to a Decision.
print("2. a supported choice")
result = validate_result(request, answer_with({"billing": 0.70, "returns": 0.30}))
decision = decision_for(request, result, "route")
print(f" value={decision.value!r} score={decision.score:.2f} provider={decision.provider!r}")
print(f" alternatives={decision.alternatives}")
print(f" evidence={decision.evidence[0]!r}")
# 3. A refusal is a typed answer without a distribution, not an exception.
print("3. the provider cannot satisfy a valid request")
refusal = Refusal(reason=RefusalReason.LABELS_UNSEEN, message="never fitted on these labels")
refused = decision_for(request, validate_result(request, DecisionResult("fake-1", {"route": refusal})), "route")
print(f" value is a Refusal: {isinstance(refused.value, Refusal)}, score={refused.score}, alternatives={refused.alternatives}")
# 4. Abstention is a value inside the declared answer space, which is a different thing.
print("4. abstention is a declared option, not a refusal")
with_abstain = {**CRITERIA, Abstention.ABSTAIN.value: "None of these fits"}
request_a = request_for(with_abstain)
picked = answer_with({"billing": 0.20, "returns": 0.10, "ABSTAIN": 0.70})
abstained = decision_for(request_a, validate_result(request_a, picked), "route")
print(f" value={abstained.value!r} (in the answer set: {abstained.value in abstained.answer_set})")
# 5. A provider that breaks the contract is a defect, and it is caught, not smoothed over.
print("5. contract violations")
attempt("probabilities sum to 1.2", lambda: answer_with({"billing": 0.70, "returns": 0.50}))
attempt("winner is not the argmax", lambda: answer_with({"billing": 0.70, "returns": 0.30}, choice="returns"))
attempt("answers a different option set", lambda: validate_result(request, answer_with({"billing": 0.6, "shipping": 0.4})))
attempt("silently drops an option", lambda: validate_result(request_a, answer_with({"billing": 0.7, "returns": 0.3})))
1. an invalid request is a precondition error
blank state -> ValueError: state must not be empty: a decision about nothing is nothing
only one option -> ValueError: choice needs 2-255 options, got 1
2. a supported choice
value='billing' score=0.70 provider='fake-1'
alternatives=(('billing', 0.7), ('returns', 0.3))
evidence='provider-returned distribution; not a truth guarantee'
3. the provider cannot satisfy a valid request
value is a Refusal: True, score=None, alternatives=()
4. abstention is a declared option, not a refusal
value='ABSTAIN' (in the answer set: True)
5. contract violations
probabilities sum to 1.2 -> ValueError: choice: probabilities sum to 1.2, not 1.0
winner is not the argmax -> ValueError: choice 'returns' is not the most probable option ('billing' is)
answers a different option set -> ValueError: answer must cover exactly the declared option set
silently drops an option -> ValueError: answer must cover exactly the declared option set
Read it against the table.
- A bad request is the caller’s bug. A blank state or a one-option question raises before any provider is asked. Nothing was refused, because nothing was asked.
- A supported choice comes back as a
Decisionwith its score, its alternatives in the request’s own order, the provider that answered, and an evidence note that says outright it is not a truth guarantee. - A refusal is an answer without a distribution. The
Decisioncarries theRefusalas its value, with no score and no alternatives. The caller learns the provider could not answer, and nothing in the result resembles a probability. - Abstention is a value inside the declared answer set. Here the request lists
ABSTAINas one more option and the provider picks it. That is an answer, and a different thing from the provider declining to answer at all. This chapter fixes how the two are represented; later chapters decide when to use them. - Violations are defects, not outcomes. Probabilities that sum to 1.2, a winner that is not the argmax, an answer over the wrong option set, and an answer that silently drops a declared option all raise. The last two are the reason
validate_resultbinds the answer to this request: a distribution that sums to one cannot reveal that it answers a different question.
What actually broke
OBSERVED — JEV-08-01: all eight existing providers at ca1e676 raised an
exception for a valid request with unsupported labels. All could answer their
supported fixture, but none expressed refusal through the first-draft result.
The hosted stub did not exist at that baseline and is excluded from this count.
The prediction of two or three failures was refuted, not softened to match.
The completed suite passed 49 tests. It caught four of four seeded defects: invalid probability mass, an out-of-set winner, wrongly attached label scores, and silent truncation. The truncating provider returned a normalized distribution over only part of the declared set; exact request binding caught it. The mutated order case needed the known-score oracle. Structural invariants alone accepted that internally consistent but semantically wrong answer.
The loose schema control accepted three malformed outcomes that the checked contract rejected. This demonstrates the value of these additional checks against that control. A stronger request-specific schema could reject some of them too; the result does not prove that a new language construct is necessary.
These numbers come from results/ch08.jsonl: first_draft_refusal_failures,
property_tests_passed, seeded_bugs_caught, and schema_control_contract_failures.
The swap, including the caller’s cost
OBSERVED — JEV-08-02: every provider saw the same 200 BANKING77 threshold texts. Its score computation was faked; the frozen train subset was selected but not used to estimate model performance. The rule provider and recorded stub refused the answer space. Refusing counts as an explicit outcome, not success at classification.
| Provider | Adapter lines | Edits per subsequent swap | P1 refusals | P2 refusals |
|---|---|---|---|---|
| Majority | 23 | 0 | 0 | 200 |
| Lookup | 26 | 0 | 0 | 200 |
| Rule | 36 | 0 | 200 | 200 |
| TF-IDF + LR | 20 | 0 | 0 | 200 |
| fastText-style | 58 | 0 | 0 | 200 |
| Embedding similarity | 46 | 0 | 0 | 200 |
| NLI | 49 | 0 | 0 | 200 |
| Option scoring | 39 | 0 | 0 | 0 |
| Hosted Jev stub | 11 | 0 | 200 | 200 |
Sources: adapter_lines_corrected, adapter_lines_by_arm, and
caller_changes_per_swap, including each row’s arm and refusals fields.
The adapter rule includes provider-specific fit bodies, rather than counting only
the final return statement. Its original absolute size predictions mostly failed.
An audit-counter repair preserved the original rows and appended corrections;
the dated amendment explains the missed helper and model-loading exclusions.
P1 lets the provider choose presentation. P2 requires bare-number markers through a shared wrapper costing 11 additional lines. Both arms use the same caller with zero edits per swap. The strict predicted caller-change advantage for P1 was refuted by a tie. P2’s coverage loss is conditional on that fixed constraint; it does not establish that all caller constraints are harmful.
The literal older runner could not consume a refusal: it reads answer.choice
unconditionally. Migrating its loop required five added nonblank lines under
the declared diff measurement. After that migration, swapping providers required
no further edits. The unchanged Chapter 1 application accepted the common
refund-outcome adapter, including escalation on refusal. Literal zero migration
changes was therefore refuted; subsequent interchangeability was observed.
Provider knowledge still reaches the original benchmark caller:
Location in run_ch04.py |
Finding | Treatment |
|---|---|---|
labels_of, line 64, and its call sites |
Caller discovers the answer space through provider.labels |
Provider detail leaks into request construction |
build, lines 114–122 |
Constructs providers by name | Recorded, excluded as factory setup |
| Seed branches, lines 143, 215, 310 | Runner knows which providers are stochastic | Setup knowledge remains in the runner |
except ValueError, lines 160, 224, 318 |
Generic fit/budget failures become provider_refused |
Refusal and failure are conflated |
decide_all, line 57 |
Reads a success-only field | Refusal is not handled |
The raw AST scan reported 12 candidates, including five excluded factory
branches; removing those leaves seven. Manual reading counted 14 leak
locations, including indirect label-helper calls and unhandled refusal. The new
common caller has zero AST and manual leaks under these rules. The original
runner remains unchanged, so its leaks remain. This also corrects the brief’s
description: its handlers catch generic ValueError; they do not match an LR-only
error string. See the ast_leaks and manual-audit rows for exact locations.
Regression evidence and its boundary
OBSERVED — JEV-08-03: recorded-score replay preserved 1,540 decisions across
11 task/provider combinations, with probability differences below the
preregistered tolerance. These are the golden_decisions_identical and
golden_max_probability_delta rows. Replay checks label binding, normalization
and winner selection. Because it injects recorded scores, it cannot detect a
changed encoder or training algorithm. The earlier end-to-end golden capture and
check belong to commits 9296ae4 and 303ddc5; this session does not claim to
have repeated them.
The five earlier self-test files passed without edits. Chapter 5’s encoder was replaced at its loading seam with a fake, so its pretrained-model smoke remains NOT_OBSERVED here. The compatibility shim measures 13 physical code lines; legacy imports, successful outputs and unsupported-label exceptions survive. NLI’s committed code was wrapped and fake-tested without editing its Chapter 6 file in this resumption. Earlier results and preregistrations were not changed.
The stub parses a transcribed response shape from evidence/jev-access.md using
the Chapter 2 parser. Its recorded choice/noul shape fits; no answer-shape change
is needed. It does not establish byte equality with a raw HTTP payload, which that
note did not retain. No Jev call occurred. TypeSafe’s inspected successful-response
types have no refusal arm, but that does not establish that the server cannot
decline through an HTTP error. Vendor SDK types.
The decision to carry forward
PROPOSED verdict: H-contract is PARTIALLY_SUPPORTED. Explicit inability is useful under the broad neutral interface. Refusal as a value is our representation choice; a shared typed exception remains an untested alternative. The swap does not establish exclusive provider ownership of presentation. Keep presentation as optional audit evidence, with feasibility checked by the provider.
I recommend accepting part (a) of Chapter 7’s pivot narrowly: make presentation feasibility visible in the chapter and refuse an impossible readout. Defer a stronger rule that presentation must be a mandatory neutral-contract field or must never be constrained by a caller. The author still decides; the proposal’s status has not been changed to accepted.
You can inspect the example and verify the recorded experiment without a model:
python examples/ch08-the-decision-contract/demo_ch08.py
python examples/ch08-the-decision-contract/verify_ch08.py
The durable artifacts are the checked contract, common adapter, property suite, golden replay, caller audit and results record. The failed fixture assumption about integer label keys is recorded, as are the presentation-arm constraint, partial paper reading and the limits of fake-model regression. The chapter stays a draft for author review. Confidence Is Not Probability now needs to distinguish a checked score from a calibrated one.
What is now permanently separated, and what could still leak?