One Contract, Many Models
Where does each kind of provider win, on the same decisions?
One Contract, Many Models
This chapter runs no new experiment. It compiles the winner table from rows in
results/ch04.jsonl,ch05.jsonlandch06.jsonland marks every cell without a measuredsplit=testrow asNOT_OBSERVED. The number of filled cells (17) is smaller than the number of empty ones, and that is a result.
The problem
Chapter 8 defined the contract every provider satisfies. This chapter asks the question the contract exists for: where does each kind of provider win, on the same decisions?
The honest scale of the answer is set by what is measured in this book: rule, majority, TF-IDF+LR, fastText-style, embedding zero-shot (bge), NLI (DeBERTa, BART) and a cross-encoder. An LLM judge and a decision model were never run here. The winner table therefore cannot be complete, and this chapter says so.
What we expect and why
Our starting hypothesis follows H4: the decision-model region is real but narrow — many changing label sets, tight latency, few labels, probabilities needed. Three papers frame the method of the benchmark itself:
-
Liang et al. (2022) — HELM argues for reporting accuracy, calibration, robustness and efficiency together, not accuracy alone. (Read at abstract level.)
-
Srivastava et al. (2022) — BIG-bench’s breadth shows that aggregate scores hide structure; a single “best provider” over all decisions is exactly such a hiding score. (Read at abstract level.)
-
Biderman et al. (2024) — reproducible evaluation is about eliminating silent variance; every row this chapter cites carries its
seed,nandnote. (Read at abstract level.)
The filling rule
The runner benchmarks/great-decision-benchmark/winner_table.py reads the results files and, for each provider and decision family, takes the best measured split=test row; if none exists, the cell is NOT_OBSERVED. The walkthrough shows the rule on a tiny hand-made table.
# 1. The filling rule: split=test rows only, best value, NOT_OBSERVED else.
print("1. the filling rule")
v = best_test(ROWS, "tfidf-lr")
print(f" tfidf-lr intent (tiny source): best test row = {v}")
# 2. Cells with no measured row stay NOT_OBSERVED.
print("2. empty cells")
for provider in ("cross-encoder", "llm-judge", "decision-model"):
v = best_test(ROWS, provider)
print(f" {provider:<16} -> NOT_OBSERVED" if v == 0.0
else f" {provider:<16} -> {v}")
1. the filling rule
tfidf-lr intent (tiny source): best test row = 0.8779
2. empty cells
cross-encoder -> NOT_OBSERVED
llm-judge -> NOT_OBSERVED
decision-model -> NOT_OBSERVED
The winner table
Best split=test accuracy per provider and family, compiled from results/ch04.jsonl, ch05.jsonl, ch06.jsonl (runner output; LLM-judge and decision-model cells are always NOT_OBSERVED):
| provider | intent ch04 | intent ch05 | intent ch06 | safety ch04 | safety ch06 |
|---|---|---|---|---|---|
| majority | 0.0130 | — | — | 0.4828 | — |
| rule | — | — | — | 0.4914 | — |
| lookup | 0.0130 | — | — | 0.4828 | — |
| tfidf-lr | 0.8779 | 0.9000 | 0.9000 | 0.8966 | 0.8966 |
| fasttext-style | 0.8526 | — | — | 0.9052 | — |
| embed-sim | — | 0.8029 | — | — | — |
| nli-deberta | — | — | 0.6735 | — | 0.5690 |
| nli-bart | — | — | 0.6706 | — | 0.5000 |
| cross-encoder | — | — | — | — | — |
| llm-judge | NOT_OBSERVED | ||||
| decision model | NOT_OBSERVED |
The table’s own caveats are visible in the rows: the Ch5 embed-sim cell is the unseen-label best (0.8029); the seen-label best is 0.7312 (Ch5). The Ch6 NLI intent cell (0.6735) is the intent-seen best; Chapter 6’s headline is that NLI loses to the embedding provider on unseen labels by 11.5 points. A max-on-test cell masks both distinctions, which is why the chapter reports the table and the chapters it cites.
Two ways the table flatters, which the rows do not show. First, each cell is the maximum over every split=test row for that provider, which includes every wording, template and seed the earlier chapters ran. A provider with many configurations (embedding similarity with four wordings, NLI with three templates) gets the best of several draws on the test split, and a provider with one configuration (tuned LR) gets one. That is selection on test, the practice this book otherwise forbids, and it favours the providers with more knobs. Second, the 17 filled cells are 15 distinct measurements: the tuned LR’s 0.9000 and 0.8966 each appear twice, because the Chapter 6 columns repeat the frozen Chapter 4 and 5 baselines that Chapter 6 refit to confirm. The columns also mix tasks (77-way intent, seen-60, unseen-17), so a value in one column is not a rival of a value in another. Read each column on its own, and treat the table as an inventory of what exists, not a leaderboard.
Three sections, each only as far as the evidence
Where the tiny classifier wins. Measured, so it goes first. TF-IDF+LR holds the intent-77 row (0.8779 tuned, Ch4) and the safety-test row is held by fastText-style (0.9052, Ch4). On Ch5’s frozen split the tuned LR reaches 0.900 on seen labels and beats the embedding provider on seen labels by 17-30 points; on Ch6 it tops both NLI models. This is the book’s strongest provider result and it is a classifier.
Where an LLM judge wins. NOT_OBSERVED. No LLM judge has been run in this book, so there is no row to put in the cell. The adjacent evidence is on the judge’s slowest sibling: both NLI models trail the embedding provider (Ch6), which is not a statement about judges. This section exists to say what it cannot say.
Where a decision model wins. NOT_OBSERVED, and there is no region to describe. Every decision-model cell in the table is empty. H4 predicted a narrow region — changing label sets, tight latency, few labels, probabilities needed — but nothing in this book measures a decision model, so the region is a prediction, not a row. Chapter 35 must classify H4 as INSUFFICIENT_EVIDENCE, not as found.
Wrong / Correct. Wrong: “The benchmark shows where each provider wins.” Correct: “The benchmark shows where a classifier wins, and that the other regions are unmeasured. A table with seventeen filled cells and the rest
NOT_OBSERVEDis the honest result.”
What to carry forward
Chapter 29 routes between providers. The router needs cost and quality profiles per provider; this chapter’s honest message is that the quality side of those profiles exists for exactly the providers above, and NOT_OBSERVED for the rest. Routing on empty cells is routing on guesses, and the router chapter must count that.
Close by
Is there any region where the decision model is the best choice? On the measured evidence, no region can be claimed: the cell is empty. The question is passed to Chapters 33 and 35 with INSUFFICIENT_EVIDENCE attached.
Limitations
- Compilation only: no new provider run. Cells are best-
split=testrows and mask seen/unseen and wording distinctions carried in the cited chapters. - LLM-judge and decision-model rows are
NOT_OBSERVEDbecause those providers were never run; a rate like “0.3% of cells filled” is a statement about this book, not about the providers. - The three papers are read at abstract level only; the full HM/IPE in Session R reading pass.
- The same SciFact dev claims served as the “test” split in Chapters 21-23 and 25; the Chapter 35 audit must state what that does to the Part VI pipeline claims (hand-off note).
- Per-unit-cost Pareto fronts are not drawn: cost and latency per provider exist in the cited chapters, and compiling them into one table was deferred to avoid mixing CPU, GPU and cache-economy figures.