The Smallest Decision
What deliberately boring baselines achieve on fixed-label intent and safety tasks — and the measured bar every more sophisticated provider must clear.
Chapters 1 through 3 argued about a product. This chapter builds the thing the rest of the book measures against: five deliberately boring providers, two real datasets, and the shared harness every later chapter imports. No new model, no hosted call, no cleverness. The question is narrow on purpose: how far do boring providers get, and with how few labels? Whatever a sophisticated provider achieves later has to beat this, by a stated margin, at a stated cost — or it has not earned its place.
That last sentence is the chapter’s principle, and it needs a definition before it is a principle. Sophistication earns its place iff it beats the best boring provider by a measured margin on a stated metric at a stated cost. The margin in this chapter is more than ~3 points of accuracy or macro-F1 — because LinguaSynth (Zhang & Mo, arXiv:2506.21848) showed that better features alone buy ~3 points inside the transparent linear family, so a smaller win cannot be attributed to modeling — or strictly lower CPU latency at parity (±1 point), with training seconds, labels and memory reported alongside. “Earn” is now a number, not an attitude. Every later chapter that claims a win will be graded against it.
The providers
Five, all implementing the DecisionProvider protocol (decide(request) -> DecisionResult) from the contract Chapter 2 built. The caller cannot tell them
apart by shape — that interchangeability is the H2 claim, and Chapter 8 tests
whether it survives contact with real callers. If you have read Models From First
Principles you know the pattern: one interface, many implementations, the
differences measured rather than asserted. This chapter does not re-teach it.
- majority — the training majority, always. Trains on nothing. The floor no trained provider may lose to without that loss being the headline.
- lookup — exact-match memory over training texts, majority fallback on anything unseen. It diagnoses memorization: near-100% train accuracy with collapsed test accuracy means lexical overlap did the work, not learning.
- rule — hand-written keyword rules, safety task only. It refuses to fit on more than a handful of labels, because rules stop being honest past that. There is no rule baseline for 77-way intent, and this chapter does not invent one.
- tfidf-lr — word unigrams and bigrams, sublinear TF-IDF, logistic regression. The regularization was tuned once on a declared grid (C × class_weight, rev2, Amendment B) after review caught it sitting at defaults; the tuning budget is the grid, fixed before seeing test, and it is spent, not open.
- fasttext-style — Joulin et al.’s architecture (arXiv:1607.01759,
§2) reimplemented from the paper: hashed word+bigram embeddings averaged,
one linear layer, full softmax NLL, mini-batch SGD. Full softmax instead of
hierarchical (77 classes need no tree), our hashing instead of theirs. Same
architecture, same loss, fewer tricks — and the
fasttextmodule itself is not installed here, so there was never a shortcut to take.
The script below produces that table. It reads the committed results/ch04.jsonl and the harness’s tuning audit (benchmarks/harness/tuning.py), so it runs in a couple of seconds and needs no model and no network.
"""The earn-ledger: what boring gets, what it costs, and whether it was tuned.
python examples/ch04-the-smallest-decision/demo_ch04.py
Reads `results/ch04.jsonl` and prints, per provider and task, the test accuracy,
the warm p95 per-item latency, and the full-train CPU seconds — then the
tuning audit (rev2 guard: a baseline with no declared grid is UNTUNED).
Nothing is measured here; everything is read. The point of the file is that
the chapter's verdict — tuned LR leads intent, the two tie on safety test
inside a wide interval, LR keeps shift — is computed from rows, not quoted
from prose.
Exp selection (rev2): TF-IDF+LR rows come from JEV-04-02R (symmetric tuning,
supersedes JEV-04-01); fastText-style rows from JEV-04-01R; every other
provider is frozen at JEV-04-01.
"""
from __future__ import annotations
import json
import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
RESULTS = ROOT / "results" / "ch04.jsonl"
exps = {"tfidf-lr": "JEV-04-02R", "fasttext-style": "JEV-04-01R"}
frozen = "JEV-04-01"
def exp_for(provider: str) -> str:
return exps.get(provider, frozen)
def value(rows: list[dict], provider: str, dataset: str, metric: str,
seed: int = 0) -> float | None:
want = exp_for(provider)
for row in rows:
if (row.get("exp") == want and row.get("provider") == provider
and row.get("dataset") == dataset
and row.get("metric") == metric and row.get("seed") == seed
and row.get("split") == "test"):
return row.get("value")
return None
def train_seconds(rows: list[dict], provider: str, dataset: str) -> float | None:
want = exp_for(provider)
for row in rows:
if (row.get("exp") == want and row.get("provider") == provider
and row.get("dataset") == dataset
and row.get("metric") == "train_seconds_cpu"
and row.get("split") == "train"
and "full" in str(row.get("note", ""))):
return row.get("value")
return None
def main() -> int:
rows = [json.loads(line) for line in
RESULTS.read_text(encoding="utf-8").splitlines() if line.strip()]
for dataset, providers in [
("banking77", ["majority", "lookup", "tfidf-lr", "fasttext-style"]),
("deepset/prompt-injections",
["majority", "lookup", "rule", "tfidf-lr", "fasttext-style"]),
]:
print(f"{dataset}:")
for provider in providers:
acc = value(rows, provider, dataset, "accuracy")
lat = value(rows, provider, dataset, "latency_warm_p95_ms")
cost = train_seconds(rows, provider, dataset)
print(f" {provider:14s} acc={acc} p95={lat} ms train={cost} s")
sys.path.insert(0, str(ROOT))
sys.path.insert(0, str(ROOT / "src"))
from arbiter.providers.linear import FastTextStyle, TfidfLogisticRegression
from arbiter.providers.lookup import LookupProvider, MajorityProvider
from arbiter.providers.rule import RuleProvider
from benchmarks.harness.tuning import audit_tuning, format_audit
print(format_audit(audit_tuning([
MajorityProvider(), LookupProvider(), RuleProvider(),
TfidfLogisticRegression(), FastTextStyle(seed=0),
])))
return 0
if __name__ == "__main__":
raise SystemExit(main())
python examples/ch04-the-smallest-decision/demo_ch04.py
banking77:
majority acc=0.013 p95=0.0259 ms train=0.0 s
lookup acc=0.013 p95=0.017 ms train=0.002 s
tfidf-lr acc=0.8779 p95=0.6128 ms train=2.722 s
fasttext-style acc=0.8526 p95=0.1368 ms train=186.698 s
deepset/prompt-injections:
majority acc=0.4828 p95=0.0034 ms train=0.0 s
lookup acc=0.4828 p95=0.0036 ms train=0.0 s
rule acc=0.4914 p95=0.0082 ms train=0.0 s
tfidf-lr acc=0.8966 p95=0.4928 ms train=0.051 s
fasttext-style acc=0.9052 p95=0.1925 ms train=7.25 s
tuning audit:
majority DECLARED grid=n/a: no hyperparameters (trains on nothing) split=n/a
lookup DECLARED grid=n/a: no hyperparameters (exact-match memory) split=n/a
rule DECLARED grid=n/a: hand-written rules (nothing fitted) split=n/a
tfidf-lr DECLARED grid=C in {0.1, 1, 10, 100, 1000} x class_weight in {None, balanced} split=threshold
fasttext-style DECLARED grid=lr in {0.05, 0.1, 0.2, 0.5} x epochs in {10, 25} (Amendment A) split=threshold
Those lines are computed from results/ch04.jsonl (3084 rows: 1877 frozen
JEV-04-01 rows, 857 JEV-04-01R fastText-style revision rows, 350 JEV-04-02R
TF-IDF+LR revision rows), not quoted from prose. The TF-IDF+LR lines read
from the second revision (symmetric tuning); the fastText-style lines from
the first; every other line is frozen from the first run. The audit says no
baseline is UNTUNED — that guard did not exist in the first two runs, which
is why this chapter needed a second revision. Read slowly; the rest of the
chapter is commentary on them.
The harness, because the numbers need a guarantor
benchmarks/harness/ implements planning/benchmark-protocol.md without
redesign — loaders with pinned revisions, fixed stratified splits cut once, a
test-touch gate that refuses a second touch of test without a recorded
violation, metrics (accuracy, per-class precision/recall/F1, ECE and AUROC
separately), CPU timing (cold and warm, load never counted), and a results store
that rejects rows with holes in their provenance. Later chapters import it; none
may fork it. If you have read PyTorch From First Principles you know why the
training loop looks the way it does; this chapter does not re-teach that either.
Two harness behaviors deserve naming because they fired during this chapter.
First, the 50- and 100-per-class learning-curve budgets do not fit
BANKING77’s tail classes — some hold 31–47 examples in our 8003-example train
cut — and the harness refused with 24 provider_refused rows instead of
truncating (plus 6 more in the rev1 re-run, same budgets). The intent curve
therefore has two points (10/class and full), not four. That refusal is the
harness working: a budget that does not fit the data is a design error, and
silent truncation would have laundered it. Second, the test gate opened each
task’s test set exactly once in the first run — and exactly once more in the
revision re-run, with a recorded reason and a test_touch_count row (value 2)
per task. The learning curves live on threshold, never on test, so
repeated evaluation during development cost nothing; the two counted touches
are the finding runs, both on the record.
Datasets, with licences and pins in datasets/README.md: banking77
(CC-BY-4.0) for 77-way intent — upstream train cut into train/calibrate/
threshold, upstream test (3,080) as test, no shift split because this chapter
fits neither calibration nor abstention; deepset/prompt-injections (Apache-2.0
on the card, CC-BY-4.0 in the metadata — the stricter applies) for safety, label
1 mapped to block per EvalHub’s config, with jackhhao/jailbreak-classification
as the shift split (different source, same binary labels, never trained on).
The exact Red Hat set ids came from EvalHub’s pinned config, not from the
article’s example links — no guessing, per the kickoff.
What boring gets
On intent (n=3080): tuned TF-IDF+LR 0.8779 accuracy / 0.8786 macro-F1. Majority and lookup both 0.013 — the floor is the class prior over 77 labels, and lookup’s perfect train memory buys exactly nothing on unseen text. fastText-style (rev1) reaches 0.8526 / 0.8524 (0.8526, 0.8516 on seeds 1–2): LR leads by 2.5 points, and the paired bootstrap interval on the difference (+0.025, 95% CI [+0.014, +0.037], seed 0; the seed-1 and seed-2 pairings agree) excludes zero. With symmetric tuning the intent crown goes back to LR — this time significantly.
On safety (n=116): tuned TF-IDF+LR 0.8966 / 0.8964. Majority 0.4828, lookup 0.4828, rule 0.4914 — one point over majority for all that hand-writing, and the strong refutation (the task needs no learning) does not fire. Against fastText-style’s 0.9052 / 0.9051 (mean 0.8908 across seeds), the rev1 twenty-point headline is gone: tuned LR trails by nine-tenths of a point, and the paired interval (−0.009, 95% CI [−0.060, +0.043]) comfortably includes zero. On this distribution, with both baselines tuned on the same terms, the two tie — and the interval says the data cannot tell them apart. Read the n=116 width honestly: the CI spans ten points, and the shift split below is the sturdier safety headline.
On shift (n=262, safety only): tuned TF-IDF+LR 0.5496 / 0.3901 against fastText-style 0.5382 / 0.3644 — eleven-tenths of a point apart, paired CI [+0.000, +0.027], touching zero: a tie with a whisper of an LR edge. But note what tuning cost: the untuned LR scored 0.6298 here, eight points higher. Tuning for threshold and test (C=1000, nearly unregularized) overfit the source distribution. The rule sits at 0.5344, level with both learners. Collapse, not transfer, is still the default across distributions — Haroon (arXiv:2607.14131) measured worse (F1 0.005 one direction), and H3’s transfer question stays genuinely open.
Learning curves (threshold split, means across seeds 0–2 except where noted): intent tuned LR climbs from a mean 0.614 at 10/class (0.623, 0.598, 0.621 per seed) to 0.875 full-train (seed 0; LR is deterministic, one fit is the fit); intent fastText-style climbs from a mean 0.472 to a mean 0.844. Safety tuned LR climbs 0.585 → 0.744 → 0.815 → 0.900 across 10/50/100/full (means); safety fastText-style climbs 0.470 → 0.711 → 0.782 → 0.867 (means). Ten labels per class buys you seven-tenths of full-train accuracy for the tuned linear model — the “how few labels” half of the chapter question, answered for both learners this time.
What boring costs, and the two surprises
Training (full, seed 0, CPU): intent — majority 0.0 s, lookup 0.002 s, tuned TF-IDF+LR 2.722 s, fastText-style 186.698 s. Safety — tuned TF-IDF+LR 0.051 s, fastText-style 7.25 s. The embedding learner costs ~70x the tuned LR fit on intent while losing by 2.5 points — and ~140x on safety for a statistical tie. Tuned LR wins on train cost and on accuracy; the embedding model answers only with inference speed (0.137 ms vs 0.613 ms per item on intent). The earn-ledger entry is per task and per split, not once.
Per-item inference (warm p95, CPU round-trip, seed 0): intent — lookup 0.017 ms, majority 0.0259, fastText-style 0.1368, tuned TF-IDF+LR 0.6128. The linear model is the slowest at inference on both tasks (safety: tuned LR 0.4928 vs fastText-style 0.1925). A 100k-feature dot product costs more than an EmbeddingBag forward. The preregistration said LR “wins on latency” — true for training cost, false per-item against every provider here, and the chapter reports the split rather than smoothing it. Anyone deploying TF-IDF+LR for its speed is optimizing the wrong half of the ledger.
Calibration versus ranking (test, seed 0): intent tuned TF-IDF+LR has ECE 0.0519 with AUROC 0.9971; fastText-style ECE 0.0408 with AUROC 0.9962. Safety: tuned LR 0.0593 / 0.9661; fastText-style 0.0642 / 0.9658. Both boring models now rank excellently and calibrate — tuning fixed LR’s calibration along with its accuracy (its frozen ECE was 0.4747, wildly overconfident). The old contrast between the two providers is gone, which is itself a finding: the overconfidence was regularization, not architecture. The protocol’s separation still does its job — its mirror image stands, when the broken fastText variant scored ECE 0.0 at chance accuracy. Near-uniform uncertainty is trivially “calibrated”, so ECE without accuracy is meaningless. Both numbers, always, or neither.
Wrong: TF-IDF+LR is the fast, cheap baseline that later providers must beat on accuracy.
Correct: TF-IDF+LR is the accurate, cheap-to-train baseline that later providers must beat on accuracy — and it is the slowest thing here at inference time. “Cheap” has two halves and they point in opposite directions.
Against the prediction
The preregistration (commit 8eefe7a, before any download or fit) predicted LR
within a few points of the best later provider with wins on latency and cost.
Against the boring five, third run, both baselines tuned on the same terms:
on intent tuned LR leads 0.8779 to 0.8526 (supported, and the paired
interval [+0.014, +0.037] excludes zero — a real lead, not noise); on safety
test the two tie at ~0.90 (tuned LR 0.8966, fastText-style 0.9052, paired CI
[−0.060, +0.043] includes zero — the rev1 twenty-point headline does not
survive symmetric tuning); on shift the two tie at ~0.54–0.55 (paired CI
[+0.000, +0.027], touching zero). Lookup collapses on test and majority
floors hold as predicted; the rule beats nothing trained, so the strong
refutation does not fire; LR trains in seconds (supported); LR does not
win per-item latency (refuted — lookup and fastText-style are both
faster).
The honest paragraph this revision owes you. The first run reported fastText-style thirty points behind on intent and collapsed on safety, with an untested story about embeddings that 366 examples cannot move. External review found the story was covering an implementation defect: mean-reduced batch loss with a fixed rate had shrunk every update ~32x, so the baseline never trained. A corrected trainer (same config, per-example-scale updates, decaying rate — grid-selected on threshold in Amendment A of the notes) closed the whole thirty-point intent gap untuned (0.5403 to 0.8526) and took safety test by a finding — the exact failure Lipton and Steinhardt describe, caught here by a reviewer asking why a standard method sat exactly on the majority line. The old rows are still in the results file, marked superseded; the chapter now cites only the new ones.
The second revision owes one more sentence: tuning one baseline and not the other is as bad as tuning neither — the rev1 safety headline (fastText-style by twenty) evaporated the moment LR got its own declared grid, and the paired intervals are the reason the chapter can say “evaporated” instead of “maybe”. Grinsztajn et al. (arXiv:2207.08815) give this chapter its principle and its warning together: boring stays on top even after large tuning budgets on tabular data — but their findings are about tabular inductive biases, and this chapter must not cite them as text findings. What transfers is the method (§3: budgeted searches, released raw baselines), not the mechanism. Lipton and Steinhardt (arXiv:1807.03341) supply the rest: error analysis, ablations and robustness checks “can be adopted by everyone”, and this harness’s per-class profiles, learning curves and shift split are those three practices. “Decision model” remains at risk of suitcase-word status (§3.4.3); every claim in this chapter names its property.
Limitations
- Intent learning curves have two points, not four. The 50/100-per-class budgets do not fit BANKING77’s tail classes; the harness refused rather than truncate (24 rows first run, 6 rev1, 6 rev2). Safety has all four points.
- The first-run fastText-style verdict was a defect, corrected in rev1. Mean-reduced batch loss with a fixed rate under-trained the baseline ~32x; the “thirty points behind” finding and the “cannot move embeddings” story are withdrawn. Corrected grid (lr × epochs, Amendment A), JEV-04-01R rows, old rows superseded in place. The selftest now includes the regression test that would have caught it (separable-data fit + batch-size invariance).
- The rev1 safety headline did not survive symmetric tuning (rev2).
fastText-style sat at defaults-versus-defaults +19 points on safety test;
with LR on its own declared grid (C × class_weight, Amendment B) the gap
is nine-tenths of a point inside a ten-point interval. Tuning one baseline
and not the other is as bad as tuning neither — the audit guard
(
tuning.py, 34 selftest checks) exists so the next asymmetry is flagged, not found. - The revisions touched
testagain. Touch count is now 3 per task: each run wrote onetest_touch_countrow (values 2 in rev1, 3 in rev2), gate violations recorded with reason each time, Amendments 2 and 3 in the protocol. Majority, lookup and rule rows are frozen from the first run and byte-identical. - fastText-style needed three training variants before rev1 (full-batch GD, decaying mini-batch SGD, constant-LR mini-batch SGD); the choice was made on train/threshold behavior only, but 200 test items were evaluated once with a broken config during diagnosis — exploratory, unrecorded, disclosed here.
- Full-train fits run twice (once in the curve loop at budget=full, once in the test section). Redundant compute, honestly recorded in duplicate rows; later chapters should drop the curve-full budget.
- No intent shift split (no calibration/abstention fitted; protocol-compliant).
- Safety test is n=116. The paired interval on the safety-test gap spans ten points ([−0.060, +0.043]) — the data cannot separate the two tuned baselines there. The shift split (n=262) is the sturdier safety headline, and there the tuned pair ties at ~0.54–0.55. Neither number travels without its interval.
- Tuning for test cost shift. Tuned LR scores 0.5496 on shift against 0.6298 untuned — eight points down (unpaired comparison, read with the interval width). C=1000 overfit the source; regularization trades test for shift, and the chapter reports both sides.
- Memory is a tracemalloc floor — native allocations escape it, and it is labeled as such in every row.
- LinguaSynth and Benchmarks-Mislead read PARTIAL (results sections); the ~3 pp featurization headroom and the 4–7 pp in-domain band are cited only for what was read.
- No live Jev behavior; no hosted calls. CPU numbers only, this machine.
What Chapter 4 leaves behind
src/arbiter/providers/— the protocol plus majority, lookup, rule, symmetrically tuned TF-IDF+LR and corrected fastText-style.benchmarks/harness/— loaders, splits (+TestGate), metrics (+paired bootstrap), timing, store, tuning audit.datasets/README.md— licences, pins, split construction, what was not used.examples/ch04-the-smallest-decision/— runner, two re-runners, selection scripts, 34-check selftest, demo (+audit), requirements, README.results/ch04.jsonl— 3084 rows (1877 JEV-04-01 frozen + 857 JEV-04-01R + 350 JEV-04-02R).evidence/notes-ch04.md(Amendments A grid selection and B),research/ch04-additions.md,evidence/ledger.md(claims 4.1–4.16, LR rows now rev2, fastText rows rev1).
What would a provider need to offer that these cannot? Three things, each naming its chapter: judgments on labels it never trained on (Chapter 5); a number with its answer that behaves like certainty under shift (Chapters 9–11); and decisions that compose — one state, many questions, answered together (Chapter 7 onward). Boring has set the final bar per task and per split: intent 0.8779 (tuned TF-IDF+LR, paired interval excludes zero — the named bar, trainable in seconds); safety test ~0.90 tied (fastText-style 0.9052 nominal, tuned LR 0.8966, interval includes zero — beat the tie by more than ~3 points); safety shift ~0.55 tied (tuned LR 0.5496 nominal, fastText-style 0.5382). Later providers must clear the stronger per split by more than ~3 points, with an interval, with the latency ledger still pointing the wrong way for LR. Everything after this chapter is trying to clear it.