Tuning an Arbiter
Define Arbiter-2 as a tuned decision model, then design the fair three-way comparison that decides whether tuning earns its cost.
Design draft: this chapter argues from the literature and from earlier measured results; its own experiment has not been run.
Up to now the book has mostly read decisions from frozen models. The First Token built an option-scoring provider without running its model sweep. Decisions as Entailment measured off-the-shelf NLI and found it losing to embeddings. Reading the Hidden State designed, but did not run, the probe comparison. This chapter asks the next question: if you are allowed to change weights, what is the cheapest change that makes a small model good at decisions?
The candidate is Arbiter-2: a small causal model tuned to map (state, question, options) directly to a decision distribution. The chapter does not claim that such tuning is worthwhile. Its job is to define Arbiter-2 sharply enough that a fair measurement can reject it. In particular, the classical baseline gets the same labels, the same splits and the same uncertainty reporting. If tuning cannot beat that baseline by a validated margin, the book keeps the cheaper provider and says so.
The run needs a GPU and cannot happen yet. What can happen now is the part that decides whether its comparison would be fair, and building that part turned up a flaw in the design as first preregistered. Both are in this chapter.
What the tuning literature promises, and conditions
LoRA: train a sliver of the model
DOCUMENTED: Hu and colleagues’ LoRA freezes pretrained weights and trains low-rank update matrices instead. For GPT-3 175B, they report roughly 10,000× fewer trainable parameters and about one-third the GPU memory versus Adam full fine-tuning, with quality on par or better on RoBERTa, DeBERTa, GPT-2 and GPT-3. Merged adapters add no inference latency by construction. [Hu et al., arXiv:2106.09685, Abstract, §§1, 4.2, 7.2]
Their rank study matters for our grid: very low rank can suffice, but the right rank is empirical and not universal. Two limits matter for us. Training still needs the base model resident, and batching different tasks with merged adapters in one forward pass is nontrivial. Those limits shape the cost ledger as much as accuracy does.
Distillation: learn from a teacher’s soft targets
DOCUMENTED: Hinton and colleagues’ distillation trains a small student on a cumbersome teacher’s soft targets, often at elevated temperature. High temperatures expose the ratios among small probabilities, and matching logits is the high-temperature limit. On MNIST, a small network drops from 146 errors to 74 using only the teacher’s soft targets. On speech, most of a ten-model ensemble’s frame-accuracy gain transfers to one distilled model. [Hinton et al., arXiv:1503.02531, Abstract, §§1–4.1]
But very negative logits may be noisy, and intermediate temperatures can win when the student is small. Distillation is therefore not “use soft targets”; it is a temperature-and-weight design problem.
Rationale distillation: the one that complicates a simple race
DOCUMENTED, and it challenges a simple three-way race: Hsieh and colleagues add teacher rationales as multi-task supervision. Step-by-step distillation beats standard finetuning and standard distillation with much less data: on e-SNLI, 12.5% of the data beats full-data standard finetuning, and a 770M T5 beats 540B PaLM few-shot CoT while using 80% of the ANLI data that standard finetuning cannot match even at 100%. [Hsieh et al., arXiv:2305.02301, Abstract, §§3–4.4, Limitations]
But rationale quality matters, smaller teachers help less, and teacher bias transfers to the student. Our preregistered comparison is therefore LoRA against standard distillation against plain supervised tuning. Rationale supervision is an explicitly deferred exploratory arm, not a hidden fourth contestant that can rescue a failed prediction.
PROPOSED: Deferral is not dismissal. If rationale supervision later beats all three preregistered methods, that result starts a new preregistration rather than amending this one after the fact. The book has already paid for post hoc rescue once, in Chapter 4’s asymmetric-tuning revision, and the protocol now charges interest: exploratory work is labelled exploratory, kept off the test split, and never presented as predicted.
The cost bar, set by a third party before any tuning begins
DOCUMENTED: AnyJev’s technical report shows label-free readout improvements doing real work: rotations lower order-flip rates and raise accuracy on 11 of 11 models, while a label-free stopping rule reads 7.3 of 18 rotations and serves 2.2× decisions per second. The same report finds per-question heads beat readout locally but shared heads fail on held-out questions. [AnyJev Technical Report, arXiv:2610.00831, Abstract, §§5–9]
Tuning must therefore beat a moving target: not an untuned baseline, but the selected readout plus stopping and calibration machinery.
PROPOSED: That moving target is why JEV-14-01 refuses to compare a tuned model against a weak default. The readout baseline must be the selected Chapter 7 cell, with rotation, prior correction and temperature already applied. Otherwise tuning can “win” by rediscovering preprocessing. The same logic applies to the classical baseline: it must be the rev2 tuned LR configuration, not an untrained default. Chapter 4 learned this the expensive way through two revisions, first an under-trained fastText arm and then asymmetric tuning, and the design keeps both mistakes visible by freezing the symmetric rule before any run.
If you have read Models From First Principles, the student, teacher and base-model vocabulary below needs no re-teaching. If you have read PyTorch From First Principles, frozen parameters, adapter updates and temperature-scaled softmax need no re-teaching either. This chapter uses those abstractions; it does not rebuild them.
Wrong: “Tuning is justified when the tuned model has the highest accuracy.”
Correct: “Tuning is justified when it wins accuracy-per-label and accuracy-per-GPU-hour against the same-label classical baseline, and when that win survives across base models.”
flowchart TD
A[same labels + splits] --> B[LoRA readout]
A --> C[teacher distillation]
A --> D[plain supervised tuning]
A --> E[classical baseline CPU]
B --> F[accuracy/ECE/cost ledger]
C --> F
D --> F
E --> F
F --> G{beat baseline by validated margin?}
G -- yes --> H[Arbiter-2 candidate]
G -- no --> I[keep cheaper provider]
What earlier chapters already established
OBSERVED: The Smallest Decision sets the bar tuning must clear. Tuned TF-IDF+LR reaches 0.8779 on intent test, with a paired interval excluding zero against fastText-style; safety test is approximately a 0.90 tie and safety shift approximately a 0.55 tie. The earn rule is explicit: beat the boring provider by more than about three points with an interval, or match it with strictly lower cost. (metadata/04-chapter.yaml, thesis and claims 4.6–4.8, 4.14, 4.17.)
OBSERVED: Chapter 6 shows that a new model family does not earn its place by existing. Its best NLI trails the embedding provider by 0.115 on unseen labels, with an interval excluding zero, while elaborated templates collapse BART onto one label. Tuning therefore enters against two warnings: mechanism novelty is not evidence, and prompt or template sensitivity can dominate model choice.
PROPOSED: Two baselines are dependencies, not gaps to fill with optimism. JEV-07-03 has not measured calibrated readout, and JEV-13-01 has not measured probes. A tuning chapter written before those runs must say what happens if they never arrive. The corresponding cells stay NOT_OBSERVED, and the surviving comparison, tuned LM versus classical baseline, still answers the chapter’s cost question. What it cannot answer without them is whether tuning beats the best readout or probe. That incompleteness is recorded, not hidden, and no friendlier comparator is substituted for either missing cell.
The design
PROPOSED: Arbiter-2 is a small tuned causal LM whose output is a distribution over the declared options. Three methods train on identical nested label budgets, with seeds 0 to 2: 100, 300, 1,000 and 3,000 intent labels, and 90, 180 and 366 safety labels.
- LoRA readout: frozen base model plus LoRA on attention query and value projections; rank grid r in {4, 8, 16} with alpha=2r and dropout=0.05; learning-rate grid {1e-4, 2e-4}; epochs=3; hard-label cross-entropy.
- Distillation: a student trained on fixed teacher soft targets plus a small hard-label term; temperature grid T in {1, 2, 4}; soft-target weight grid {0.7, 0.9}; student learning-rate grid {1e-4, 2e-4}; epochs=3. The preregistered teacher is Chapter 4’s rev2 tuned LR plus its Chapter 10 primary temperature map: strong, available and boring by design.
- Plain supervised tuning: full-parameter finetuning on hard labels; learning-rate grid {1e-5, 3e-5}; epoch grid {2, 3}.
The same selection rule applies to every method on the threshold split: highest threshold accuracy, then lower ECE, then smaller rank, lower temperature and fewer epochs. Each tuned model then receives one post-hoc temperature from the fixed Chapter 7 grid on the calibrate split. Test is evaluated once. Safety shift is reported separately.
PROPOSED: The teacher is boring on purpose. A stronger but exotic teacher would entangle two questions: whether distillation works and whether that particular teacher is good. Tuned LR plus its primary temperature map is already the book’s fixed-label champion, with known accuracy, known ECE behaviour and no prompt sensitivity. The distillation temperature grid preserves Hinton’s warning: T=1 keeps the teacher’s confident ranking, while higher temperatures expose the small-probability structure a small student may need, or may find noisy. The grid, not intuition, decides. Teacher calls are counted because each soft target is a dependence: a method that needs the teacher forever is less deployable than one that needs it once.
The base-model axis separates method effects from model effects: Qwen3-0.6B, Qwen3-1.7B-Instruct and Llama-3.2-1B-Instruct, three sizes across two families, reusing Chapter 7’s pinned revisions. The smallest model tests the compression question most directly; the two larger models test whether a tuning ranking survives scale and tokenizer family. A CPU-only Chapter 4-style classical baseline trains on the same label budgets. It is both the fixed-label comparator and the cost comparator: it needs no GPU, no teacher calls and no adapter storage.
Calibration is reported, not used as rescue. Every tuned distribution is evaluated for accuracy first; ECE is reported both raw and after one fixed temperature fit. A method that gains accuracy while destroying calibration has not won the decision question; it has moved the failure from the label to the probability. Likewise, a method that needs extensive temperature search has a hidden label cost, because calibration labels are labels. The fixed Chapter 7 temperature grid keeps that cost visible and equal.
The cost ledger is a first-class result. Labels are the scarcest resource in the chapter’s question; teacher calls price distillation’s dependence; GPU-hours price the tuning itself; adapter size prices deployment and task switching. A method can therefore lose in two distinct ways: by being less accurate at equal cost, or by being equally accurate at greater cost. P4 exists to force the second comparison. If the classical baseline ties on accuracy, its CPU-only ledger wins without further argument.
Seeds mean the same thing everywhere here. Tuned inference is deterministic; seeds drive nested label sampling, solver initialisation where applicable, and bootstrap resampling. Three seeds therefore do not create three runs of a stochastic method; they create three comparable draws of the scarce-label regime the chapter is about. Identical nested subsets across methods and models keep the comparison paired.
The design, made executable
The checks are src/arbiter/tuning_design.py, tested by 8 tests in tests/tuning/. Everything below is examples/ch14-tuning-an-arbiter/walkthrough_ch14.py, which you can run as it stands. It needs no model and trains nothing. Section 3 reads the real BANKING77 training labels from the harness’s local cache, which is a dataset statistic and not a model result; the rest is arithmetic on the preregistered design.
The setup
import dataclasses
from pathlib import Path
import sys
ROOT = Path(__file__).resolve().parents[2]
sys.path[:0] = [str(ROOT), str(ROOT / "src")]
from arbiter.tuning_design import (
MethodSpec,
SelectionRule,
TuningPlan,
missing_classes,
nested_label_subsets,
verify_plan,
)
from benchmarks.harness import loaders, splits
PLAN = TuningPlan(
methods=(
MethodSpec("lora", {"rank": (4, 8, 16), "lr": (1e-4, 2e-4)}, {"epochs": 3}),
MethodSpec("distillation",
{"temperature": (1, 2, 4), "soft_weight": (0.7, 0.9), "lr": (1e-4, 2e-4)},
{"epochs": 3}),
MethodSpec("supervised", {"lr": (1e-5, 3e-5), "epochs": (2, 3)}),
),
models=("qwen3-0.6b", "qwen3-1.7b-instruct", "llama-3.2-1b-instruct"),
seeds=(0, 1, 2),
budgets={"intent": (100, 300, 1000, 3000), "safety": (90, 180, 366)},
)
def main() -> None:
1. The plan, checked and counted
verify_plan returns the ways a plan would compare methods unfairly, and an empty list passes. fit_count is arithmetic: grid cells per method, times budgets, models and seeds.
# 1. The preregistered design, checked and counted.
print("1. the preregistered plan")
print(f" violations: {verify_plan(PLAN)}")
for m in PLAN.methods:
print(f" {m.name:<13} {m.n_cells:>2} grid cells ({', '.join(f'{k}={len(v)}' for k, v in m.grid.items())})")
cells = sum(m.n_cells for m in PLAN.methods)
budgets = sum(len(b) for b in PLAN.budgets.values())
print(f" {cells} cells x {budgets} budgets x {len(PLAN.models)} models x {len(PLAN.seeds)} seeds = {PLAN.fit_count()} grid fits")
for minutes in (1, 5, 15):
print(f" at {minutes:>2} min per fit (an assumption, not a measurement): {PLAN.fit_count() * minutes / 60:,.0f} GPU-hours")
The per-fit times are an assumption for sizing and not a measurement; this book has no measured fit time on this hardware. The point is the order of magnitude. A full grid of this design is on the order of a thousand fits, which is why the stage plan shards it into resumable pieces and why a GPU run cannot be improvised.
2. One selection rule for every method
# 2. One selection rule for every method.
print("2. the shared selection rule")
rule = SelectionRule()
candidates = [
{"name": "A", "threshold_accuracy": 0.85, "ece": 0.09, "rank": 16},
{"name": "B", "threshold_accuracy": 0.85, "ece": 0.04, "rank": 8},
{"name": "C", "threshold_accuracy": 0.85, "ece": 0.04, "rank": 4},
{"name": "D", "threshold_accuracy": 0.80, "ece": 0.01, "rank": 4},
]
print(f" accuracy, then ECE, then rank: picks {rule.select(candidates)['name']} (D has the best ECE but loses on accuracy)")
The accuracy-first rule picks C. D has the best ECE of the four but loses on accuracy, which is the rule working as written: calibration is reported beside accuracy and does not rescue it.
3. Nested label budgets, and what a small budget never shows
The budgets are nested prefixes of one seeded ordering, so every method and model sees the same labels. Reading the real training labels shows what the smallest budget costs.
# 3. Nested label budgets, and what a small budget never shows.
print("3. nested label budgets on the real BANKING77 training labels")
bank = loaders.load_banking77()
cut = splits.stratified_cut(bank["train_texts"], bank["train_labels"],
train_n=8003, calibrate_n=1000, threshold_n=1000, seed=0)
classes = sorted(set(cut.train_labels))
index = {c: i for i, c in enumerate(classes)}
labels = [index[c] for c in cut.train_labels]
print(f" {len(labels)} training items, {len(classes)} classes")
print(" budget missing classes (seeds 0, 1, 2)")
for budget in (100, 300, 1000, 3000):
row = [missing_classes(labels, nested_label_subsets(len(labels), [budget], s)[budget], len(classes)) for s in (0, 1, 2)]
print(f" {budget:>6} {row}")
nested = nested_label_subsets(len(labels), [100, 300, 1000, 3000], seed=0)
print(f" 100 inside 300 inside 3000: {set(nested[100]) <= set(nested[300]) <= set(nested[3000])}")
4. The verifier earns its place
# 4. The verifier earns its place by catching the mistakes Chapter 4 made.
print("4. the verifier catches deliberate mistakes")
no_lr = dataclasses.replace(PLAN, methods=(MethodSpec("lora", {"rank": (4, 8, 16)}), *PLAN.methods[1:]))
print(f" LoRA without a learning rate -> {verify_plan(no_lr)}")
empty = dataclasses.replace(PLAN, methods=(MethodSpec("supervised", {"lr": (), "epochs": (2, 3)}),))
print(f" an empty learning-rate grid -> {verify_plan(empty)}")
reordered = dataclasses.replace(PLAN, budgets={"intent": (300, 100)})
print(f" budgets that cannot nest -> {verify_plan(reordered)}")
What it prints
1. the preregistered plan
violations: []
lora 6 grid cells (rank=3, lr=2)
distillation 12 grid cells (temperature=3, soft_weight=2, lr=2)
supervised 4 grid cells (lr=2, epochs=2)
22 cells x 7 budgets x 3 models x 3 seeds = 1386 grid fits
at 1 min per fit (an assumption, not a measurement): 23 GPU-hours
at 5 min per fit (an assumption, not a measurement): 116 GPU-hours
at 15 min per fit (an assumption, not a measurement): 346 GPU-hours
2. the shared selection rule
accuracy, then ECE, then rank: picks C (D has the best ECE but loses on accuracy)
3. nested label budgets on the real BANKING77 training labels
8003 training items, 77 classes
budget missing classes (seeds 0, 1, 2)
100 [19, 24, 22]
300 [2, 2, 2]
1000 [0, 0, 0]
3000 [0, 0, 0]
100 inside 300 inside 3000: True
4. the verifier catches deliberate mistakes
LoRA without a learning rate -> ["lora: knob 'lr' is neither gridded nor fixed"]
an empty learning-rate grid -> ["supervised: knob 'lr' has an empty grid"]
budgets that cannot nest -> ['intent: budgets must be strictly increasing so they can nest']
Read it in order.
- Section 1: the plan verifies. The three methods have 6, 12 and 4 grid cells, 22 in all, and the design implies 1,386 grid fits.
- Section 2: the rule picks
Cand not the best-ECE candidate. - Section 3: on the real intent training labels, a 100-label budget misses 19, 24 and 22 of the 77 classes on seeds 0 to 2. At 300 labels two classes are missing on every seed, and from 1,000 labels up none. Missing classes at the smallest budget are a property of the data and not of any method, which is why they are reported and not smoothed.
- Section 4: a LoRA grid without a learning rate, an empty learning-rate grid and budgets that cannot nest are each caught by name. The first is the kind of incomplete grid this book’s own Chapter 4 once shipped.
A flaw the checks surfaced
Section 3 of the output has a consequence the original design missed. The preregistered teacher is Chapter 4’s tuned LR, trained at the maximum available budget, which means on all 8,003 intent training labels. At a student budget of 100 labels, the student sees 100 items, but each item carries a soft target from a teacher that has seen every one of the 77 classes. A 100-label subset misses about a quarter of the classes; the teacher’s soft targets quietly supply them.
So distillation’s “100 labels” is not 100 labels. The student is fed information from 8,003. Prediction P1, that distillation leads at low budgets, would then hold close to by construction, and holding by construction is not a finding.
PROPOSED, as a dated amendment to the preregistration before any run, stricter and not softer: distillation runs with two teachers.
- T-full is the preregistered teacher. Its label cost is charged honestly: the teacher’s 8,003 training labels plus the student’s. It answers a different question: how well can a small model compress a known strong teacher?
- T-matched is the same classical model trained only on the student’s own n nested labels. Its label cost is n. It answers the label-budget question, and P1 is evaluated on T-matched.
The amendment is recorded in metadata/14-chapter.yaml and the ledger. It also reflects a general lesson of Chapters 4 and 5: a comparison is only as fair as its accounting, and the accounting is easiest to get wrong where something looks free.
The experiment, designed but not run
JEV-14-01 is preregistered with status: NOT_RUN; the full block is in metadata/14-chapter.yaml, including the amendment above.
Predictions (HYPOTHESIS):
- P1. At 100 and 300 intent labels, distillation with the budget-matched teacher leads on accuracy and ECE in most model-budget cells.
- P2. At 1,000 and 3,000 labels, plain supervised tuning leads on accuracy in most cells.
- P3. No method wins on all three base models at full budget.
- P4. At full intent budget, no tuned LM beats the same-label classical baseline by more than 3 points with a paired interval excluding zero. A larger validated win refutes P4 and adopts Arbiter-2.
The cascade rule is preregistered with the prediction: if tuning does not earn its place, the classical provider stays in the cascade. That is a result, not a failure.
PENDING_RUN: result for JEV-14-01, P1–P4 Command:
python examples/ch14-tuning-an-arbiter/run_ch14.py --models qwen3-0.6b,qwen3-1.7b-instruct,llama-3.2-1b-instruct --budgets 100,300,1000,3000 --seeds 0,1,2 --teachers full,matchedFills:results/ch14.jsonl,static/figures/ch14-accuracy-per-cost-*.pngStage plan (each ≤15 min, GPU-bound, resumable; adapters outside git): S1 prepare identical nested budgets (nested_label_subsets) and cache teacher targets for both teachers; S2 tune LoRA/distillation/SFT shards; S3 select cells with the shared rule and fit post-hoc temperature; S4 evaluate test once plus safety shift; S5 record cost ledger and write rows/figures.
What would change your mind
A universal tuning winner across all three base models would refute P3 and reopen the “one best way to tune” question. A large validated win over the classical baseline would refute P4 and make Arbiter-2 the default. A rationale-supervised exploratory arm that beats all three preregistered methods would not refute P1 or P2; it would motivate a new preregistered comparison, because rationale supervision was never in this race.
If T-full beats T-matched by a wide margin at 100 labels, that is not a distillation win. It is a measurement of how much the teacher’s extra labels were worth, which is useful, and different.
No pivot is proposed: a design supplies no new measured evidence. H3 stays INSUFFICIENT_EVIDENCE; transfer is Chapter 15’s question, not this chapter’s verdict.
Limitations
JEV-14-01 is NOT_RUN, and results/ch14.jsonl does not exist. The GPU-hour figures in section 1 are arithmetic on an assumed fit time, not measurements.
Required papers are PARTIAL by section; appendices, hyperparameters and code were not reproduced. AnyJev material is third-party, not vendor documentation. Two of the three named baselines (JEV-07-03 readout and JEV-13-01 probes) are deferred dependencies, and unavailable cells stay NOT_OBSERVED. Pinned weights may be unavailable at run time; substitution needs a dated amendment.
Sparse 100-label budgets on 77 classes miss classes, as section 3 measures; missing-class behaviour is reported, not smoothed. ECE beside accuracy does not certify calibration. Rationale supervision, teacher quality and cross-task transfer remain outside the preregistered test. The safety label budgets are smaller than the intent budgets, so safety intervals will be wider; reporting them wide is part of the acceptance rule, not an apology for it.
The design checks catch an incomplete grid, an inconsistent selection rule and budgets that cannot nest. They would not have caught the teacher-leak flaw on their own; that came from reading the missing-class counts against the preregistered teacher. Other flaws of that kind may remain.
What the next chapters inherit
Chapter 15 inherits the tuning machinery and the cost ledger: transfer must be tested with tuned models whose price is already known. Chapters 16 to 19 inherit Arbiter-2 only conditionally. If P4 holds, the language work builds on the cheaper provider, not on a tuned model that failed to earn its place. Chapter 28 inherits the accuracy-per-cost framing: a provider wins by occupying the Pareto front, not by winning accuracy alone.
The prompt’s closing question is therefore conditional by design. Arbiter-2 exists as a specification. Whether it is better than the cheapest thing that already worked is exactly what JEV-14-01 is built to decide, and the receipt, labels, teacher calls and GPU-hours, always matters more than the method. A later success with rationale supervision, larger teachers or cross-task transfer would be interesting, but it would not retroactively validate LoRA, standard distillation or plain supervised tuning. Each claim needs its own preregistered race.