The First Token
What tokenisation does to option scoring: which readouts exist at 77 options, which are impossible, and what that does to the claim that a decision model is just a way of asking a model.
Chapter 4 fixed a world with 77 labels and put a number on it. Chapter 5 loosened one joint: the label set could be described at run time instead of trained in. Both chapters compared providers on accuracy. This one asks a question that comes before any provider runs.
Here is the shortest version of it. A decision model that returns a probability
over a set of options must, somehow, turn “which of these?” into a number. The
most common way is to write the options into a prompt and read the model’s
distribution over whatever comes next. If the options are printed as letters,
that is a probability over 26 possible answers. If they are printed as words, it
is a probability over strings of different lengths. If they are printed as
numbers, it is a probability over digits — and digits are not the same size in
every tokenizer. The third-party account of how Jev works that this chapter set
out to rebuild (research/primary-sources.md, Victor Dibia’s write-up — not
vendor documentation; the vendor has published no internals) describes exactly
this: a prompt that ends where the answer goes, then each option scored by its
log-probability, then a softmax across options. That is the mechanism. It is
also, as we are about to see, at least four different mechanisms depending on how
the options are printed.
The claim under test is the book’s own: H1, that decision models reduce to known techniques plus packaging. If “option scoring” turns out to be one thing, H1 wins its easiest clause. If it turns out to be a family whose members disagree about whether they exist at all at 77 options, then the clause was never a claim about one thing, and the useful question becomes which member you picked and why.
What we expected, and why
Robinson & Wingate (arXiv:2210.12353) are the source of the technique this chapter builds: present the options as symbol-enumerated lines and read the probability of the symbol at the answer position. Their §3 lists four problems with the older alternative — scoring each option’s text separately — and the third is the one that matters here: with cloze-style scoring you must choose a normalisation, and that choice “often incur[s] a computational cost or depend[s] on choice of tokenization scheme”. Their Table 3 measures it: randomly re-casing or re-spacing the option text costs the cloze approach 12.4% and 10.3% accuracy, and costs the letter approach 1.3% and 0.5%. Surface form is a big deal for one readout and nearly irrelevant for the other. That is a strong hint that the printing of the options is not a formatting detail.
Zheng et al. (arXiv:2309.03882) qualify the
technique in a way that a later chapter will have to fight: the bias that makes
option order matter is, in §2.4, mostly token bias rather than position bias —
the model a priori assigns more mass to some ID tokens — and removing the IDs
reduces it while degrading accuracy. Their §2.4 also reports that swapping the
symbol alphabet (a/b/c/d, 1/2/3/4, (A)/(B)/(C)/(D)) does not help. If the
choice of symbol does not matter for bias, then our chapter’s question is
somewhere else. Zhao et al. (arXiv:2102.09690)
supply the correction the third-party account assumes: divide out a prior
estimated from a content-free input.
Then, in the last twelve months, the prior art moved. LLM-as-Jev
(arXiv:2610.02076) presents bracketed numeric
identifiers [1] … [K] and argues in §3.2 that they are prefix-free because
every suffix ends in ], so the option count is unbounded where letters stop at
26. Its Appendix C observes that Qwen tokenises digits individually, so “options
1 and 10 share their first token” — and its Appendix H reports a benchmark where
fine-tuning moved answers from option 10 to option 1, with the honest admission
“We have not identified the cause”. Four searches of arXiv (recorded with dates
in evidence/notes-ch07.md) found no paper that measures first-token collision
rates over a real label set. So the measurement below is not a reproduction; it is
the missing measurement.
So we predicted, and committed before running (metadata/07-chapter.yaml,
commit a19ff6b), twelve token-level predictions. The honest summary of what we
expected: letters would be collision-free and would stop existing above 26
options; bare digits 1–77 would be both non-prefix-free and collision-heavy;
bracketed identifiers would be prefix-free but otherwise unchanged; and the 77
label names would collide heavily enough to block a first-token readout of the
strings. Six held. We report the six refutations with the same prominence.
The build
src/arbiter/providers/option_scoring.py, written from scratch on the Chapter 2
contract. Nothing was forked. It has two readouts, named for what they are
rather than for what they are supposed to be:
- letter readout — the probability of a single-token symbol at the answer
position, restricted to the option symbols and renormalised. Robinson &
Wingate’s MCP; AnyJev’s
raw. One token per option, so one forward pass. - full-option-string log-probability — the sum of the option’s token log-probabilities, softmaxed across options. Robinson & Wingate’s CP; LLM2Jev’s Eq. (1).
Two decisions were declared before any run and are visible in the code. The
multi-token rule is the sum, not the mean, because a sum is the joint
likelihood of strings of unequal length and a mean is not a probability of
anything; the cost is that the score structurally prefers short options. (The
same confound appears independently in 2607.27421’s calibration section, which
warns that the sequence-level log-probability it uses is “sensitive to output
label length”.) And the letter readout is capped at 26 options, raising at
fit rather than at the three-thousandth item — a ceiling we took from AnyJev’s
own code, MAX_OPTIONS = 26, because 26 is where the letters run out and a
provider that truncates a 77-way question is worse than one that refuses it.
The tokenisation finding
Run now, on this machine, CPU only, tokenizer only. 6001 rows in
results/ch07.jsonl, all tagged "mode": "observed-cpu". Three pinned
tokenizers: Qwen3-1.7B (instruct), Qwen3-1.7B-Base, and Llama-3.2-1B-Instruct as
a second family. Twelve option sets (the 77 label names, four Chapter 5 wordings
of all 77, the 17 held-out labels under four wordings, the two safety labels
under four wordings) and six ways of printing them.
The script below produces that table. It reads the committed results/ch07.jsonl, so it runs in a fraction of a second and needs no model: the tokenizers were the only thing consulted when the rows were written.
"""The chapter's code block: the tokenisation ledger, computed from
results/ch07.jsonl. No model, no GPU, no network, under a second.
python examples/ch07-the-first-token/demo_ch07.py
Everything printed is read out of the results file or recomputed from it. If a
number here is not in that file, this script is wrong and the chapter is wrong
with it; that is the point of generating the block rather than typing it.
"""
from __future__ import annotations
import json
import sys
from collections import defaultdict
from pathlib import Path
ROOT = Path(__file__).resolve().parents[2]
sys.path.insert(0, str(ROOT))
sys.path.insert(0, str(ROOT / "src"))
RESULTS = ROOT / "results" / "ch07.jsonl"
PRIMARY = "qwen3-1.7b-instruct"
SECOND = "llama-3.2-1b-instruct"
rows = [json.loads(l) for l in RESULTS.open(encoding="utf-8")]
def cell(tokenizer: str, option_set: str, presentation: str,
metric: str) -> float | None:
for r in rows:
if (r.get("provider") == f"tokenizer/{tokenizer}"
and r.get("option_set") == option_set
and r.get("presentation") == presentation
and r.get("metric") == metric):
return r["value"]
return None
def infeasible(tokenizer: str, option_set: str, presentation: str) -> bool:
for r in rows:
if (r.get("provider") == f"tokenizer/{tokenizer}"
and r.get("option_set") == option_set
and r.get("presentation") == presentation
and r.get("metric") == "presentation_feasibility"):
return r["value"] == 0.0
return False
def depth_groups(tokenizer: str, option_set: str, presentation: str,
d: int) -> float | None:
for r in rows:
if (r.get("provider") == f"tokenizer/{tokenizer}"
and r.get("option_set") == option_set
and r.get("presentation") == presentation
and r.get("metric") == "exploratory_prefix_groups_at_depth"
and r.get("read_depth") == d):
return r["value"]
return None
def verdict(pid: str) -> tuple[str, str]:
for r in rows:
if r.get("metric") == f"prediction_{pid}":
return r["verdict"], r["value"]
return "MISSING", -1.0
print("JEV-07-02 what tokenisation does to option scoring")
print(f"tokenizer: Qwen/Qwen3-1.7B @ 70d244cc (primary), "
f"meta-llama/Llama-3.2-1B-Instruct @ 92131767 (second family)")
print(f"option sets: 77 BANKING77 label names, 4 wordings x 77, "
f"17 held-out x 4, 2 safety x 4")
print(f"rows in results/ch07.jsonl: {len(rows)}\n")
print("77 options, one presentation per line")
print(" presentation readout distinct collide pairs "
"pfx_tok depth_to_separate")
for pres, name in (("letter", "letter readout"),
("number_bare", "bare digits"),
("number_bracketed", "bracketed [k]"),
("full_string", "full option strings")):
if infeasible(PRIMARY, "banking77_names_w0", pres):
print(f" {pres:15s} {name:17s} IMPOSSIBLE by construction "
f"(26 letters < 77 options)")
continue
d1 = cell(PRIMARY, "banking77_names_w0", pres, "distinct_first_tokens")
col = cell(PRIMARY, "banking77_names_w0", pres,
"options_first_token_collision")
pairs = cell(PRIMARY, "banking77_names_w0", pres,
"first_token_collision_pairs")
pfx = cell(PRIMARY, "banking77_names_w0", pres, "strict_prefix_pairs_token")
depth = cell(PRIMARY, "banking77_names_w0", pres,
"exploratory_max_depth_to_separate_all")
depth_s = "1" if depth == 1 else f"{depth:.0f}"
print(f" {pres:15s} {name:17s} {d1:8.0f} {col:8.0f} {pairs:6.0f} "
f"{pfx:8.0f} {depth_s}")
print("\n17 held-out options and 2 safety options, same tokenizers")
print(" set pres distinct collide pairs pfx_tok")
for oset in ("unseen17_w0", "safety_w0"):
for pres in ("letter", "number_bare", "number_bracketed", "full_string"):
if infeasible(PRIMARY, oset, pres):
print(f" {oset:14s} {pres:17s} IMPOSSIBLE")
continue
print(f" {oset:14s} {pres:17s} "
f"{cell(PRIMARY, oset, pres, 'distinct_first_tokens'):8.0f} "
f"{cell(PRIMARY, oset, pres, 'options_first_token_collision'):8.0f} "
f"{cell(PRIMARY, oset, pres, 'first_token_collision_pairs'):6.0f} "
f"{cell(PRIMARY, oset, pres, 'strict_prefix_pairs_token'):8.0f}")
print("\nsurface form moves the first token (77 label names, Qwen3)")
print(f" one leading space changes the first token of "
f"{cell(PRIMARY, 'banking77_names_w0', 'full_string', 'leading_space_changes_first_token'):.0f} of 77")
print(f" capitalising the first letter changes it for "
f"{cell(PRIMARY, 'banking77_names_w0', 'full_string', 'capitalisation_changes_first_token'):.0f} of 77")
lens = [cell(PRIMARY, "banking77_names_w0", "full_string", k) for k in
("token_len_min", "token_len_median", "token_len_max")]
print(f" option token length min/median/max: {lens[0]:.0f}/"
f"{lens[1]:.0f}/{lens[2]:.0f}")
print("\nwhich presentation collides, per wording (77 options)")
print(" wording distinct collide largest class")
for w in ("w0", "w1", "w2", "w3"):
oset = "banking77_names_w0" if w == "w0" else f"banking77_{w}"
print(f" {w:8s} "
f"{cell(PRIMARY, oset, 'full_string', 'distinct_first_tokens'):8.0f} "
f"{cell(PRIMARY, oset, 'full_string', 'options_first_token_collision'):8.0f} "
f"{cell(PRIMARY, oset, 'full_string', 'largest_first_token_class'):14.0f}")
print("\ntwo tokenizer families, bare digits (77 options)")
for tok, label in ((PRIMARY, "Qwen3 "), (SECOND, "Llama3.2")):
print(f" {label} distinct first tokens: "
f"{cell(tok, 'banking77_names_w0', 'number_bare', 'distinct_first_tokens'):.0f}"
f" strict prefix pairs: "
f"{cell(tok, 'banking77_names_w0', 'number_bare', 'strict_prefix_pairs_token'):.0f}")
print("\nthe largest colliding first-token classes (Qwen3, w0 label names)")
for r in rows:
if (r.get("option_set") == "banking77_names_w0"
and r.get("presentation") == "full_string"
and r.get("metric") == "largest_first_token_class_members"):
print(f" {r['note']}")
break
print("\npredictions T1-T12, as preregistered")
holds = 0
total = 0
for i in range(1, 13):
pid = f"T{i}"
v, val = verdict(pid)
if v == "NOT_APPLICABLE":
continue
total += 1
holds += 1 if val == 1.0 else 0
print(f" {pid:4s} {v}")
print(f" {holds} of {total} applicable predictions hold")
print("\nJEV-07-03 model run: DEFERRED")
print(" no accuracy, no calibration, no flip rate, no latency was measured.")
print(" commands in examples/ch07-the-first-token/STAGE_PLAN.md;")
print(" every model number in the chapter is PENDING_RUN.")
JEV-07-02 what tokenisation does to option scoring
tokenizer: Qwen/Qwen3-1.7B @ 70d244cc (primary), meta-llama/Llama-3.2-1B-Instruct @ 92131767 (second family)
option sets: 77 BANKING77 label names, 4 wordings x 77, 17 held-out x 4, 2 safety x 4
rows in results/ch07.jsonl: 6001
77 options, one presentation per line
presentation readout distinct collide pairs pfx_tok depth_to_separate
letter letter readout IMPOSSIBLE by construction (26 letters < 77 options)
number_bare bare digits 9 75 366 68 2
number_bracketed bracketed [k] 1 77 2926 0 3
full_string full option strings 43 48 94 0 5
17 held-out options and 2 safety options, same tokenizers
set pres distinct collide pairs pfx_tok
unseen17_w0 letter 17 0 0 0
unseen17_w0 number_bare 9 9 36 8
unseen17_w0 number_bracketed 1 17 136 0
unseen17_w0 full_string 14 6 3 0
safety_w0 letter 2 0 0 0
safety_w0 number_bare 2 0 0 0
safety_w0 number_bracketed 1 2 1 0
safety_w0 full_string 2 0 0 0
surface form moves the first token (77 label names, Qwen3)
one leading space changes the first token of 77 of 77
capitalising the first letter changes it for 76 of 77
option token length min/median/max: 2/3/8
which presentation collides, per wording (77 options)
wording distinct collide largest class
w0 43 48 10
w1 29 59 12
w2 54 36 5
w3 47 43 11
two tokenizer families, bare digits (77 options)
Qwen3 distinct first tokens: 9 strict prefix pairs: 68
Llama3.2 distinct first tokens: 77 strict prefix pairs: 0
the largest colliding first-token classes (Qwen3, w0 label names)
'card':[7, 11, 12, 17, 18, 21, 33, 43, 47, 57]; 'top':[20, 26, 27, 30, 34, 42, 67]; 'transfer':[5, 28, 41, 74]; 'pending':[8, 16, 59, 62]; 'exchange':[3, 19, 52]
predictions T1-T12, as preregistered
T1 HOLDS
T2 REFUTED
T3 REFUTED
T4 REFUTED
T5 REFUTED
T6 REFUTED
T7 HOLDS
T8 HOLDS
T9 HOLDS
T10 REFUTED
T11 HOLDS
T12 HOLDS
6 of 12 applicable predictions hold
JEV-07-03 model run: DEFERRED
no accuracy, no calibration, no flip rate, no latency was measured.
commands in examples/ch07-the-first-token/STAGE_PLAN.md;
every model number in the chapter is PENDING_RUN.
Read the first block row by row, because the three surviving presentations fail in three different ways.
Letters do not exist at 77 options. Not “perform badly” — do not exist. There
are 26 and the task needs 77. AnyJev’s code says the same thing
(MAX_OPTIONS = 26), which is why that number is a constant in ours rather than
a guess. This is the one row in the table that is a hard impossibility, and it is
a property of the alphabet, not of any model.
Bare digits are a trap, and the trap is bigger than we predicted. Qwen3 gives
77 identifiers only 9 distinct first tokens: 1, 10, 11, 12 … 19 all start
with 1, and so on, so 75 of the 77 options collide and there are 366
colliding pairs with a largest class of 11. There are also 68 strict token-prefix
relations. A readout that reads one token cannot separate option 1 from option
11, and a readout that reads the full string has to decide what to do about the
prefix nesting. Both failures come from the same fact: 10 is two tokens on this
tokenizer.
Bracketing fixes the trap by creating a worse one. [1] … [77] has zero
prefix relations, exactly as LLM2Jev §3.2 argues — the bracket is the trick, and
it works. But every bracketed identifier now begins with the single token [, so
the 77 options share one first token and all 77 collide. Reading depth 1 is
not merely uninformative here, it is identically uninformative. Separation takes
depth 3 on Qwen3 and depth 2 on Llama-3.2. The prefix-free design did not remove
the collision; it moved it from a scattered pattern to a uniform one and charged
one extra token of read depth for it. Our prediction T3 said the collision count
would be unchanged. It went from 75 to 77. That refutation is the finding.
Full option strings collide less than we predicted, and are still blocked. We
predicted at least 55 of 77 label names would share a first token; 48 do, with a
largest class of 10 (the card … family: arrival, acceptance, linking, not
working, payment fee, delivery estimate and so on). 48 of 77 is a majority, so a
single-token readout of the strings is still unusable for this label set, but our
threshold was too high and the direction was right. Wording moves the number a
lot: the four Chapter 5 wordings give 43, 29, 54 and 47 distinct first tokens, so
the same options presented four ways differ by 25 distinct tokens — more than
the number of options lost to collisions in the best wording.
Surface form is not a detail. One leading space changes the first token of 77 of 77 label names. Capitalising the first letter changes it for 76 of 77. This is Robinson & Wingate’s Table 3 result (“Caps” and “Space” corruptions) arriving from the other direction: they measured that the cloze approach loses accuracy under those corruptions; we measured that the representation itself changes for almost every option, which is the mechanism behind their number. And the two tokenizer families agree exactly, on all four wordings — this is not a quirk of one vocabulary.
The finding that changes the most
The bare-digit trap does not exist on Llama-3.2. Llama tokenises 10 as a
single token; Qwen3 splits it into 1, 0. So the same 77 options, printed the
same way, under the same prompt, with the same readout, give:
- Qwen3: 9 distinct first tokens, 75 colliding, 68 prefix pairs, separation at depth 2.
- Llama-3.2: 77 distinct first tokens, 0 colliding, 0 prefix pairs, separation at depth 1.
LLM2Jev reports the Qwen behaviour as a fact about its tokenizer (Appendix C) and then finds a benchmark effect it cannot explain (Appendix H: answers moving from option 10 to option 1, “We have not identified the cause”). This chapter says what the cause is, at least structurally: on a digit-splitting tokenizer, options 1 and 10 are the same token until the second one. The effect is reproducible from the tokenizer files alone, with no model, no weights and about forty seconds of CPU — and it exists on one tokenizer family and not on the other. That is a stronger and much more actionable statement than “small models are biased”, and it is not in the literature we read.
What this does to “Jev is just option scoring”
Wrong: a decision model is an inference strategy — build the prompt, read the option probabilities, normalise — and what remains of Jev after you remove that is branding.
Correct: “read the option probabilities” is at least four different mechanisms. At 77 options one of them does not exist, one of them collapses 75 options onto 9 tokens, one of them collapses all 77 onto a single token, and the fourth needs up to 5 tokens of read depth — and which mechanism you get depends on which tokenizer you load, as the Qwen/Llama split above shows with everything else held constant. The interesting engineering is not “score the options”. It is choosing a presentation the tokenizer can read, refusing the ones it cannot, and knowing the read depth you are buying.
That is a claim about the shape of the family, and it is why this chapter ends
up proposing a change to how the book states H1
(planning/pivot-proposals.md, 2026-10-07, status PROPOSED — the author
decides). We did not edit planning/hypotheses.md.
The cost accounting (analytical, not measured)
One forward pass over the shared prompt, with the KV cache reused for every
option, against the naive one-pass-per-option version: both are implemented in
HFScorer (path="shared_prefix" and path="naive"), and the provider exposes
the token counts through count_tokens. Those counts are ANALYTICAL and every
row that carries them says so in its own accounting field. At 77 options the
naive path re-reads the prompt 77 times; the shared path reads it once plus the
option tokens. That is arithmetic. No wall-clock number has been measured for this
provider and the chapter claims none.
PENDING_RUN: measured tokens and latency, both execution paths Command:
python examples/ch07-the-first-token/run_ch07.py --stage S7 --metric tokens --device cudathen--stage S8 --metric latency --device cudaFills:results/ch07.jsonl, rows withmetric: tokensandmetric: latencyNote: the latency stage needs an idle machine, 200 items, warm, 3 repeats.
What the tests cover
python examples/ch07-the-first-token/selftest_ch07.py — 87 checks, no model, no
GPU, about fifteen seconds.
- A fake model with known logits. Every probability, normalisation, rotation,
prior division and temperature is checked against arithmetic done by hand in the
test file. The letter readout’s output is compared to a hand-computed softmax;
the prior division to
p/priorrenormalised by hand; the batch prior to a running mean; the content-free prior to the mean of three probes. Nothing is trusted because it ran. - A trivially separable task (RUN.md standing lesson 2, which exists because
Chapter 4 shipped a revision over an under-trained baseline). Six options with
disjoint characters; the provider must get it right through the full
DecisionRequest→ChoiceAnswerpath, and it does. - The invariance that rotation has to preserve. Under an option-invariant fake, all cyclic rotations give zero flips and the same distribution; rotating and un-rotating the answer names one option across all rotations. Under a deliberately order-dependent fake, flips appear. Both directions are tested, because a rotation test that only ever passes proves nothing.
- The mathematical claim behind log-space marginalisation. Under
logit(i at position j) = c_i + b_j, the log-mean over all rotations returnssoftmax(c)exactly; the arithmetic mean does not. That is AnyJev’s argument and it is checked numerically rather than cited. - Two seeded bugs. (1) A tokenizer that merges the prompt’s last token with the continuation’s first — the trap LLM2Jev’s Appendix C warns about — must make the prefix invariant raise, and it does, both through the helper and through the provider. (2) The rotation permutation machinery is checked to be genuine: shift 0 is the identity, every shift is a true permutation, and every option visits every position exactly once — the properties whose absence would silently corrupt every answer while looking like an improvement.
- A real code path. The last block builds a randomly initialised 2-layer
Qwen3 in-process on the real Qwen3 tokenizer and runs the provider through
it. This proves the transformers calls, the cache path and the prompt
construction execute. Its numbers are meaningless by construction and are never
written to
results/.
Against the prediction
Clause by clause, six of twelve hold.
| # | Prediction | Verdict |
|---|---|---|
| T1 | bare digits, 77 options: 9 distinct first tokens, 75 colliding, 366 pairs, largest 11 | HOLDS — exactly |
| T2 | bare digits: 68 strict token-prefix pairs, 0 character-prefix pairs | REFUTED — token half right (68), character half wrong: 1 is a character prefix of 10. Our own arithmetic error |
| T3 | bracketed: 0 prefix pairs, collision count unchanged from bare | REFUTED, substantively — 0 prefix pairs confirmed, but collisions go 75 → 77: the bracket makes all options share [ |
| T4 | letters: 26 distinct at K≤26, impossible at 77 | REFUTED on the number, HOLDS on the substance — at 17 options there are 17 letters, not 26; we predicted the ceiling where we meant the count. Letters are collision-free (17/17 distinct, 0 pairs) and 77 is impossible |
| T5 | label names: ≥55 of 77 share a first token, largest ≥15 | REFUTED — 48 collide, largest 10. Still a majority, so the readout is still blocked; our threshold was too high |
| T6 | label names: 0 character-prefix pairs, ≥1 token-prefix pair | REFUTED — both zero. Sharing a first word is a collision, not a nesting: card arrival → [4951, 18647] and card not working → [4951, 537, 3238] diverge at the second token |
| T7 | one leading space changes ≥70 of 77 first tokens | HOLDS — 77 of 77 |
| T8 | capitalising changes ≥70 of 77 | HOLDS — 76 of 77 |
| T9 | median 3–6 tokens, max ≥8, ≥95% multi-token | HOLDS — median 3, max 8, 100% multi-token |
| T10 | Llama-3.2 shows the same picture as Qwen3 | REFUTED, and it is the chapter’s sharpest finding — Llama gives 77 distinct first tokens, 0 collisions, 0 prefix pairs |
| T11 | safety labels: 2 distinct, no collisions or prefixes, all four wordings | HOLDS |
| T12 | unseen-17: 2–8 of 17 share a first token | HOLDS — 6 |
Two of the six failures (T2, T6) were arithmetic and reasoning errors of ours,
diagnosable in a few lines (examples/ch07-the-first-token/diagnose_refutations.py
reproduces each). Two more (T4, T5) were the right direction with thresholds in
the wrong place. Two (T3, T10) are findings about the world, and T10 is the one
that changes the chapter’s conclusion.
The local run is deferred, and why
Local GPU work on this machine is paused by the author’s decision of 2026-10-07: five GPU runs had failed or been too slow, and a Chapter 6 run was still occupying the card. So Chapter 7’s model sweep did not happen. What that costs is specific, and it is the half of the chapter that would have said how often these structural limits are reached in practice.
What would change the book’s answer: if option scoring on frozen Qwen3 lands close to the Chapter 4 bar of 0.8779 on 77-way intent, H1’s “reduces to known techniques” clause is strongly supported and the packaging is most of it. If it lands far below, as the prior art suggests — 2607.27421 finds no model above 80% on Banking77 across a 41-model cohort, and LLM2Jev’s own frozen 4B reaches 69.0 — then “cheap” is the only thing option scoring has, and the interesting question becomes where in the accuracy-cost plane the intersection is. And if the readout collisions measured here turn out to predict accuracy across models, the readout’s availability is a first-order constraint on any system with a large label set, which is a stronger claim than H1 makes in either direction.
PENDING_RUN: all of JEV-07-03. Accuracy, macro-F1, ECE (15 equal-mass bins), Brier, AUROC, flip rate in option space, answer mass, tokens and latency, for Qwen3 at 0.6B / 1.7B / 1.7B-Base / 4B on 77-way intent, 17-way unseen and the two safety splits, raw and corrected, against the frozen Chapter 4 bar and the Chapter 5 embedding provider with paired bootstrap intervals. Command: stages S1–S9 in
examples/ch07-the-first-token/STAGE_PLAN.md, one per model/task/split, each ≤15 minutes, each resumable, each writing a.donemarker; merge withrun_ch07.py --mergeonly when all nine markers exist. Fills:results/ch07.jsonl, then the results and prediction sections above. Estimated total: about 1 h 45 min of GPU time.
What we got wrong about our own predictions
We wrote twelve predictions expecting to be right about most of them and wrote them in a form precise enough to be wrong in a diagnosable way. Six held. Of the six that failed, two were arithmetic slips that took a minute each to find, two were threshold errors in the right direction, and two were discoveries about tokenisation that we could not have got right by reasoning alone — the bracketed scheme’s total depth-1 collision, and the fact that the digit trap exists on one tokenizer family and not another. That ratio is the argument for preregistering even the mechanical parts: the mechanical predictions were the ones that made the interesting failures legible, because a failed arithmetic claim is fast to diagnose and a failed intuition is not.
One more thing we would flag for the next chapter: the exploratory
depth-of-read measurement (how many tokens before options separate) was not
preregistered. It was added after T1 showed that bracketed identifiers share
their first token, because a collision count with no statement of what to do
instead is not actionable. It is labelled exploratory_ in every row, and it is
what turned “there is a collision” into “you need depth 3, and here is the depth
for each presentation”.
Limitations
- No model ran. Every accuracy, calibration, flip-rate and latency number in
this book for Chapter 7 is
NOT_OBSERVED. The provider exists, is tested against a fake model with known logits, and has never been evaluated. - Three tokenizers, two families. Qwen3 (instruct and base) and Llama-3.2. Digit grouping is known to vary across families — Phi-4-style tokenizers group up to three digits, per LLM2Jev Appendix C — and we did not measure one. The family-dependence claim rests on two points, not a survey.
- Two of the twelve predictions failed for our own arithmetic. We report the ratio honestly, but a preregistration with more slack would have predicted fewer failures and taught us less.
- The label sets are one author’s. The four wordings come from Chapter 5’s frozen fixture, hand-written by the author without seeing dataset text. The collision counts are properties of these wordings; a different vocabulary would give different numbers, though the two tokenizer families agreeing exactly suggests the shape is robust to the tokenizer, not the author.
- Only BANKING77 and the two safety labels. One 77-way set, one 17-way set, one binary set. 2610.07716 shows that the candidate menu itself moves accuracy 26 points on CLINC150, wider than a model-size step, so the menu is a first-order variable this chapter deliberately held fixed rather than varied.
- The exploratory depth measurement is not preregistered. Labelled as such in every row; it is a lead, not a result.
- No hosted Jev.
src/arbiter/providers/hosted_jev.pydoes not exist in this repository and there is no response cache, so the optional comparison of Jev’s behaviour on the tokenisation-collision cases — does Jev separate the tencard …labels that a single token cannot? — is skipped, not deferred quietly. It is the most interesting question this chapter raises for a reader with an API key.
What Chapter 7 leaves behind
src/arbiter/providers/option_scoring.py— two readouts, declared multi-token rule, log-space rotation, batch and content-free priors, 26-option refusal, both execution paths, ANALYTICAL token accounting. Untested against a real model, and labelled as such in its own docstring.examples/ch07-the-first-token/— the tokenisation analysis (the chapter’s one real result), the 87-check selftest with two seeded bugs, the demo that prints this chapter’s code block, the refutation diagnosis, and the nine-stage plan for the deferred run.results/ch07.jsonl— 6001 rows,"mode": "observed-cpu".- A dated entry in
planning/pivot-proposals.mdproposing that H1’s list of constituent techniques stop treating “first-token LM inference” as one thing.
The prompt ends with a question, and this chapter can only answer half of it. Is
this a model, or a way of asking a model? The tokenisation answer is clear: at
77 options, whether you have a way of asking at all is settled by the
alphabet and the tokenizer, not by the model. What would have to be learned for
it to be a model? Nothing in this chapter tests that — the deferred run is the
test. Chapter 8 takes the other half, which we can answer now: whatever the
mechanism is, the program should not be able to ask for a decision the provider
cannot represent, and the refusal in fit is the smallest honest version of
that.