Reading the Hidden State
Ask whether a linear probe can read a decision from hidden activations more cheaply than readout, then require controls before believing it.
Design draft: this chapter argues from the literature and from earlier measured results; its own experiment has not been run.
You have been reading decisions from model outputs: letters, digits, brackets, full strings, entailment scores and calibrated probabilities. Those are all downstream of the same expensive object: a forward pass that turns a state into hidden activations, followed by whatever projection turns those activations into an answer. This chapter asks whether you should stop earlier. If the hidden state already contains the decision in linearly accessible form, a small probe may answer more accurately and more cheaply than the output readout. If it does not, or if the probe merely memorises its training labels, then hidden-state reading is an attractive illusion.
The question has two halves. Accuracy asks whether a probe beats readout on the question it was fitted for. Transfer asks what survives when the question changes. The literature says those halves point in opposite directions, and the most useful third-party report on this exact problem agrees. That tension is the chapter.
The real run needs a GPU and cannot happen yet. What can happen now is the part that decides whether its results would be believable: the checks that separate a probe that found a decision from a probe that found capacity or a nuisance feature. This chapter builds those checks, tests them, and shows them working on synthetic activations that are labelled as such.
What probes are, and what they are not
The original idea: a thermometer, not a control
DOCUMENTED: Alain and Bengio define a probe as a linear classifier trained independently on a layer’s features to predict the original classes. Probes do not affect model training; they are thermometers, not controls. Their headline observation is that linear separability increases almost monotonically with depth in the image models they test. [Alain & Bengio, arXiv:1610.01644, Abstract, §§1, 3.2–3.4, 5.1, 7]
They also warn that probes can diagnose failure: in a pathological 128-layer network with a long skip connection, half the model stays unused despite successful loss minimisation. They explicitly prefer validation behaviour to training behaviour, and warn that very wide early layers may need feature subsampling because a probe can otherwise become larger than the model it measures.
Those cautions transfer directly. A probe is a measurement instrument with its own parameters, data requirements and failure modes. It is not a way of asking the model what it believes. If you need that distinction later, this is where it enters the book.
Probing for truth: topic transfer built in
DOCUMENTED: Azaria and Mitchell’s SAPLMA trains a classifier on hidden-layer activations to predict whether a statement is true or false, including statements the model itself generated. For OPT-6.7B, the 20th of 32 layers averages 0.8060 across held-out topics, while few-shot prompting stays near chance; for LLaMA2-7B, the middle layer performs best. [Azaria & Mitchell, arXiv:2304.13734, Abstract, §§4–5, 7–8]
Training uses all topics except the held-out test topic and repeats each fit three times, so the method is already a transfer test: topic-specific memorisation is not enough. But the best layer depends on the model, and generated sentences are weaker and more threshold-sensitive than curated true/false sentences. Their limitations section is unusually relevant here: it distinguishes detecting truth from detecting certainty, and warns that long-response activations mix correct and incorrect information over time.
Removing supervision, and what that costs
DOCUMENTED: Burns and colleagues remove supervision. Contrast-Consistent Search finds a linear direction satisfying negation consistency: a statement and its negation should receive probabilities summing to one, while avoiding the degenerate all-0.5 solution. Across six models and ten QA datasets, CCS averages 71.2% against 67.2% for calibrated zero-shot. When a misleading prefix drops UnifiedQA zero-shot from 80.4% to 70.9%, CCS moves from 82.1% to 83.8%. It transfers across unrelated tasks and often works with little data. [Burns et al., arXiv:2212.03827, Abstract, §§2.2, 3.2–3.3, 5.1]
Their stated limitation is the one our controls enforce: CCS assumes that a supervised probe could in principle succeed, and it does not establish that the discovered direction is the model’s knowledge rather than another prominent feature.
DOCUMENTED, and it challenges unsupervised optimism: Farquhar and colleagues prove that arbitrary binary features can be optimal under the CCS loss, then show empirically that unsupervised probes can learn distractors, simulated-character opinions or prompt artefacts instead of knowledge. Their conclusion is blunt: existing unsupervised methods are insufficient for discovering latent knowledge, although contrastive activations remain useful interpretability tools. They also warn that future consistency-based methods may inherit the same identification problem. [Farquhar et al., arXiv:2312.10029, Abstract, §§1–3, 6]
This is why JEV-13-01 treats supervised probes and CCS as different claims. A supervised win says the signal is linearly accessible under labels. An unsupervised win, without controls, says only that something prominent was found.
A recent study that sharpens both the hope and the boundary
DOCUMENTED: Cencerrado and colleagues train a difference-of-means linear probe on residual-stream activations before any answer token is generated. The direction generalises across factual datasets and beats black-box question-embedding baselines, but fails on GSM8K mathematical reasoning. Separability saturates in intermediate layers. [Cencerrado et al., arXiv:2509.10625, Abstract, §§3–4.4, 5–6]
Their cost accounting is directly useful. Activation collection dominates, at about 60 A100 hours plus 100 A40 hours in their runs, while probe training on 10,000 cached activations takes under three minutes on CPU. A probe can therefore be cheap at decision time and expensive to prepare. Their limitations also bind our design: correctness is a single-sample binary label, larger-model coverage is thin, and layer selection comes from one dataset.
If you have read PyTorch From First Principles, the implementation shape below will feel familiar: frozen activations, no gradient into the base model, a small convex head. If you have read Embeddings From First Principles, the representation warning will feel familiar too: linear accessibility is a property of a representation under a task, not proof that the representation “understands” the task. Neither book is re-taught here.
Wrong: “A probe that beats readout has found where the model keeps the decision.”
Correct: “A supervised probe has found a linearly accessible correlate of the label under the fitted distribution. Transfer and controls decide whether it found anything else.”
flowchart TD
A[shared state + question] --> B[one forward pass]
B --> C[hidden state per layer]
C --> D[linear probe]
D --> E[decision distribution]
B --> F[first-token readout]
F --> E
D --> G[random-label control]
D --> H[unseen-question transfer]
What you have already measured
The First Token measured the readout side without running a model decision. Its tokenizer-only run wrote 6,001 rows: Qwen3 gives nine distinct first tokens for bare digits 1 to 77, bracketed identifiers share one first token, full option strings need up to depth five, and letter readout is impossible at 77 options. (Rows: results/ch07.jsonl, JEV-07-02 demo block.)
The model sweep, JEV-07-03, is deferred. That matters because Chapter 13 cannot compare against a measured calibrated readout until it exists. The design therefore names JEV-07-03 as a dependency and does not pretend its baseline is already in hand.
Chapter 7 also pinned the models this chapter reuses: Qwen3-1.7B instruct and base, Qwen3-0.6B, Qwen3-4B and Llama-3.2-1B-Instruct, with exact revisions in metadata/07-chapter.yaml. No weights were downloaded for Chapter 7’s tokenisation run, and none are downloaded here. The future probe run must either find those pinned revisions locally or record a dated amendment; silent substitution would repeat the oldest baseline error in this book.
The third-party head experiment
DOCUMENTED: AnyJev’s technical report fits a closed-form head on one question’s gold labels and finds it scores 0.13 above debiased readout at the same depth. But leave-one-question-out shared heads (four extraction variants, five depths, four functional forms) all score below raw readout on held-out questions. The best reaches 0.580 against 0.635 debiased and 0.628 raw, with reversal flip rates of 0.17 to 0.30 versus 0.07 for raw readout. [AnyJev Technical Report, arXiv:2610.00831, §8]
Their sentence is the chapter hypothesis in miniature: “The signal is present per question; a shared linear rule that reaches an unseen question is not.”
That report also supplies two warnings our design copies. Its 20-option results have no intervals because they record aggregates rather than per-item outcomes, so JEV-13-01 stores per-item rows. And its serving threshold was selected and bounded on overlapping data; the honest deployment bound needs disjoint selection and verification splits. Our threshold and operating-point discipline does the same work for probes.
A current documentation change must be said loudly. DOCUMENTED: AnyJev’s present docs/levels.md, read 2026-10-08, says version 0.3.0 removed label-trained L1 and L2 levels; they remain in anyjev==0.2.0. Its limitations page adds that every published decision is scored in isolation, not inside an agent loop. So the prompt’s “matching AnyJev’s published limitation” now needs a version pin. The limitation is real in the technical report and in 0.2.0; it is not a claim about the current live levels. JEV-13-01 pins anyjev==0.2.0 for historical-method comparison only. This is a third-party documentation state, not vendor documentation and not our measurement.
The design
PROPOSED: JEV-13-01 extracts the final-prompt-token hidden state at every transformer block for three pinned models (Qwen3-1.7B-Instruct, Qwen3-4B and Llama-3.2-1B-Instruct) on 77-way BANKING77 intent and binary safety decisions. Each layer gets a linear softmax probe.
- Probe budgets are 100, 300 and 1,000 labels per question.
- Regularisation uses one shared L2 grid, C in {0.01, 0.1, 1.0, 10.0}, selected on the threshold split.
- Seeds 0 to 2 drive label subsampling and any stochastic solver initialisation.
- Accuracy and ECE are reported together, because Chapter 9 showed that ranking and calibration are different properties.
Controls are acceptance conditions, not decorations.
- Every probe has a random-label twin trained on the same items and budgets. A 77-way control must stay within [0.00, 0.05] on test. High training accuracy with chance test accuracy is the expected signature of memorisation, so the control checks whether the reporting pipeline could mistake capacity for knowledge.
- A second control varies superficial prompt format while holding labels fixed. If probe accuracy follows formatting rather than labels, the transfer claim fails even when in-question accuracy looks strong. That is Farquhar’s test translated into our harness: prominent features, not only random labels, are the rival hypothesis.
- Held-out option wording uses Chapter 5’s
w1tow3descriptions after fitting onw0. - Cross-question transfer uses the binary safety arm, where the output dimensionality permits the same head to be evaluated on a new question.
- A CCS or unsupervised arm may be reported, but Farquhar’s result forbids interpreting it as knowledge discovery without the same transfer and distractor checks.
Baselines are named, not assumed. The primary baseline is JEV-07-03’s raw and temperature-calibrated readout in its selected template, rotation and prior cell. The calibration-reporting convention follows Chapter 10. AnyJev 0.2.0 L2 is a historical comparator where its question and model scope applies. If JEV-07-03 has not run when JEV-13-01 is scheduled, the probe results wait; comparing a new probe against an unmeasured readout would manufacture a finding.
Cost is a measurement, not arithmetic. The probe path uses one shared forward pass plus a linear projection; the readout path uses its declared number of passes. The run records forward passes, prompt and continuation tokens by path, cold and warm p50 and p95 latency, GPU time, labels and peak memory exactly. The shared-state question for Chapter 27 is built in: one cached state can serve several probe heads, so the marginal cost of the second question is the measurement that matters.
The checks, made executable
The checks are src/arbiter/probe_design.py, tested by 8 tests in tests/probes/. They work on any activation arrays, so they will run unchanged on real hidden states later. Everything below is examples/ch13-reading-the-hidden-state/walkthrough_ch13.py, which you can run as it stands.
Everything it feeds the checks is synthetic. Twenty classes, 64 dimensions, and a signal planted in a seeded generator, strongest in the middle layers because that is how it was built. No model produced a single number below, so none of it is evidence about a real hidden state. It shows that the checks do their job. (The planted signal is deliberately weak. A first, stronger setting saturated at 100% accuracy in every layer and showed nothing, so it was weakened until the tables varied. That is tuning a demonstration to display a mechanism, which is fine for a demonstration and would be fraud for a result.)
The setup
import numpy as np
ROOT = Path(__file__).resolve().parents[2]
sys.path[:0] = [str(ROOT), str(ROOT / "src")]
from arbiter.probe_design import (
control_verdict,
evaluate,
memorisation_gap,
random_label_twin,
select_c_on_threshold,
select_layer,
signal_over_control,
)
N_CLASSES, DIM = 20, 64
SIGNAL_BY_LAYER = [0.05, 0.12, 0.20, 0.27, 0.30, 0.27, 0.23, 0.20] # peaks mid-depth, by construction
C_GRID = [0.01, 0.1, 1.0, 10.0]
CENTRES = np.random.default_rng(123).normal(0, 1, (N_CLASSES, DIM)) # one class geometry for every split
WORDING = np.random.default_rng(456).normal(0, 1, DIM) # a nuisance direction that marks the wording
def make_split(n_per_class, signal, seed, spurious=None):
"""Synthetic activations. `spurious` is the chance that the wording marker agrees with the label."""
rng = np.random.default_rng(seed)
y = np.repeat(np.arange(N_CLASSES), n_per_class)
X = signal * CENTRES[y] + rng.normal(0, 1, (len(y), DIM))
if spurious is not None:
agrees = rng.random(len(y)) < spurious
marker = np.where(agrees, y % 2, rng.integers(0, 2, len(y)))
X = X + 3.0 * marker[:, None] * WORDING
return X, y
1. Choose the layer without seeing test
One probe per layer. The layer is chosen by threshold accuracy, and select_layer receives threshold accuracy only, so it cannot select on test.
chance = 1 / N_CLASSES
print(f"ILLUSTRATIVE synthetic activations: {N_CLASSES} classes, {DIM} dims, chance {chance:.2f}")
# 1. Choose the layer on the threshold split. Test is touched once, afterwards.
print("1. one probe per layer, layer chosen on threshold")
per_layer, fits = [], []
for layer, signal in enumerate(SIGNAL_BY_LAYER):
tr, th, te = (make_split(30, signal, s) for s in (10 + layer, 20 + layer, 30 + layer))
fit = evaluate(*tr, *th, *te, C_GRID)
per_layer.append(round(fit.threshold, 3))
fits.append((tr, th, te, fit))
best = select_layer(per_layer)
print(f" threshold accuracy by layer: {per_layer}")
print(f" chosen layer {best}; test accuracy {fits[best][3].test:.3f} (chosen without seeing test)")
2. The random-label twin
The twin has the same inputs and the same label counts, with the link between them destroyed. If it scores well on held-out data, the pipeline is manufacturing accuracy. signal_over_control is the number that says whether the real probe found anything.
# 2. The random-label twin: same inputs, same label counts, no link between them.
print("2. the twin must not find signal")
tr, th, te, real = fits[best]
twin = evaluate(tr[0], random_label_twin(tr[1], 1), *th, *te, C_GRID)
print(f" real probe: train {real.train:.3f}, test {real.test:.3f}")
print(f" twin probe: train {twin.train:.3f}, test {twin.test:.3f} control verdict {control_verdict(twin.test, band=(0.0, 0.10))}")
print(f" signal over control: {signal_over_control(real, twin):.3f}")
3. Too few labels, and the twin shows it
At a few labels per class a probe memorises: it fits the training data and the gap to held-out accuracy is large. The preregistered budgets of 100 to 1,000 labels per question on a 77-way task are sparse by construction, so this is the regime that matters.
# 3. With too few labels the probe memorises, and the twin shows it.
print("3. labels per class (layer " + str(best) + ")")
print(" per class real test twin test real train-test gap")
for n in (3, 10, 30, 100):
a = make_split(n, SIGNAL_BY_LAYER[best], 40 + n)
b, c = make_split(30, SIGNAL_BY_LAYER[best], 50 + n), make_split(30, SIGNAL_BY_LAYER[best], 60 + n)
r = evaluate(*a, *b, *c, C_GRID)
t = evaluate(a[0], random_label_twin(a[1], 1), *b, *c, C_GRID)
print(f" {n:>9} {r.test:>9.3f} {t.test:>9.3f} {memorisation_gap(r):>10.3f}")
4. Right for the wrong reason
In-wording accuracy cannot distinguish a decision from a prominent nuisance feature. Here a wording marker agrees with the label 90% of the time during fitting, then stops predicting anything at the held-out wording. This is Farquhar’s rival hypothesis in miniature, and the reason held-out wording is an acceptance condition.
# 4. A probe can be right for the wrong reason: accuracy on the fitted wording says nothing.
print("4. the wording shift")
strong = SIGNAL_BY_LAYER[best]
a = make_split(30, strong, 70, spurious=0.9)
b = make_split(30, strong, 71, spurious=0.9)
same = make_split(30, strong, 72, spurious=0.9)
held_out = make_split(30, strong, 73, spurious=0.0) # the marker no longer predicts the label
fit = evaluate(*a, *b, *same, C_GRID)
from arbiter.probe_design import accuracy, fit_probe
probe = fit_probe(a[0], a[1], fit.c)
print(f" fitted wording (marker agrees with label 90%): test {fit.test:.3f}")
print(f" held-out wording (marker uninformative): test {accuracy(probe, *held_out):.3f}")
5. Selection ignores test
Scramble every test label and the chosen layer and the chosen regularisation do not change. A selection rule that could be moved by test labels would be tuning on test.
# 5. Selection cannot see test: scramble every test label and nothing chosen changes.
print("5. selection ignores test labels")
chosen = []
for scramble in (False, True):
picks, cs = [], []
for layer, signal in enumerate(SIGNAL_BY_LAYER):
tr, th, te = (make_split(30, signal, s) for s in (10 + layer, 20 + layer, 30 + layer))
y_test = np.random.default_rng(0).permutation(te[1]) if scramble else te[1]
fit = evaluate(*tr, *th, te[0], y_test, C_GRID)
picks.append(round(fit.threshold, 3))
cs.append(fit.c)
chosen.append((select_layer(picks), cs))
print(f" layer and C per layer, real test labels: layer {chosen[0][0]}, C {chosen[0][1]}")
print(f" layer and C per layer, scrambled test labels: layer {chosen[1][0]}, C {chosen[1][1]}")
print(f" identical: {chosen[0] == chosen[1]}")
What it prints
ILLUSTRATIVE synthetic activations: 20 classes, 64 dims, chance 0.05
1. one probe per layer, layer chosen on threshold
threshold accuracy by layer: [0.08, 0.13, 0.295, 0.478, 0.568, 0.508, 0.365, 0.238]
chosen layer 4; test accuracy 0.607 (chosen without seeing test)
2. the twin must not find signal
real probe: train 0.823, test 0.607
twin probe: train 0.372, test 0.065 control verdict PASS
signal over control: 0.542
3. labels per class (layer 4)
per class real test twin test real train-test gap
3 0.250 0.060 0.750
10 0.450 0.047 0.490
30 0.570 0.053 0.235
100 0.597 0.045 0.148
4. the wording shift
fitted wording (marker agrees with label 90%): test 0.647
held-out wording (marker uninformative): test 0.342
5. selection ignores test labels
layer and C per layer, real test labels: layer 4, C [10.0, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01]
layer and C per layer, scrambled test labels: layer 4, C [10.0, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01]
identical: True
Read it in order.
- Section 1: threshold accuracy rises to layer 4 and falls after it, and the chosen layer’s test accuracy is 0.607. This is the planted mid-depth peak, by construction, and not a finding about any model. Alain and Bengio saw near-monotone growth with depth in image models; Azaria and Mitchell found mid-depth peaks in language models. A real run decides which pattern a decision task shows.
- Section 2: the twin reaches 0.372 on training data and 0.065 on test, so the control passes its band, and the real probe’s margin over its twin is 0.542.
- Section 3: with 3 labels per class the real probe’s train-to-test gap is 0.75, and it closes to 0.148 at 100 labels per class. The twin stays near chance at every budget.
- Section 4: accuracy at the fitted wording is 0.647 and falls to 0.342 where the marker stops helping, a drop of about 30 points that in-wording accuracy alone would never reveal.
- Section 5: identical layer and identical regularisation with real and scrambled test labels.
The experiment, designed but not run
JEV-13-01 is preregistered with status: NOT_RUN; the full block is in metadata/13-chapter.yaml.
Predictions (HYPOTHESIS):
- P1. At 1,000 labels per question, the best-layer supervised probe beats calibrated readout by at least 3 accuracy points on 77-way intent for at least 2 of 3 models.
- P2. The same probe loses at least 5 points on held-out option wording for at least 2 of 3 models.
- P3. Random-label controls stay within [0.00, 0.05] on 77-way test while true-label probes exceed 0.30.
- P4. Cross-question safety transfer falls at least 5 points below in-question accuracy.
- P5. Measured probe latency is lower than the selected readout path on identical hardware.
Refutation: P1 fails if no probe beats readout. P2 and P4 fail if probes transfer. P3 fails if random labels also predict. P5 fails if one shared forward pass plus projection is not cheaper in wall time. The allowed negative result is explicit: probes may learn the dataset rather than the decision, in which case the chapter recommends the readout and records the probe as a diagnostic, not a provider.
PENDING_RUN: result for JEV-13-01, P1–P5 Command:
python examples/ch13-reading-the-hidden-state/run_ch13.py --models qwen3-1.7b-instruct,qwen3-4b,llama-3.2-1b-instruct --budgets 100,300,1000 --seeds 0,1,2Fills:results/ch13.jsonl,static/figures/ch13-per-layer-*.pngStage plan (each ≤15 min, GPU-bound, resumable cached activations): S1 extract hidden states by model/task/layer shard; S2 fit supervised and random-label probes (withprobe_design.evaluateandrandom_label_twin); S3 evaluate in-question and unseen-wording transfer; S4 evaluate cross-question transfer and unsupervised sanity checks; S5 measure cost/shared-state reuse and write rows/figures.
What would change your mind
A shared probe that survives held-out wording and a new question would overturn the transfer half. A probe that beats readout only by memorising random labels would overturn the accuracy half by invalidating the instrument. A cheaper readout path would overturn the cost half even if accuracy held.
No pivot is proposed, because a design supplies no new measured evidence. The compression-thesis question in planning/purpose.md stays open: smaller models become viable if probes work, but only the GPU run can say whether “work” means accuracy, transfer, cost, or none of them. H3 therefore stays INSUFFICIENT_EVIDENCE.
Limitations
JEV-13-01 is NOT_RUN, and results/ch13.jsonl does not exist. Every number in the walkthrough is synthetic.
Required papers are PARTIAL by section; proofs, supplements, appendices and code were not reproduced. AnyJev numbers are third-party claims from its report and repository, not vendor documentation and not our measurements. The present-tense AnyJev limitation must carry its version pin: L1 and L2 were removed in 0.3.0 and remain in 0.2.0.
The design assumes pinned weights remain available; otherwise a dated amendment is required. Probe budgets of 100 labels on a 77-way task are sparse by construction, and missing-class behaviour must be reported, not smoothed. ECE beside accuracy does not make a probe calibrated. Shared-state reuse is a first look for Chapter 27, not its verdict.
The checks catch memorisation, test leakage in selection and a planted nuisance feature. They cannot catch every confound a real model can hide, and passing them would make a probe credible, not correct.
What the next chapters inherit
Chapter 14 inherits the probe as a tuning baseline: any LoRA, distillation or supervised-tuning claim must beat both readout and the probe, not merely beat readout.
Chapter 15 inherits the transfer machinery: same-question accuracy without cross-question evaluation is not evidence for general decision capability.
Chapter 27 inherits the cost question: one cached forward pass serving several heads is the shared-state hypothesis in miniature.
The prompt’s closing question stays with the reader: what does a probe know that readout does not, and what does it forget when the question changes? This chapter’s answer is designed, not measured: probably the local label geometry, and probably only that.