Did the Context Help?

Explain this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Apply this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

The staged compiler passed every test its contract set. Forty-two fixture-budget combinations, forty-two correct success-or-failure verdicts, zero illegal admissions of any class, required-item recall at 1.0 on every feasible combination. And the same frozen run admits five distractors and two evaluator-labelled harmful items into staged bundles β€” information the hidden ledger marks as useless or worse, sitting inside policy-valid, budget-compliant, fully traced context. Nothing malfunctioned. The compiler cannot see hidden labels, and legality never promised usefulness. So here is the question the previous chapter earned but could not answer: the bundle is legal, but did it help?

A legal bundle is not necessarily a useful bundle.

That sentence opens the chapter because everything in it follows from the gap between constructing context and proving context mattered. Make the gap concrete with one frozen case before any theory. On the heterogeneous-basic fixture at the medium budget, the staged bundle is fully policy-valid: mandatory constraints, the required rule, the retained memory note, all present and correctly ordered. It also contains, by trace, one misleading instruction-shaped note the hidden ledger marks harmful β€” eligible, in scope, fresh enough, relevant at 0.55, admitted under uncertainty. The bundle is beyond reproach and possibly beyond use in the same breath. Now ask the only question that matters: does the reader act on it?

Book result β€” compiler construction

Before the gap, the promotion, stated with its full provenance chain so the evidence grade is checkable rather than asserted:

Book result β€” compiler construction. On the synthetic compiler-v1 suite (14 fixtures Γ— 3 budget regimes Γ— 6 strategies = 252 compilations), the deterministic staged compiler produced correct success/failure status on all 42 fixture-budget combinations, required-item recall of 1.0 on all 34 feasible combinations, zero scope/freshness/authority/floor/dependency violations, and explicit compile failures with exact reason codes on every infeasible combination. Experiment compiler-v1, run run-001, implementation commit ec642dc, artifact validated clean.

The numbers below are read from that frozen artifact, not reconstructed from any summary:

Strategy Status correct Dependency Group Floor Distractor Scope Freshness
dump/truncate 34/42 1 0 3 18 5 4
top-k 34/42 1 1 3 18 6 6
weighted 34/42 1 0 3 17 6 6
hard-gated greedy 34/42 1 1 0 9 0 0
staged compiler 42/42 0 0 0 5 0 0
oracle 42/42 0 0 0 0 0 0

The promotion covers bundle construction only: hard gates eliminated the exercised violations, dependency-aware assembly prevented the cheap-reference failure, representation choice respected floors while changing cost, infeasibility reported itself, budget slack survived. Read the baselines as the contrast that earns those verbs their keep. Dump, top-k, and weighted strategies each admit a dozen or more illegal items and dozens of scope, freshness, and floor violations across the same 42 combinations; even hard-gated greedy, which shares the staged compiler’s gates, still breaks a required group and a dependency closure and recovers only two-thirds of the stale-cheap must-evidence that form-aware staging preserves. The gates earn their keep against the ungated packers; the staging earns its keep against gates alone. The staged compiler’s five distractor admissions (two qualification-adjacent roadmap notes, one large archive note twice across budgets, one plausible note in the heterogeneous pool) and two harmful over-admissions (one misleading instruction-shaped note, admitted at medium and roomy budgets where it fits the earn rule) are part of the promoted result, not footnotes to it β€” a legal admission under uncertainty is rational even when hidden truth disagrees, and the price of uncertainty is exactly what the next experiment must measure. Oracle token gaps (βˆ’210 to +676, mean +51) are reported as diagnostics: negative gaps do not mean the compiler beat anything, since under-fidelity, uncertainty handling, and ledger definitions all move the difference before behaviour is even involved.

What the run does not prove needs equal ink, because the temptation runs the other way: nothing here shows a reader reasoning better, coding better, answering better, or acting more safely. No reader has seen these bundles. Bundle quality is not behavioural quality, and the chapter treats any sentence crossing that boundary as a different claim requiring different evidence.

Influence is not utility

The sibling Memory book earned the distinction this chapter reuses rather than re-derives: behavioural influence asks whether changing the supplied context changed what the model did; behavioural utility asks whether that change improved the externally scored outcome. Context can control behaviour and make it worse β€” the memory experiments showed wrong memory driving harmful actions β€” and that movement is evidence of influence, never evidence of success. The causal chain under test runs bundle into reader into observable action into task outcome, with the bundle as intervention and behaviour as outcome, and every measurement below respects that direction.

Three non-equivalences guard the reasoning throughout. A correct answer does not prove the context helped: the model may have used the bundle, known enough already, guessed, substituted another item, or carried unnecessary decisive context. Mention, citation, and quotation prove even less β€” a model can echo a candidate it never depended on. And a clean bundle proves nothing downstream at all: must-recall, scope correctness, dependency closure, floor legality, and budget compliance describe the input artefact, never what the reader noticed, ignored, obeyed, or was interfered with by. Consider the heterogeneous harmful note under this discipline. Suppose the reader, shown the medium bundle, issues the action the misleading note suggests while the no-context reader holds the release. Behaviour changed and the score fell β€” influence without utility, recorded as exactly that rather than averaged into a middling success number. Suppose instead both readers hold. Same bundle, same labels, zero influence: the hidden harmfulness never reached behaviour, and the ledger’s warning stays a hypothesis about the bundle rather than a fact about the system. Attribution requires counterfactual intervention β€” remove the item and watch the action change, restore it and watch the action return β€” with everything below that standard honestly labelled context-consistent output rather than causal use. The ladder runs from present, to mentioned, to consistent, to removal-sensitive, to restoration-confirmed, and tasks are labelled by how far down they reach rather than assumed at the bottom.

What outside evidence already says

The measurement philosophy arrives with company. The Memory book’s behavioural chapter ran the same matched discipline β€” task, reader, prompt, and decoding fixed while supplied context varied β€” and showed remove/restore attribution with negative controls and wrong-context positives, including the finding that a ledger-minimum oracle can lose to a non-oracle bundle on a particular reader. That last result travels here as a standing warning: the compiler-v1 oracle knows minimum fixture-defined evidence, never which extra support a reader needs, so it is a bundle oracle rather than a behavioural optimum, and any staged-over-oracle behavioural win would demand investigation rather than celebration.

Three external studies sharpen distinct edges. ContextBench, a 2026 preprint, augments coding-agent issue resolution with human-annotated gold contexts and trajectory-level recall, precision, and efficiency β€” and reports the gap this chapter is built around: sophisticated scaffolding barely moves retrieval, models favour recall over precision, and explored context substantially exceeds utilised context. Intermediate quality and final success are different measurements, on 1,136 tasks the authors measured rather than assumed. An ETH Zurich study of repository-level context files finds the sharper version: across agents including Claude Code, Codex, and Qwen Code, supplied AGENTS.md-style files generally do not improve task success while raising inference cost by over twenty per cent β€” more context with more agent activity and no better outcome, published as an ICLR 2026 workshop paper with its population attached. That last pairing deserves emphasis because it is the cost story this chapter must tell separately from the quality story: tokens spent, steps taken, and latency incurred are real prices even when success stands still, and a compiler that saves bundle tokens while multiplying reader calls has moved cost rather than removed it. NoLiMa, peer-reviewed at ICML 2025, supplies the capacity moral: nominal context windows do not establish effective use, with most tested models falling below half their short-context performance at length under latent-association probes. A 2026 preprint on experience reuse in coding tasks adds the selection qualifier the compiler needs: compact correctly-selected prior experience helps effectiveness and efficiency, while autonomous reuse without reliable selection does little or harms β€” selected versus unfiltered being precisely the staged-versus-dump distinction. None of these populations transfer to the book’s fixtures; each is cited for its narrow lesson, and the book’s own matched intervention remains the only evidence that can settle its own question.

The instrument: hold everything fixed except context

Candidate Pool
     ↓
  Compiler
     ↓
ContextBundle
     ↓
   Reader
     ↓
Structured Action
     ↓
Deterministic Grader
     ↓
Behavioural Outcome

      ↑
only ContextBundle changes
across matched conditions

The behavioural ladder reuses Stage 3 outputs directly, never re-approximated: no-context floor, dump, top-k, weighted, hard-gated greedy, staged compiler, and oracle bundles, byte-identical to the frozen run-001 renders. Each rung has a job beyond filling a table. Dump is the naive capacity answer the whole book argues against. Top-k and weighted test whether relevance alone, scored or blended, approximates governance. Hard-gated greedy is the key rival: if it matches staged behaviourally, the dependency, representation, and group machinery beyond gates has no behavioural case on these tasks, whatever its construction virtues. The oracle diagnoses rather than competes, marking what perfect knowledge could achieve. Recompiling silently during inference is forbidden β€” any re-render gets digest-checked against the frozen bytes first, and a changed bundle is a new condition under a new name. Fixtures qualify for behavioural use only under pre-registered rules: the staged condition must emit a model-readable bundle or an intentional failure, a gradable observable action must exist, hidden truth must define correct behaviour, at least one strategy difference must plausibly matter, and the task must not leak its answer independently of context. Every compiler fixture meeting the rule enters; none is cherry-picked for flattering the staged condition, and compile-failure fixtures stay in the analysis as compiler-level evidence without a reader run manufactured for them.

Budgets do not multiply the matrix blindly. One primary regime is chosen from bundle properties alone β€” feasible for staged on enough tasks, genuinely differentiating across strategies β€” before any reader output exists, with tight and roomy reruns confined to a pre-registered subset where budget is itself the mechanism. The reader contract is model-agnostic and deliberately spare: frozen prompt plus frozen bundle in, structured action out, with no tools, no cross-condition memory, no prior conversation, and no network unless some later experiment studies exactly those. Every task-condition-repeat runs isolated, interleaved by task rather than batched by condition, because the condition is the intervention and contamination would be the result. One primary reader is chosen for stable access, pricing visibility, structured output, deterministic decoding, and sufficient capacity β€” documented by criterion, never by flattery β€” with a capability-distant second reader pre-registered to test whether effects belong to bundles or to bundle-reader interactions, each reported separately since stronger readers may absorb distractors and weaker ones may amplify them. That interaction is a first-class result either way: a bundle whose advantage survives only on suggestible readers has a different architectural status from one that moves a robust reader, and the chapter refuses to average the two into a single number that describes neither.

Tasks use arbitrary synthetic project facts β€” chosen backends, blocked migrations, forbidding constraints, fixture-local identifiers and owners β€” so pretraining cannot answer around the bundle, and they demand application over repetition: hold or release the deployment, select the file, ground the parameter, preserve the constraint, reject the wrong-project setting, prefer the fresh value. The arbitrariness is load-bearing rather than decorative: where the facts could come from training data, every condition including no-context might succeed, and the experiment would measure prior knowledge with context as decoration. Actions arrive as compact schemas, parsed deterministically and graded deterministically, with parse failures scored on their own ledger rather than repaired into existence β€” unparseable output is itself a behavioural observation, and context conditions may move it. Graders, prompts, fixtures, and bundle sources are versioned before inference and never rescored after reading outputs.

Controls that can actually bite

The attribution arm runs remove, restore, and two sharper variants on pre-registered decisive items named from hidden truth before any model call: the staged bundle, the same bundle minus exactly the decisive unit, the bundle with it restored, a token-matched variant restoring volume without evidence, and a wrong-context variant substituting a plausible fixture-local falsehood β€” wrong backend, stale version, dropped qualification, wrong identifier. Each variant answers a different sceptic. Removal asks whether the action depended on the item at all; restoration asks whether the dependence replicates or was noise; token-matching asks whether volume rather than information did the work; wrong-context asks whether the reader is even listening, since a reader unmoved by falsehoods cannot credit truths either. All fixtures stay inert, with a deterministic simulator scoring proposed actions and nothing consequential executing. Decisive items are frozen by identity before inference β€” candidate, representation, and reason recorded where the reader can never see them β€” because choosing the removal target after results would turn attribution into storytelling. A no-context negative control anchors every task: behaviour achievable from the bare request calibrates both interference (context harming what needed nothing) and necessity, and no-context success is investigated as leakage, guessability, or grader weakness rather than discarded. The full raw pool may appear diagnostically but never in the headline ladder unless it obeys the same budget β€” the canonical naive condition remains dump/truncate, matched token for token.

Measurement stays dimensional by the placeholder’s standing rule: task success on frozen deterministic scores, constraint and scope and freshness and grounding dimensions only where exercised, harmful-task and harmful-action counts first-class, parse rates, bundle and total-input tokens with output and reasoning tokens beside them, provider cost, latency, and cache telemetry where exposed β€” never a composite, never averaged across readers, never ranked into model winners, since readers are instruments and the unit of study is one reader under different context. The no-context condition earns its place in this ledger twice over: a strong reader may solve history-free tasks from pretraining alone, which calibrates necessity, and supplied context may actively interfere where none was needed, which prices the harm side every other condition assumes away β€” including the subtle case where context merely slows the reader without changing its answer, a latency cost with zero behavioural benefit that only separate cost columns can see. Small samples report task-level outcomes with exact denominators and paired directional counts; no decorative significance machinery. The future results table is designed now with its behavioural columns explicitly pending β€” success, harm, action-change, tokens, input, cost unfilled until a run exists β€” while the completed compiler table above stands beside it under an unmissable separate label. Two tables, two evidence grades, never one.

What the diagnostics are for

The five staged distractor admissions become a diagnostic set rather than a defect list: for each, the experiment asks whether it was behaviourally inert, merely costly, action-altering, or interfering β€” with the live possibility, already demonstrated by the sibling book, that the hidden ledger was stricter than behaviour requires. The two harmful-labelled admissions are mandatory probes wherever valid reader tasks can host them: hidden harmfulness must translate into observed harmful action before the label means anything behavioural, and a reader ignoring them is itself a finding about suggestibility rather than a clean bill of health. Oracle gaps get their behavioural reading only from runs: extra compiler context helping, doing nothing, or interfering; under-fidelity sufficing; oracle minima too strict; readers inferring around gaps. Mechanism ablations arrive as explicit policy versions with digests and traces β€” full staged against staged-without-dependency-costing, without representation alternatives, without group handling, without coverage preference β€” on the minimal fixture subsets isolating each, with hard governance ablated only in synthetic fixtures and every ablation deterministic rather than hand-edited. Aggregate staged wins never promote individual mechanisms; Chapter 24 keeps only what targeted evidence supports. The oracle needs the same discipline applied in reverse. If the staged bundle beats the oracle bundle behaviourally on some task, the ledger is not wrong β€” the reader needed something the minimum left out, perhaps redundant support the fixture author deemed unnecessary, perhaps an anchor the oracle stripped as non-minimal. If the oracle wins, the gap measures avoidable uncertainty overhead rather than reader failure. Either direction refines the fixture as much as the compiler, which is why oracle gaps are diagnostics first and verdicts never.

Falsification is accepted in advance across the full space the prompt requires: staged matching no-context or hard-gated greedy, top-k winning despite worse bundles, oracle failing to beat staged, extra legal context harming readers, distractors proving useful, harmful labels proving inert, remove/restore showing nothing, token-matched recovery matching decisive recovery, effects reversing across readers, bundle metrics failing to predict utility. Each simplifies the final architecture rather than failing the experiment β€” staged complexity unneeded for behaviour may survive for legality, auditability, and failure handling, and that split verdict belongs in Chapter 24. Work one scenario to show what acceptance looks like in practice: if top-k matches staged behaviourally across the feasible set despite its fifteen illegal admissions elsewhere, the conclusion is not that legality is worthless but that this reader on these tasks does not punish the violations top-k commits β€” the gates stay for auditability and for readers that do punish them, while the behavioural case for staging narrows to the fixtures where violations actually bite. Evidence grades the mechanism per population, never in general.

The Stage 4 contract

The implementation handoff is a pipeline, not a sketch: consume frozen compiler bundles with full provenance (experiment, run, commit, strategy, fixture, budget, bundle digest); invoke the fixed reader; parse the structured action; run the deterministic grader; record observations; execute the matched interventions; freeze the behavioural result β€” reusing the existing invocation, observation, and manifest records, adding only what the bundle-to-behaviour link strictly requires. Each reuse is concrete rather than aspirational: the reader call becomes a model invocation with its telemetry discipline intact; the graded outcome becomes an observation with metric, value, verdict, and evidence; the frozen behavioural set becomes a manifest naming the source run it extends. What must be genuinely new is small β€” the link from bundle digest to invocation, and the condition ladder identity per observation β€” and anything beyond that needs justification against the existing structures before a single new record type is admitted. Bundles are never altered during inference; the behavioural run freezes separately with its own experiment, run, reader, prompt, grader, and commit identities while the Stage 3 artifact stays immutable. Synthetic fixtures answer assembly and behavioural causality; real traces will later answer prevalence and realism; real tasks will eventually test external validity β€” each rung a different question, none skipped, none merged.

The evidence hierarchy this chapter leaves behind may be its most durable page: bundle-construction evidence from the frozen run, behavioural correlation from matched reader outcomes, causal attribution from remove/restore and targeted ablation, ecological prevalence from future genuine traces, production benefit from nowhere yet. Nothing in the chapter advances a rung without its run. The hierarchy also disciplines ambition in both directions. Upward: no production claim follows from synthetic causality without prevalence, realism, and external validity, each earned separately. Downward: no stipulation about what “good context” means in the abstract survives contact with the ladder β€” goodness is always goodness-for-a-reader-on-a-task, measured by movement and improvement rather than asserted from bundle properties. What remains unknown is stated plainly β€” whether legal bundles move behaviour, which bundle dimensions predict movement, which mechanisms earn their keep downstream β€” and the next repository to speak is not this one. The compiler has said everything it can say about its own output. A reader must now answer back.

References

  • Project Context. compiler-v1 / run-001, implementation commit ec642dc. Frozen internal result, validated clean, local-only artifact. Consumed: strategy status table, violation counts, distractor and harmful admission identities, oracle token gaps, per-combination bundle bytes. Bundle-construction evidence only; no behavioural claim.
  • Memory book (sibling manuscript, ernanhughes/memory, in development; frozen runs are internal book evidence). Consumed: Ch12 matched-intervention method (fixed task/reader/prompt/decoding, influence/utility split, remove/restore, negative and wrong-context controls) and the ledger-oracle lesson. https://github.com/ernanhughes/memory
  • Li, H., Zhu, L., Zhang, B., et al. “ContextBench: A Benchmark for Context Retrieval in Coding Agents.” Preprint, arXiv:2602.05892, February 2026. 1,136 issue tasks with gold contexts; intermediate quality versus end-to-end success; explored-versus-utilised gap. https://arxiv.org/abs/2602.05892
  • Gloaguen, T., MΓΌndler, N., Mueller, M. N., et al. “Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?” ICLR 2026 Workshop on Memory for LLM-Based Agentic Systems. Context files generally without success gain at 20%+ cost across tested agents. Population-bound. https://arxiv.org/abs/2602.11988
  • Modarressi, A., et al. “NoLiMa: Long-Context Evaluation Beyond Literal Matching.” Peer-reviewed, ICML 2025. Nominal capacity versus effective use under latent-association probes. Measurement-philosophy use only. https://arxiv.org/abs/2502.05167
  • SWE Context Bench authors. “SWE Context Bench: A Benchmark for Context Learning in Coding.” Preprint, arXiv:2602.08316, February 2026. Compact correctly-selected experience helping; unfiltered autonomous reuse limited or harmful. Narrow selection qualifier only. https://arxiv.org/html/2602.08316