← Context From First Principles

Did the Context Help?

Evaluate context by what the model does with it: influence versus utility, a matched design that changes only the bundle, and what a small set of behavioural runs does and does not show.

Two prerequisites have been met, on small cases. The compiler builds a legal bundle and can show its working (Chapters 23 and 24). The bundle can be delivered to a running model and independently seen to have arrived (Chapter 25). Neither says anything about the model.

This chapter asks a different question. Given that the intended context was constructed and delivered, did it improve what the model did?

The book now has three kinds of evidence, and this is the third.

Kind Asks Chapters What it cannot say
Structural is the bundle legal, reproducible and traced? 23, 24 that it arrived, or helped
Transport did that exact bundle reach the model-context boundary? 25 that it helped
Behavioural did the model behave differently, and better? 26 how often it happens in real work

A fourth kind, ecological evidence about how often any of this occurs in real sessions, does not exist yet. The corpus of real sessions is empty. Everything below is controlled evidence on synthetic tasks.

One thing needs saying before the results, because it is easy to let the strength of the transport proof leak into them. The behavioural runs reported below did not go through the live path. The bundles were supplied to the reader as its input through a plain chat interface, with its output constrained to a fixed set of actions. They did not travel through the coding agent’s runtime. Delivery was therefore not in question in those runs, and the transport work of Chapter 25 says nothing about them, or they about it. They are evidence about what a reader does with a bundle. They are not evidence about the live system.

Chapter 24 ended with a table in which the compiler admitted five distractors and two items that a hidden ledger labels harmful. The bundles were legal, within budget and fully traced. Whether those extra items would matter to a reader is not a question the compiler can see. Here is one case in the mixed pool. The bundle is policy-valid, and it contains a misleading, instruction-shaped note that is in scope, fresh, above the relevance threshold, and fits. The compiler can see nothing wrong with it. Does the reader act on it?

That is the only question that matters here, and it can only be answered by running a reader.

Influence is not utility

Two ideas, kept apart.

Influence asks whether changing the supplied context changed what the model did. Utility asks whether that change improved the outcome, as scored by something other than the model.

They come apart. Context can control behaviour and make it worse: wrong context drove harmful actions in the project’s earlier memory experiments, and it does so again below. A change of action is evidence of influence. It is never, alone, evidence of success.

Three non-equivalences guard the reasoning.

A correct answer does not prove the context helped. The model may have known enough already, guessed, or used another item.

Mentioning, quoting or citing a piece of context proves less still. A model can echo material it did not depend on.

And a clean bundle proves nothing downstream. That a bundle has every required item, closes its dependencies and respects every gate describes the input. It says nothing about what the reader noticed, ignored, obeyed or was misled by.

Attribution needs an intervention. Remove the item and watch the action change; restore it and watch the action return. Anything short of that is consistent with the context having been used, not shown to have been. The ladder runs from present, to mentioned, to consistent with, to sensitive to removal, to confirmed by restoration, and a result is labelled by how far down it reaches.

The instrument: hold everything else fixed

The design is a matched one.

candidate pool
      ↓
  compiler
      ↓
ContextBundle  ← only this changes between conditions
      ↓
    reader
      ↓
structured action
      ↓
deterministic grader
      ↓
behavioural outcome

The task, the reader, the prompt and the decoding stay fixed. Only the bundle changes. Seven bundles are compared on each task: no context at all; four simpler ways of building a bundle (dump, top-k, weighted, hard-gated greedy); the compiler’s; and an oracle bundle built from the hidden evidence. They are the frozen outputs of the construction experiment in Chapter 24, byte for byte. Nothing is recompiled during a run.

Each rung has a job. Dump is the naive answer to a large window. Top-k and the weighted packer test whether relevance alone, scored or blended, does the work of governance. Hard-gated greedy is the important rival: if it matches the compiler on behaviour, the extra machinery the compiler carries beyond the gates has no behavioural case on these tasks, whatever its virtues in construction. The oracle diagnoses. It marks what perfect knowledge of the hidden evidence achieves.

The tasks use arbitrary facts invented for the purpose: a chosen backend, a blocked migration, a forbidding rule, a fixture-local identifier. Nothing in pretraining can supply them. Where the facts could come from training data, every condition, including no context, might succeed, and the experiment would measure the model’s prior knowledge with context as decoration.

The reader sees a frozen prompt and a frozen bundle and answers with one action from a fixed set. A deterministic program parses the answer and grades it against hidden truth. No model judges anything. Output that does not parse is scored on its own ledger and never repaired, because unparseable output is itself an observation that context can move.

The measurements are kept in separate columns: task score, harmful action, parse rate, tokens, cost. Nothing is combined into one number, and readers are never ranked, since they are instruments and the unit of study is one reader under different context.

What was run

The evidence is small, and its size is the first thing to say.

Six synthetic tasks, each run at the tightest of the three budgets. Seven bundle conditions per task. One mid-sized local model, at temperature zero, reading each bundle in isolation, its output constrained to a fixed action set. One task, the qualification trap, also received four surgical variants of the compiler’s bundle: with its decisive item removed, restored, replaced by volume without evidence, and replaced by a plausible falsehood. Fifty-four calls in all, forty-six matched cases and eight planned repeats. Every answer parsed.

A second, smaller local model from a different family read a further slice of one task as a transfer probe. It is reported separately and never averaged in.

Two limits apply to everything below. The main model was referred to by a name that moves, and its exact weights were not recorded, so the run cannot be repeated from the name: a rerun would be a new run. And the transfer probe changed the reader and the bundle together, because the main ladder used the tight budget and the probe used the medium one, so it cannot show that readers differ on identical bundles.

Book result — matched behaviour. On the six-task tight ladder, changing the frozen bundle changed the full parsed action on all six tasks between no context and the compiler’s bundle (the action name on four of six). It improved the scored outcome on two, degraded it on one, and left three unchanged. From hard-gated greedy to the compiler, four of six response records changed in the wording of target, value or reason, none changed the decided action, and none changed the score. From the compiler to the oracle, two of six records changed in wording, none changed the action, none changed the score. The evidence register lists this result as Matched behaviour, and the transfer probe as Reader-transfer probe.

Read the three comparisons separately, because they say different things.

No context to compiler: context can influence behaviour without improving utility. Six tasks were influenced; the net gain was one task.

Hard-gated greedy to compiler: the extra machinery changed no decision and no score on this set. That is not a tie to break with a larger sample. It is the finding.

Compiler to oracle: the hidden-evidence minimum changed no decision and no score either. A richer bundle matched the minimum sufficient one exactly.

The task-level rows carry the weight, because aggregates hide direction:

Task No context With supplied context What it shows
stale cheap candidate abstains, 0.0 proceeds, 1.0 in every supplied condition context necessary
mixed pool abstains, 0.0 refuses correctly, 1.0 context moves abstention to the right refusal
qualification trap holds, 1.0 releases, 0.0, harmful, in every supplied condition including the compiler’s and the oracle’s supplied context systematically causes the failure
dependency trap correct correct no scored difference
wrong scope correct correct no scored difference; the no-context reader already succeeds
calibration abstains, 0.0 0.0: one bundle abstains; the others echo bundle-derived wrong values instead of the identifier the task states the negative control failed

No sentence of the form more selected context is better survives that table. Context helped, hurt and did nothing, sometimes on adjacent rows of the same run.

The compiler did not win

Say it plainly, because the temptation runs toward the larger machine. Nothing in the behavioural run promotes the compiler over hard-gated greedy. The comparison moved no action and no score on any of the six tasks.

The construction experiment established that the compiler is far more reliable than the simpler assemblers at building legal bundles. That stands untouched. But reliability of construction and utility of behaviour are different objectives, measured by different runs, and the second does not repeat the first’s verdict.

The right synthesis is a separation, not a ranking. The compiler earned construction-level machinery the simpler packer lacked. The extra machinery did not demonstrate behavioural utility over hard-gated greedy on this small experiment. That is a better guide to engineering than a declared winner, because it says what kind of claim each mechanism can carry. A fixture on which the two decide differently, with the compiler’s decision scoring higher, would promote the compiler behaviourally on that population. Until one exists, the promotion stays where the runs put it.

The results that do not flatter

The qualification trap stays in the foreground. The reader with no context holds. Every reader given any supplied bundle releases, harmfully, including the compiler’s and the oracle’s, and the surgical variants and the repeats agree. Adding context can create the failure, and it does so with bundles that are legal, budgeted, traced and fully delivered in the sense that they were the input. The compiler saw eligible, useful-looking items. The reader read them damagingly. That is the argument for evaluating behaviour after assembly, and it is not an argument for a new harmfulness classifier, and none is proposed.

The failed negative control stays equally visible. On the calibration task every condition scored zero, because the reader followed content in the bundle over the identifier the task itself stated. That reads two ways at once: as evidence of strong context sensitivity, and as a limitation of the instrument, since a benchmark that cannot keep its own negative control does not get to make sweeping claims from its positive conditions.

Removal and restoration did not resolve. The preferred standard is to remove the decisive item and watch the action change, then restore it and watch the action return. On the qualification task the parent bundle already sat at the behavioural floor, so removal could show nothing. The restored bundle was byte-equal to its parent, which shows the restoration path has integrity. The volume-matched bundle was well formed, and the falsehood bundle showed the reader responding to misleading framing without being rescued by the true item. None of that attributes the harm to any candidate. The result supports the weaker claim that the bundle condition caused a matched difference. Which candidate caused it is open.

Reader dependence is open too. On the mixed pool the main reader refused correctly for every supplied bundle at the tight budget. The transfer probe ran the second reader on the same task at the medium budget and it abstained on all seven cases. Reader and bundle changed together. Only the no-context condition is identical across the two runs, and both readers abstained there. So the probe shows that a second reader abstained on bundles the first never saw. It does not show that readers behave differently on identical bundles. That run does not exist.

The oracle lesson follows. The bundle oracle knows the hidden minimum sufficient evidence. On the qualification trap it releases harmfully like every other supplied bundle, and on the calibration task it echoes wrong content like the rest. That is not the oracle failing. It optimises a different objective, hidden evidence sufficiency, from reader behaviour. Minimum semantic evidence may be reader-relative, since one model may need redundancy the ledger never encoded and another may ignore support it was given. This run tested one reader and does not isolate that.

What the failed wave is and is not

Chapter 25 opened on a behavioural wave that ran 24 times without a valid observation. It does not appear in the results above, and it should not. It is an instrumentation result. Nothing in it can be read as a finding about whether context helps, or as a failure of the hypothesis. Twenty-four excluded runs contribute no behavioural evidence, in either direction.

That wave was designed to ask the next question, and it is worth stating, because it is the experiment the working live path now makes possible. Eight tasks, three conditions: no context, a distractor of the same size and shape, and the exact minimal context the task requires, delivered through the live runtime and independently observed. Its success condition is narrow: the context condition succeeds where the no-context condition fails, with transport verified on every run, and the distractor behaves like no context. It is a leverage check on a small set, not an estimate of effect size. And it starts only behind the liveness gates: it will spend no inference until a live canary has shown, for the scheduled model, that injected material is observed once and that a compiled bundle survives the whole chain.

It has not been run. The behavioural evidence in this book is therefore exactly what is reported above.

Other people’s evidence

The wider evidence is narrow and comes from other populations, and it is cited for what it shows and no further.

A 2026 benchmark that augments coding-agent issue resolution with human-annotated gold context reports that the context an agent explores far exceeds the context it uses, and that intermediate retrieval quality and final success are different measurements. A study of repository-level context files across several coding agents found that supplying them generally did not improve task success while raising inference cost by more than a fifth: more context, more agent activity and no better outcome. Long-context work shows nominal windows overstating effective use. A preprint on reusing experience in coding tasks finds that compact, correctly selected prior experience helps and unfiltered autonomous reuse does little or harms, which is the selected-against-unfiltered distinction this book has drawn since Chapter 14. None of these populations transfers to the synthetic tasks here. The book’s own matched design is the only evidence that can settle the book’s own question.

What the evidence amounts to

Follow the hierarchy honestly, and each step is a different size.

Structural evidence is strong for what it covers. Transport evidence is a single qualified pass on a trivial case, with a stated boundary. Behavioural evidence is six synthetic tasks, one reader whose exact identity was not recorded, one budget, and a transfer probe that does not isolate what it was meant to. Ecological evidence does not exist.

The method is the part that is sound, and it is what the book can hand on. Fix the model, the task, the decoding and everything except the bundle. Vary the bundle deliberately. Grade deterministically. Refuse every inference the comparison does not support. On that discipline the result is modest and clear: a legal, delivered bundle can help, hurt or do nothing, and the extra machinery of the compiler has not yet been shown to change what the reader does.

That leaves the whole system to be put together, and an honest account of what the project has shown and what it has not.

What does the finished system look like, and what is still open?

References

  • Project Context evidence register (companion repository, evidence/README.md). Lists the identifiers, code revision and published artefacts behind each measured result. Consumed here: Matched behaviour (the seven-condition ladder on six tasks, its two influence granularities, the harm decomposition, the removal and restoration series; real reader behaviour on synthetic tasks, the main model recorded only by a moving name), Reader-transfer probe (seven abstentions at the medium budget; does not isolate reader dependence) and Instrumentation failure (an inconclusive wave, discussed and excluded from the behavioural evidence). No ecological claim.
  • Memory book (sibling manuscript, unpublished). Method background, not evidence for this book: the influence and utility split, the remove-and-restore attribution standard with negative and wrong-context controls, and the lesson that a minimum-evidence oracle can lose to another bundle on a particular reader.
  • Li, H., Zhu, L., Zhang, B., et al. “ContextBench: A Benchmark for Context Retrieval in Coding Agents.” Preprint, arXiv:2602.05892, February 2026. 1,136 issue tasks with gold contexts; intermediate quality versus end-to-end success; explored-versus-utilised gap. Population-bound. https://arxiv.org/abs/2602.05892
  • Gloaguen, T., Mündler, N., Mueller, M. N., et al. “Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?” ICLR 2026 Workshop on Memory for LLM-Based Agentic Systems. Context files generally without success gain at 20%+ cost across tested agents. Population-bound. https://arxiv.org/abs/2602.11988
  • Modarressi, A., et al. “NoLiMa: Long-Context Evaluation Beyond Literal Matching.” Peer-reviewed, ICML 2025. Nominal capacity versus effective use under latent-association probes. Measurement-philosophy use only. https://arxiv.org/abs/2502.05167
  • SWE Context Bench authors. “SWE Context Bench: A Benchmark for Context Learning in Coding.” Preprint, arXiv:2602.08316, February 2026. Compact correctly-selected experience helping; unfiltered autonomous reuse limited or harmful. Narrow selection-qualifier use only. https://arxiv.org/html/2602.08316