Chapter 12 of 17

Does Better Context Change Behaviour?

Concepts

Chapter 12 β€” Does Better Context Change Behaviour?

Source: 12-chapter.md

What this chapter is really about

Underneath the ladders and interventions, this chapter is about the moment the book is forced to honour its own definition. Every earlier chapter measures what the context contains; this one removes the past and watches what the system does. Its deepest move is splitting the behavioural question in two: influence (did memory change the action?) from utility (was the change better?). A system can be highly influenced and worse, uninfluenced and correct, or correct in both conditions β€” and only the paired intervention tells those apart. The hidden thesis: answer quality is a proxy for memory, and proxies have now been caught disagreeing with the thing itself (0.72 recall with 0.955 coverage), so the book needs an instrument that measures the thing itself.

Current thesis

Explicit claims (now measured in ch12-20260920T204414Z-behavior, grader v2, prompt v2)

  • Matched seven tasks: no-memory 0.226 β†’ RAG 0.393 β†’ selected 0.357 β†’ assembled 0.488 β†’ oracle 0.524; restore tops all (0.778 over 6 ablation tasks).
  • Undifferentiated history hurts: full 68-unit history 0.048 < no memory 0.226; project-only history (B1P) recovers to 0.179 β€” contamination explains roughly half the loss, unselected volume the rest.
  • Remove/restore tracks: 0.250 removed β†’ 0.778 restored; three full Mβ†’MAβ†’MR traces (fix-store, cite-rule, corpus-cleanup).
  • Wrong memory is influence with teeth: BW 0.042 with harmful tasks in 2 of 4 (facade deletion, stale-backend configuration, both B6); full history and decisive-removal also delete the facade once each.
  • Strict influence table (73 pairs): 14 full-success contributions, 2 harmful, 0 not-necessary, 57 failed-to-repair; directional table beside it: actions change on 73/73, scores improve on 32, degrade on 8, unchanged on 33.
  • Second reader agrees on direction, differs on profile: ministral matched 0.000 β†’ 0.524 β†’ 0.452 β†’ 0.643 β†’ oracle 0.536; deletes the facade even with no memory.
  • Transfer (5 tasks): memory 0.95 vs no-memory 0.35; one task non-diagnostic (promote-decision self-answers).
  • Type A on fixture evidence with five demotion clauses (synthetic/few tasks, one primary reader, two behaviourally flat tasks, repeats on 3 tasks Γ— B4 only and deterministic, 73 single-shot remainder).

Implied claims

  • Behavioural quality is irreducibly multi-dimensional: constraint, avoidance, continuation, currency and harm do not move together, so no single success number can certify memory.
  • The ledger is conservative about necessity: unlabelled/distractor items can carry actionable content (loop record substituting for the experiment-pending MUST), which the anomaly investigation promotes to a versioned finding rather than a silent edit.
  • Readers differ in exploitability: the same memory helps both readers, but floors, ceilings and suggestibility differ β€” the instrument measures the memory–reader interaction.
  • Layer effects are non-monotonic: frame selection (B3 0.357) trails strong RAG (B2 0.393) on both readers while assembly (B4 0.488) recovers the advantage at fewer tokens β€” better selection metrics do not imply better behaviour.
  • An evidence oracle is not automatically a behavioural oracle: ministral’s assembled context (0.643) beats the auditable oracle (0.536), and restoration (0.778) beats it on the primary reader β€” the oracle strips SHOULD-grade material readers actually use.

Not yet established

  • That the nine fixtures represent real project decisions (five transfer tasks say consistent, at small-n, with one self-answering).
  • That the dimension set is complete: efficiency, calibration and multi-step behaviour are secondary or absent.
  • That results survive vague, evolving tasks: all fixtures are crisp, standalone decisions.

What the chapter already gives us

  • The influence/utility split. Two measurements where the literature often reports one; a system can be influential and harmful, and the table shows both cells occupied.
  • The remove/restore intervention. The strongest controlled demonstration available: behaviour breaks when the decisive item leaves and recovers when it returns, under otherwise matched conditions.
  • The full-history negative. Preservation without selection scores below nothing β€” the Paradox as a number, and the reason selection chapters precede behaviour chapters.
  • The ledger-conservatism finding. Missing-MUST/high-coverage cases traced to substituting evidence, recorded as a benchmark-version finding rather than relabelled.
  • A reusable evaluator. Versioned tasks, graders, simulator and paired-comparison machinery that Chapters 13–17 inherit as their downstream judge.

Where the current treatment stops

  • Grader contracts freeze the author’s normativity in place (what counts as full credit is asserted before running, never derived from outcomes). Different engineers might grade the prose/architecture split differently; the contracts don’t encode profiles.
  • Name-matching on free-text mechanism names proved brittle (credit-rule rejects under B4/BO/BA/BW): four verdicts lost to a missing name field rather than a wrong verdict. The contract limitation is recorded, not repaired post hoc.
  • One fixture is non-diagnostic (review-prose: 0.0 under every condition including oracle) and one intervention likewise; both are published as instrument failures.
  • Repeats cover 3 tasks Γ— B4 only (9 outcomes, deterministic); the remaining 73 outcomes are single-shot at temperature zero.

The deeper territory

  • Profiles and normativity. Grader contracts are utility functions in disguise. Making risk profiles explicit (cautious vs fast-moving) would turn grading disagreements into testable profile differences rather than author taste.
  • Vague-task behaviour. Real tasks arrive underspecified and evolving; the current fixtures are crisp one-shot decisions. Behaviour under uncertainty β€” hedging, clarifying, updating β€” needs task shapes this instrument doesn’t have.
  • Harm asymmetry. Two harmful actions under wrong-memory against zero elsewhere suggests memory safety is a tail phenomenon: means hide it, and only the separate harmful-action rate catches it. The measurement lesson generalises beyond this chapter.
  • Reader suggestibility as a trait. The second reader benefits more and follows poison more readily. If exploitability and suggestibility correlate across readers, memory architecture and reader hardening are joint problems.

Concepts worth developing

Intervention-graded memory quality

Idea. Score memory systems by the behaviour delta they cause under matched remove/restore interventions, not by retrieval or answer metrics. Quality becomes a causal quantity: the improvement attributable to the specific remembered item.

Why it matters. It ends the proxy regress β€” every earlier metric is validated against whether moving the memory moves the behaviour.

Connection to the current chapter. Generalises the M3→MA→MR pattern from three traces to the standard certification for any memory claim.

Broader implication. Later chapters (consolidation, forgetting, procedures) inherit their admission test: no intervention delta, no mechanism.

What remains unresolved. Cost (model calls per intervention); nondeterminism handling at scale; which tasks qualify as diagnostic.

The ledger as a hypothesis about necessity

Idea. Treat MUST labels as claims about behavioural necessity that interventions can falsify: if removing X never changes behaviour while alternatives substitute, X was sufficient, not necessary. Support-group semantics (Ch 7 OR-groups) absorb the correction without relabelling.

Why it matters. It converts the anomaly (0.72 β†’ 0.955) from an embarrassment into a mechanism for improving the benchmark: necessity is discovered, not stipulated.

Connection to the current chapter. The loop-record-for-experiment-pending substitution is the first recorded instance.

Broader implication. Every ledger in the book becomes revisable under version rules rather than authoritative β€” evidence discipline applied to the instrument itself.

What remains unresolved. How many interventions falsify a label (one reader? two?); whether necessity is reader-relative.

Important distinctions

  • Influence (did behaviour change?) vs utility (was it better?) β€” the chapter’s master cut.
  • Memory-as-answer (recall the fact) vs memory-as-resource (apply it to decide).
  • Correct-with-memory vs correct-because-of-memory (only the counterfactual separates them).
  • Evidence sufficiency (ledger) vs reader sufficiency (behaviour) β€” different quantities, honestly diverged.
  • Ordinary errors (averaged) vs harmful actions (never averaged away).
  • Diagnostic fixtures (separate conditions) vs non-diagnostic ones (published, not dropped).

What mechanism would make this work?

Nine BehaviourTasks with hidden ledgers and frozen grader contracts + deterministic ProjectWorld simulator + frozen-context ladder (B0/B1/B2/B3/B4/BO) + remove/restore/scramble/wrong-memory interventions + fixed-reader paired runs with repeats + second-reader transfer + five time-locked transfer tasks. Missing: vague-task shapes, multi-step behaviour, efficiency dimensions, profile-parameterised grading, name-robust verdict scoring.

Connections to the rest of the book

  • Consumes Chapters 3/10/14 as frozen conditions (C0, C5, A6@768, CO-auditable) β€” the pipeline on trial as a whole.
  • Gives Chapter 11 its first action-level validation (facade trap must not fire; caller loops must steer the plan).
  • Hands Chapter 13 the wrong-memory premise with numbers (harmful actions under BW/B1) plus the reusable evaluator.
  • Hands Chapters 15–17 the downstream judge: later mechanisms earn themselves by moving this instrument.
  • Returns Chapter 1’s definition operationalised, and Chapter 2’s failure taxonomy extended (B0–B8).

Beyond the current book

  • Interventionist causation in evaluation: remove/restore as the behavioural analogue of ablation in systems work and of counterfactual fairness probes.
  • Safety evaluation practice: tail-harm measurement separated from mean performance β€” the harmful-action rate as a general pattern.
  • Benchmark design: non-diagnostic fixtures published rather than dropped; contracts frozen before running as preregistration.
  • Shen et al., Mem2ActBench (ACL 2026) β€” the memory-as-answer vs memory-as-resource cut with reverse-generated memory-dependent tasks; motivates memory-dependence-by-construction. Status: peer-reviewed. Verified via ACL Anthology (2026.acl-long.370).
  • He et al., MemoryArena (2026) β€” Memory-Agent-Environment loops with interdependent subtasks; saturated recall collapsing in agentic settings parallels the 0.72/0.955 divergence. Status: preprint. Verified via project site and arXiv record (2602.16313).
  • Wang et al., Agent Workflow Memory (ICML 2025) β€” remembered structure moving task success and step efficiency on web navigation; precedent for action-level memory claims. Status: peer-reviewed. Verified via PMLR (v267).
  • Tan et al., MemBench (ACL Findings 2025) β€” effectiveness/efficiency/capacity with participation/observation scenarios; background for cost dimensions, no decision metric of its own. Status: peer-reviewed. Verified via ACL Anthology (2025.findings-acl.989).
  • HagstrΓΆm et al., CUB (ACL 2026) β€” context-present β‰  context-used; synthetic wins inflate. Status: peer-reviewed. Verified via ACL Anthology (2026.acl-long.1151).
  • HagstrΓΆm et al., DRUID / Reality Check (ACL 2025) β€” real retrieved contexts vs synthetic; motivates the real-project transfer. Status: peer-reviewed. Verified via ACL Anthology (2025.acl-long.968).
  • Chang and Chen, LOCOMO-CONV (2026) β€” strong retrieval not translating into response quality; silent grounding as a discrepancy class. Status: preprint. Verified via arXiv (2609.03467).

Possible future claims

Already supportable

  • Structured project memory improves the behaviour of a fixed reader over no memory and over RAG on memory-dependent tasks (measured, fixture-scoped).
  • Removing decisive memory regresses behaviour; restoring it recovers it (three full traces).
  • Undifferentiated history is worse than none (measured).
  • Harmful memory causes harmful actions the clean conditions never take (four primary-reader instances: BWΓ—2, B1Γ—1, BAΓ—1, separated metric).

Plausible but needs development

  • Assembled memory holds its edge over RAG on new tasks and readers (both readers: B4 > B2; the B3 dip does not replicate into B4).
  • Real-project decisions show the same deltas (transfer consistent at n=5).
  • Substitution-aware ledgers (OR-groups from intervention data) replace MUST conservatism.

Speculative

  • Suggestibility correlates with exploitability across readers generally.
  • The instrument suffices for consolidation/forgetting/procedure admission without modification.

Claims worth attacking

  • “Memory improves behaviour.” Counter: on 57 of 73 pairs neither condition reaches full success; the tasks where memory helps may be the ones engineered to need the exact supplied items β€” circularity by fixture design. The defence is the remove/restore within-task control plus the irrelevant-task control, but the charge has force wherever B0 already succeeds.
  • “Full history hurts.” Counter: the full-history condition dumps unsectioned text through a render the assembler never built; the harm may be formatting, not content. B1P answers it partway: project-only history is still sectioned-away-from-nothing yet recovers roughly half the loss, so contamination and unselected volume both contribute.
  • “Restoration tops oracle.” Counter: BR contexts are larger than BO (1,297 vs 1,329 tokens β€” comparable) and contain redundant SHOULD material; the win may be coverage-by-volume, not decisiveness. Check token-matched comparisons before repeating this claim.

Tensions and counterarguments

  • Hard tasks vs diagnostic tasks: the 57 never-fully-successful pairs argue the fixtures are too hard; the 14 full-success (32 improved directional) pairs argue they are diagnostic. Both can be true of different tasks β€” the per-task table, not the aggregate, is the honest unit.
  • Reader-relative necessity: if MUST-falsification depends on the reader, ledgers become reader-indexed and the benchmark loses a fixed target. The version-rules finding needs a same-reader replication policy.
  • Safety means vs tails: BW scores far below B0 on average (0.042 vs 0.226) yet the harms also appear under B1 and BA, never under B0/selection/assembly/oracle. Any future chapter reporting means without the harmful-action rate repeats the exact error this chapter was built to end.

Examples and thought experiments

  • The sqlite trap across readers: deliberately wrong memory drives both readers to configure SQLite (B6 recorded in both); framed conditions choose PostgreSQL. The same trap, two readers, same failure β€” the book’s isolation method applied to action, with the harm located at the wrong-memory extreme rather than between RAG and framing.
  • The facade under two listings: staged triples protect it, bare listings delete it. The difference between B4-cleanup (facade intact) and BW-cleanup (B6) is exactly Chapter 11’s licence, expressed in behaviour.
  • The honest baseline: no-memory fix-store emits unknown_backend with a misspelled parameter key (0.333) β€” without demonstrations the reader hedges rather than copies. What does the tightened goal dimension reward, and what would a demonstration have bought?

Potential demonstrations or experiments

Add: (1) vague-task conditions (underspecified present state, clarification allowed, scored on eventual action); (2) multi-step behaviour (plan then execute with world feedback between steps); (3) profile-parameterised grading (same actions, cautious vs fast-moving contracts); (4) name-robust verdict scoring (paraphrase-tolerant mechanism identity). Proposed; none run.

Research questions this chapter creates

  • Which tasks are diagnostic for memory behaviour, and can diagnosticity be pre-registered rather than observed?
  • Do exploitability and suggestibility correlate across readers and scales?
  • What is the cheapest intervention that preserves attribution (single remove? matched paraphrase? counterfactual prompt)?
  • How do support-group semantics absorb intervention-falsified MUST labels under benchmark version rules?

Architectural implications

  • The evaluator is shared infrastructure for Chapters 13–17: wrong-frame fallback, consolidation, compression, forgetting, procedures and the capstone all face this judge.
  • Grader contracts are frozen before running, versioned in manifests β€” preregistration as architecture.
  • Harmful-action rate stays separate from every mean, book-wide from here on.

How would we know this works?

The chapter works if matched memory/no-memory interventions separate behaviour with the gains tracing to decisive items (remove regresses, restore recovers) while wrong-memory and irrelevant-memory controls behave as designed. It fails usefully if no-memory already succeeds everywhere (fixtures non-diagnostic) or if memory never moves behaviour (pipeline behaviourally inert) β€” either published as such.

The chapter at its highest level

The ideal version would teach: influence/utility with the table fully populated; remove/restore as the certification standard; the ledger-as-hypothesis with versioned falsification; reader-relative necessity faced directly; profiles in grading; vague tasks admitted. The current version builds the instrument, populates it on nine fixtures, and records exactly where it is blind.

Discussion

Start here

  • 57 of 73 pairs never reach full success under either condition. Are the fixtures too hard, or is the reader too small β€” and which hypothesis would a repeat with a stronger reader test?
  • The v1 B0 baseline copied example shapes (goal credit without constraint credit). Answered in v2: prompts carry no action demonstrations and goal credit requires a valid backend β€” B0 fix-store fell from 0.667 to 0.333 on the tightened dimension.
  • review-prose is dead under all conditions including oracle, and credit-rule is flat at 0.0 likewise: the render contains the checklist units yet both readers report experiment markers instead (prose), and reject without naming the mechanism everywhere including oracle (credit-rule). Render-content audit says the evidence is present; the failure is reader-side prioritisation and contract over-specificity respectively. Both stay published as instrument failures.

Push the idea further

  • If necessity is reader-relative, do ledgers become reader-indexed β€” and does the benchmark then need a reference reader pinned in the contract?
  • Harm lives in tails that means hide. What other book metrics (precision, recall, coverage) might be hiding tail failures behind acceptable means?
  • After the capstone, does the instrument retire the key-claim scorer, or do the two proxies keep each other honest?

Decisions we need to make

  • Whether profile-parameterised grading enters E-13 or stays future work.
  • Whether the loop-record/or-alternative finding triggers a benchmark-version proposal now or waits for replication.
  • Whether vague-task shapes join the instrument before Chapter 13 uses it.

Claims worth attacking

  • “Remove/restore proves the memory mattered.” Counter: restoration adds tokens and SHOULD material alongside the decisive item β€” coverage-by-volume confounds decisiveness. Token-matched restore control needed.
  • “Full history hurts.” Counter: the B1 render is unsectioned and unmarked; assembly confounds selection. B1P answers it partway β€” project-only history still unsectioned recovers roughly half the loss β€” but a sectioned full-history control would separate formatting from content fully.
  • “Two readers agree.” Counter: n=2 with opposite failure modes (cite-rule, cleanup) is divergence as much as agreement β€” report the interaction, not the consensus.

New ideas worth exploring

  • Diagnosticity prediction: pre-registering which tasks should separate conditions before running.
  • Intervention-cheap attribution: single-remove or paraphrase controls instead of full ladders.
  • The tail-metric programme: harmful-action-style separated rates for every book mean.

Explain this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Apply this chapter with AI

Copy this prompt into ChatGPT, Claude, Gemini, a local model, or another AI.

Chapter 1 defined memory behaviourally: remove the past, rerun the present task, and ask whether behaviour changed and whether the change was an improvement. Eleven chapters later, that test has never actually been run. Chapters 3 to 11 built a pipeline β€” retrieval, structure, association, routing, lineage, temporal state, open loops, frames, derived consequences β€” and measured each stage against ledgers of what the context contains. During development, a later assembly experiment, reported in Chapter 14, exposed the measurement anomaly that motivates the behavioural experiment reported here: at a 768-token budget the assembled context holds required-evidence recall of only 0.72 yet the reader answers at 0.955 key-claim coverage, better than the full raw context at 0.879. Either the ledger overstates what the task needs, or the reader answers without the evidence the book calls required, or the metric gives false confidence. Answer quality alone cannot distinguish those cases. This chapter builds the instrument that can.

A correct answer is not yet memory

A correct answer is compatible with four states of the world: the system used memory well, the system knew the answer already, the system guessed, or the system needed only part of what the ledger demands. The key-claim scorer cannot tell them apart. Neither can citation: a model can emit source identifiers it did not reason from. The only evidence that survives is interventional β€” hold the model, the task, the prompt and the decoding fixed, change only the memory supplied, and watch whether the structured behaviour changes and in which direction.

That gives the chapter its two separated measurements. Behavioural influence asks whether memory changed what the system did. Behavioural utility asks whether the change was better. A system can show influence without improvement, improvement without influence is impossible by definition, and a correct answer under both conditions demonstrates nothing about memory at all.

Tasks that require the past

Nine controlled tasks carry hidden behavioural ledgers written before any model ran, each with a frozen grader contract stating what earns full, partial and zero credit, what is forbidden, and what counts as equivalent. At least half depend on project-specific facts no model could know parametrically: the pgvector/HNSW/tsvector configuration of the baseline store, the reversibility rule that refuses a measured 0.474-to-0.464 mechanism, the four open loops blocking release, the contracted facade that must survive the corpus_import cleanup. And the tasks require application, not repetition β€” choosing a backend, holding a chapter, planning a migration β€” scored through a tiny deterministic project simulator whose transitions (including which actions count as harmful) are fixed in advance. The ninth fixture is the negative control: three present-state labels to echo exactly, where any behaviour change under memory is intrusion, not help.

The ladder

Every task runs the same reader (llama3.1:8b, temperature zero) under paired conditions that differ only in supplied memory. No-memory (B0) is the mandatory floor. Full visible history (B1, all 68 units across three projects) asks whether undifferentiated past helps, and a project-only control (B1P, the 57 Memory-project units) separates foreign-project contamination from unselected volume. Strong RAG (B2, frozen Chapter 10 C0) and frame-conditioned selection (B3, frozen C5) reprise the book’s own pipeline. Assembled memory (B4, the frozen A6 assembly output at 768 tokens whose construction is reported in Chapter 14) is the practical system. The auditable oracle (BO) is the diagnostic ceiling. Three interventions complete the ladder: decisive memory removed (BA), restored (BR), and deliberately wrong memory (BW) as a positive control for influence. A scrambled control (BS) replaces decisive evidence with distractors at matched shape. Frozen contexts are hashed; the runtime never sees a hidden label; repeats on three headline tasks test determinism.

Book result (frozen run ch12-20260920T204414Z-behavior, grader v2, prompt v2, 9 tasks, 82 controlled outcomes, 3 repeats each on 3 tasks, plus a 76-outcome second-reader run and 10 time-locked transfer outcomes). Headline means are task-averaged over the seven tasks that run the full B0/B2/B3/B4/BO ladder, repeats averaged within task first β€” the only comparison in which every condition faces the same tasks:

condition            success   harmful tasks   mean tokens
B0  no memory          0.226      0/7               β€”
B1  full history (68)  0.048      1/7            6242
B1P project-only (57)  0.179      0/7            5495
B2  strong RAG         0.393      0/7            2029
B3  frame-selected     0.357      0/7            1909
B4  assembled (768)    0.488      0/7            1180
BO  oracle-auditable   0.524      0/7            1329
BA  decisive removed   0.250*     1/6            1365
BR  decisive restored  0.778*     0/6            1480
BW  wrong memory       0.042*     2/4             120
BS  scrambled          0.167*     0/2            1015

Starred rows run on intervention subsets (6, 6, 4, 2 tasks), not the matched seven. Parse rate is 75 of 82; all 7 failures sit in long unsectioned contexts (B1/B1P/B3/BA).

Four things in that table matter more than the headline.

Undifferentiated history hurts β€” and contamination is only half the story. Full 68-unit history across three projects scores 0.048 against 0.226 with nothing. Project-only history (57 Memory units, the B1P control) recovers to 0.179 β€” better than full history, still worse than nothing. Both history conditions sit below no-memory, so the earlier reading that only the full dump fails is false. Removing foreign projects repairs roughly half the loss; the remainder is volume without selection. So the Perfect Memory Paradox from Chapter 1 arrives as a number with its decomposition: preservation without selection is not memory but interference, and foreign-project contamination is one component of the interference, not the whole of it. Even perfectly scoped accumulation remains harmful without selection.

Better selection is not yet better behaviour. Strong RAG reaches 0.393, frame selection falls back to 0.357, assembly recovers to 0.488 at roughly three-fifths the tokens. Frame conditioning, which Chapter 10 showed improves context-selection metrics, does not improve downstream behaviour by itself here β€” on either reader (the second reader shows the same non-monotonicity: 0.524 β†’ 0.452 β†’ 0.643). That is a valuable finding, not a failure: better context-selection metrics do not automatically produce better downstream behaviour. The stronger end-to-end result is structured assembled memory over strong RAG over no memory, for both readers. Assembly recovers the behavioural advantage while using substantially less context β€” roughly 39% fewer tokens than selection on matched tasks. The layers interact non-additively: RAG establishes a strong retrieval baseline, frame-conditioned selection improves evidence properties without improving behaviour by itself, and bounded assembly over the framed evidence produces the best practical behavioural result. That non-additivity matters for later chapters. The oracle leads at 0.524, and restoration (0.778, over the 6 ablation-applicable matched tasks) tops even the oracle β€” returning the decisive item to an otherwise assembled context beats the ledger-minimum set, because the assembled context carries useful SHOULD-grade material the oracle strips.

Removal and restoration track. Across ablation-applicable tasks, removing decisive memory drops success to 0.250 and restoring it lifts success to 0.778. Three full Mβ†’MAβ†’MR traces show the pattern individually: fix-store holds 1.0 across all three repeats, collapses to 0.333 with the postgres record stripped, and returns to 1.0 on restore; cite-rule goes 0.0 β†’ 0.0 β†’ 1.0, the restored decisive unit producing the book’s own rule identifier against two neighbour distractors; corpus-cleanup goes 0.833 β†’ 0.0 β†’ 0.833, the ablated run updating docs immediately and deleting the contracted facade. The reader’s behaviour depends on the specific remembered item, in both directions.

Wrong memory is influential and destructive. BW averages 0.042 with two harmful tasks out of four: the bare facade-deletion listing deletes the contracted facade (B6), and the stale SQLite record drives a SQLite configuration (B6 under grader v2, which records acting on superseded operational state as a harmful event). Full history also deletes the facade once. Most strikingly, removing decisive caller evidence (BA) lets the trap win too: the ablated cleanup run deletes the facade with B6 recorded. On fix-store the ladder is progressive rather than divergent: no-memory emits an unusable unknown backend (0.333), strong RAG and frame selection configure PostgreSQL partially (0.833 each), assembled and restored contexts configure it fully (1.0) β€” while deliberately wrong memory configures SQLite with B6 recorded. The failure-avoidance mechanism Chapter 8 built now has behavioural consequences, but in this run they appear at the wrong-memory extreme rather than between RAG and framing. The two directions together are the chapter’s result in miniature: correct memory can improve behaviour, and wrong memory can also control behaviour. A memory layer that decides what the model may see can cause the harmful action, not merely fail to prevent it. That is the positive control doing its job, and it is the bridge to Chapter 13: the system now has enough leverage to hurt the reader, so the next question is when it should trust its own framing that strongly.

Dimensions, not one number

Task success averages hide where memory helps. Reported separately, task-averaged over all tasks: constraint adherence moves 0.083 (B0) β†’ 0.500 (B2/B3) β†’ 0.583 (B4) β†’ 0.600 (BO), and removal reverses it (0.250) while restoration completes it (0.750); open-work continuation moves 0.000 β†’ 0.625 (B4) β†’ 0.667 (BO), with restoration reaching 1.000 from 0.000 removed; failure avoidance is 1.000 everywhere except under project-only history, where a parse failure zeroes the single task, and under wrong memory, where the stale record wins (0.000); goal adherence moves 0.333 β†’ 0.500 β†’ 0.429, restoration lifting it to 0.714. Harmful actions are zero under no-memory, selection, assembly and oracle β€” and nonzero under full history, decisive-removed cleanup, and wrong memory. The book’s pipeline does not raise a single score; it raises constraint-following, continuation and avoidance while leaving the reader’s goal-directedness intact β€” and the negative control confirms the boundary: on the echo task, no-memory scores 1.0 while RAG and frame-selected context both go silent, abstaining (0.0 each) on the grounds that the labels are “not found in the evidence”. Memory that helps eight tasks hurts the ninth by displacing obedience to the present task: with evidence present, the reader treats the memory store as the source of truth and withholds the trivially correct echo.

The strict influence table over 73 paired comparisons reads: 14 full-success contributions, 2 harmful, 0 memory-not-necessary, 57 failed-to-repair. The last number needs the directional table beside it, because the strict table credits only perfect treated scores: across the same 73 pairs, structured actions change on all 73, scores improve on 32, degrade on 8, and stay equal on 33. A 0.00 β†’ 0.75 move counts as failed-to-repair in the strict table and as improvement in the directional one; both are published, labelled for what they are. The honest majority is partial movement, not full success β€” memory moves behaviour on every pair, and moves it upward four times as often as down.

The 0.72 / 0.955 paradox, resolved

The anomaly that motivated this chapter dissolves into two findings, one about the ledger and one about the reader.

First, the ledger overstates necessity. On T1-architecture at A6@768, four of six MUST units are absent yet coverage is 1.0 β€” because the open-loop record mb-loop-ch10-experiment, unlabelled in that task’s ledger and therefore a distractor by default, carries the same actionable content (“no committed run”) as the missing MUST unit, alongside SHOULD-grade support for the remaining claims. The same substitution pattern covers the other high-coverage/low-recall tasks. This chapter confirms it behaviourally: review-arch under the identical 0.333-recall context holds the chapter for exactly the experiment-pending reason, scoring 1.0. The evidence was sufficient for the behaviour; the ledger was conservative about which units count. That is recorded as a ledger finding for the next benchmark version β€” alternative support between the loop record and the experiment-pending unit β€” not as a silent label edit, which the version rules forbid.

Second, the reader sometimes answers from partial evidence. At 768 tokens the composed policy protects disagreement and licences at the cost of required recall (0.72), yet answers better than the full context. That is not a contradiction once influence is measured: the protected structure carries the decision-relevant content even when MUST-counted units fall. Evidence sufficiency and reader sufficiency are different quantities, and the key-claim scorer measures only the second. Chapter 14’s headline stays valid within its bounds; it cannot be read as proof the missing evidence was unnecessary in general.

Two readers, one direction

A second reader (ministral-3:8b, 76 outcomes, no repeats) shows the same direction with different magnitudes on matched tasks: 0.000 β†’ 0.524 (RAG) β†’ 0.452 (selected) β†’ 0.643 (assembled), oracle 0.536. Three reader-dependence findings stand out. First, the smaller reader starts from zero: it abstains without memory almost everywhere, so memory’s measured delta is larger even where its ceiling matches. Second, it is more suggestible: under wrong memory it follows the stale record, ships the unready chapter, and deletes the contracted facade twice over β€” and it deletes the facade even with no memory at all, where the primary reader holds. Memory teaches this reader restraint it does not have (assembled cleanup: facade intact, though migrations incomplete). Third, the ledger oracle is not a behavioural oracle for this reader: assembled context (0.643) beats the oracle (0.536) β€” the oracle strips SHOULD-grade material this reader actually uses. The instrument measures the memory–reader interaction, not memory alone: the architecture supplies better evidence, but readers differ in floors, ceilings and suggestibility. The two readers are reported separately and never averaged.

Real-project transfer

Five time-locked repository questions at the frozen commit, manually adjudicated, scored separately from fixtures: which chapters have canonical runs, the fair-oracle gap in tokens, whether to promote the rejected drop revision, what the frame needs for misread tasks, and the repaired null-denominator rule. Memory-supplied answers average 0.95 against 0.35 without, per task: run-report 0.75/0.25 (one field wrong β€” the answer denies Chapter 14 a canonical run it has), fair-gap 1.0/0.0, promote-decision 1.0/1.0, frame-remedy 1.0/0.5, null-rule 1.0/0.0. The promote-decision task is non-diagnostic: its question states the breach, so rejection needs no memory (both conditions score 1.0). The null-rule task now works as designed: copying the 0.0 demonstration scores 0 without memory, stating explicit nulls scores 1.0 with it. Transfer is consistent with control, at small-n and with the disagreements that exist preserved in the frozen file.

What this chapter earns, and what it does not

Against the outcomes declared before the run, this is a Type A result on fixture evidence: on matched tasks, correct memory causes meaningful beneficial behavioural changes over no memory (0.226 β†’ 0.488 assembled) and over strong RAG (0.393 β†’ 0.488), removal drops success to 0.250 and restoration lifts it to 0.778, and the attribution traces name the decisive items. Full history scoring below no-memory (0.048 against 0.226) is the complementary Type A finding in the other direction: selection is what makes the past usable, and its absence is measurably harmful.

Five demotion clauses apply, and each is load-bearing. The fixtures are synthetic and few (9 controlled, 5 transfer); the primary reader is one small model with a second as transfer check only. One fixture is non-diagnostic (review-prose scores 0.0 under every condition including the oracle β€” the task does not separate conditions and is kept as a published instrument failure, not dropped). One task is behaviourally flat in v2 (credit-rule scores 0.0 under every condition including the oracle β€” without action demonstrations the reader abstains or omits the mechanism name everywhere, so the task measures format compliance rather than memory use; kept, not repaired after the fact). Repeats are deterministic on all three headline tasks (fix-store 3Γ—1.0, ship-ch10 3Γ—0.667, corpus-cleanup 3Γ—0.833). Nothing here claims that memory improves behaviour in general; it claims that on memory-dependent controlled tasks, this reader acts better with structured project memory, and that targeted interventions attribute part of the improvement to specific remembered evidence.

What remains unsolved. The instrument now exists, and the first thing it reveals is how much behaviour it cannot yet explain: 57 of 73 pairs never reach full success under either condition, wrong memory is far worse on average (0.042) while being catastrophically worse in two cases, and the reader that benefits most is also the most suggestible. Chapter 12 establishes that memory can change behaviour beneficially, and that removing specific memories can reverse those gains. It also establishes the danger in the other direction: stale or inappropriate memory can produce destructive actions. Once memory has behavioural power, the next question is no longer merely which memories are relevant. It is when the system should trust its own framing strongly enough to let those memories control behaviour. That is Chapter 13 β€” When the Frame Is Wrong: whether memory can be made safer without being made weaker β€” the abstention path Chapter 10 specified but never built β€” and it now has a way to be measured.

Research foundations

The outside literature converges on the instrument rather than the mechanism. Mem2ActBench (Shen and colleagues, ACL 2026) draws the operative cut between remembering information and using memory to act, with tool tasks reverse-generated so the past is necessary by construction; this chapter’s memory-dependent fixtures follow the same rule in project form. MemoryArena (He and colleagues, 2026, preprint) couples acquisition with later use across interdependent sessions and reports that saturated recall collapses in agentic settings β€” the same direction as the 0.72/0.955 divergence, approached from the opposite side. CUB (HagstrΓΆm and colleagues, ACL 2026) supplies the warning this chapter heeds throughout: context present does not imply context used, and simple synthetic wins inflate. DRUID (HagstrΓΆm and colleagues, ACL 2025) motivates the real-project transfer: synthetic utilisation results overstate, so a small adjudicated set anchors the fixtures. LOCOMO-CONV (Chang and Chen, 2026, preprint) names the discrepancy class this chapter classifies as substitution β€” strong retrieval not translating into response quality, with silent grounding where memory helps without surfacing the gold fact. Agent Workflow Memory (Wang and colleagues, ICML 2025) and MemBench (Tan and colleagues, ACL Findings 2025) are precedent and background respectively: the former that remembered structure can move task success, the latter that effectiveness without a decision metric is not behaviour.

References

  • Yiting Shen and colleagues, Mem2ActBench: A Benchmark for Evaluating Long-Term Memory Utilization in Task-Oriented Autonomous Agents (ACL 2026).
  • Zexue He and colleagues, MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks (2026, preprint).
  • Lovisa HagstrΓΆm and colleagues, CUB: Benchmarking Context Utilisation Techniques for Language Models (ACL 2026).
  • Lovisa HagstrΓΆm and colleagues, A Reality Check on Context Utilisation for Retrieval-Augmented Generation (ACL 2025).
  • Wen-Yu Chang and Yun-Nung Chen, When Users Don’t Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents (2026, preprint).
  • Zora Zhiruo Wang and colleagues, Agent Workflow Memory (ICML 2025).
  • Haoran Tan and colleagues, MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents (ACL Findings 2025).