Decisions Cannot Read Everything
What happens to a decision when the state is larger than the model can use well?
Design draft: this chapter argues from the literature and from earlier measured results; its own provider experiment has not been run. What did run is a token count of a real corpus and a toy model whose output is labelled as the consequence of its assumptions.
You are building a decision system, and the evidence it needs lives in a corpus: documents, passages, records. The decision model can read only a finite amount of text. The contract you have been using hands the provider a single state string, so the easiest design is also the most tempting one: put the whole corpus in the state and let the model sort it out.
This chapter asks whether that works, and what to do if it does not. The question is at what state size does a decision degrade, and does selecting what to read restore it? The first half is a question about models and the literature answers it better than we can. The second half is a question about arithmetic and about design, and part of it can be settled on this machine without running a model.
What the long-context literature found
Three papers set the expectation. All three are about language models reading long inputs, not about decision providers, so they bear on this chapter as a warning and not as a measurement of anything in this book.
Position matters
DOCUMENTED: Liu and colleagues study multi-document question answering and key-value retrieval. They find that performance can degrade significantly when the position of the relevant information changes: it is often highest when the information sits at the beginning or end of the input, and degrades significantly when models must use information in the middle of a long context, even for models built for long contexts. [Liu et al., arXiv:2307.03172, Abstract]
The authors also report the size of the effect for one model: GPT-3.5-Turbo’s multi-document performance can drop by more than 20% when the relevant document moves to the middle. That figure comes from the paper’s body and not its abstract.
The claimed window is not the usable window
DOCUMENTED: Hsieh and colleagues build RULER, a benchmark that goes beyond the needle-in-a-haystack test with more kinds of needles, multi-hop tracing and aggregation tasks. Despite near-perfect accuracy on the vanilla needle test, almost all of the 17 models they evaluate show large drops as context length grows. All claim context sizes of 32K tokens or more, and only half can maintain satisfactory performance at 32K. [Hsieh et al., arXiv:2404.06654, Abstract]
Length alone hurts
DOCUMENTED: Levy and colleagues isolate the effect of input length by extending the same samples with padding of different lengths, types and locations. They find a notable degradation in reasoning performance at input lengths much shorter than the models’ technical maximum, in every version of their dataset though at different intensities, and that the usual next-word-prediction metric correlates negatively with performance on their reasoning task. [Levy et al., arXiv:2402.14848, Abstract]
What these do not tell us
None of the three measures a decision provider, a calibrated probability, or a classifier. They say that stuffing a context is not free for the models they tested. They do not say how much it costs a provider in this book, and they say nothing about calibration, which this chapter’s second prediction concerns. Both are open until a real provider is run.
Wrong: “Our simulation reproduces the lost-in-the-middle curve, so the effect is confirmed.”
Correct: “A simulation returns its assumptions. A U-shaped curve out of a model that was given a U-shaped bias is not a finding. The curve is evidence only when a real model produces it.”
flowchart TD
A[corpus too large for the window] --> B{read everything?}
B -- yes --> C[cost grows with the corpus; use degrades with length and position]
B -- no --> D[select what to read]
D --> E[short state]
E --> F[decision]
What you can count
Whether a corpus fits is a measurement, and it needs no model. The evidence in the Part VI experiments is the SciFact corpus: 5,183 scientific abstracts, with claims that each need a few of them. Everything below is examples/ch20-decisions-cannot-read-everything/walkthrough_ch20.py, which you can run as it stands. It uses two cached tokenizers and does arithmetic on the counts. No language model runs.
The setup
import json
import statistics
from pathlib import Path
import sys
ROOT = Path(__file__).resolve().parents[2]
sys.path[:0] = [str(ROOT), str(ROOT / "src")]
from transformers import AutoTokenizer
from benchmarks.long_state.long_state import PositionBias, accuracy_vs_position
DATA = ROOT / "datasets" / "scifact" / "data"
TOKENIZERS = {
"bge-small (WordPiece)": ("BAAI/bge-small-en-v1.5", "5c38ec7c405ec4b44b94cc5a9bb96e735b38267a"),
"Qwen3 (byte-level BPE)": ("Qwen/Qwen3-1.7B", "70d244cc86ccca08cf5af4e1e306ecf908b1ad5e"),
}
WINDOWS = (4_096, 32_768, 131_072)
def load_jsonl(path: Path) -> list[dict]:
return [json.loads(line) for line in path.open(encoding="utf-8")]
def document_text(doc: dict) -> str:
return f"{doc['title']}. {' '.join(doc.get('abstract', []))}".strip()
def retrievable_claims(path: Path) -> int:
"""Claims with at least one supporting or contradicting evidence document."""
count = 0
for claim in load_jsonl(path):
evidence = claim.get("evidence") or {}
if any(e.get("label") in ("SUPPORT", "CONTRADICT") for evs in evidence.values() for e in evs):
count += 1
return count
def main() -> None:
texts = [document_text(d) for d in load_jsonl(DATA / "corpus.jsonl")]
claims = retrievable_claims(DATA / "claims_dev.jsonl")
1. How big is the evidence?
Two unrelated tokenizers agree to within about two percent, which is the useful thing to know: the answer does not depend on which tokenizer you happen to use.
# 1. How big is the evidence, in tokens?
print(f"1. the SciFact corpus: {len(texts)} abstracts, counted with two tokenizers")
totals = {}
for name, (repo, revision) in TOKENIZERS.items():
tokenizer = AutoTokenizer.from_pretrained(repo, revision=revision)
lengths = [len(ids) for ids in tokenizer(texts, add_special_tokens=False)["input_ids"]]
totals[name] = (sum(lengths), statistics.mean(lengths))
print(f" {name:<24} total {sum(lengths):>9,} mean {statistics.mean(lengths):>6.1f} "
f"median {statistics.median(lengths):>5.0f} longest {max(lengths):>5}")
2. How much fits in a window?
# 2. How much of it fits in a window?
print("2. the share of the corpus one context window can hold")
name = "bge-small (WordPiece)"
total, mean = totals[name]
for window in WINDOWS:
print(f" window {window:>7,} tokens holds {window / total:>6.1%} of the corpus, about {window / mean:>5.0f} abstracts")
3. What does reading everything cost?
There are 188 labelled dev claims. Reading the whole corpus for each claim, against reading a ten-abstract candidate list, is arithmetic on the counts above.
# 3. What reading everything costs, against reading a short candidate list.
print(f"3. tokens to read for the {claims} labelled dev claims (arithmetic on the counts above)")
everything = claims * total
top_ten = claims * 10 * mean
print(f" reading the whole corpus for each claim: {everything:>13,.0f}")
print(f" reading ten candidate abstracts instead: {top_ten:>13,.0f}")
print(f" ratio {everything / top_ten:.0f} to 1; the corpus holds {len(texts)} abstracts and the list holds 10")
4. A toy, and what it can and cannot say
The toy picks the one relevant passage among 11 by an attention score made of relevance, an assumed U-shaped position bias, and noise. The bias is a formula, 4(x − 0.5)², put into the code. The simulation therefore returns a U-shaped curve for the same reason a calculator returns 4 for 2 + 2.
# 4. A toy: what an ASSUMED position bias implies, and what sets how strong it looks.
print("4. a toy attention model with an assumed U-shaped position bias (ILLUSTRATIVE)")
print(" accuracy of picking the one relevant passage among 11, by where it sits")
print(" noise position 0 position 5 position 10")
for noise in (0.0, 0.05, 0.1, 0.5):
acc = accuracy_vs_position(11, PositionBias.U_SHAPED, noise, 200, seed=0)
print(f" {noise:<5} {acc[0]:>9.3f} {acc[5]:>9.3f} {acc[10]:>10.3f}")
flat = accuracy_vs_position(11, PositionBias.UNIFORM, 0.1, 200, seed=0)
print(f" no bias at all (noise 0.1): position 0 {flat[0]:.3f}, position 5 {flat[5]:.3f}, position 10 {flat[10]:.3f}")
print(" the shape comes from the assumed bias; its depth comes from the assumed noise")
What it prints
1. the SciFact corpus: 5183 abstracts, counted with two tokenizers
bge-small (WordPiece) total 1,742,496 mean 336.2 median 315 longest 1938
Qwen3 (byte-level BPE) total 1,707,618 mean 329.5 median 303 longest 1928
2. the share of the corpus one context window can hold
window 4,096 tokens holds 0.2% of the corpus, about 12 abstracts
window 32,768 tokens holds 1.9% of the corpus, about 97 abstracts
window 131,072 tokens holds 7.5% of the corpus, about 390 abstracts
3. tokens to read for the 188 labelled dev claims (arithmetic on the counts above)
reading the whole corpus for each claim: 327,589,248
reading ten candidate abstracts instead: 632,046
ratio 518 to 1; the corpus holds 5183 abstracts and the list holds 10
4. a toy attention model with an assumed U-shaped position bias (ILLUSTRATIVE)
accuracy of picking the one relevant passage among 11, by where it sits
noise position 0 position 5 position 10
0.0 1.000 0.000 1.000
0.05 1.000 0.085 1.000
0.1 1.000 0.085 1.000
0.5 0.640 0.085 0.660
no bias at all (noise 0.1): position 0 1.000, position 5 1.000, position 10 1.000
the shape comes from the assumed bias; its depth comes from the assumed noise
Read it in order.
- Section 1: the corpus is 1,742,496 tokens by the bge tokenizer and 1,707,618 by the Qwen3 tokenizer. An abstract averages about 336 tokens and the longest is about 1,900.
- Section 2: even a 131,072-token window holds 7.5% of the corpus, about 390 abstracts. A 4,096-token window holds twelve.
- Section 3: reading the whole corpus for each of 188 claims is about 328 million tokens. Reading ten candidates each is about 0.6 million, a ratio of 518 to 1. The corpus is 5,183 abstracts and the list is 10, so the ratio is just the ratio of the two counts. These are token counts and say nothing about money or time; they set the scale of the problem.
- Section 4: accuracy at the edges is 1.000 and at the middle position it is near the 1-in-11 chance level (0.091), for any noise from 0.05 up. At noise 0.0 the middle scores 0.000, but that is the code’s tie rule: every passage in the middle of the toy scores exactly zero, and ties go to the first passage. At noise 0.5 even the edges fall to about 0.65. With no bias at all, every position is found. So the shape is the assumed bias, and the depth is the assumed noise.
The right use of the toy is as a statement of what follows if a position bias of that shape exists, and a reminder that the size of any such effect in a real model has to be measured. It supports no claim that such a bias exists, and it does not show that failure is “selection, not capacity”.
The argument that does not depend on the toy
PROPOSED: The case for selecting what to read rests on the count, not the toy. A decision about one claim needs a handful of abstracts; the corpus holds more than five thousand. Even a large window holds under a tenth of it. And the literature above says that using the window you do have is not free. Together these say that something must choose what the provider reads, and that something is a design decision with its own error rate. Chapter 21 measures that error rate.
The word “retrieval” is deliberately absent from this argument. The need to select is a consequence of the numbers. What does the selecting (BM25, dense embeddings, a human, a rule) is a separate question.
The experiment, designed but not run
JEV-20-01 is preregistered with status: NOT_RUN; the full block is in metadata/20-chapter.yaml. It needs a real provider that reads long states, which is a language-model run and is deferred.
Predictions (HYPOTHESIS, about a real provider):
- P1. Decision accuracy falls as state size grows, and depends on where the evidence sits.
- P2. Calibration degrades before accuracy does: ECE rises while accuracy is still high.
- P3. Showing only the relevant passages (oracle selection) recovers accuracy to within a stated margin of the short-state accuracy.
- P4. The failure is in finding the evidence and not in using it: the provider decides correctly when the evidence is presented in a short state.
Refutation: P1 fails if accuracy is flat across state sizes. P2 fails if calibration improves with size. P3 fails if oracle selection does not recover accuracy. P4 fails if the provider cannot use the evidence even when it is presented alone.
PENDING_RUN: result for JEV-20-01, P1–P4 Needs: a provider that reads long states (language-model run, deferred) and the Chapter 9 calibration measures. Stage plan: S1 build states of increasing size around a planted relevant passage at controlled positions; S2 run the provider on each state; S3 compute accuracy and ECE by size and position with paired intervals; S4 run the oracle and truncated arms; S5 write rows.
What would change your mind
If a real provider’s accuracy is flat in state size and position, the warning in the literature does not apply to it, and the argument for selection weakens to the cost argument alone, which the count still supports. If oracle selection does not help, the problem is elsewhere. Neither outcome would affect the token counts above, which are facts about the corpus.
No pivot is proposed.
Limitations
The provider experiment is NOT_RUN, and results/ch20.jsonl does not exist. The only measurements here are token counts for one corpus with two tokenizers, and arithmetic on them. The counts are not a cost in money or time, and a different corpus would give a different scale.
The toy simulation is an illustration of its own assumptions. It is not evidence about any model, and a previous draft of this chapter read it as confirming the literature; that reading is withdrawn. The three required papers are cited at the level of their abstracts above; the one body figure (the 20% drop) comes from the session’s full-text notes and should be re-checked against the paper before quotation. The simulation does not model duplicate padding, calibration, or any provider.
What the next chapter inherits
Chapter 21 inherits the need: something must choose what the decision reads. It also inherits the numbers that make the choice concrete (a 5,183-abstract corpus, a ten-abstract candidate list) and the instruction that retrieval error and decision error must be counted separately. It then measures both on real data, on a CPU.