Decisions Over Time
Do several questions over one state share work, and does that matter?
Decisions Over Time
Design draft: this chapter computes ANALYTICAL token counts under assumed costs. No model and no inference engine runs. Every number is a consequence of the assumed state length, question length, per-call overhead and batch-scaling exponent, which are swept and stated next to the results. None of it is a measurement of Prompt Cache, SGLang, or any engine.
The problem
A decision request in the vendor’s model is one state, several questions, answered in one call. The shared state is the selling point: ask several questions over one state, and the work is shared. Chapter 26 recorded the graph around this. This chapter asks the arithmetic question the claim rests on: do several questions over one state share work, and does that matter?
The tempting answer — “yes, because the state is encoded once” — is exactly what must be decomposed. Most of it may be a property of caching, which any provider can use, and only the remainder is a property of decision models.
What we expect and why
Our starting hypothesis is that prefix caching delivers most of the saving for any provider, and shared-state heads add a smaller further saving at the cost of generality.
Three sources frame it:
-
Gim et al. (2023) — Prompt Cache reuses attention states of shared prompt modules, cutting time-to-first-token 8× on GPU and 60× on CPU inference with no model changes. The mechanism is a cache, not a model. (Read at abstract level.)
-
Zheng et al. (2023) — SGLang’s RadixAttention automatically reuses KV cache across a structured program’s calls, up to 6.4× higher throughput. The reuse is an inference-engine optimisation. (Read at abstract level.)
-
Caruana (1997) — multitask learning trains one shared representation for several tasks (heads). The shared-heads idea behind Chapter 13’s probes is old and is a modelling arrangement, distinct from inference caching. (Verified via Springer DOI 10.1023/A:1007379606734, Machine Learning 28, 41-75.)
The accounting
benchmarks/shared_state/accounting.py counts tokens under four schemes. The assumptions: state length L, question length q, per-call overhead, and (for shared heads) a head_cost.
from benchmarks.shared_state.accounting import account # type: ignore[attr-defined]
L, Q, K, HEAD = 1000, 64, 4, 8
# 1. Token counts for k questions over one state (ANALYTICAL).
print("1. tokens processed for k questions over one state (ANALYTICAL)")
print(f" assumptions: state L={L}, question q={Q}, k={K}, head_cost={HEAD}")
for r in account(K, L, Q, HEAD, rate=1.0, overhead=0.0):
print(f" {r.scheme:<14} {r.tokens:>6} tokens ({r.ratio_vs_independent:.2f}x)")
# 2. The saving reproduced by prefix caching, without any shared-state model.
print("2. what prefix caching already buys (ANALYTICAL)")
a1 = account(K, L, Q, HEAD, 1.0, 0.0)
ratio = {r.scheme: r.ratio_vs_independent for r in a1}
print(f" prefix_cached is {ratio['prefix_cached']:.2f}x; shared_heads is {ratio['shared_heads']:.2f}x")
print(" -> most of the saving is prefix caching; heads add a small increment")
The walkthrough prints:
1. tokens processed for k questions over one state (ANALYTICAL)
assumptions: state L=1000, question q=64, k=4, head_cost=8
independent 4256 tokens (1.00x)
prefix_cached 1256 tokens (0.30x)
shared_heads 1032 tokens (0.24x)
batched 1256 tokens (0.30x)
2. what prefix caching already buys (ANALYTICAL)
prefix_cached is 0.30x; shared_heads is 0.24x
-> most of the saving is prefix caching; heads add a small increment
The sweep
run_ch27.py sweeps k over {1, 4, 16} and L over {1000, 8000, 32000} tokens and writes results/ch27.jsonl (mode: analytic; assumptions in every row). At k=16:
| L | independent | prefix_cached | shared_heads | batched |
|---|---|---|---|---|
| 1,000 | 17,024 (1.00×) | 2,024 (0.12×) | 1,128 (0.07×) | 2,024 (0.12×) |
| 32,000 | 513,024 (1.00×) | 33,024 (0.06×) | 32,128 (0.06×) | 33,024 (0.06×) |
What it says
-
Prefix caching captures most of the saving. At L=1,000 the cached scheme is 0.12× and the shared-heads scheme 0.07×; at L=32,000 both are ≈0.06×. The gap between them shrinks as the state grows, because the state (processed once either way) dominates. Most of the “shared state” saving is a cache any provider can use. P3, the chapter’s central claim, holds under this model.
-
The shared-heads increment is small and assumption-bound. At L=1,000 it buys 0.05× more (0.07 vs 0.12), and only because
head_cost=8was assumed againstq=64. Set head cost equal to the question length and the two schemes coincide. The shared-heads advantage is an architecture position (Caruana multitask; Chapter 13’s probes), not an inference fact. P2 holds only under its own assumption. -
Batching and prefix caching coincide in token count here. They differ in scheduling, not in how many tokens pass through the model; that difference is an engine detail outside this accounting.
-
The trade is generality. Prefix caching needs no model change; shared heads require the multi-headed architecture. The vendor’s “one state, several questions” is on the cache side of the line, and on the modelling side only to the extent heads exist (P1-P4, all ANALYTICAL).
Wrong / Correct. Wrong: “The shared-state advantage proves decision models share compute.” Correct: “Under this analytic model, most of the saving is prefix caching, which any provider can use; the shared-heads increment is a small, architecture-dependent remainder.”
The distinction this chapter keeps
Modelling idea vs inference-engine idea. Shared heads are a modelling arrangement (one representation, several heads — Caruana); prefix caching is an engine optimisation (Prompt Cache, SGLang). The vendor’s claim straddles the line; this chapter’s accounting places most of the saving on the engine side and says the modelling remainder is exactly the part that must be measured, not assumed.
What to carry forward
What part of the Jev interface is a model idea and what is an inference-engine idea? Under this accounting: processing the shared state once is cache (engine); deciding several questions from one representation is model (heads). Chapter 28 builds the benchmark that would measure which holds for real providers — and its cells for this saving are NOT_OBSERVED, because no engine ran.
Close by
What part of the Jev interface is a model idea and what is an inference-engine idea? The chapter’s answer: the split above, ANALYTICAL, pending measurement.
Limitations
- This is ANALYTICAL arithmetic.
results/ch27.jsonlismode: analytic; the latency proxy includes a batch-scaling exponent as an assumption and says nothing about a real engine. - No real provider, cache engine, or GPU ran; the Chapter 13 probe-heads shared-state hypothesis is deferred (GPU).
- Prompt Cache, SGLang and Caruana are cited at abstract/summary level; the full-text reading is scheduled for the Session R reading pass.
- The shared-heads increment depends on the assumed
head_cost; swept values are L=1,000 / 8,000 / 32,000 and k=1 / 4 / 16, stated in the rows.