Compaction
How Pi decides when the context window is full, what it replaces with a summary, what survives untouched, and where an extension can take the summarization over.
Two hours into a refactor, the footer says you are at 87% context. Nothing is wrong. Then it says the agent compacted, and the model asks you to re-explain the schema it was told about an hour ago.
That re-asking is the real behaviour of compaction, not a bug in your prompt. The old messages were not deleted — they are all still in the session file — but they are no longer in the next request. What the model now has instead is a paragraph written by another model about what those messages said.
This chapter is about that mechanism: when it fires, what it replaces, and where you can intervene.
Two mechanisms, one idea
| Mechanism | Trigger | Purpose |
|---|---|---|
| Compaction | Context exceeds threshold, or /compact |
Summarize old messages to free up context |
| Branch summarization | /tree navigation |
Preserve context when switching branches |
“Both use closely related structured formats and track file operations cumulatively. Summarization requests disable prompt-cache writes because these one-off prompts are unlikely to be reused.”
That last clause is worth internalizing. Compaction deliberately breaks your prompt cache, because the request it sends is unlike any request that came before. Cache hits resume right after.
When it fires
The trigger is a subtraction:
contextTokens > contextWindow - reserveTokens
reserveTokens defaults to 16384 and exists “to leave room for the LLM’s response.” The model has to be able to write a long answer even when the conversation is nearly full; without the reserve, the first response that overflows would be the one that pushes past the limit.
The check does not happen at one point. “During a multi-turn agent run, Pi checks the canonical projected context after tools finish and their results are appended, before starting the next assistant response.” Pi “also checks before a new user prompt and performs final-attempt overflow recovery after the low-level run ends.”
It skips the between-turn check in exactly one case: “when the completed tool batch terminates the run and no queued message requires another response.” No next request, no need to free space for it.
There is a third trigger that is not a threshold at all: “A provider context-overflow error or an early final stopReason: "length" can select one compact-and-retry recovery attempt.” One attempt. Not a loop.
What compaction actually does
Five steps, from compaction.md:
- Find cut point — walk backwards through the finalized session projection, accumulating token estimates until
keepRecentTokens(default 20000) is reached. - Extract messages — collect projected messages from the previous kept boundary, or session start, up to the cut point.
- Generate summary — call the LLM with a structured format, passing the previous summary as iterative context when one exists.
- Append entry — save a
CompactionEntrywith the summary andfirstKeptEntryId. - Rebuild context — the session rebuilds context for the next request using summary plus messages from
firstKeptEntryIdonwards.
Step 3 is the one that surprises people. The previous summary is not discarded; it is passed in as context for the new one. Each compaction summarizes what the last one summarized, plus what has happened since. That is why the “re-asking” degrades gradually rather than all at once.
Drawn out, using the entry numbering from the documentation’s own example. Entries 0 to 9 exist before compaction; entry 10 is the compaction entry appended after it:
flowchart TD
subgraph FILE["Session file — every entry stays"]
direction LR
e0["0 hdr"] --> e1["1 usr"] --> e2["2 ass"] --> e3["3 tool"] --> e4["4 usr"] --> e5["5 ass"] --> e6["6 tool"] --> e7["7 tool"] --> e8["8 ass"] --> e9["9 tool"]
end
e9 --> c10["10 compaction"]
subgraph NEXT["Next request — rebuilt context"]
direction LR
sys["system"] --> sm["summary of 0 to 3"] --> k4["4 usr"] --> k5["5 ass"] --> k6["6 tool"]
end
c10 --> sm
e4 --> k4
The first box is the session file, and compaction only ever appends to it. The second box is what the model is given, and it is a different thing: system prompt, the summary, then the entries from firstKeptEntryId — entry 4 in this example — onwards.
Wrong: “compaction deletes the old messages.” Correct: it changes what the next request is built from.
session-format.md: Pi “replaces older summarized entries with the compaction summary and keeps the range beginning at this entry.”compaction.mdadds that “Omitted raw entries remain stored but do not affect cut selection, summaries, checkpoints, or token estimates.” Nothing is deleted from the file, and what the model lost is real anyway.
Cut point rules
Valid cut points are user messages, assistant messages, BashExecution messages, and custom messages. Then the rule that prevents half a tool call from existing:
“Never cut at tool results (they must stay with their tool call).”
A cut between a tool call and its result would leave the model asking about a tool invocation that never returns. The budget of keepRecentTokens can therefore be overshot slightly to land on a legal boundary.
There is a second case called a split user-message span. “A user-message span starts with a user message and includes all turns until the next user message.” Normally compaction cuts at one of those boundaries. But when a single span exceeds keepRecentTokens on its own — one long turn with many tool calls — the cut lands inside it, at an assistant message. The doc marks this isSplitTurn = true, with messagesToSummarize empty and the early part kept as turnPrefixMessages, and Pi “generates two summaries and merges them”: one for previous context, one for the user-message-span prefix.
The entry it writes
interface CompactionEntry<T = unknown> {
type: "compaction";
id: string;
parentId: string | null;
timestamp: string;
summary: string;
firstKeptEntryId: string;
tokensBefore: number;
usage?: Usage; // LLM usage that generated the summary
fromHook?: boolean; // true if provided by extension (legacy field name)
details?: T; // implementation-specific data
}
// Default compaction uses this for details:
interface CompactionDetails {
readFiles: string[];
modifiedFiles: string[];
}
firstKeptEntryId is the whole mechanism. “Pi replaces older summarized entries with the compaction summary and keeps the range beginning at this entry.” Change that field and you change which history the model sees.
tokensBefore is recalculated “from the rebuilt, context-edited session projection before writing the new CompactionEntry, so the token count reflects the actual pre-compaction context being replaced.” Note what it does not do: “Omitted raw entries remain stored but do not affect cut selection, summaries, checkpoints, or token estimates.”
The summary format
Both compaction and branch summaries use structured output. Compaction summaries have sections Goal, Constraints & Preferences, Progress, Key Decisions, Next Steps, and Critical Context. Branch summaries stop after Next Steps. “Pi appends file lists to either format when relevant.”
## Goal
[What the user is trying to accomplish]
## Constraints & Preferences
- [Requirements mentioned by user]
## Progress
### Done
- [x] [Completed tasks]
### In Progress
- [ ] [Current work]
## Key Decisions
- **[Decision]**: [Rationale]
## Next Steps
1. [What should happen next]
## Critical Context
- [Data needed to continue]
Two deliberate choices here. Critical Context exists because the summary’s job is not narrative completeness — it is carrying the facts the next request needs. And file lists are appended because “which files did I read and which did I change” is exactly the state a summarizing model drops first, and exactly what the next turn needs to pick up work.
Before summarizing, messages are serialized to plain text rather than sent as a conversation: [User]:, [Assistant thinking]:, [Assistant]:, [Assistant tool calls]:, [Tool result]:. “This prevents the model from treating it as a conversation to continue.” Tool results are truncated to 2000 characters, “since tool results (especially from read and bash) are typically the largest contributors to context size.”
Configuring it
{
"compaction": {
"enabled": true,
"reserveTokens": 16384,
"keepRecentTokens": 20000
}
}
| Setting | Default | Description |
|---|---|---|
enabled |
true |
Enable auto-compaction |
reserveTokens |
16384 |
Tokens to reserve for LLM response |
keepRecentTokens |
20000 |
Recent tokens to keep (not summarized) |
“Disable auto-compaction with "enabled": false. You can still compact manually with /compact.”
Per-model overrides exist because a 1M-context model and a 200K model should not compact at the same point:
{
"compaction": {
"modelOverrides": {
"some-provider/big-model": {
"reserveTokens": 400000
}
}
}
}
Keys are “exact, case-sensitive provider/modelId values, including any slashes within the model ID.” Each value “falls back independently from the model override to the ordinary setting to the built-in default.” One subtlety: “reserveTokens also influences summarization output limits, capped by the model’s maximum output tokens; it is not solely a trigger threshold.” A 400000 reserve on a model with a 64K output cap does not make summaries longer — the output cap binds first.
Taking it over with an extension
session_before_compact fires before auto-compaction or /compact, and can cancel or supply your own summary:
pi.on("session_before_compact", async (event, ctx) => {
const { preparation, branchEntries, customInstructions, reason, willRetry, signal } = event;
// preparation.messagesToSummarize — messages to summarize
// preparation.previousSummary — previous compaction summary
// preparation.fileOps — extracted file operations
// preparation.firstKeptEntryId — where kept messages start
// reason — "manual" (/compact), "threshold", or "overflow"
// willRetry — whether the aborted turn is retried after compaction
return { cancel: true };
// or
return {
compaction: {
summary: "Your summary...",
firstKeptEntryId: preparation.firstKeptEntryId,
tokensBefore: preparation.tokensBefore,
usage: summaryResponse.usage,
details: { readFiles: preparation.fileOps.readFiles },
}
};
});
If you want to generate the summary with a different model, serializeConversation is the documented path: convert AgentMessage[] with convertToLlm, serialize to text, send it wherever you want. session_compact_failed is the terminal counterpart — “useful for telemetry extensions that need to pair session_before_compact attempts with terminal outcomes” — carrying reason, errorMessage, aborted, willRetry, and fromExtension.
Failure is not a dead end
“Compaction can fail if the provider is unavailable or cannot accept the summarization request. Correct the provider problem and run /compact again.” The failure mode is visible in the event stream as compaction_end with aborted: false and an errorMessage, and separately as summarization_retry_scheduled when a retry is scheduled.
One asymmetry worth knowing: “Pi does not automatically carry file lists from extension-generated summaries whose fromHook field is true; extensions manage their own details format.” If your custom summary drops the file lists, later compactions will not recover them — cumulative tracking only accumulates across Pi-generated summaries.
So the model forgetting your schema after a compaction is not a failure of the mechanism. It is what a paragraph about a schema does. The remedy is on your side — put durable facts in a context file (chapter 28) so they are rebuilt from disk every request rather than carried in history.
The other exit from a full context is not compression at all. When a session takes a wrong turn, /tree, /fork, and /clone are three different answers to the same question.