The Session File Format
One JSON object per line, a tree of entries, and the three-pass walk that turns a stored session into exactly the messages a model request contains.
You open a session file to find out what the agent actually saw. There are four hundred lines. Somewhere in them is the system prompt, most of a conversation that has been summarized away, a branch you abandoned last week, and a custom entry from an extension storing its own counter.
Which of those reached the model is not obvious from the file. This chapter is about the file and the rule that decides.
The container
“Sessions are stored as JSONL (JSON Lines) files. Each line is a JSON object with a type field. Session entries form a tree structure via id/parentId fields, enabling in-place branching without creating new files.”
Three properties in one sentence. Append-only line format, typed objects, and a tree carried entirely in the file. No separate index, no rewriting of history when you branch.
Location:
~/.pi/agent/sessions/--<path>--/<timestamp>_<session-id>.jsonl
“For <path>, Pi removes the leading path separator and replaces /, \, and : with -.” So C:\work\api becomes --C-work-api--. sessions.md adds that sessions are “grouped by working directory” and that sessionDir, PI_CODING_AGENT_SESSION_DIR and --session-dir change the root, with the CLI option winning.
Deletion is deliberately boring. “Sessions can be removed by deleting their .jsonl files,” and /resume in the picker does the same thing, preferring the trash CLI “when available.”
Versions
The header carries a version, and there are three:
- Version 1 — linear entry sequence (legacy, auto-migrated on load)
- Version 2 — tree structure with
id/parentIdlinking - Version 3 — renamed
hookMessagerole tocustom
“Existing sessions are automatically migrated to the current version (v3) when loaded.”
Version 3 is the smallest change and the easiest to trip over: if your extension reads entries by the old hookMessage type name, it silently sees nothing on any session created by a current Pi.
The header
{"type":"session","version":3,"id":"uuid","timestamp":"2024-12-03T14:00:00.000Z","cwd":"/path/to/project"}
“First line of the file. Metadata only, not part of the tree (no id/parentId).”
With a parent:
{"type":"session","version":3,"id":"uuid","timestamp":"2024-12-03T14:00:00.000Z","cwd":"/path/to/project","parentSession":"/path/to/original/session.jsonl"}
parentSession appears “For sessions with a parent (created via /fork, /clone, or newSession({ parentSession })).” This is the only durable record of where a forked session came from — the fork’s branch summary does not exist there, because the new session has no record of what it left.
Every other entry extends:
interface SessionEntryBase {
type: string;
id: string; // Usually an 8-char hex ID; may fall back to a full UUID
parentId: string | null; // Parent entry ID (null for a root entry)
timestamp: string; // ISO timestamp
}
Note “Usually an 8-char hex ID; may fall back to a full UUID.” Do not assume a fixed id length.
The entry types
There are eleven. What follows is what each is for, not a restatement of the field list.
message — a conversation message. The message field holds an AgentMessage. System messages among these carry the prompt and tool loadout, and that is the interesting case: “the first request of a session persists one with every prompt section and tool declaration, and later changes persist as system messages that patch sections by name (null removes one) and list toolsAdded/toolsRemoved. Replaying them in order yields the current prompt and tools; there is no separate prompt state entry.”
model_change — “Emitted when the user switches models mid-session. The latest entry is the selected model, which may be a virtual model; assistant messages then name the physical model that answered.”
thinking_level_change — same idea for reasoning level. The example carries "thinkingLevel":"high".
usage — “Records model-attributed usage that is not an assistant message and does not participate in LLM context.” kind is “an arbitrary string identifying the operation; for example, cache warming uses "cache_warm".” Two rules: “Usage entries contribute to session token and cost totals. Pi hides them from the conversation tree.” And: “Consumers should treat unknown kind values as normal usage rather than rejecting them.” That second sentence is a compatibility instruction — do not write a closed enum.
compaction — from chapter 23. Required firstKeptEntryId, plus optional systemMessage, usage, details, fromHook.
context_edit — “Append-only edit of one earlier context-producing entry. It changes only future model context; the target entry and its metadata remain unchanged in raw history, UI, exports, and session accounting.”
{"type":"context_edit","id":"g6h7i8j9","parentId":"f6g7h8i9","timestamp":"2024-12-03T14:11:00.000Z","targetId":"c3d4e5f6","replacement":null}
Four rules here. “Targets may be user, assistant, tool-result, or custom-message entries.” “replacement: null omits the target from model context.” “String replacements for assistant and tool-result entries are normalized to one text block because those roles require content arrays.” And “If several edits target the same entry, the latest edit on the active branch wins.”
branch_summary — from chapter 24.
custom — extension state, “Does NOT participate in LLM context.”
{"type":"custom","id":"h8i9j0k1","parentId":"g7h8i9j0","timestamp":"2024-12-03T14:20:00.000Z","customType":"my-extension","data":{"count":42}}
“Use customType to identify your extension’s entries on reload.” And an aside worth knowing: “Pi stores virtual model router state as custom entries with customType pi.virtual-model-state.” So a namespace you did not choose can occupy your entry types. Prefix yours.
custom_message — extension content that does reach the model. content, display, and “Optional extension-specific metadata (not sent to LLM).”
The custom / custom_message pair is the one to internalize. Two extension entry types with confusingly similar names and opposite visibility:
| In LLM context | Rendered in TUI | |
|---|---|---|
custom |
No | Yes, via registerEntryRenderer |
custom_message |
Yes | display: true shows it, false hides it |
label — “User-defined bookmark/marker on an entry.” targetId plus label. “Set label to undefined to clear a label.”
session_info — the session display name, set via /name, --name / -n, or pi.setSessionName(). “The session name is displayed in the session selector (/resume) instead of the first message when set.”
Three passes to a model request
This is the part that answers your opening question. buildContextEntries() walks from the current leaf to the root and produces the active entry list while honoring compaction:
- Collects all entries on the path.
- If one or more
CompactionEntryvalues are on the path, uses the latest one, includes the compaction entry first, includes non-system entries fromfirstKeptEntryIdup to but not including the compaction entry, and includes entries after the compaction entry. - Preserves non-message entries in the selected range so interactive mode can render them.
Then buildSessionProjection() “applies the latest context_edit for each selected target. It returns the model-visible messages together with their source entries. Omitted targets produce no message; replacements retain the source entry’s role and metadata while changing only content. The raw selected entries are not modified.”
Then buildSessionContext() builds on that projection. It first “extracts current model and thinking level settings from the full path” — note full path, not the compacted range — and then converts selected entries to messages.
flowchart TD
F["session.jsonl tree of entries"] --> A["active path: leaf to root"]
A --> B["buildContextEntries: latest compaction picks the range"]
B --> C["buildSessionProjection: latest context_edit per target"]
C --> D["buildSessionContext: settings from full path, then entries to messages"]
D --> E["messages sent to the model"]
F -.-> X["entries that never arrive"]
C -.-> X
Wrong: “A stored entry is a message the model was sent.”
Correct: the file is the record. Three ordered stages reduce it to the request, and two of those stages have no on-disk representation at all.
One consequence of the ordering: because edits apply after compaction selects the range, an edit can target an entry that compaction already excluded. Such an edit has nothing to act on for that request, but the raw entry stays in the file.
Another: “Pre-compaction system messages are folded into the complete checkpoint rather than replayed from the retained range.” So after a compaction, the prompt comes from the compaction’s systemMessage, not from the system entries you can still see in the file.
Reading a session yourself
import { readFileSync } from "fs";
const lines = readFileSync("session.jsonl", "utf8").trim().split("\n");
for (const line of lines) {
const entry = JSON.parse(line);
switch (entry.type) {
case "session":
console.log(`Session v${entry.version ?? 1}: ${entry.id}`);
break;
case "message":
console.log(`[${entry.id}] ${entry.message.role}: ${JSON.stringify(entry.message.content)}`);
break;
case "compaction":
console.log(`[${entry.id}] Compaction: ${entry.tokensBefore} tokens summarized`);
break;
case "branch_summary":
console.log(`[${entry.id}] Branch from ${entry.fromId}`);
break;
case "usage":
console.log(`[${entry.id}] Usage (${entry.kind}): ${entry.usage.totalTokens} tokens`);
break;
case "custom":
console.log(`[${entry.id}] Custom (${entry.customType}): ${JSON.stringify(entry.data)}`);
break;
case "custom_message":
console.log(`[${entry.id}] Extension message (${entry.customType}): ${entry.content}`);
break;
case "label":
console.log(`[${entry.id}] Label "${entry.label}" on ${entry.targetId}`);
break;
case "model_change":
console.log(`[${entry.id}] Model: ${entry.provider}/${entry.modelId}`);
break;
case "thinking_level_change":
console.log(`[${entry.id}] Thinking: ${entry.thinkingLevel}`);
break;
}
}
This walks the file, not the branch. To see what the model saw, filter to the active branch yourself and apply the compaction and edit rules, or ask over RPC — get_entries gives you everything with a cursor, get_messages gives you the conversation as the model has it.
Practical rules for anything that writes this format
Append, never rewrite. SessionManager owns the tree and the file is its record; the SDK is the supported path, and hand-editing is not.
Use a namespaced customType. Pi already owns pi.virtual-model-state, and your type is the key everything else matches on.
Expect version 1 sessions. A linear legacy file migrates on load, so a parser that assumes a tree can fail on old files and work on new ones.
The file is the record, and it is the only copy. Everything the session told the model can be summarised away or edited out of the next request — but nothing in this chapter is rebuilt from scratch when Pi restarts, which is exactly what the next one is about.