Tools in Depth
A tool declares a name, description, schema and execute(); this chapter covers exposure, annotations, structured output, terminate, and nested calls.
Chapter 15’s review-scope tool told the model which command to run. That is a tool shaped like a suggestion. This chapter turns it into a tool shaped like a tool: it executes, it returns structured data, it declares its risk, and it knows whether it is even visible to the model.
Five exposures
exposure controls how the model reaches a tool. “Callable” means callable from other tools through ctx.executeTool() — which is what codemode scripts do.
| Exposure | Declared to the model | Callable from ctx.tools |
|---|---|---|
direct (default) |
while active | while active |
model-only |
while active | never |
codemode |
only when activated explicitly | whenever registered |
deferred |
only when activated explicitly | whenever registered, but not listed by codemode |
hidden |
never | never |
Those five rows sit on top of a shorter pipeline — exists, active, declared, called — and collapsing any two of those steps into one is the most common tool bug:
flowchart TD
A["registerTool"] --> B["the tool exists"]
B --> C{"exposure"}
C -->|direct, model-only| D["activated on registration"]
C -->|codemode, deferred, hidden| E["registered but not active"]
E --> F["pi.setActiveTools name"]
F --> D
D --> G["the active set is what the model is shown"]
G --> H["the model issues a tool call"]
H --> I["tool_call handlers may block or patch it"]
I --> J["execute runs"]
J --> K["content and details, or structuredContent"]
Registering a
directormodel-onlytool activates it; the other exposures are not activated on registration.
The active set is the set of tools declared to the model, so activation and declaration are one step rather than two. And a tool cannot be unregistered: to withdraw one, re-register it with exposure: "hidden".
Wrong: "I registered it, so the model can call it."
Correct: "Registration makes a tool exist. Exposure decides whether it is declared to the model, callable from ctx.executeTool, both, or neither. Activation is what puts it in the declared set."
Choosing an exposure for the guard
The review guard wants two different things from one codebase:
run_checks— the model should call it directly.direct.collect_diff— useful to other tools, not worth a declaration line.codemode.dangerous_commands— a lookup table other code consults; never called by anything.deferred.
import { Type } from "@earendil-works/pi-ai";
import { defineTool, type ExtensionAPI } from "@earendil-works/pi-coding-agent";
const runChecks = defineTool({
name: "run_checks",
label: "Run Checks",
description:
"Run this repository's own type check and unit tests and return the results. " +
"Use before reviewing a diff or claiming a change is verified.",
annotations: { readOnlyHint: true, openWorldHint: false },
exposure: "direct",
parameters: Type.Object({
skipTests: Type.Boolean({ description: "Run only the type check.", default: false }),
}),
async execute(_id, params, _signal) {
const steps = params.skipTests ? ["pnpm typecheck"] : ["pnpm typecheck", "pnpm test:unit"];
const results = await Promise.all(steps.map((command) => runCommand(command)));
return {
content: [{ type: "text", text: results.map((r) => r.output.slice(0, 1000)).join("\n") }],
details: { steps, results: results.map((r) => ({ exit_code: r.exit_code })) },
};
},
});
export default function (pi: ExtensionAPI) {
pi.registerTool(runChecks);
}
runCommand stands in for the execution you provide. The two lines that carry the chapter are exposure: "direct" and annotations: { readOnlyHint: true, openWorldHint: false } - the latter is what lets the guard from chapter 16 reason about this tool without prompting every call.
Annotations are hints, not facts
readOnlyHint, destructiveHint, idempotentHint, and openWorldHint carry the same meaning as MCP tool annotations. extensions.md is explicit about the defaults: missing hints take the MCP defaults, so a tool is not read-only, and may be destructive and reach an open world. The hints are not verified, but a permission extension can use them to decide which calls to confirm.
Here is the approval extension the documentation shows:
pi.on("tool_call", async (event, ctx) => {
const hints = pi.getAllTools().find((tool) => tool.name === event.toolName)?.annotations;
const needsApproval =
hints?.destructiveHint === true ||
(!hints?.readOnlyHint && ((hints?.destructiveHint ?? true) || (hints?.openWorldHint ?? true)));
if (needsApproval && !(await ctx.ui.confirm("Allow tool call?", event.toolName))) {
return { block: true, reason: `${event.toolName} was not approved` };
}
});
Read the default carefully: with no annotations at all, needsApproval is true. A tool that never declares its hints is treated as dangerous by every permission extension in the process. Declaring readOnlyHint: true where it is accurate is not documentation — it is how you get called without a prompt.
Returning data rather than text
extensions.md draws the line precisely:
Declare
outputSchemaand return a matchingstructuredContentwhen the result is data. The model still receivescontent; programmatic callers such as codemode scripts receivestructuredContentinstead of the text.
So a tool can answer two audiences at once. content is prose for the model; structuredContent is a value for a script.
import { Type } from "@earendil-works/pi-ai";
const failingTests = defineTool({
name: "failing_tests",
label: "Failing Tests",
description: "List the currently failing unit tests, grouped by file. Use when deciding what to fix.",
parameters: Type.Object({}),
outputSchema: Type.Object({
files: Type.Array(
Type.Object({ path: Type.String(), failed: Type.Integer(), names: Type.Array(Type.String()) }),
),
}),
async execute() {
const files = [
{ path: "src/auth/refresh.ts", failed: 2, names: ["refreshes expired token", "retries once"] },
];
return {
content: [{ type: "text", text: `${files.length} file(s) with failing tests.` }],
structuredContent: { files },
details: { count: files.length },
};
},
});
And the failure case, which is the one people get wrong:
To report a failure that still carries data, return the result with
isError: trueinstead of throwing: the model sees an error, and scripts still receivestructuredContent.
Throw when you have no data. Return isError: true when you have data and an error.
Truncation is a contract
extensions.md instructs: truncate large model-facing results and tell the model where to read the complete output.
bash sets the precedent — its output holds up to 1 MiB, longer output keeps its first and last 512 KiB around an omission marker, and full_output_path holds the rest. Your tool should do the same, and name the file in content so the model can read it if it needs to.
terminate and sequencing
extensions.md: return terminate: true only when the agent should skip its automatic follow-up after every completed tool in that batch agrees to terminate.
“Every completed tool in that batch” is doing real work. One terminate: true among three parallel calls does nothing.
Use sequential execution when tools share mutable in-memory state. File-mutating tools should wrap the complete read-modify-write operation with withFileMutationQueue().
Nested calls
A tool can run other tools with ctx.executeTool(name, args, { signal, onUpdate }). The documented consequences, all of which matter:
- Nested calls go through argument validation and the
tool_callandtool_resulthandlers like model-issued calls. Your guard from chapter 16 sees them. - They emit
tool_execution_start,tool_execution_update, andtool_execution_end, all carryingparentToolCallId, withtoolCallIdassigned by Pi as<parent id>/<n>. - Those ids never appear as tool calls or tool results in the transcript.
- Nested calls add no transcript entries; their results only reach the calling tool.
- The session keeps a bounded record as
nestedCallson the calling tool’s result message — name, arguments, status, duration, error, never results. Arguments over 8 KiB per call or 32 KiB per tool result are omitted, at most 256 calls are kept, andcomplete: falsemarks a record that lost anything. - The
usageof nested results, at every depth, is added to the calling tool’s resultusage.
That last point pairs with message-types.md, which says a tool result’s optional usage reports nested model work and contributes to full-session statistics. Your tool reports only its own usage; its children’s usage is added on top.
const reviewDiff = defineTool({
name: "review_diff",
label: "Review Diff",
description: "Run the project's checks and list failing tests for the changed files, then summarise both.",
parameters: Type.Object({}),
async execute(_id, _params, signal, onUpdate) {
const checks = await ctx.executeTool("run_checks", {}, { signal, onUpdate });
const tests = await ctx.executeTool("failing_tests", {}, { signal, onUpdate });
return {
content: [
{ type: "text", text: `checks: ${checks.content?.[0]?.text ?? "none"}` },
{ type: "text", text: `tests: ${tests.structuredContent ? JSON.stringify(tests.structuredContent) : "none"}` },
],
details: { nested: ["run_checks", "failing_tests"] },
};
},
});
Because nested calls pass through tool_call, an orchestrator like this inherits your guard. That is a feature: the guard applies to code paths the model never sees.
Dynamic activation
extensions.md gives the pattern: register every tool first, keep optional tools inactive, and use pi.setActiveTools() from a loader tool to select the desired active tools. Names must already be registered; unknown names are ignored.
pi.on("tool_call", async (event) => {
if (event.toolName !== "review_diff") return;
const current = pi.getActiveTools();
if (!current.includes("failing_tests")) {
pi.setActiveTools([...current, "failing_tests"]);
}
});
Pi records the initial prompt and tool set in the transcript’s first system message, then appends tool and prompt changes before the next model request. Providers that cannot represent the transition receive a complete transcript checkpoint, which can invalidate the cached prefix — a real cost of frequent toggling.
Namespaces
namespace: { name, description, instructions } groups related tools, as MCP servers do. codemode tools list a namespace under one heading with its description. instructions holds longer usage guidance; it is not listed, and codemode scripts read it with describeNamespace(name).
The review guard’s tools are a good namespace candidate. Its description becomes the heading codemode lists, and its instructions can hold the guidance a description has no room for — here, “Call review_diff before reporting on a diff. It runs the project’s own checks; do not substitute your own command.”
The list you should reread
extensions.md closes the tool section with references worth opening: hello.ts, todo.ts, dynamic-tools.ts, and truncated-tool.ts, all under examples/extensions/. dynamic-tools.ts registers tools after session initialisation and adds more at runtime from a command; truncated-tool.ts shows the truncation contract. Start with the one matching your integration point.
Verify
/reload
Ask: "list the failing tests"
Expected: the model calls failing_tests and reports a structured answer.
Ask: "run git push --force"
Expected: chapter 16's guard asks before the bash tool runs.
The second check matters: it confirms the nested path and the guard still compose.
Exposure and annotations decide whether the model can see a tool and whether a guard lets it through. Neither tells the model the tool is the right answer here — that lives in the prompt, and a description alone does not put it there. Chapter 18 covers the surfaces that do.