Learning Path 04

Advanced Agents From First Principles

Engineer autonomous systems for production: evidence, reliability, authority, distributed execution, durable workflows, security, coordination and explicit control planes.

A production-engineering path for moving beyond agent demos into measurable, constrained, recoverable and trustworthy autonomous systems.

Production AI 46 chapters Sequential or reference
Steps 00–11

01 — Reasoning Architectures

Generation, candidates, search, specialists and architecture selection.

  1. 00When Should You Use an Advanced Agent Architecture?
  2. 01Does Your AI Agent Fail on Complex Reasoning Tasks? Treat Chain of Thought as Computation, Not Proof
  3. 02Why Does My Reasoning Agent Give a Different Answer Every Time? Use Self-Consistency Without Confusing Consensus With Truth
  4. 03Does Your Agent Commit to a Bad Reasoning Path Too Early? Build a Tree of Thoughts
  5. 04Does Your Agent Prune Good Ideas Too Early? Use Monte Carlo Tree Search for Long-Horizon Reasoning
  6. 05Is One Model Doing Everything? Build a Mixture of Experts at the Agent Level
  7. 06Does One Agent Plan, Execute and Judge Its Own Work? Build a Planner-Executor-Critic Architecture
  8. 07Do Your Agents Agree Too Easily? Use Adversarial Review and Multi-Agent Debate Without Confusing Debate With Truth
  9. 08Is Your Agent Spending the Same Compute on Every Task? Build Adaptive Agents That Escalate Only When Needed
  10. 09Can Your Agent Actually Learn From Previous Runs?
  11. 10Are You Combining Every Agent Technique Into One Monster? Build a Mixture-of-Agents Runtime
  12. 11Which Advanced Agent Architecture Should You Use? A Practical Selection Guide
Steps 12–18

02 — Evidence & Optimization

Benchmarking, observability, uncertainty, compute allocation and information value.

  1. 12Is Your Advanced Agent Actually Better? Benchmark It Under Equal Budgets
  2. 13How Do You Debug an Agent That Made the Wrong Decision? Add Trajectory Observability
  3. 14Can Your Agent Learn From Its Own Trajectories Without Learning the Wrong Lessons?
  4. 15How Do You Optimize an Agent Policy Without Turning It Into Another Black Box?
  5. 16Where Should an Agent Spend Its Compute? Build a Dynamic Budget Scheduler
  6. 17What Is Your Agent Actually Uncertain About?
  7. 18What Should Your Agent Observe Next? Use Expected Value of Information
Steps 19–22

03 — Distributed Execution

Parallelism, leases, fencing, backpressure and failure containment.

  1. 19Can Your Agent Explore in Parallel Without Creating Chaos? Use Speculative Execution and Early Cancellation
  2. 20Can Your Agent Coordinate Across Machines Without Duplicating Work? Use Leases, Idempotency and Fencing
  3. 21What Happens When Too Many Agents Compete for the Same Resources? Add Admission Control, Quotas and Backpressure
  4. 22What Happens When One Dependency Starts Failing? Add Circuit Breakers, Bulkheads and Graceful Degradation
Steps 23–26

04 — Behavioral Production Engineering

Drift, releases, replay, provenance and incident forensics.

  1. 23Your Infrastructure Is Healthy. Why Is the Agent Getting Worse? Detect Behavioral Drift and Roll Back Safely
  2. 24How Do You Release Agent Behavior Safely? Add Behavioral Contracts, Compatibility Checks and Promotion Gates
  3. 25Can You Reproduce an Agent Run Months Later? Add Deterministic Replay and Provenance
  4. 26Why Did the Agent Fail? Build an Incident Forensics Pipeline
Steps 27–28

05 — Reliability Engineering

SLOs, error budgets and evidence-driven reliability investment.

  1. 27How Reliable Does an Agent Need to Be? Define SLOs and Error Budgets
  2. 28Where Should You Spend the Next Engineering Hour? Prioritize Reliability by Risk and Expected Return
Steps 29–31

06 — Authority & Competence

Escalation, competence envelopes and sandboxed capability acquisition.

  1. 29When Should an Agent Stop and Ask a Human? Design Authority Boundaries and Escalation
  2. 30Is This Task Outside Your Agent’s Competence? Build Competence Envelopes and OOD Detection
  3. 31How Can an Agent Learn New Capabilities Without Expanding Its Own Authority? Use Sandboxed Capability Acquisition
Steps 32–34

07 — Capability Architecture

Capability portfolios, dependency graphs and execution placement.

  1. 32Which Capabilities Are Actually Worth Building? Design a Capability Portfolio
  2. 33Which Shared Components Actually Unlock More Capability? Build a Capability Dependency Graph
  3. 34Where Should This Task Actually Run? Build Capability-Aware Placement Across Models, Providers and Resource Pools
Steps 35–37

08 — Continuity & Temporal Correctness

Portable state, freshness and versioned intent across long-running work.

  1. 35How Do You Move a Running Agent Between Workers Without Losing Meaning? Build Portable Execution State and Safe Handoff
  2. 36Is Your Agent Acting on Stale State? Build Temporal Consistency, Freshness Budgets and Conflict Detection
  3. 37Is Your Agent Still Solving the Right Task? Build Intent Versioning, Supersession and Cancellation
Steps 38–43

09 — Durable Autonomous Systems

Commitments, durable workflows, recovery, trust, coordination and the control plane.

  1. 38A Plan Is Not a Commitment — Model Goals, Commitments and Executable Work
  2. 39How Do You Make an Agent Survive for Days? Build Durable Long-Running Workflows
  3. 40Your Agent Changed the World. What Happens When Step Two Fails? Build Transactions, Compensation and Reconciliation
  4. 41What Should Your Agent Trust? Build Explicit Security and Trust Boundaries
  5. 42How Do Multiple Agents Coordinate Without Becoming a Distributed Argument?
  6. 43Who Controls the Agent? Build an Explicit Agent Control Plane
Steps 44–45

10 — Synthesis

The complete reference architecture—and the case for removing what you do not need.

  1. 44Build a Production AI Agent From First Principles: The Complete Reference Architecture
  2. 45You Probably Don't Need All of This: Build the Minimum Production Agent Architecture