Agent Architectures - Complete Deep Dive
Prerequisites: Tool Calling, Structured Outputs, Evals Used in: Agent Memory and State, Model Context Protocol, Durable Execution Build it: Lesson 4 - The Agent Loop From Scratch implements this as runnable, tested code you can execute offline.
What Is an Agent Architecture?
Start with the arithmetic, because it governs every other decision on this page.
Success rates multiply. If one step succeeds 95 percent of the time and the steps are roughly independent, five steps in a row succeed about 77 percent of the time. Ten steps land near 60 percent. Twenty steps are a coin flip you lose more often than you win. None of that is an LLM fact β it is what happens to any chain of fallible operations β but it is precisely why agent demos look magical and agent products feel broken. The demo runs the three-step happy path. Production runs the eleven-step path with a malformed API response in the middle.
An agent architecture is therefore not βhow do I let the model be autonomous.β It is how do I arrange model calls, tool calls, and ordinary code so that the number of fallible steps stays small, every hop is validated, and no loop can run away. Fewer steps, validation at each hop, and bounded loops are not polish you add before launch. They are the architecture.
Real-world analogy: a relay race. Individual runners are fast and the race is won or lost at the baton handoffs, so coaching obsesses over handoff mechanics rather than top speed. An agentβs handoffs are the moments where model output becomes a tool call and a tool result becomes model input. That is where agents drop the baton, and that is where the engineering goes.
The Reliability Arithmetic
End-to-end success for a chain of independent steps, each at the stated per-step rate:
| Per-step success | 3 steps | 5 steps | 10 steps | 20 steps |
|---|---|---|---|---|
| 99% | 97.0% | 95.1% | 90.4% | 81.8% |
| 95% | 85.7% | 77.4% | 59.9% | 35.8% |
| 90% | 72.9% | 59.0% | 34.9% | 12.2% |
| 80% | 51.2% | 32.8% | 10.7% | 1.2% |
Read the 95 percent row twice. A step quality most teams would call good β one failure in twenty β produces a system that fails roughly a quarter of the time over five steps. The only ways out are the three the table itself suggests: raise per-step reliability, cut the number of steps, or add recovery so a failed step does not end the run.
Two honest caveats about the independence assumption, because it cuts both ways.
It is optimistic when failures correlate. A bad first step poisons everything downstream β a wrong plan, a misread document, a tool called with a wrong entity ID β and the later steps then fail at a much higher rate than their standalone numbers. Cascades are the normal failure shape of long agent runs, not the exception.
It is pessimistic when the loop can observe and recover. A ReAct-style loop that sees a tool error and retries with corrected arguments has turned one fatal step into a step with a second chance, which is exactly the point of the loop.
π‘ ReAct is just a loop - the model says what it wants done next, your code does it, the result goes back in, and round it goes until the model says it is finished. That is real, and it is the strongest argument for loops over straight-line chains β but it only holds when the observation genuinely carries the information needed to correct, and when the retry is bounded. See Retries and Backoff for the mechanics you should reuse here rather than reinvent.
The practical consequence: validate at every hop so a failure is caught at the step that caused it, not five steps later when a confident summary arrives built on a silent error.
The Progression
Four architectures, in the order you should try them. Each one buys something specific and costs something specific.
1. A single call with tools
One model call, a set of tool definitions, the model picks zero or more tools, your code executes them, results go back, the model answers. One or two round trips, bounded by construction.
π‘ A tool definition is the written description plus the argument list you hand the model, so it knows a function exists and what to pass it.
Start here, and stay here longer than feels exciting. A large share of problems labelled βwe need an agentβ are this: answer a question using a lookup, file a ticket from a description, classify and route an inbound message, pull three fields from a document and write them somewhere. Per the table above, a two-step path at 95 percent is a 90 percent system; the same task expanded into a nine-step agent loop is a coin flip. The architecture that does the job in the fewest fallible steps wins, and this is that architecture. Mechanics live in Tool Calling.
2. ReAct-style reason - act - observe loops
The model alternates between reasoning about what to do next, taking an action, and observing the result, repeating until it decides it is done. The loop is what makes multi-step work possible at all, and what it actually buys is narrow and worth naming precisely:
- Adaptivity. The next action depends on what the last one returned, so the path is not fixed at design time. Necessary when step three genuinely cannot be known until step two answers.
- Error recovery. A tool error is an observation, so the model gets a chance to correct rather than the run dying.
- Unbounded-shape exploration. Search-like tasks where the number of lookups is not known in advance.
What it costs: every iteration re-sends the accumulated history, the number of model calls is data-dependent rather than fixed, and the model is now deciding control flow β which is the least testable place to put a decision.
flowchart LR
U[User goal] --> C[Controller - checks step and time and cost budget]
C --> R[Model reasons and proposes next action]
R --> V[Validate action name and arguments against allowlist]
V --> T[Execute tool]
T --> O[Append observation to history]
O --> C
R --> F[Model emits final answer]
C --> B[Budget exhausted - return partial result and state why]
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class U client
class C edge
class R,V service
class T async
class O data
class F,B client
The controller node is the part people skip when they write their first loop, and it is the only thing standing between you and an agent that spends an afternoon calling the same search twice.
3. Planner - executor separation
Produce a plan once, then execute it. The planning model call emits an ordered list of steps with expected inputs and outputs; a separate executor β ideally plain deterministic code β walks that list, calling tools and feeding results forward.
The separation buys four things a free-running loop cannot:
- The plan is an inspectable artifact. You can log it, diff it, show it to a user for approval, and eval it directly. βWas the plan rightβ and βwas the execution rightβ become two separate questions with two separate scorecards.
- The plan is validatable before anything happens. Check it against a schema, an allowlist of permitted steps, a step cap, and a policy check for destructive operations β all before the first side effect.
- The plan is cacheable. Many user goals are structurally identical with different parameters. Cache on the goal shape rather than the literal text and you skip the planning call entirely, which is often the most expensive call in the run. See Caching.
- Execution is deterministic and cheap. If the executor is code, its steps do not carry model non-determinism at all, so the per-step reliability of most of the chain jumps to ordinary software levels.
The cost: a plan fixed up front cannot adapt to a surprise. Real systems therefore allow bounded replanning β if a step fails or returns something the plan did not anticipate, return to the planner once or twice with the failure attached, and count replans against a budget.
flowchart LR
G[User goal] --> P[Planner model call - one shot]
P --> S[Plan - ordered steps with expected outputs]
S --> V[Plan validator - schema and allowlist and step cap]
V --> E[Executor - deterministic loop over steps]
E --> T[Tools and APIs]
T --> E
E --> A[Assemble final answer]
V --> X[Reject plan - replan once within budget]
X --> P
S --> K[Plan cache keyed on goal shape]
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class G client
class P service
class S,K data
class V,E service
class T async
class A client
class X edge
4. Multi-agent with specialised roles and a supervisor
Several agents, each with a narrow role and its own tool set, coordinated by a supervisor that routes work and assembles results. A researcher, a writer, a critic. A triage agent and three domain specialists.
Be skeptical of this, and be skeptical early. It is the most-reached-for and least-justified pattern in the space, for a reason that is entirely structural rather than aesthetic:
- It multiplies cost and latency. Every sub-agent runs its own loop with its own history, so you are paying for several agentsβ worth of context regrowth, usually serially.
π‘ Context regrowth means every turn re-sends everything that came before it, so the bill for step ten includes steps one through nine all over again. - Communication between agents is lossy. The only channel between them is natural language. Agent A compresses its findings into prose, Agent B re-interprets that prose, and information degrades at every handoff β a new class of failure that has no analogue in a single agent and is miserable to debug because nothing errored.
- It multiplies the failure surface. Go back to the arithmetic. Each agent is itself a multi-step chain, so a supervisor coordinating four agents is a chain of chains. Failures also become non-local: the writer produced a bad paragraph because the researcher quietly returned a thin result three handoffs ago.
- It usually substitutes for work not done. Most teams reach for multi-agent before exhausting a single well-built agent with good tool descriptions, validated arguments, and a tight loop. Splitting an agent that fails for prompt and tool reasons produces several agents that fail for the same reasons, plus coordination bugs.
Reach for it when concerns are genuinely separable with a clear interface β different tool sets with no overlap, different permission scopes that should not share a context, genuinely parallelisable independent subtasks, or a critic whose value depends on not having seen the writerβs reasoning. Those cases are real. They are rarer than the enthusiasm suggests. When you do build it, treat inter-agent messages as an API contract with a schema, not as chat, and read Saga Pattern for how multi-participant workflows handle partial failure and compensation.
Choosing an Architecture
| Architecture | When it fits | Cost and failure profile |
|---|---|---|
| Single call with tools | The task needs one or two lookups and the shape is known in advance | Cheapest and most predictable. Fails by choosing the wrong tool or wrong arguments - caught by validation |
| Deterministic workflow with model calls at decision points | The steps are known but individual judgements need a model | Near-software reliability. Fails only at the judgement steps, each independently testable |
| ReAct loop | The next step genuinely depends on the last result and the path length is unknown | Data-dependent cost, context regrows every iteration. Fails by looping, drifting from the goal, or spending the budget |
| Planner-executor | Multi-step work with a mostly knowable shape and a need for approval or audit | One expensive planning call plus cheap execution. Fails by producing a bad plan - visible before execution |
| Multi-agent with a supervisor | Separable concerns, disjoint tool sets or permission scopes, real parallelism | Highest cost and latency. Fails by lossy handoffs and non-local errors that never raise an exception |
Who Decides the Next Step
This is the decision that determines whether your system is debuggable, and it deserves to be made explicitly rather than inherited from whichever framework you installed.
Model-driven control flow: the model picks the next action each turn. Maximum flexibility, minimum testability. The execution path differs run to run, so a failure may not reproduce, and there is no code path to set a breakpoint on.
Code-driven control flow: your code decides what happens next, calling the model for specific judgements β classify this, extract that, is this good enough, which of these three branches. The path is a code path. It has a diff, a test, and a stack trace.
Push as much control flow as possible into ordinary deterministic code, and use the model only for the judgement steps. A workflow with model calls at its decision points is cheaper, faster, more reliable, and dramatically more debuggable than a free-running agent, and it is what most shipped βAI agentsβ actually are once you read their source. The reason is the arithmetic again: deterministic steps do not contribute model-level failure rates to the product, so moving a step from model-driven to code-driven does not just simplify it, it removes it from the risk chain.
The honest test for whether you need model-driven flow: can you enumerate the decision points? If you can β even as twenty branches β write the branches. Reserve model-driven control flow for cases where the space of next actions genuinely cannot be enumerated at design time. And note that this is a spectrum, not a binary: a deterministic outer workflow containing one small bounded loop at the single genuinely open-ended step is usually the right answer, and is far better than either extreme.
Termination
An agent without explicit termination is not an agent, it is an unbounded spend authorisation. Every one of these is mandatory, not a hardening task for later.
class AgentBudget:
max_steps = 12 # hard ceiling on loop iterations
max_wall_clock_s = 60 # user-facing patience, enforced across the whole run
max_cost_usd = 0.50 # cumulative token plus tool spend for this run
max_tool_errors = 3 # consecutive failures before giving up on a path
max_replans = 2 # planner re-entries
def step_allowed(state, budget):
if state.steps >= budget.max_steps:
return False, "step budget exhausted"
if state.elapsed_s >= budget.max_wall_clock_s:
return False, "time budget exhausted"
if state.cost_usd >= budget.max_cost_usd:
return False, "cost budget exhausted"
if state.action_fingerprint in state.recent_fingerprints:
return False, "loop detected - identical action repeated"
return True, None
- Step budget. A hard iteration cap. Choose it from the task, not from optimism β if the job needs four steps, twelve is generous and fifty is an accident waiting to bill you.
- Wall-clock ceiling. Steps can be slow independently of how many there are. A run that has consumed the userβs patience should stop even with steps remaining.
- Cost ceiling. Track cumulative tokens and tool spend per run and stop at the limit. Without this, one pathological run can cost more than a thousand normal ones. See Tokens and Cost Math.
- Loop detection. Fingerprint each action as tool name plus normalised arguments. A repeated identical action means the model is stuck, and a stuck agent will stay stuck until the budget dies. Break immediately rather than waiting for the step cap. Also watch for two-step oscillation, which the naive check misses.
- An explicit failure path. This is the one most often missing. When any budget trips, the agent must surface βI could not complete thisβ with what it did accomplish, what it tried, and what blocked it. Silent truncation and a confidently fabricated final answer are far worse than an honest failure, because the caller cannot tell them apart from success.
Repeated tool failures against the same dependency deserve the same treatment you would give any flaky downstream call β stop hammering it. Circuit Breaker applies to agent tool calls unchanged.
Context Growth Is the Dominant Cost
Every ReAct iteration re-sends the system prompt, the tool definitions, the full reasoning and action history, and every observation returned so far. The prompt regrows on each pass, so cost and latency per step rise as the run progresses. A ten-step run does not cost ten times a one-step run; it costs substantially more, because the later steps carry the weight of all the earlier ones.
This has a blunt consequence: the single most effective agent optimisation is usually not a better prompt, it is returning less from your tools. A tool that dumps a full API response into the transcript pays for that payload on every subsequent step. Return the fields the model needs and keep the rest behind a handle it can ask for.
Four other levers are worth pulling. Cap observation size and truncate with a marker so the model knows the result was trimmed. Summarise or drop stale early history once it stops being decision-relevant, and keep the stable prefix stable so prompt caching can apply. Then put large intermediate artifacts in storage, passing references rather than contents. Storage and eviction strategy is the subject of Agent Memory and State.
π‘ Prompt caching lets the provider charge less for the opening chunk of your prompt, but only while that chunk is character-for-character the same as last time.
Watch cost per successful run, not cost per call. An agent that halves its per-call cost while doubling its step count got worse.
Human-in-the-Loop Checkpoints
Autonomy is a dial, and for anything with real consequences it should not start at maximum.
Put a checkpoint where an action is irreversible, externally visible, or expensive: sending a message on a userβs behalf, moving money, deleting or overwriting records, filing something with a third party, writing to a primary data store. Cheap reversible reads need no gate.
Three patterns worth distinguishing. Approve the plan β the highest-leverage placement, because one approval covers a whole run and planner-executor gives you the artifact to approve. Approve the action β gate specific tool calls by category rather than every call, since a confirmation on everything trains users to click through. Review the output β the agent drafts, a human accepts, and the edit distance on acceptance is a free quality signal for Evals.
Design the waiting properly. A run paused on a human is a run that must survive minutes to days, which means checkpointed state rather than a held-open request, and it means a timeout policy for approvals that never arrive.
π‘ Checkpointing means writing the runβs progress to storage after each step, so a crash picks up where it stopped instead of starting the whole job again. Treat the agentβs proposed action as untrusted input to your authorization layer, not as an already-authorised instruction β the model may be acting on injected content, which is the subject of Prompt Injection and Jailbreaks. The state machinery is in Durable Execution.
Bad to Good to Great
Bad - one free-running loop with every tool attached
A single prompt, twenty tools, βkeep going until the task is done.β No step cap, no argument validation, no cost tracking, no failure path.
It will work in the demo and fail in ways you cannot reproduce. Tool selection degrades because twenty descriptions compete for attention. Context grows unbounded until the run is expensive or truncated. A stuck model loops until something external stops it. And when it fails it produces a fluent final answer built on a silently broken middle step, which is the worst possible failure mode because it is indistinguishable from success.
Good - a bounded loop with validated tools
A step cap, schema validation on tool arguments, a scoped tool set for the task, retries with backoff on transient tool errors, and traces you can replay.
This genuinely ships. Its remaining limits are specific. Control flow still lives in the model, so the execution path is not a code path you can test. The plan exists only implicitly in the transcript, so you cannot approve or cache it. Context still regrows every iteration, and a crash mid-run loses everything because there is no persisted state.
Great - a deterministic workflow with bounded model judgement
- Control flow in code. The outer path is an ordinary workflow with explicit branches. Model calls sit at the judgement points only.
- The fewest fallible steps that do the job. Steps are merged or moved into deterministic code wherever possible, because every removed step multiplies back into the success rate.
- A plan as a first-class artifact where the task warrants it β validated against a schema and an allowlist before execution, cached on goal shape, and approvable by a human.
- Validation at every hop. Tool arguments type-check, tool results are schema-checked, and a failure is attributed to the step that caused it.
- Full budget enforcement β steps, wall clock, cost, consecutive errors, replans, and loop-fingerprint detection.
- An explicit failure path that reports partial progress and the blocking reason instead of fabricating completion.
- Checkpointed state with idempotent steps, so a crash resumes rather than restarts and a resumed run does not duplicate side effects.
- Per-step evals and traces. Plan quality and execution quality are scored separately, and every run is replayable.
The difference between Good and Great is not autonomy. It is that Great has fewer places to fail, and every remaining one is named, bounded, and observable.
When to Use
β Reach for an agent loop when:
- The next step genuinely depends on what the previous step returned, and the path cannot be enumerated at design time
- The number of tool calls varies by input in a way no fixed workflow captures
- Error recovery mid-task has real value, because tools fail in ways the model can correct from
- The task decomposes into steps you can validate individually
β Do not reach for one when:
- A single call with tools, or a fixed workflow with model calls at decision points, would do the job β this covers most cases
- The steps are always the same, in which case you are paying model prices for an if-statement
- Each stepβs output cannot be validated, so failures compound invisibly
- Latency or cost budgets cannot absorb a data-dependent number of model calls
- You are considering multi-agent before a single agent with good tools has been genuinely exhausted
Common Interview Questions
Q1: Why do multi-step agents fail so much more than single calls, and what do you do about it?
Because independent step success rates multiply. At 95 percent per step, five steps is roughly 77 percent end to end and ten steps is around 60 percent, and correlated failures make it worse than that in practice since a wrong early step poisons everything downstream. There are only three levers: raise per-step reliability with validated arguments and schema-checked results, cut the number of fallible steps by moving control flow into deterministic code, or add recovery so a failed step does not end the run. I reach for the second first, because a step moved into ordinary code leaves the risk chain entirely rather than just getting better.
Q2: When would you choose planner-executor over a ReAct loop?
When I need the plan to be an artifact rather than an emergent property of a transcript. That buys four concrete things. I can validate the whole plan against a schema, an allowlist, and a step cap before any side effect happens. I can show it to a human for one approval instead of gating every action, and I can cache it on the goal shape and skip the most expensive call in the run. Execution also becomes deterministic code, so most of the chain stops carrying model-level failure rates. I stay with ReAct when the path genuinely cannot be known in advance, and in practice I often combine them - a validated plan with a small bounded loop at the one step that is actually open-ended.
Q3: A team proposes five specialised agents with a supervisor. What do you push back on?
I ask what a single agent with well-scoped tools has already failed at, because splitting an agent that fails for prompt and tool-description reasons yields five agents that fail for the same reasons plus coordination bugs. Then I name the structural costs: each sub-agent runs its own loop so cost and latency multiply, the only channel between agents is natural language so information degrades at every handoff with nothing raising an error, and a chain of chains compounds the reliability arithmetic. I would support it where concerns are genuinely separable - disjoint tool sets, different permission scopes that should not share a context, real parallelism, or a critic that must not see the writerβs reasoning - and only with schema-defined messages between agents rather than freeform chat.
Q4: How do you stop an agent from running forever or spending unbounded money?
Layered budgets, all enforced by the controller rather than requested in the prompt. A hard step cap, plus a wall-clock ceiling because slow steps are independent of step count. Then a cumulative cost ceiling across tokens and tool spend, a consecutive-error limit per dependency, and a replan cap. On top of those, loop detection that fingerprints tool name plus normalised arguments and breaks immediately on a repeat instead of waiting for the step cap, plus an oscillation check for two-step cycles. The part teams forget is the explicit failure path - when a budget trips the agent must surface that it could not complete the task, with partial progress and the blocking reason, because a fabricated confident answer is indistinguishable from success to the caller.
Q5: What dominates the cost of an agent loop, and how do you reduce it?
Context regrowth. Every iteration re-sends the system prompt, the tool definitions, and the entire accumulated history and observations, so per-step cost rises as the run proceeds and a ten-step run costs far more than ten times a one-step run. The highest-leverage fix is usually not prompt wording but returning less from tools - a tool that dumps a full API response pays for that payload on every subsequent step, so return the needed fields and keep the rest behind a handle. After that: cap and mark truncated observations, drop or summarise history that is no longer decision-relevant, keep the stable prefix stable so prompt caching applies, and pass references to stored artifacts instead of their contents. I track cost per successful run, since halving per-call cost while doubling step count is a regression.
Build it in code: Agentic AI Course · Fundamentals: Core Concepts