Lesson 1 - What an Agent Is and When Not to Build One
Run it:
python3 demo.pyConcept: Agent Architectures covers the theory and the interview framing, without code.
What you will build
- A working mental model with three names in it: chatbot, workflow, agent β and a test for telling them apart in one question.
- The reliability arithmetic that every later lesson inherits, worked out exactly rather than hand-waved.
- The same task written twice against the real
agenticAPI: once as a workflow, once as an agent, with the step counts compared. - A decision rule for when a loop earns its cost, so lessons 2 to 24 are building something you can justify.
The idea
Three things get called βan agentβ and only one of them is.
A chatbot takes text and returns text. It has no tools, so it cannot touch the world. It can be wrong, but it cannot do anything.
A workflow is code you wrote, with model calls inside it. Your code decides the order of steps. The model fills in blanks that are hard to write rules for β classify this message, extract these fields, summarise this in two sentences. The control flow lives in your source file, where you can read it, diff it, and unit test it.
An agent is a loop in which the model decides the next step. You hand it a goal and a set of tools; it picks a tool, you run it, you give it the result, and it picks again until it decides it is done. The control flow lives in the modelβs weights, where you cannot read it.
The analogy that has never let me down: a workflow is a recipe, an agent is a cook. A recipe that says βsautΓ© the onions for six minutesβ is reproducible, auditable, and boring. A cook tastes the sauce and decides what it needs next, which is worth paying for when you cannot know what the sauce will need β and is pure overhead when the answer was always βsix minutesβ.
So the recommendation, stated plainly: most problems described as needing an agent are workflows. A workflow is cheaper (fewer model calls), faster (no round trip to decide the obvious), more debuggable (the path is in your code, not inferred from a transcript), and more reliable (fewer fallible steps). The burden of proof is on the agent. If you cannot say what the loop buys you in one sentence, you want a workflow.
The arithmetic that governs everything downstream
Independent step success rates multiply. That is not an LLM fact; it is what happens to any chain of fallible operations. But it is exactly why agent demos look magical and agent products feel broken.
>>> from decimal import Decimal
>>> Decimal("0.95") ** 5
Decimal('0.7737809375')
>>> 0.95 ** 5 # the same number, in binary floating point
0.7737809374999998
The exact value is 0.7737809375, because 955 over 1005 terminates cleanly in decimal. The float result is the same number with the usual binary rounding, shown so you are not surprised by it in your own REPL.
Five steps, each succeeding 95 percent of the time β a rate most teams would happily call good β is a system that fails roughly one run in four. Ten steps at 95 percent lands near 60 percent. Twenty steps is a coin flip you lose more often than you win.
| Per-step success | 3 steps | 5 steps | 10 steps | 20 steps |
|---|---|---|---|---|
| 99% | 97.0% | 95.1% | 90.4% | 81.8% |
| 95% | 85.7% | 77.4% | 59.9% | 35.8% |
| 90% | 72.9% | 59.0% | 34.9% | 12.2% |
Independence cuts both ways, so two honest caveats. The assumption is optimistic when failures correlate: a wrong entity id in step one poisons every step after it, and cascades are the normal shape of a long agent run rather than the exception. It is pessimistic when the loop can observe and recover, because an agent that sees a tool error and retries with corrected arguments has turned one fatal step into a step with a second chance β the strongest argument there is for loops over straight lines, and you will watch it happen in lesson 3βs test_tool_error_is_returned_to_the_model_which_can_retry.
Either way the design consequence is the same: the architecture that does the job in the fewest fallible steps wins. Every lesson in this course is downstream of that sentence.
The same task, both ways
A support question: can I still refund order 4471? Two facts are needed β the order, and the refund policy.
As a workflow. Your code fetches both, then asks the model one question.
from agentic import FakeModel, Message
ORDERS = {"4471": {"id": "4471", "status": "delivered", "delivered_days_ago": 12}}
POLICY = "Refunds are accepted within 30 days of delivery."
model = FakeModel(["Yes. Delivered 12 days ago, inside the 30-day window."])
def refund_eligible(order_id: str) -> str:
order = ORDERS[order_id] # step 1 - your code chose this
policy = POLICY # step 2 - your code chose this
completion = model.complete([ # step 3 - the model fills one blank
Message(role="system", content="Answer only from the facts given."),
Message(role="user", content=f"Order {order}. Policy: {policy}. Refundable?"),
])
return completion.text
print(refund_eligible("4471"))
print(f"model calls: {model.call_count}") # 1
One model call. The path is three lines of Python you can read. If it breaks, you know which line.
As an agent. You hand over both lookups as tools and let the model sequence them.
from agentic import Agent, Budget, FakeModel, Registry, tool, tool_call
from agentic.tools import ToolError
@tool(description="Look up an order by id.", order_id="The order id, digits only.")
def get_order(order_id: str) -> dict:
if order_id not in ORDERS:
raise ToolError(f"no order {order_id!r}")
return ORDERS[order_id]
@tool(description="Search company refund policy.", q="A short search phrase.")
def search_docs(q: str) -> str:
return POLICY
model = FakeModel([
tool_call("get_order", {"order_id": "4471"}),
tool_call("search_docs", {"q": "refund window"}),
"Yes. Delivered 12 days ago, inside the 30-day window.",
])
run = Agent(model, Registry([get_order, search_docs]), budget=Budget(max_steps=6)).run(
"Can I still refund order 4471?"
)
run.output # the same answer
run.stop # Stop.ANSWERED
run.trajectory() # ["get_order", "search_docs"] - tools that actually RAN
run.step_count # 3 model calls, not 1
Same answer. Three model calls instead of one, three chances to pick the wrong tool, and a control flow you can only reconstruct from run.steps after the fact. On this task the agent version is strictly worse, and it is the version most people build first.
Workflow versus agent, in one picture
flowchart LR
subgraph Workflow
W1[Your code runs step one]
W2[Model fills one blank]
W3[Your code runs step two]
W4[Answer]
end
subgraph Agent
A1[Model sees the goal and the tools]
A2[Model picks the next tool]
A3[Your code runs it and appends the result]
A4[Answer or a named stop reason]
end
W1 --> W2
W2 --> W3
W3 --> W4
A1 --> A2
A2 --> A3
A3 --> A2
A2 --> A4
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class W1,W3,A3 service
class W2,A2 async
class A1 edge
class W4,A4 data
The only structural difference is the arrow from A3 back to A2. That one arrow is the entire subject of this course: it is what makes an agent adaptive, and it is what makes it unbounded until you bound it.
When an agent genuinely earns its cost
One condition, and it is narrow:
The path cannot be known ahead of time. Step three depends on what step two returned, and no amount of upfront thinking lets you write the branch. A support triage that must follow wherever the runbook points. A debugging session where the next log to read depends on the last one. A research question where you do not know how many lookups it takes until you start looking.
Three supporting reasons that are real but weaker on their own: error recovery (a tool error becomes an observation the model can act on), unbounded fan-out (you cannot write a loop over a list whose length nobody knows yet), and the long tail (a hundred rare variants that each deserve a branch, and nobody will write a hundred branches).
And the honest counter-test, which I would ask in an interview: could you replace the loop with a switch statement and lose nothing? If yes, write the switch statement. It is faster, cheaper, and you can put a breakpoint in it.
What you will build over 24 lessons
A small, readable agent library with no framework in it β because the fastest way to understand what LangGraph does for you is to have written it once yourself.
Lessons 2 to 6 build the machinery: the model boundary, tools, the loop, structured outputs, memory. You have a working agent by the end of lesson 4. Lessons 7 to 12 are the half that separates a demo from something deployable β retries, budgets, tracing, evals, a judge, guardrails. Lessons 13 to 16 add knowledge: embeddings, vector search, hybrid retrieval, retrieval as a tool. Lessons 17 to 22 scale the pattern and run it in production. Lessons 23 and 24 are a capstone and the interview.
Everything runs offline, with no API key and no pip install, for one reason you will meet in the next lesson.
Exercise
Pick a task from your own work that someone has proposed building an agent for. Write it twice against the real API: once as a workflow (your code sequences the steps, the model fills in blanks) and once as an agent (tools plus a loop). Then compare.
Success criterion: your script prints the model-call count for both shapes and asserts the workflow uses strictly fewer. Run it with python3 compare.py from inside agentic-course/. Then answer, in one sentence, what the extra calls bought you β and if the answer is βnothingβ, you have just found a workflow.
Worked solution
```python """Same task, two shapes. Save as compare.py inside agentic-course/ and run it.""" from agentic import Agent, Budget, FakeModel, Message, Registry, Stop, tool, tool_call ORDERS = {"4471": {"id": "4471", "status": "delivered", "delivered_days_ago": 12}} POLICY = "Refunds are accepted within 30 days of delivery." # -- workflow: your code decides the path --------------------------------- wf_model = FakeModel(["Yes, it is inside the 30-day window."]) def workflow(order_id: str) -> str: order = ORDERS[order_id] return wf_model.complete([ Message(role="system", content="Answer only from the facts given."), Message(role="user", content=f"Order {order}. Policy: {POLICY}. Refundable?"), ]).text wf_answer = workflow("4471") # -- agent: the model decides the path ----------------------------------- @tool(description="Look up an order by id.", order_id="The order id.") def get_order(order_id: str) -> dict: return ORDERS[order_id] @tool(description="Search company refund policy.", q="A short search phrase.") def search_docs(q: str) -> str: return POLICY ag_model = FakeModel([ tool_call("get_order", {"order_id": "4471"}), tool_call("search_docs", {"q": "refund window"}), "Yes, it is inside the 30-day window.", ]) run = Agent(ag_model, Registry([get_order, search_docs]), budget=Budget(max_steps=6)).run( "Can I still refund order 4471?" ) assert run.stop is Stop.ANSWERED assert run.trajectory() == ["get_order", "search_docs"] assert wf_model.call_count < run.step_count print(f"workflow {wf_model.call_count} model call -> {wf_answer}") print(f"agent {run.step_count} model calls -> {run.output}") print(f"agent ran {run.trajectory()} to reach the same answer") ``` Both print the same answer. The agent spent three model calls and three opportunities to go wrong. On a task where the path is knowable, that is a pure loss β which is the point of the exercise.Checkpoint
What is the one-question test for agent versus workflow? Who decides the next step. If your code decides the order and the model only fills in blanks, it is a workflow. If the model decides the order, it is an agent.
Five steps at 95 percent each β what is the end-to-end success rate, exactly?
0.7737809375, about 77 percent. Roughly one run in four fails even though every individual step is βgoodβ. UseDecimal("0.95") ** 5if you want the exact digits rather than the float.
Why does the independence assumption understate risk in practice? Because agent failures correlate. A wrong id or a misread document in an early step raises the failure rate of everything downstream, so real long runs do worse than the multiplication suggests.
Name the one condition under which an agent loop is clearly worth it. The path cannot be known ahead of time β step three depends on what step two returned, and no upfront branch can capture it. Error recovery and unbounded fan-out are supporting arguments, not sufficient ones.
What does
run.trajectory()report, and what does it deliberately leave out? The tools that actually executed successfully, in order. It leaves out calls the model requested but that were denied, invalid, or failed β those live inrun.requested_tools()andrun.blocked_tools().
Theory and interview framing: Become an AI Engineer