Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 13 min read

Lesson 1 - What an Agent Is and When Not to Build One

Run it: python3 demo.py Concept: Agent Architectures covers the theory and the interview framing, without code.


What you will build


The idea

Three things get called β€œan agent” and only one of them is.

A chatbot takes text and returns text. It has no tools, so it cannot touch the world. It can be wrong, but it cannot do anything.

A workflow is code you wrote, with model calls inside it. Your code decides the order of steps. The model fills in blanks that are hard to write rules for β€” classify this message, extract these fields, summarise this in two sentences. The control flow lives in your source file, where you can read it, diff it, and unit test it.

An agent is a loop in which the model decides the next step. You hand it a goal and a set of tools; it picks a tool, you run it, you give it the result, and it picks again until it decides it is done. The control flow lives in the model’s weights, where you cannot read it.

The analogy that has never let me down: a workflow is a recipe, an agent is a cook. A recipe that says β€œsautΓ© the onions for six minutes” is reproducible, auditable, and boring. A cook tastes the sauce and decides what it needs next, which is worth paying for when you cannot know what the sauce will need β€” and is pure overhead when the answer was always β€œsix minutes”.

So the recommendation, stated plainly: most problems described as needing an agent are workflows. A workflow is cheaper (fewer model calls), faster (no round trip to decide the obvious), more debuggable (the path is in your code, not inferred from a transcript), and more reliable (fewer fallible steps). The burden of proof is on the agent. If you cannot say what the loop buys you in one sentence, you want a workflow.


The arithmetic that governs everything downstream

Independent step success rates multiply. That is not an LLM fact; it is what happens to any chain of fallible operations. But it is exactly why agent demos look magical and agent products feel broken.

>>> from decimal import Decimal
>>> Decimal("0.95") ** 5
Decimal('0.7737809375')
>>> 0.95 ** 5              # the same number, in binary floating point
0.7737809374999998

The exact value is 0.7737809375, because 955 over 1005 terminates cleanly in decimal. The float result is the same number with the usual binary rounding, shown so you are not surprised by it in your own REPL.

Five steps, each succeeding 95 percent of the time β€” a rate most teams would happily call good β€” is a system that fails roughly one run in four. Ten steps at 95 percent lands near 60 percent. Twenty steps is a coin flip you lose more often than you win.

Per-step success 3 steps 5 steps 10 steps 20 steps
99% 97.0% 95.1% 90.4% 81.8%
95% 85.7% 77.4% 59.9% 35.8%
90% 72.9% 59.0% 34.9% 12.2%

Independence cuts both ways, so two honest caveats. The assumption is optimistic when failures correlate: a wrong entity id in step one poisons every step after it, and cascades are the normal shape of a long agent run rather than the exception. It is pessimistic when the loop can observe and recover, because an agent that sees a tool error and retries with corrected arguments has turned one fatal step into a step with a second chance β€” the strongest argument there is for loops over straight lines, and you will watch it happen in lesson 3’s test_tool_error_is_returned_to_the_model_which_can_retry.

Either way the design consequence is the same: the architecture that does the job in the fewest fallible steps wins. Every lesson in this course is downstream of that sentence.


The same task, both ways

A support question: can I still refund order 4471? Two facts are needed β€” the order, and the refund policy.

As a workflow. Your code fetches both, then asks the model one question.

from agentic import FakeModel, Message

ORDERS = {"4471": {"id": "4471", "status": "delivered", "delivered_days_ago": 12}}
POLICY = "Refunds are accepted within 30 days of delivery."

model = FakeModel(["Yes. Delivered 12 days ago, inside the 30-day window."])

def refund_eligible(order_id: str) -> str:
    order = ORDERS[order_id]                  # step 1 - your code chose this
    policy = POLICY                           # step 2 - your code chose this
    completion = model.complete([             # step 3 - the model fills one blank
        Message(role="system", content="Answer only from the facts given."),
        Message(role="user", content=f"Order {order}. Policy: {policy}. Refundable?"),
    ])
    return completion.text

print(refund_eligible("4471"))
print(f"model calls: {model.call_count}")     # 1

One model call. The path is three lines of Python you can read. If it breaks, you know which line.

As an agent. You hand over both lookups as tools and let the model sequence them.

from agentic import Agent, Budget, FakeModel, Registry, tool, tool_call
from agentic.tools import ToolError

@tool(description="Look up an order by id.", order_id="The order id, digits only.")
def get_order(order_id: str) -> dict:
    if order_id not in ORDERS:
        raise ToolError(f"no order {order_id!r}")
    return ORDERS[order_id]

@tool(description="Search company refund policy.", q="A short search phrase.")
def search_docs(q: str) -> str:
    return POLICY

model = FakeModel([
    tool_call("get_order", {"order_id": "4471"}),
    tool_call("search_docs", {"q": "refund window"}),
    "Yes. Delivered 12 days ago, inside the 30-day window.",
])

run = Agent(model, Registry([get_order, search_docs]), budget=Budget(max_steps=6)).run(
    "Can I still refund order 4471?"
)

run.output         # the same answer
run.stop           # Stop.ANSWERED
run.trajectory()   # ["get_order", "search_docs"] - tools that actually RAN
run.step_count     # 3 model calls, not 1

Same answer. Three model calls instead of one, three chances to pick the wrong tool, and a control flow you can only reconstruct from run.steps after the fact. On this task the agent version is strictly worse, and it is the version most people build first.


Workflow versus agent, in one picture

flowchart LR
    subgraph Workflow
        W1[Your code runs step one]
        W2[Model fills one blank]
        W3[Your code runs step two]
        W4[Answer]
    end
    subgraph Agent
        A1[Model sees the goal and the tools]
        A2[Model picks the next tool]
        A3[Your code runs it and appends the result]
        A4[Answer or a named stop reason]
    end
    W1 --> W2
    W2 --> W3
    W3 --> W4
    A1 --> A2
    A2 --> A3
    A3 --> A2
    A2 --> A4
    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000
    class W1,W3,A3 service
    class W2,A2 async
    class A1 edge
    class W4,A4 data

The only structural difference is the arrow from A3 back to A2. That one arrow is the entire subject of this course: it is what makes an agent adaptive, and it is what makes it unbounded until you bound it.


When an agent genuinely earns its cost

One condition, and it is narrow:

The path cannot be known ahead of time. Step three depends on what step two returned, and no amount of upfront thinking lets you write the branch. A support triage that must follow wherever the runbook points. A debugging session where the next log to read depends on the last one. A research question where you do not know how many lookups it takes until you start looking.

Three supporting reasons that are real but weaker on their own: error recovery (a tool error becomes an observation the model can act on), unbounded fan-out (you cannot write a loop over a list whose length nobody knows yet), and the long tail (a hundred rare variants that each deserve a branch, and nobody will write a hundred branches).

And the honest counter-test, which I would ask in an interview: could you replace the loop with a switch statement and lose nothing? If yes, write the switch statement. It is faster, cheaper, and you can put a breakpoint in it.


What you will build over 24 lessons

A small, readable agent library with no framework in it β€” because the fastest way to understand what LangGraph does for you is to have written it once yourself.

Lessons 2 to 6 build the machinery: the model boundary, tools, the loop, structured outputs, memory. You have a working agent by the end of lesson 4. Lessons 7 to 12 are the half that separates a demo from something deployable β€” retries, budgets, tracing, evals, a judge, guardrails. Lessons 13 to 16 add knowledge: embeddings, vector search, hybrid retrieval, retrieval as a tool. Lessons 17 to 22 scale the pattern and run it in production. Lessons 23 and 24 are a capstone and the interview.

Everything runs offline, with no API key and no pip install, for one reason you will meet in the next lesson.


Exercise

Pick a task from your own work that someone has proposed building an agent for. Write it twice against the real API: once as a workflow (your code sequences the steps, the model fills in blanks) and once as an agent (tools plus a loop). Then compare.

Success criterion: your script prints the model-call count for both shapes and asserts the workflow uses strictly fewer. Run it with python3 compare.py from inside agentic-course/. Then answer, in one sentence, what the extra calls bought you β€” and if the answer is β€œnothing”, you have just found a workflow.

Worked solution ```python """Same task, two shapes. Save as compare.py inside agentic-course/ and run it.""" from agentic import Agent, Budget, FakeModel, Message, Registry, Stop, tool, tool_call ORDERS = {"4471": {"id": "4471", "status": "delivered", "delivered_days_ago": 12}} POLICY = "Refunds are accepted within 30 days of delivery." # -- workflow: your code decides the path --------------------------------- wf_model = FakeModel(["Yes, it is inside the 30-day window."]) def workflow(order_id: str) -> str: order = ORDERS[order_id] return wf_model.complete([ Message(role="system", content="Answer only from the facts given."), Message(role="user", content=f"Order {order}. Policy: {POLICY}. Refundable?"), ]).text wf_answer = workflow("4471") # -- agent: the model decides the path ----------------------------------- @tool(description="Look up an order by id.", order_id="The order id.") def get_order(order_id: str) -> dict: return ORDERS[order_id] @tool(description="Search company refund policy.", q="A short search phrase.") def search_docs(q: str) -> str: return POLICY ag_model = FakeModel([ tool_call("get_order", {"order_id": "4471"}), tool_call("search_docs", {"q": "refund window"}), "Yes, it is inside the 30-day window.", ]) run = Agent(ag_model, Registry([get_order, search_docs]), budget=Budget(max_steps=6)).run( "Can I still refund order 4471?" ) assert run.stop is Stop.ANSWERED assert run.trajectory() == ["get_order", "search_docs"] assert wf_model.call_count < run.step_count print(f"workflow {wf_model.call_count} model call -> {wf_answer}") print(f"agent {run.step_count} model calls -> {run.output}") print(f"agent ran {run.trajectory()} to reach the same answer") ``` Both print the same answer. The agent spent three model calls and three opportunities to go wrong. On a task where the path is knowable, that is a pure loss β€” which is the point of the exercise.

Checkpoint

What is the one-question test for agent versus workflow? Who decides the next step. If your code decides the order and the model only fills in blanks, it is a workflow. If the model decides the order, it is an agent.

Five steps at 95 percent each β€” what is the end-to-end success rate, exactly? 0.7737809375, about 77 percent. Roughly one run in four fails even though every individual step is β€œgood”. Use Decimal("0.95") ** 5 if you want the exact digits rather than the float.

Why does the independence assumption understate risk in practice? Because agent failures correlate. A wrong id or a misread document in an early step raises the failure rate of everything downstream, so real long runs do worse than the multiplication suggests.

Name the one condition under which an agent loop is clearly worth it. The path cannot be known ahead of time β€” step three depends on what step two returned, and no upfront branch can capture it. Error recovery and unbounded fan-out are supporting arguments, not sufficient ones.

What does run.trajectory() report, and what does it deliberately leave out? The tools that actually executed successfully, in order. It leaves out calls the model requested but that were denied, invalid, or failed β€” those live in run.requested_tools() and run.blocked_tools().


Theory and interview framing: Become an AI Engineer

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access