Agentic AI, End to End
A hands-on course that builds a production-shaped AI agent from scratch โ the loop, the tools, the memory, the evals, the guardrails โ with no framework at all. When you later pick up LangGraph or CrewAI, you will already know what it is doing for you, because you will have written it.
Everything runs offline. No API key, no pip install, no virtualenv. The code ships with 116 passing tests and a working demo you can run in the next sixty seconds.
This is the workshop. The map is next door. Twenty-four lessons that build one thing deeply, with code you execute and exercises you check. For the breadth of the role โ how LLMs work, prompting, model selection, tokens and cost, chunking, hallucination and grounding, fine-tuning, inference serving, GPU cost, multimodal โ read the AI Engineer track. None of those are covered here. Every lesson links back to its concept page, marked Concept at the top. Read the concept if you want the framing; skip it if you want to build.
Start in sixty seconds
Grab just the course folder. It is 184 KB โ do not clone the whole site.
npx degit iluzn/HLD-Designs/agentic-course agentic-course
cd agentic-course
No Node? Use a sparse checkout, which skips the site entirely:
git clone --filter=blob:none --sparse https://github.com/iluzn/HLD-Designs.git
cd HLD-Designs
git sparse-checkout set agentic-course
cd agentic-course
Then, either way:
python3 demo.py # full stack, 6 scenarios
python3 -m unittest discover -s tests -t . # 116 tests
Python 3.10+. Standard library only. That is the whole setup.
Why it runs with no API key. Everything talks to a model through one method,
Model.complete. That single seam is the trick:FakeModelis scripted and deterministic, so the agent loop, the traces, and the eval suite behave identically whether a real provider is on the other end or not. You learn the mechanics for free, then swap one class to go live.
What you actually build
from agentic import Agent, FakeModel, Registry, tool, tool_call
@tool(description="Look up an order by id.", order_id="The order id.")
def get_order(order_id: str) -> dict:
return {"id": order_id, "status": "delivered"}
model = FakeModel([
tool_call("get_order", {"order_id": "4471"}),
"Order 4471 was delivered.",
])
run = Agent(model, Registry([get_order])).run("Where is order 4471?")
run.output # "Order 4471 was delivered."
run.trajectory() # ["get_order"] - tools that actually RAN
run.stop # Stop.ANSWERED
By the end you have added budgets, tracing, an eval suite, a judge, hybrid retrieval, injection containment, and a capstone you can put in front of an interviewer.
Three ideas the whole course turns on
A loop with no budget is not an agent. It is a while loop with a credit card attached. Every run in this course ends with a named stop reason โ answered, step budget, cost budget, repeated call, fatal tool error. There is no unnamed terminal state, because โit just stoppedโ is how you get a surprise invoice.
Requested is not executed. run.trajectory() lists tools that actually ran. run.requested_tools() lists what the model asked for. Conflating them makes a working guardrail look like a breach โ a blocked injection would fail a never_used("refund") assertion even though the defence held. Writing the tests for this course is what surfaced that bug in my own code.
Prompt injection is contained, not prevented. Instructions and data share one channel, so no phrasing makes a documentโs text non-instruction. The course therefore removes capability rather than trying to detect attacks: a context that has read untrusted text loses its mutating tools. Demo scenario 3 shows a fully compromised model that still cannot issue the refund.
Agentic AI Course โ Lesson Tracker
Track your progress through all 24 lessons. Sign in to save across devices.
Part 1 โ Build an Agent From Scratch
No frameworks. You write the loop, so you understand every layer above it.
| # | Lesson | Code | Read |
|---|---|---|---|
| 1 | What an Agent Is (and When Not to Build One) | โ | Read โ |
| 2 | The Model Boundary โ the seam that makes agents testable | model.py |
Read โ |
| 3 | Tools โ schemas, validation, dispatch | tools.py |
Read โ |
| 4 | The Agent Loop โ ReAct from scratch | loop.py |
Read โ |
| 5 | Structured Outputs โ contracts a program can consume | tools.py |
Read โ |
| 6 | Memory โ history windows, summarization, a fact store | memory.py |
Read โ |
Part 2 โ Make It Not Break
The half of agent engineering that separates a demo from something you can deploy.
| # | Lesson | Code | Read |
|---|---|---|---|
| 7 | Failure Handling โ retries, timeouts, idempotent tools | model.py |
Read โ |
| 8 | Budgets and Termination โ step, cost, and loop guards | loop.py |
Read โ |
| 9 | Tracing โ nested spans, replay, redaction | trace.py |
Read โ |
| 10 | Agent Evals โ golden tasks and trajectory scoring | evals.py |
Read โ |
| 11 | LLM-as-Judge โ rubrics, kappa, judge validation | evals.py |
Read โ |
| 12 | Guardrails โ injection, policy, containment | guard.py |
Read โ |
Part 3 โ Give It Knowledge
| # | Lesson | Code | Read |
|---|---|---|---|
| 13 | Embeddings โ cosine similarity from scratch | retrieval.py |
Read โ |
| 14 | Vector Search โ build a store, then learn what ANN buys | retrieval.py |
Read โ |
| 15 | Hybrid and Rerank โ BM25, reciprocal rank fusion | retrieval.py |
Read โ |
| 16 | Retrieval as a Tool โ agentic search | retrieval.py |
Read โ |
Part 4 โ Scale the Pattern
| # | Lesson | Code | Read |
|---|---|---|---|
| 17 | Planning โ planner-executor and decomposition | loop.py |
Read โ |
| 18 | Multi-Agent โ supervisors, specialists, and when it is theatre | loop.py |
Read โ |
| 19 | Durable Execution โ checkpoints, resumption, idempotency | memory.py |
Read โ |
| 20 | MCP and Frameworks โ porting what you built | โ | Read โ |
Part 5 โ Run It in Production
| # | Lesson | Code | Read |
|---|---|---|---|
| 21 | Latency and Cost โ caching, parallelism, budgets | trace.py |
Read โ |
| 22 | Deployment and Rollout โ versioning, shadow, canary | evals.py |
Read โ |
Part 6 โ Ship Something
| # | Lesson | Code | Read |
|---|---|---|---|
| 23 | Capstone โ build and benchmark an agent | all | Read โ |
| 24 | The Agentic AI Interview | โ | Read โ |
The library you build
Eight modules, each short enough to read in one sitting.
| File | What it is | The lesson it teaches |
|---|---|---|
model.py |
Model boundary, FakeModel, HTTP adapter |
Why one seam makes everything testable |
tools.py |
Schemas, validation, authorization, dispatch | This is the security boundary |
loop.py |
The loop, budgets, named stop reasons | What every framework wraps |
memory.py |
History, summarization, fact store | Memory poisoning is the real risk |
trace.py |
Spans, token and cost accounting, redaction | The trace is the bug report |
evals.py |
Golden cases, trajectory scoring, judge | The skill that gets you hired |
guard.py |
Fencing, policy, egress, confirmation | Containment when prevention fails |
retrieval.py |
Cosine, vector store, BM25, RRF, rerank | Why hybrid beats dense-only |
Tests as the curriculum
The test names tell you what the course is really about:
test_step_budget_stops_a_runaway_agent
test_repeated_identical_call_is_detected
test_internal_exception_does_not_leak_details_to_the_model
test_unknown_tool_does_not_reveal_hidden_tools
test_missing_confirm_handler_crashes_rather_than_degrading
test_permission_filter_runs_before_ranking
test_rrf_combines_by_rank_not_score
test_always_pass_judge_has_zero_kappa_despite_high_agreement
test_injected_document_cannot_reach_a_mutating_tool
test_protected_keys_cannot_be_written_by_an_agent
Each encodes a failure mode that costs real money. Two of them exist because writing them found genuine design bugs in the library.
Honest limitations
Stated up front, because a course that hides its simplifications teaches you to trust the wrong things.
HashingEmbedderis a bag-of-words hash, not a semantic model. It cannot tell โcheapโ and โinexpensiveโ are related. It exists so the retrieval mechanics run offline. Swap in a real embedding model for anything real โ and note that doing so means re-indexing your whole corpus, never a config flip.VectorStoreis exact brute-force search, linear in corpus size. Real systems use approximate indexes and treat recall as a tunable SLO. Brute force is the baseline you measure those against.guard.scanis a pattern matcher and will always be incomplete โ an attacker needs one phrasing you did not think of. Use it for triage, never as the thing standing between an attacker and a tool.- No async. The library stays synchronous so the control flow is readable.
Prerequisites
Working Python and comfort with HTTP APIs. You do not need a maths background, a PhD, an ML background, or a GPU. If you want the conceptual grounding first, the AI Engineer track covers the vocabulary; this course is the hands-on half.
Related on SystemCraft
- Become an AI Engineer โ the 36-topic reference track this course implements
- Design ChatGPT โ the system-design view of an AI product
- System Design Concepts โ idempotency, queues, caching, observability
- DSA Problemset โ the coding round is still an ordinary coding round
Navigation
| AI Engineer Track | All Concepts |