Become an AI Engineer
A complete path from βI write backend codeβ to βI ship production AI features.β Thirty-six deep dives, ordered, free, no signup.
This track exists because most AI learning material stops at the demo. Calling a model API is the easy part and takes an afternoon. The job is everything after that: getting the right context in front of the model, proving the output is actually correct, stopping it being talked into something harmful, and keeping it fast and affordable. That is the material below.
This is the map. The workshop is next door. Thirty-six topics covering the breadth of the role β read it to understand the landscape and to prepare for interviews. It is prose, not code. If you would rather build an agent and run tests against it, take the Agentic AI Course β 24 lessons over a real Python library with 116 tests that run offline with no API key. Twenty-one topics here have a paired lesson there, marked Build it at the top of the page. Use both: read the concept, then implement it.
What an AI Engineer Actually Is
An AI Engineer works at the product layer β you take a foundation model you did not train and build reliable systems on it. An ML Engineer works at the model layer β training, tuning, and serving the models themselves. Both are real engineering; they are different jobs, and the first one is where most hiring currently sits.
The good news for anyone already writing backend code: you own most of the skill surface already. API design, idempotency, retries, queues, caching, observability, cost control β all of it transfers. What is new is that one of your dependencies is non-deterministic, bills per token, and will return a confidently wrong answer in a perfectly valid response format.
You do not need a PhD, the ability to derive backpropagation, or to train a model from scratch. You do need the six bits of maths in topic 3, an accurate mental model of how these systems fail, and the discipline to measure instead of guess.
New here? Start with The AI Engineer Role, then How LLMs Actually Work. If you want the fastest useful path, take the Crash Course below.
The Five Skills That Get Screened
Job postings vary wildly in wording but converge on five clusters. This track is weighted toward them, which is why evaluation gets five topics and prompt tricks get one.
| Skill cluster | Why it matters | Start here |
|---|---|---|
| Evals | The load-bearing skill, and the one most self-taught candidates skip. Without a measurement you cannot tell an improvement from a regression | Evals |
| Retrieval architecture | Retrieval quality caps answer quality. Most βthe model is dumbβ bugs are retrieval bugs | RAG End to End |
| Agent orchestration | Letting a model act without the system becoming unbounded or unpredictable | Tool Calling |
| Cost and latency | Model calls are the slowest, most expensive dependency most products have shipped | Tokens and Cost Math |
| AI security | Prompt injection has no complete fix, so the discipline is bounding the damage | AI Security |
Curated Playlists
Three ordered routes through the same 36 topics. Pick one based on what you need next β the full tracker is further down.
Crash Course
10 topics, about a week. The shortest path to shipping something real and knowing whether it works. Skips agents, tuning, and self-hosted serving entirely.
- The AI Engineer Role
- How LLMs Actually Work
- LLM APIs and SDKs
- Tokens and Cost Math
- Prompt Engineering
- Structured Outputs
- Embeddings
- RAG End to End
- Evals
- AI Security
GenAI Product Track
16 topics. For building user-facing features on foundation models. Heavy on retrieval and evaluation, which is where these products actually succeed or fail.
Topics 1 and 2, then 5 through 18 in order β orientation, the full model-interaction surface, all six retrieval topics, then evals, judges, and grounding.
Start at The AI Engineer Role and read straight through to Hallucination and Grounding, skipping topics 3 and 4 if your Python and maths are already solid.
Production AI Track
12 topics. For engineers who already have something working and now have to operate it at scale, on a budget, on call.
- Evals
- AI Security
- Tracing and Observability
- Tool Calling
- Agent Architectures
- Agent Memory and State
- Inference Serving
- GPU Cost Engineering
- Latency Engineering
- Prompt and Semantic Caching
- Deploying AI Features
- The AI Engineer Interview
AI Engineer Track β Learning Tracker
Track your progress through all 36 topics. Sign in to save across devices.
Stage 1 β Orientation
Where the role sits, how the models behave, and the small amount of maths and Python you actually need.
| # | Topic | Deep Dive |
|---|---|---|
| 1 | The AI Engineer Role (vs ML Engineer, vs Data Scientist) | Read β |
| 2 | How LLMs Actually Work (Next-Token Prediction, Sampling) | Read β |
| 3 | The Math You Actually Need (and What to Skip) | Read β |
| 4 | Python for AI Engineers (Async, Bounded Concurrency) | Read β |
Stage 2 β Working With Models
The request surface, what it costs, and how to get output your program can actually consume.
| # | Topic | Deep Dive |
|---|---|---|
| 5 | LLM APIs and SDKs (Roles, Streaming, Error Classes) | Read β |
| 6 | Tokens, Context Windows, and Cost Math | Read β |
| 7 | Prompt Engineering, Systematically | Read β |
| 8 | Structured Outputs (JSON Mode, Constrained Decoding) | Read β |
| 9 | Model Selection and Routing (Cascades, Self-Hosting) | Read β |
Stage 3 β Retrieval
Getting the right context in front of the model. Retrieval quality caps answer quality, so this stage is six topics deep.
| # | Topic | Deep Dive |
|---|---|---|
| 10 | Embeddings and Semantic Similarity | Read β |
| 11 | Vector Databases and ANN Search (HNSW, IVF-PQ) | Read β |
| 12 | Document Parsing and Chunking | Read β |
| 13 | RAG End to End | Read β |
| 14 | Hybrid Search and Reranking (BM25, Cross-Encoders) | Read β |
| 15 | Advanced Retrieval (Multi-Hop, GraphRAG, Agentic Search) | Read β |
Stage 4 β Evaluation and Reliability
The professional core of the discipline. Without this stage you are not engineering, you are guessing.
| # | Topic | Deep Dive |
|---|---|---|
| 16 | Evals: The Core Discipline | Read β |
| 17 | LLM-as-Judge and Automated Grading | Read β |
| 18 | Hallucination, Grounding, and Citations | Read β |
| 19 | Prompt Injection, Jailbreaks, and the OWASP LLM Top 10 | Read β |
| 20 | Tracing and Observability for LLM Apps | Read β |
Stage 5 β Agents
Letting a model take actions, without the system becoming unbounded, unpredictable, or dangerous.
| # | Topic | Deep Dive |
|---|---|---|
| 21 | Tool Calling Fundamentals | Read β |
| 22 | Agent Architectures (ReAct, Planner-Executor, Multi-Agent) | Read β |
| 23 | Model Context Protocol and Tool Ecosystems | Read β |
| 24 | Agent Memory, State, and Durable Execution | Read β |
Stage 6 β Customization
What to reach for after prompting, retrieval, and evals are exhausted. Reaching here first is the classic mistake.
| # | Topic | Deep Dive |
|---|---|---|
| 25 | Fine-Tuning: When, Why, and How (LoRA, QLoRA) | Read β |
| 26 | RLHF, DPO, and Preference Tuning | Read β |
| 27 | Distillation and Small Models | Read β |
| 28 | Synthetic Data Generation | Read β |
Stage 7 β Production
Serving, cost, latency, caching, and rollout. This is where backend experience pays off most directly.
| # | Topic | Deep Dive |
|---|---|---|
| 29 | Inference Serving (Continuous Batching, KV Cache, vLLM) | Read β |
| 30 | GPU Capacity and Cost Engineering | Read β |
| 31 | Latency Engineering for LLM Apps | Read β |
| 32 | Prompt and Semantic Caching | Read β |
| 33 | Deploying and Rolling Out AI Features | Read β |
| 34 | Multimodal (Vision, Audio, Documents) | Read β |
Stage 8 β Getting Hired
| # | Topic | Deep Dive |
|---|---|---|
| 35 | Portfolio Projects That Get You Hired | Read β |
| 36 | The AI Engineer Interview | Read β |
How Long This Takes
An estimate, not a promise. It assumes you already code professionally, you are studying part time, and you are building rather than only reading.
| Period | What you cover | What you can build |
|---|---|---|
| Weeks 1-4 | Stages 1 and 2 | A useful single-call feature with validated structured output |
| Months 2-3 | Stage 3 and the start of Stage 4 | A full RAG pipeline plus your first real eval suite |
| Months 3-5 | Rest of Stage 4, Stage 5, parts of Stage 7 | A traced, guarded, tool-using system inside a latency and cost budget |
| Months 5-8 | Stage 6 or deeper Stage 7, then Stage 8 | A measured portfolio project and interview readiness |
Roughly six months to solid interview readiness. The curve is steepest in months 2 and 3, because retrieval quality is genuinely fiddly and that is where most people quit. Anyone selling a weekend transition is selling a demo, not the job.
The One Habit That Matters
Write an eval set early, even a bad one. Fifty labelled cases from real inputs will teach you more than fifty hours of reading, because it converts every later change from a guess into a measurement. Teams that ship reliable AI features are not the ones with the cleverest prompts β they are the ones whose eval set grows every week from real failures.
Related on SystemCraft
- Design ChatGPT β the worked system design for an AI chat product: streaming, GPU scaling, context management
- System Design Concepts β caching, queues, idempotency, observability. The fundamentals underneath every AI system
- HLD Problems β system design practice for the architecture round
- DSA Problemset β the coding round is still an ordinary coding round
Navigation
| All Concepts | HLD Designs |