Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 20 min read

Portfolio Projects That Get You Hired - Complete Deep Dive

Stage 8 - Getting Hired Lesson 35 of 36

Prerequisites: Evals, RAG End to End, Deploying AI Features Used in: The AI Engineer Interview Build it: Lesson 23 - Capstone - Build and Benchmark an Agent implements this as runnable, tested code you can execute offline.


What is a Portfolio Project?

A portfolio project is evidence submitted to a skeptical reader. Not a demo, not a tutorial artifact, not a screenshot. Evidence.

That framing settles almost every decision on this page, because it forces the question: evidence of what? And here is the thesis this entire page defends: a demo that works on three inputs proves nothing, because the entire difficulty of AI engineering is reliability. Anyone can get a model to answer one question well. The job β€” the part that is actually hard, the part companies are paying for β€” is getting it to answer the tenth thousandth question well, inside a latency budget, at a cost that does not sink the feature. A demo cannot demonstrate that. Only a measurement can.

So the rule is: a project is only evidence if it measures something.

The most differentiating artifact you can produce is not the most impressive-sounding system. It is a project with a recorded baseline, three deliberate changes, and the measured effect of each change on quality, p95 latency, and cost per request. That is the organising idea of this page, and everything below is in service of it. A modest grounded question-answering system with that table beats an ambitious multi-agent framework without it, every single time, because the table is the only part of either project that proves you can tell improvement from regression.

Real-world analogy: two mechanics describe the same engine tune. The first says β€œit feels much quicker now.” The second hands you a dyno sheet: before, after, three changes, what each one did to power and fuel consumption, plus the one change that made it worse and got reverted. The second mechanic is the one you trust with your car, and it is not because their work was fancier. It is because they measured.

This is the artifact that The AI Engineer Role calls the Great tier, and it happens to exercise all five skill clusters at once β€” evals, retrieval, agents, cost and latency, security β€” which is why one good project outperforms five shallow ones.


Why the Common Portfolio Fails

Look at what most candidates submit. The pattern is remarkably consistent, and every element of it is a signal that the candidate has not met the real difficulty of the work.

What is submitted What the interviewer concludes
A chat wrapper built from a tutorial, same stack and same structure as thousands of others The candidate followed instructions. Nothing shows they could have made a decision.
No eval set anywhere in the repository Quality was judged by eyeballing. The candidate has no instrument, so every claim is unverifiable.
No failure analysis - the README lists only what works Either the system never broke, which is false, or the candidate never looked.
No cost or latency numbers The candidate has not thought about the two constraints that kill most AI features in production.
A README that lists features rather than tradeoffs Feature lists describe scope. Hiring decisions turn on judgement, and judgement only shows in rejected alternatives.
Five half-finished projects instead of one finished one Breadth without depth. The interviewer cannot go deep on any of them, so the conversation dies.
A framework tour - four orchestration libraries wired together Tooling familiarity, not engineering. The first mechanism question exposes it.

The common thread: none of it is falsifiable. There is nothing in the repository a reader could check, disagree with, or be surprised by. A portfolio that cannot be wrong cannot be evidence.

The sharpest version of this failure is the project that does work. It really does answer questions about the documents. And the first interview question is β€œhow do you know it works?” β€” and the honest answer is β€œI tried it and it seemed good.” At that point the project has stopped helping and started hurting, because it established that the candidate does not measure.


The Anatomy of a Project That Lands

Eight properties. They are not a wish list; each one closes off a specific way the interviewer could dismiss the work.

1. A real domain you personally know. Tax rules for your country, the regulations in your last job’s industry, the rulebook of a sport you play, your company’s internal runbooks, a hobby with a dense corpus. This matters for one non-obvious reason: you cannot build an eval set for a domain where you cannot judge correctness. Pick a generic domain and you will be unable to label your own test cases, which means you cannot measure, which means you cannot produce the artifact that matters. Domain knowledge is not decoration here β€” it is the enabling condition.

2. Real messy data, not a clean sample. Actual PDFs with two-column layouts, scanned pages, tables that span page breaks, inconsistent headings, near-duplicate documents with different versions. The clean sample dataset skips every problem that makes Document Parsing and Chunking a discipline. Mess is where the interesting decisions live.

3. A deployed endpoint. Something the reader can call. Deployment forces you to confront timeouts, concurrency, key management, cold starts, and rate limits β€” all of which turn into conversation. A repository that only runs on your machine leaves the reader trusting your description of behaviour instead of observing it.

4. An eval set of 50 to 200 labelled cases, sourced from real inputs. This is the spine. Fifty carefully chosen cases with recorded expected behaviour, deliberately loaded with the awkward ones β€” ambiguous questions, questions your corpus genuinely cannot answer, multi-part questions, false premises, near-duplicate entities. Build it the way Evals describes: versioned in the repository, each case annotated with why it exists, with a held-out slice you only touch at the end.

5. A recorded before-and-after table. The baseline, then each deliberate change, then the delta on quality, p95 latency, and cost per request. This is the single highest-value object in the whole repository.

6. An honest write-up of what broke and what you rejected. Cluster your failures and report the clusters. Name the change you tried that made things worse, and say by how much. Nothing else in a portfolio buys as much credibility per sentence.

7. Cost and latency numbers. Cost per request, p50 and p95 end to end, and where the time actually goes. Most candidates cannot produce any of these, so having them is disproportionately differentiating.

8. One finished thing. Scoped small enough that everything above is actually true of it.


The Loop That Produces the Artifact

The table does not appear at the end as documentation. It is a byproduct of working in a loop.

flowchart LR
    A[Pick a real task you can judge] --> B[Build the simplest thing that works end to end]
    B --> C[Freeze an eval set of 50 to 200 labelled cases]
    C --> D[Record the baseline - quality and p95 and cost per request]
    D --> E[Make exactly one deliberate change]
    E --> F[Re-run the same suite and read the failures]
    F --> G[Write the delta into the results table - keep or revert]
    G --> E
    G --> H[Publish the README leading with the measured result]

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class A client
    class B,E service
    class C,D data
    class F async
    class G,H edge

Two properties of this loop do the work. The suite is frozen before the first change, so the comparison is honest. And exactly one change per iteration, so each row of the table attributes a delta to a cause. Bundle three changes together and you have measured their sum, which tells you nothing about which one to keep β€” and the interviewer will ask.

Here is the shape of the output. The numbers are invented to show the format, not results from a real system.

change                          quality    p95      cost/req   verdict
baseline - dense top 5           0.61     1.9 s      1.00x      -
+ hybrid retrieval and rerank    0.74     2.4 s      1.05x      kept
+ cheap model for easy routes    0.73     1.4 s      0.42x      kept
+ multi hop query decomposition  0.75     4.8 s      2.30x      reverted

Four lines. That is the entire portfolio, compressed. The reverted row is the most valuable one, because it shows a decision being made against a budget rather than a preference for the more sophisticated design.


Six Project Briefs

Increasing difficulty. Pick one and finish it. Each brief states what it demonstrates, the specific hard part, and the measurement that makes it credible.

1. Grounded question answering with verified citations

Answer questions over a corpus you know, where every claim carries a citation and the citation is checked against what was actually retrieved.

Demonstrates: the full retrieval path, grounding discipline, abstention behaviour. The hard part: not retrieval, but refusal. Making the system reliably say β€œthe corpus does not answer this” instead of producing a fluent paragraph is the real work, and it is what Hallucination, Grounding, and Citations exists to solve. Verifying that each cited span supports the sentence attached to it is harder still. The measurement: grounded answer rate, citation validity rate, and abstention rate on a deliberately unanswerable slice of your eval set. An abstention rate of zero means the system is bluffing.

2. Messy documents to validated structured records

Turn a pile of real documents into typed records that pass schema validation, with a dashboard of schema failures over time.

Demonstrates: Structured Outputs, parsing under mess, repair loops, field-level measurement. The hard part: partial extraction. A record where nine fields are right and the date is wrong is worse than a clean failure, because it passes validation and poisons downstream data. You need field-level scoring and a confidence path that routes uncertain records to review. The measurement: per-field precision and recall against a labelled set, schema failure rate before and after your repair strategy, and the share of records that needed human review.

3. An eval harness and judge for an existing open source AI app

Take a popular open source AI application, build the eval suite it does not have, and validate your judge against your own human labels.

Demonstrates: the skill employers screen hardest for, applied to somebody else’s code. The hard part: the judge. An unvalidated judge produces authoritative-looking numbers that mean nothing. You label a couple of hundred cases by hand, measure agreement between your labels and the judge, and iterate the rubric until agreement is defensible β€” the process in LLM-as-Judge. The measurement: judge-to-human agreement, plus the failure clusters your suite found in a project that shipped without one. A submitted issue or pull request makes this concrete.

4. A cost and latency optimisation study

Take one working feature from a single frontier model call to a routed cascade, and measure what the change actually cost you in quality.

Demonstrates: Model Selection and Routing, caching, budget thinking β€” the cluster most candidates skip entirely. The hard part: routing without quality collapse. Deciding which requests are easy is itself a classification problem, and the interesting finding is usually the quality you gave up on the hard tail rather than the money you saved on the easy head. The measurement: cost per request and p95 before and after, with the quality delta on the same frozen suite, reported per route rather than only in aggregate.

5. A bounded tool-using agent

An agent with a hard step budget, idempotent tools, and a documented failure taxonomy.

Demonstrates: Tool Calling, loop control, blast-radius thinking. The hard part: bounding it. Unbounded agents loop, burn budget, and occasionally take an action twice. The engineering is the step budget, the idempotency keys on every mutating tool, the retry and backoff policy, and the confirmation gate on anything irreversible. The measurement: task completion rate, median steps per task, budget-exhaustion rate, and a taxonomy of failures counted by category β€” wrong tool, malformed arguments, premature stop, loop.

6. A retrieval quality study

One corpus, three retrieval strategies β€” dense, hybrid, hybrid plus reranking β€” measured properly.

Demonstrates: that you understand retrieval quality caps answer quality, and that you can measure a component in isolation. The hard part: building the relevance labels. You need per-query judgements of which chunks are actually relevant, which is slow manual work, and it is exactly the work that makes the result trustworthy. Then keeping generation frozen so the retrieval delta is not confounded. The measurement: recall at k and precision at k per strategy, rank of the first relevant chunk, added latency and cost per strategy, and the end-to-end answer quality delta. Details in Hybrid Search and Reranking.


Project to Skills to Measurement

Project Skills demonstrated The measurement that makes it credible
Grounded question answering with citations Retrieval, grounding, abstention, eval design Grounded answer rate, citation validity, abstention on an unanswerable slice
Documents to validated records Structured outputs, parsing, repair loops Per-field precision and recall, schema failure rate before and after
Eval harness and validated judge Evals, judge calibration, error analysis Judge-to-human agreement, failure clusters found
Cost and latency optimisation study Routing, caching, budget discipline Cost per request and p95 before and after, with quality delta per route
Bounded tool-using agent Tool calling, loop control, idempotency Completion rate, steps per task, failure taxonomy counts
Retrieval quality study Embeddings, hybrid search, reranking, component isolation Recall at k and precision at k per strategy, with latency and cost added

Every row in the right-hand column is something a reader can check. That is the test for whether a project belongs in a portfolio at all.


The Write-Up Is the Interview Artifact

Most of your readers will read the README and never run the code. Write it accordingly.

Lead with the measured result. First screen: what the system does in one sentence, and what the numbers say. Not architecture, not setup instructions, not a feature list. If a reader learns the headline result in ten seconds, they will keep reading.

Show the before-and-after table high up. It is the reason the project exists. Include the reverted row.

State what you rejected and why. β€œI tried multi-hop decomposition; quality moved by a point and p95 more than doubled, so I dropped it” is a complete demonstration of engineering judgement in one sentence. Candidates systematically hide this, believing rejected work looks like failure. It reads as the opposite.

Describe the failure clusters. β€œForty percent of remaining failures are the retriever surfacing a superseded version of a policy document” tells a reader you did error analysis and know what you would build next. Interviewers often go straight at this, because it is the hardest thing to fake.

Be explicit about limitations. Corpus size, the domains you did not cover, the load you never tested, the eval set’s biases, what you would need before real users touched it. Stating limits raises credibility rather than lowering it, and it pre-empts the interviewer’s best attack.

Say how to reproduce it. The eval suite should run with one command against the committed golden set. If a reader can reproduce your baseline, your numbers stop being claims.

Include one trace. A single annotated end-to-end trace β€” input, retrieved chunks, assembled prompt, output, latency breakdown, token counts β€” teaches a reader more about your engineering than the architecture diagram. It also shows you instrumented the thing, which is the point of Tracing and Observability for LLM Apps.


What to Leave Out

Fake or unmeasured metrics. Never write a number you did not compute. β€œNinety-five percent accuracy” with no suite behind it is the fastest way to lose an interview, because the follow-up is always β€œon what set, graded how” and there is no recovery.

Unverifiable claims. β€œProduction ready”, β€œenterprise grade”, β€œhighly scalable” with no load test. Strong readers treat these as noise at best.

Half-finished features. A menu item that opens an empty page costs you more than the feature would have earned. Cut it and say the scope was deliberate.

Framework tours. Wiring four orchestration libraries together demonstrates configuration, not engineering. One library you can explain mechanically beats four you imported.

Leaderboard scores as your own evidence. Public benchmark numbers describe a model somebody else trained. They say nothing about your system.

Secrets and other people’s data. Committed keys are an instant negative signal, and real customer records in a public repository is worse than a missing project. Synthesise or redact, and say which.


Bad to Good to Great

Bad - the tutorial clone

A chat interface over some documents, built by following a video, deployed nowhere, no eval set, README listing features. It works on the inputs you tried. It dies on the first question about reliability, and it makes every other claim on your resume less believable.

Good - one real system, end to end, deployed

Your own domain, real messy documents, a live endpoint, a README that explains the architecture and the decisions. This is genuinely respectable and puts you ahead of most submissions. Its ceiling is specific: every quality claim in it rests on your word. When the interviewer asks whether your chunking change helped, the honest answer is still β€œit seemed to.”

Great - the same system with a recorded baseline and measured changes

Add the spine: a versioned eval set of 50 to 200 labelled cases from real inputs, a held-out slice, a recorded baseline across quality and p95 and cost per request, three deliberate changes each measured in isolation, one of them reverted with the reason stated, a clustered failure analysis, an annotated trace, and a README that leads with the result and closes with the limitations.

The distance between Good and Great is perhaps two weekends of unglamorous work, and it changes what the project is. Good is a description of a system. Great is evidence about an engineer. Only one of those survives a depth conversation, which is exactly what The AI Engineer Interview is built to prepare you for.


When to Use

βœ… Build a measured portfolio project when:

❌ Do not invest here when:


Common Interview Questions

Q1: Walk me through your project. What did you build?

I would lead with the task and the result rather than the architecture. Something like: the system answers questions over a corpus of regulations I know well, every answer carries a verified citation, and it abstains when the corpus cannot support an answer. I built a 120-case eval set from real questions, recorded a baseline, then made three changes - hybrid retrieval with reranking, a cheap model on the easy route, and query decomposition. The first two improved quality and cost and I kept them; decomposition moved quality by roughly a point while more than doubling p95, so I reverted it. Then I would offer the biggest remaining failure cluster, because that is usually where the real conversation starts.

Q2: How do you know your project actually works?

Because I can re-run the measurement in front of you. There is a versioned eval set in the repository, sourced from real inputs and deliberately weighted toward the awkward cases - ambiguous questions, questions the corpus cannot answer, near-duplicate entities. Deterministic checks cover schema validity, citation validity against what was actually retrieved, refusal on the must-refuse cases, and latency and cost ceilings. Open-ended quality goes through a rubric judge I validated against my own labels. I kept a held-out slice I only ran at the end, so the numbers are not tuned to the set I iterated on.

Q3: What was the hardest problem you hit, and how did you solve it?

The honest answer is usually retrieval rather than anything about the model. In my case near-identical documents from different years, where the retriever would confidently surface a superseded version and the model would faithfully answer from it - so the generation looked grounded and was still wrong. I found it by reading failures rather than watching the score: sampled about thirty, wrote a line on each, and one cluster was almost half of them. The fix was metadata filtering on effective date plus reranking, and I verified it by scoring retrieval separately from generation so I could see recall move without confounding it with prompt changes.

Q4: How much does your system cost to run, and how fast is it?

I track cost per request and p50 and p95 end to end, and I can break the latency down by stage - embedding, retrieval, reranking, generation. Reranking added meaningful latency and I kept it because the quality delta justified it against my budget. Routing easy requests to a smaller model cut cost per request by well over half with a quality delta small enough to accept, and I can show that per route rather than only in aggregate, because the aggregate hid a real drop on the hard tail. Treating latency and cost as correctness criteria rather than afterthoughts was the thing that changed the design most.

Q5: If you had another two weeks on this project, what would you do?

I would work the largest failure cluster rather than add a feature, and I would say which cluster and why. After that, two things the current version is weak on. First, the eval set is mine alone, so it encodes my idea of a correct answer - I would get a second person to label a slice and measure how often we disagree, because that bounds how much my scores can be trusted. Second, I have never run it under concurrency, so my p95 is a single-user number; I would load it and expect to find the retrieval tier, not the model, as the first constraint. Adding a third capability would be the wrong call while the measurement is the weak part.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access