Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
โฑ๏ธ 22 min read

Tracing and Observability for LLM Apps - Complete Deep Dive

Prerequisites: LLM APIs and SDKs, Observability, Evals Used in: Tool Calling Fundamentals, Latency Engineering, Deploying AI Features Build it: Lesson 9 - Tracing - Spans, Replay, Redaction implements this as runnable, tested code you can execute offline.


What is LLM Observability?

LLM observability is the practice of recording one user action as a tree of timed, causally-linked steps โ€” a trace โ€” rather than as a line in a log file, so that a failure you cannot reproduce can still be examined after the fact.

The distinction is not pedantic. In an ordinary CRUD service a request really is a request: one handler, one query, one response, and a log line naming the route, status, and duration captures most of what happened. An LLM feature is not shaped like that. One click fans out into an embedding call, a vector search, a rerank, a first model call, a tool call your own code executed, a second model call to interpret the tool result, a schema validation, and a retry because the first attempt emitted unparseable JSON. POST chat 200 4.2s describes all of that equally well, which is another way of saying it describes none of it.

Real-world analogy: a taxi receipt versus the GPS track. The receipt confirms the trip happened, what it cost, and how long it took. When the passenger complains the route was wrong, the receipt is useless, because every good route and every bad route produces an identical receipt. The GPS track shows the driver missed a turn early and then spent eight minutes on a road they should never have been on. Ordinary request logging is the receipt. A trace is the track.

So the unit of observability shifts. Not the request โ€” the trace, with nested spans, where a span is one step with a start time, a duration, a parent, and a payload. Standard distributed-tracing vocabulary carries over directly from Observability; what changes is that the interesting payload is no longer just a status code, it is text, and a lot of it.


Why Ordinary Request Logging Fails Here

Property of a normal service How an LLM feature breaks it
One request maps to one unit of work One request fans out into retrieval, reranking, several model calls, tool calls, and retries
Failures return non-200 A failure returns 200 with a fluent, confident, wrong answer
The same input reproduces the bug Sampling is stochastic - the same input may succeed on the next run
Duration is one number Duration splits into time-to-first-token and total, and users perceive the first one
Cost is roughly fixed per request Cost varies by an order of magnitude with prompt length and output length
The log line fits the failure The failure is a paragraph of text and the twenty chunks that produced it

The last row is the crux. When someone reports โ€œit gave a wrong answer to this question,โ€ the useful artifact is not a status code. It is the rendered prompt that actually went over the wire, the chunks that were retrieved and their scores, the model and sampling parameters in force, and the exact bytes that came back. If you did not record those, the report is unactionable and the conversation ends at โ€œI cannot reproduce it.โ€


The Shape of One Trace

Here is a single user question through a retrieval-backed assistant with one tool. Durations are illustrative, not measurements.

flowchart TD
    R[Root span - answer question - total 2400 ms] --> A[Span - embed query - 40 ms]
    R --> B[Span - vector search top 20 - 60 ms]
    R --> C[Span - rerank to top 5 - 180 ms]
    R --> D[Span - model call one - TTFT 320 ms]
    D --> E[Span - tool call - get order status - 240 ms]
    R --> F[Span - model call two - TTFT 280 ms]
    R --> G[Span - validate schema - 5 ms]
    G --> H[Span - retry model call two after invalid JSON]

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class R client
    class A,C,G service
    class B data
    class D,F,H async
    class E edge

Read the structure, not the numbers. The retry hangs off validation, so you can count retries as a first-class event instead of inferring them from duplicate log lines. The tool call is a child of the model call that requested it, so โ€œwhich model call caused this database writeโ€ is a lookup rather than a reconstruction. And every span carries a session_id and trace_id, so a multi-turn complaint becomes a query.


What to Capture Per Span

Be specific, because the fields teams skip are the fields they later need.

Field group Record Why you will want it
Identity trace_id, span_id, parent_span_id, session_id, user_id or a pseudonym, feature, tenant_id Grouping, per-customer cost, multi-turn reconstruction
Inputs The raw user input and any structured parameters The failure often starts here
Prompt The rendered prompt as sent, plus prompt_version A template plus variables is not enough - rendering is where bugs hide
Model model_id, pinned version identifier, provider, region A silent provider-side model update is invisible without this
Sampling temperature, top-p, max tokens, stop sequences, seed if supported Non-reproducibility is often just an unrecorded parameter
Retrieval Retrieved chunk ids, scores, rank, index snapshot version, filters applied Distinguishes a retrieval failure from a generation failure
Tools Tool name, arguments as sent, result or error, duration Tool misuse is a distinct failure class
Output The raw completion bytes, finish_reason finish_reason of length truncation explains a whole family of bugs
Accounting Prompt tokens, completion tokens, cached tokens, computed cost Per-feature and per-customer cost attribution
Timing Queue time, time-to-first-token, total duration, tokens per second Users feel TTFT - dashboards usually show total
Verdicts Schema valid, citations verified, guardrail triggered, sampled grade Turns a trace store into a quality dataset

A span record, with illustrative values:

{
  "trace_id": "8f2c1a",
  "span_id": "c4",
  "parent_span_id": "root",
  "span_type": "llm_call",
  "session_id": "s-4471",
  "feature": "support_answer",
  "prompt_version": "support_answer@v17",
  "model_id": "vendor-model-family",
  "model_version_pinned": "2026-04-01",
  "params": { "temperature": 0.2, "top_p": 1, "max_tokens": 800 },
  "index_snapshot": "kb-2026-04-12",
  "retrieved": [{ "chunk_id": "policy-refunds-v3#2", "score": 0.81, "rank": 1 }],
  "finish_reason": "stop",
  "prompt_tokens": 2140,
  "completion_tokens": 190,
  "cached_prompt_tokens": 1800,
  "ttft_ms": 320,
  "total_ms": 1180,
  "schema_valid": true,
  "citations_verified": true,
  "payload_ref": "blob://traces/8f2c1a/c4"
}

Two deliberate choices in that record. The heavy text lives behind payload_ref rather than inline, so metrics queries stay cheap and the sensitive part has its own lifecycle. And prompt_version is an explicit string, because โ€œwhich prompt produced thisโ€ is the first question of every investigation and reconstructing it from a deploy timestamp is guesswork.


Replayability Is the Whole Point

Say it plainly: a trace you cannot replay is an anecdote. Non-determinism means you cannot go back and make the bug happen again, so the recording has to be complete enough to stand in for the bug itself. The trace is the bug report.

Replay means you can take a stored trace and re-execute any single span against a new prompt, a new model, or a new retrieval configuration, holding everything else fixed. That capability is what makes the following possible:

Practical requirement: to replay, you need the index snapshot version, not only the chunk ids. If the corpus has been reindexed since, the same ids may point at different text and your replay silently tests a different system.


Token and Cost Accounting

Cost belongs on the same dashboard as latency, at the same granularity, reviewed by the same person. Teams that split them discover a tenfold spend increase at the end of the billing period rather than in the deploy that caused it.

Record tokens per span and roll up three ways:

Derive cost per successful request, not cost per request. A request that failed validation and retried twice cost three calls to produce one answer, and the retry-inflated figure is the one that reflects reality. Also track the ratio of prompt tokens to completion tokens: a prompt-heavy ratio drifting upward is the signature of context bloat, which usually means a retrieval change quietly raised top_k or someone appended a few more few-shot examples. The arithmetic behind all of this is in Tokens, Context Windows, and Cost Math.


The Metrics That Matter

Metric Read it as
p50 and p95 time-to-first-token Perceived responsiveness for a streaming UI - the number users actually feel
p50 and p95 total duration Whether the full answer arrives inside the taskโ€™s budget
Tokens per request, split prompt and completion Context bloat and output verbosity, the two cost drivers
Cost per successful request Unit economics including the waste from retries
Cache hit rate Whether the prompt prefix is stable enough to be reused
Schema validation failure rate Contract health - a rising rate often means a model or prompt change
Tool error rate per tool Which integration is flaky, and whether the model is calling it wrongly
Retry rate Hidden cost and hidden latency, plus a leading indicator of provider trouble
Abstention rate Zero means the system is bluffing - a spike means retrieval broke
Escalation rate The most honest quality signal you have, because the user voted with their feet

Use percentiles, not means. A mean happily hides a tail where one request in twenty takes fifteen seconds, and in a streaming interface the tail is the whole user experience. The general percentile discipline is in Performance Metrics; the tactics for moving these numbers are in Latency Engineering.


Online Quality Without Blocking the Request

Offline evals measure what you thought to ask. Traces let you measure what users actually sent โ€” but you cannot grade every response inline, because grading costs a model call and would add its latency to every request.

The pattern: sample asynchronously. A small percentage of completed traces is enqueued, graded out of band by a validated judge and by cheap deterministic checks, and the verdict is written back onto the trace. Nothing blocks. Then:

  1. The graded stream becomes a continuous online quality metric, segmented by feature, prompt version, and model version.
  2. Low-scoring traces are queued for human review rather than found by chance.
  3. Reviewed failures are promoted into the versioned golden set, where they stay forever.

Bias the sample rather than taking it uniformly: over-sample sessions with a thumbs-down, a regeneration, a heavy edit before acceptance, an escalation, a guardrail trigger, or a schema failure. Uniform sampling of mostly-fine traffic spends your grading budget confirming that easy requests are easy.


Drift Detection on Both Sides

Quality can fall without a single line of your code changing. Three causes, and you need signals for all of them.

Input drift. Users start asking different things โ€” a new product launched, a season turned, a competitorโ€™s outage sent you unfamiliar traffic. Monitor the distribution of inputs, not just the outputs: request volume by topic cluster, input length distribution, language mix, and the rate of questions whose best retrieval score falls below a floor. That last one is a genuinely good early warning, because it detects questions your corpus cannot answer before users tell you.

Corpus drift. The index changed. Documents were added, edited, or deleted, so the same question now retrieves different text. Version the index and record the snapshot on every span so a quality step-change lines up against a reindex.

Provider drift. The model behind a floating alias was updated, or its serving stack changed. You detect this only if you pin versions and record them, and only if you keep a small canary suite running on a schedule against production configuration. Without it, a provider-side change is indistinguishable from your own regression.


Alert on Ratios, Not Absolutes

Absolute thresholds break on normal traffic movement. Total spend rising is meaningless when usage doubled; error count rising is meaningless when request count rose with it. Alert on rates and ratios, which are stable under load changes:

One more that catches real incidents: alert on a change in the distribution of finish reasons. A rise in length-truncation means outputs are being cut off mid-answer, which users experience as the product being broken while every dashboard shows 200s.


PII and Retention - The Compliance Surface You Just Built

This is where an observability project becomes a data-protection project, and skipping it is how a well-intentioned tracing rollout turns into an incident.

Prompts and completions contain user data by construction. A support assistantโ€™s traces hold names, addresses, order numbers, and whatever the user pasted; a coding assistantโ€™s traces hold proprietary source; a health or finance assistantโ€™s traces hold the most regulated categories there are. Logging them wholesale means you have created a second copy of your most sensitive data, in a system that was designed for debugging convenience and is typically readable by every engineer on the team. Controls, in order of what they buy:


Bad to Good to Great

Bad - log the request and response

An access log line with route, status, and duration, plus perhaps the final answer printed to stdout.

When a user reports a wrong answer you have the answer and nothing that produced it. You cannot tell whether retrieval missed, the prompt changed, the model version moved, or a tool errored and the model improvised around it. Cost is a monthly surprise. Latency is one number that hides the only part users feel. And because the failure is not reproducible, the investigation ends at โ€œcannot reproduceโ€ โ€” which, repeated enough times, teaches the team that quality complaints are not actionable.

Good - structured logs per model call

JSON lines per call with prompt, response, token counts, model name, and duration, searchable in your existing log tool.

A real improvement: you can find a bad response and see its prompt. The limits are specific. There is no parent-child structure, so you cannot see that this call had a tool call under it or that a retry occurred, and reconstructing one user action means joining on a request id by hand. Retrieved chunk ids and scores are usually missing, so retrieval failures and generation failures remain indistinguishable. Model version is a floating alias, so a provider update is invisible. And prompts are stored in full, in the log tool, with the same access rules as everything else โ€” which is a compliance problem nobody has noticed yet.

Great - traces with spans, accounting, sampled grading, and a retention policy

  1. Trace per user action, spans for retrieve, rerank, each model call, each tool call, and each validation, with parent links and a session_id spanning turns.
  2. Full replay inputs recorded โ€” rendered prompt, prompt version, pinned model version, sampling parameters, chunk ids with scores, and the index snapshot version.
  3. Token and cost on every span, rolled up per feature, per model, and per customer, with cost per successful request on the latency dashboard.
  4. Percentile metrics on TTFT and total duration, plus cache hit rate, schema failure rate, tool error rate, retry rate, abstention rate, and escalation rate.
  5. Asynchronous sampled grading, biased toward negative signals, writing verdicts back onto traces and feeding failures into the versioned golden set.
  6. Drift monitors on the input distribution and the retrieval score floor, not only on output quality, with a scheduled canary suite to catch provider-side change.
  7. Ratio-based alerts with two-sided bounds where a collapse is as bad as a spike.
  8. Redaction at the edge, metadata and payloads stored separately, field-level access control, short payload retention, sampled capture, and a working deletion path.

The difference between Good and Great is not tooling spend. It is that in Great, a user complaint becomes a specific, replayable, permanently-tested case, and the data you collected to get there is governed.


When to Use

โœ… Invest in tracing when:

โŒ Keep it light when:


Common Interview Questions

Q1: Why is a log line per request not enough for an LLM feature?

Because one user action is not one unit of work. It fans out into embedding, retrieval, reranking, several model calls, tool calls your own code executed, validation, and retries, and a single line collapses all of that into a route, a status, and a duration. Worse, failures here mostly return 200 with fluent wrong content, so status codes carry almost no signal, and duration needs splitting into time-to-first-token and total because streaming users only feel the first. The unit of observability has to be a trace with nested spans, where the tool call is a child of the model call that requested it and the retry is a child of the validation that rejected the output. That structure is what lets you ask which component failed instead of guessing.

Q2: A user reports a wrong answer from last Tuesday. What do you need to have recorded?

Everything required to re-execute it, because the failure is stochastic and will not reproduce on demand. Concretely: the raw input, the rendered prompt as sent along with its prompt version, the model id with a pinned version rather than a floating alias, the sampling parameters, the retrieved chunk ids with their scores and the index snapshot version, any tool calls with their arguments and results, the raw completion, and the finish reason. With those I can replay generation against the frozen context to separate a retrieval failure from a generation failure, test a candidate fix against the actual failure rather than a paraphrase, and promote the trace into the golden set so it becomes a permanent regression test. Without them the report is an anecdote and the investigation ends at cannot reproduce.

Q3: How do you track LLM cost without waiting for the invoice?

Record prompt, completion, and cached token counts on every span, compute cost at write time, and roll it up per feature, per model, and per customer. Then put cost on the same dashboard as latency, because the two trade against each other and splitting them is how teams discover a tenfold increase a month late. The metric to alert on is cost per successful request, not total spend and not cost per request - a request that failed validation and retried twice consumed three calls to produce one answer, and only the per-success figure reflects that. I also watch the prompt-to-completion token ratio, since a ratio drifting prompt-heavy is the signature of context bloat, usually a raised retrieval top-k or extra few-shot examples that nobody costed.

Q4: What would you alert on, and why not on absolute numbers?

Absolute thresholds break the moment traffic moves - total spend and error counts both rise with usage, so they page you for growth and stay quiet during a real regression. I alert on ratios: schema validation failure rate, tool error rate broken out per tool so one flaky integration does not hide in a healthy aggregate, retry rate, cost per successful request, and p95 time-to-first-token against a defined budget. Abstention rate gets two-sided bounds, because a collapse to zero means the system started bluffing and a spike means retrieval broke. I would also alert on a shift in the distribution of finish reasons, since a rise in length truncation means answers are being cut off mid-sentence while every dashboard still shows 200s.

Q5: You want to log every prompt and completion for debugging. What is the objection?

That prompts and completions are user data by construction, so wholesale capture creates a second copy of your most sensitive content inside a system built for debugging convenience and usually readable by the whole team. The fix is not to stop tracing, it is to govern it. Redact high-risk fields at the edge before the payload is written, because once raw text lands it has been replicated and backed up. Split low-sensitivity metadata from the bulky payloads so they can carry different access rules and different lifetimes. Put raw payloads behind an explicit, audited, time-boxed grant while leaving metrics open. Retain payloads for days or weeks and metadata for much longer, since trend analysis only needs the metadata. Sample payload capture rather than taking everything, while always capturing errors, guardrail triggers, and negative-feedback sessions. And make sure a user deletion request actually reaches the trace store and anything derived from it.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access