Tracing and Observability for LLM Apps - Complete Deep Dive
Prerequisites: LLM APIs and SDKs, Observability, Evals Used in: Tool Calling Fundamentals, Latency Engineering, Deploying AI Features Build it: Lesson 9 - Tracing - Spans, Replay, Redaction implements this as runnable, tested code you can execute offline.
What is LLM Observability?
LLM observability is the practice of recording one user action as a tree of timed, causally-linked steps โ a trace โ rather than as a line in a log file, so that a failure you cannot reproduce can still be examined after the fact.
The distinction is not pedantic. In an ordinary CRUD service a request really is a request: one handler, one query, one response, and a log line naming the route, status, and duration captures most of what happened. An LLM feature is not shaped like that. One click fans out into an embedding call, a vector search, a rerank, a first model call, a tool call your own code executed, a second model call to interpret the tool result, a schema validation, and a retry because the first attempt emitted unparseable JSON. POST chat 200 4.2s describes all of that equally well, which is another way of saying it describes none of it.
Real-world analogy: a taxi receipt versus the GPS track. The receipt confirms the trip happened, what it cost, and how long it took. When the passenger complains the route was wrong, the receipt is useless, because every good route and every bad route produces an identical receipt. The GPS track shows the driver missed a turn early and then spent eight minutes on a road they should never have been on. Ordinary request logging is the receipt. A trace is the track.
So the unit of observability shifts. Not the request โ the trace, with nested spans, where a span is one step with a start time, a duration, a parent, and a payload. Standard distributed-tracing vocabulary carries over directly from Observability; what changes is that the interesting payload is no longer just a status code, it is text, and a lot of it.
Why Ordinary Request Logging Fails Here
| Property of a normal service | How an LLM feature breaks it |
|---|---|
| One request maps to one unit of work | One request fans out into retrieval, reranking, several model calls, tool calls, and retries |
| Failures return non-200 | A failure returns 200 with a fluent, confident, wrong answer |
| The same input reproduces the bug | Sampling is stochastic - the same input may succeed on the next run |
| Duration is one number | Duration splits into time-to-first-token and total, and users perceive the first one |
| Cost is roughly fixed per request | Cost varies by an order of magnitude with prompt length and output length |
| The log line fits the failure | The failure is a paragraph of text and the twenty chunks that produced it |
The last row is the crux. When someone reports โit gave a wrong answer to this question,โ the useful artifact is not a status code. It is the rendered prompt that actually went over the wire, the chunks that were retrieved and their scores, the model and sampling parameters in force, and the exact bytes that came back. If you did not record those, the report is unactionable and the conversation ends at โI cannot reproduce it.โ
The Shape of One Trace
Here is a single user question through a retrieval-backed assistant with one tool. Durations are illustrative, not measurements.
flowchart TD
R[Root span - answer question - total 2400 ms] --> A[Span - embed query - 40 ms]
R --> B[Span - vector search top 20 - 60 ms]
R --> C[Span - rerank to top 5 - 180 ms]
R --> D[Span - model call one - TTFT 320 ms]
D --> E[Span - tool call - get order status - 240 ms]
R --> F[Span - model call two - TTFT 280 ms]
R --> G[Span - validate schema - 5 ms]
G --> H[Span - retry model call two after invalid JSON]
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class R client
class A,C,G service
class B data
class D,F,H async
class E edge
Read the structure, not the numbers. The retry hangs off validation, so you can count retries as a first-class event instead of inferring them from duplicate log lines. The tool call is a child of the model call that requested it, so โwhich model call caused this database writeโ is a lookup rather than a reconstruction. And every span carries a session_id and trace_id, so a multi-turn complaint becomes a query.
What to Capture Per Span
Be specific, because the fields teams skip are the fields they later need.
| Field group | Record | Why you will want it |
|---|---|---|
| Identity | trace_id, span_id, parent_span_id, session_id, user_id or a pseudonym, feature, tenant_id |
Grouping, per-customer cost, multi-turn reconstruction |
| Inputs | The raw user input and any structured parameters | The failure often starts here |
| Prompt | The rendered prompt as sent, plus prompt_version |
A template plus variables is not enough - rendering is where bugs hide |
| Model | model_id, pinned version identifier, provider, region |
A silent provider-side model update is invisible without this |
| Sampling | temperature, top-p, max tokens, stop sequences, seed if supported | Non-reproducibility is often just an unrecorded parameter |
| Retrieval | Retrieved chunk ids, scores, rank, index snapshot version, filters applied | Distinguishes a retrieval failure from a generation failure |
| Tools | Tool name, arguments as sent, result or error, duration | Tool misuse is a distinct failure class |
| Output | The raw completion bytes, finish_reason |
finish_reason of length truncation explains a whole family of bugs |
| Accounting | Prompt tokens, completion tokens, cached tokens, computed cost | Per-feature and per-customer cost attribution |
| Timing | Queue time, time-to-first-token, total duration, tokens per second | Users feel TTFT - dashboards usually show total |
| Verdicts | Schema valid, citations verified, guardrail triggered, sampled grade | Turns a trace store into a quality dataset |
A span record, with illustrative values:
{
"trace_id": "8f2c1a",
"span_id": "c4",
"parent_span_id": "root",
"span_type": "llm_call",
"session_id": "s-4471",
"feature": "support_answer",
"prompt_version": "support_answer@v17",
"model_id": "vendor-model-family",
"model_version_pinned": "2026-04-01",
"params": { "temperature": 0.2, "top_p": 1, "max_tokens": 800 },
"index_snapshot": "kb-2026-04-12",
"retrieved": [{ "chunk_id": "policy-refunds-v3#2", "score": 0.81, "rank": 1 }],
"finish_reason": "stop",
"prompt_tokens": 2140,
"completion_tokens": 190,
"cached_prompt_tokens": 1800,
"ttft_ms": 320,
"total_ms": 1180,
"schema_valid": true,
"citations_verified": true,
"payload_ref": "blob://traces/8f2c1a/c4"
}
Two deliberate choices in that record. The heavy text lives behind payload_ref rather than inline, so metrics queries stay cheap and the sensitive part has its own lifecycle. And prompt_version is an explicit string, because โwhich prompt produced thisโ is the first question of every investigation and reconstructing it from a deploy timestamp is guesswork.
Replayability Is the Whole Point
Say it plainly: a trace you cannot replay is an anecdote. Non-determinism means you cannot go back and make the bug happen again, so the recording has to be complete enough to stand in for the bug itself. The trace is the bug report.
Replay means you can take a stored trace and re-execute any single span against a new prompt, a new model, or a new retrieval configuration, holding everything else fixed. That capability is what makes the following possible:
- Localise a failure. Re-run generation against the frozen retrieved context. If it now answers correctly, retrieval was the problem, not the prompt.
- Test a fix against the actual failure. Not against a paraphrase you typed from memory.
- Promote the trace into the eval set. A trace already contains the input, the context, and the observed bad output, which is exactly a golden case. This is the loop described in Evals, and traces are the supply line.
- Evaluate a model migration on real traffic by replaying a week of traces through the candidate.
Practical requirement: to replay, you need the index snapshot version, not only the chunk ids. If the corpus has been reindexed since, the same ids may point at different text and your replay silently tests a different system.
Token and Cost Accounting
Cost belongs on the same dashboard as latency, at the same granularity, reviewed by the same person. Teams that split them discover a tenfold spend increase at the end of the billing period rather than in the deploy that caused it.
Record tokens per span and roll up three ways:
- Per feature. One feature is almost always the majority of spend, and it is often not the one anybody was optimising.
- Per model. Makes a cheap-model path silently falling back to an expensive one visible immediately.
- Per customer or tenant. In B2B this is the difference between a healthy account and one being served at a loss. It is also the input to any usage-based pricing you offer.
Derive cost per successful request, not cost per request. A request that failed validation and retried twice cost three calls to produce one answer, and the retry-inflated figure is the one that reflects reality. Also track the ratio of prompt tokens to completion tokens: a prompt-heavy ratio drifting upward is the signature of context bloat, which usually means a retrieval change quietly raised top_k or someone appended a few more few-shot examples. The arithmetic behind all of this is in Tokens, Context Windows, and Cost Math.
The Metrics That Matter
| Metric | Read it as |
|---|---|
| p50 and p95 time-to-first-token | Perceived responsiveness for a streaming UI - the number users actually feel |
| p50 and p95 total duration | Whether the full answer arrives inside the taskโs budget |
| Tokens per request, split prompt and completion | Context bloat and output verbosity, the two cost drivers |
| Cost per successful request | Unit economics including the waste from retries |
| Cache hit rate | Whether the prompt prefix is stable enough to be reused |
| Schema validation failure rate | Contract health - a rising rate often means a model or prompt change |
| Tool error rate per tool | Which integration is flaky, and whether the model is calling it wrongly |
| Retry rate | Hidden cost and hidden latency, plus a leading indicator of provider trouble |
| Abstention rate | Zero means the system is bluffing - a spike means retrieval broke |
| Escalation rate | The most honest quality signal you have, because the user voted with their feet |
Use percentiles, not means. A mean happily hides a tail where one request in twenty takes fifteen seconds, and in a streaming interface the tail is the whole user experience. The general percentile discipline is in Performance Metrics; the tactics for moving these numbers are in Latency Engineering.
Online Quality Without Blocking the Request
Offline evals measure what you thought to ask. Traces let you measure what users actually sent โ but you cannot grade every response inline, because grading costs a model call and would add its latency to every request.
The pattern: sample asynchronously. A small percentage of completed traces is enqueued, graded out of band by a validated judge and by cheap deterministic checks, and the verdict is written back onto the trace. Nothing blocks. Then:
- The graded stream becomes a continuous online quality metric, segmented by feature, prompt version, and model version.
- Low-scoring traces are queued for human review rather than found by chance.
- Reviewed failures are promoted into the versioned golden set, where they stay forever.
Bias the sample rather than taking it uniformly: over-sample sessions with a thumbs-down, a regeneration, a heavy edit before acceptance, an escalation, a guardrail trigger, or a schema failure. Uniform sampling of mostly-fine traffic spends your grading budget confirming that easy requests are easy.
Drift Detection on Both Sides
Quality can fall without a single line of your code changing. Three causes, and you need signals for all of them.
Input drift. Users start asking different things โ a new product launched, a season turned, a competitorโs outage sent you unfamiliar traffic. Monitor the distribution of inputs, not just the outputs: request volume by topic cluster, input length distribution, language mix, and the rate of questions whose best retrieval score falls below a floor. That last one is a genuinely good early warning, because it detects questions your corpus cannot answer before users tell you.
Corpus drift. The index changed. Documents were added, edited, or deleted, so the same question now retrieves different text. Version the index and record the snapshot on every span so a quality step-change lines up against a reindex.
Provider drift. The model behind a floating alias was updated, or its serving stack changed. You detect this only if you pin versions and record them, and only if you keep a small canary suite running on a schedule against production configuration. Without it, a provider-side change is indistinguishable from your own regression.
Alert on Ratios, Not Absolutes
Absolute thresholds break on normal traffic movement. Total spend rising is meaningless when usage doubled; error count rising is meaningless when request count rose with it. Alert on rates and ratios, which are stable under load changes:
- Schema validation failure rate, not failure count
- Tool error rate per tool, so one flaky integration does not hide inside a healthy aggregate
- Cost per successful request, not daily spend, and p95 TTFT against a defined budget rather than yesterdayโs mean
- Retry rate and the prompt-to-completion token ratio, both as trend alerts rather than hard thresholds
- Abstention rate with two-sided bounds, since a collapse to zero and a spike are both failures
One more that catches real incidents: alert on a change in the distribution of finish reasons. A rise in length-truncation means outputs are being cut off mid-answer, which users experience as the product being broken while every dashboard shows 200s.
PII and Retention - The Compliance Surface You Just Built
This is where an observability project becomes a data-protection project, and skipping it is how a well-intentioned tracing rollout turns into an incident.
Prompts and completions contain user data by construction. A support assistantโs traces hold names, addresses, order numbers, and whatever the user pasted; a coding assistantโs traces hold proprietary source; a health or finance assistantโs traces hold the most regulated categories there are. Logging them wholesale means you have created a second copy of your most sensitive data, in a system that was designed for debugging convenience and is typically readable by every engineer on the team. Controls, in order of what they buy:
- Redact at the edge. Detect and mask high-risk fields before the payload is written, not in a nightly cleanup job. Once raw text lands in the store it has been retained, replicated, and backed up. Keep a structural placeholder so the trace remains readable and replayable in shape.
- Split metadata from payloads. Metrics, ids, token counts, timings, and verdicts are low-sensitivity and cheap. The prompt and completion bodies are high-sensitivity and bulky. Store them separately, referenced by id, so they can carry different access rules and different lifetimes.
- Field-level access control on trace storage. Metrics open to the team; raw payloads behind an explicit grant, ideally time-boxed and audited. โWho read this customerโs conversation, and whyโ should be answerable.
- Asymmetric retention. Short retention for raw payloads โ days to weeks, long enough to debug. Long retention for metadata and aggregates, which is what you need for trend analysis anyway. There is no debugging reason to hold raw prompts for a year.
- Sample instead of capturing everything. Full payload capture on every request is rarely necessary and always the largest part of both the bill and the risk. Capture payloads on a percentage of traffic plus all errors, all guardrail triggers, and all negative-feedback sessions. You get the failures, which is the point, without a complete shadow copy of production.
- Honour deletion. A user data-deletion request has to reach the trace store, and any eval set derived from it. Store enough linkage to find the records; do not store so much that the linkage itself is the exposure.
- Never trace credentials. Tokens and keys must not be in the context, so they must not be in the payload either. See Prompt Injection, Jailbreaks, and the OWASP LLM Top 10.
Bad to Good to Great
Bad - log the request and response
An access log line with route, status, and duration, plus perhaps the final answer printed to stdout.
When a user reports a wrong answer you have the answer and nothing that produced it. You cannot tell whether retrieval missed, the prompt changed, the model version moved, or a tool errored and the model improvised around it. Cost is a monthly surprise. Latency is one number that hides the only part users feel. And because the failure is not reproducible, the investigation ends at โcannot reproduceโ โ which, repeated enough times, teaches the team that quality complaints are not actionable.
Good - structured logs per model call
JSON lines per call with prompt, response, token counts, model name, and duration, searchable in your existing log tool.
A real improvement: you can find a bad response and see its prompt. The limits are specific. There is no parent-child structure, so you cannot see that this call had a tool call under it or that a retry occurred, and reconstructing one user action means joining on a request id by hand. Retrieved chunk ids and scores are usually missing, so retrieval failures and generation failures remain indistinguishable. Model version is a floating alias, so a provider update is invisible. And prompts are stored in full, in the log tool, with the same access rules as everything else โ which is a compliance problem nobody has noticed yet.
Great - traces with spans, accounting, sampled grading, and a retention policy
- Trace per user action, spans for retrieve, rerank, each model call, each tool call, and each validation, with parent links and a
session_idspanning turns. - Full replay inputs recorded โ rendered prompt, prompt version, pinned model version, sampling parameters, chunk ids with scores, and the index snapshot version.
- Token and cost on every span, rolled up per feature, per model, and per customer, with cost per successful request on the latency dashboard.
- Percentile metrics on TTFT and total duration, plus cache hit rate, schema failure rate, tool error rate, retry rate, abstention rate, and escalation rate.
- Asynchronous sampled grading, biased toward negative signals, writing verdicts back onto traces and feeding failures into the versioned golden set.
- Drift monitors on the input distribution and the retrieval score floor, not only on output quality, with a scheduled canary suite to catch provider-side change.
- Ratio-based alerts with two-sided bounds where a collapse is as bad as a spike.
- Redaction at the edge, metadata and payloads stored separately, field-level access control, short payload retention, sampled capture, and a working deletion path.
The difference between Good and Great is not tooling spend. It is that in Great, a user complaint becomes a specific, replayable, permanently-tested case, and the data you collected to get there is governed.
When to Use
โ Invest in tracing when:
- Anything LLM-generated reaches users, at any volume โ this is day-one work, not a maturity milestone
- One user action triggers more than one model, retrieval, or tool call
- You are spending real money on inference and cannot say which feature spends it
- You are about to change a model, prompt, or retriever and want your eval set fed from reality
- Multiple people can modify prompts, so โwhich version produced thisโ has to be answerable
โ Keep it light when:
- A single-call, low-volume internal tool with no retrieval and no tools โ structured logs with token counts are proportionate
- You would be building payload capture before deciding redaction and retention, which builds the liability before the control
- The data is highly regulated and you have no redaction path yet โ start with metadata-only traces, which still give you cost, latency, and failure rates
- You are collecting traces nobody reads, in which case fix the review habit before widening the capture
Common Interview Questions
Q1: Why is a log line per request not enough for an LLM feature?
Because one user action is not one unit of work. It fans out into embedding, retrieval, reranking, several model calls, tool calls your own code executed, validation, and retries, and a single line collapses all of that into a route, a status, and a duration. Worse, failures here mostly return 200 with fluent wrong content, so status codes carry almost no signal, and duration needs splitting into time-to-first-token and total because streaming users only feel the first. The unit of observability has to be a trace with nested spans, where the tool call is a child of the model call that requested it and the retry is a child of the validation that rejected the output. That structure is what lets you ask which component failed instead of guessing.
Q2: A user reports a wrong answer from last Tuesday. What do you need to have recorded?
Everything required to re-execute it, because the failure is stochastic and will not reproduce on demand. Concretely: the raw input, the rendered prompt as sent along with its prompt version, the model id with a pinned version rather than a floating alias, the sampling parameters, the retrieved chunk ids with their scores and the index snapshot version, any tool calls with their arguments and results, the raw completion, and the finish reason. With those I can replay generation against the frozen context to separate a retrieval failure from a generation failure, test a candidate fix against the actual failure rather than a paraphrase, and promote the trace into the golden set so it becomes a permanent regression test. Without them the report is an anecdote and the investigation ends at cannot reproduce.
Q3: How do you track LLM cost without waiting for the invoice?
Record prompt, completion, and cached token counts on every span, compute cost at write time, and roll it up per feature, per model, and per customer. Then put cost on the same dashboard as latency, because the two trade against each other and splitting them is how teams discover a tenfold increase a month late. The metric to alert on is cost per successful request, not total spend and not cost per request - a request that failed validation and retried twice consumed three calls to produce one answer, and only the per-success figure reflects that. I also watch the prompt-to-completion token ratio, since a ratio drifting prompt-heavy is the signature of context bloat, usually a raised retrieval top-k or extra few-shot examples that nobody costed.
Q4: What would you alert on, and why not on absolute numbers?
Absolute thresholds break the moment traffic moves - total spend and error counts both rise with usage, so they page you for growth and stay quiet during a real regression. I alert on ratios: schema validation failure rate, tool error rate broken out per tool so one flaky integration does not hide in a healthy aggregate, retry rate, cost per successful request, and p95 time-to-first-token against a defined budget. Abstention rate gets two-sided bounds, because a collapse to zero means the system started bluffing and a spike means retrieval broke. I would also alert on a shift in the distribution of finish reasons, since a rise in length truncation means answers are being cut off mid-sentence while every dashboard still shows 200s.
Q5: You want to log every prompt and completion for debugging. What is the objection?
That prompts and completions are user data by construction, so wholesale capture creates a second copy of your most sensitive content inside a system built for debugging convenience and usually readable by the whole team. The fix is not to stop tracing, it is to govern it. Redact high-risk fields at the edge before the payload is written, because once raw text lands it has been replicated and backed up. Split low-sensitivity metadata from the bulky payloads so they can carry different access rules and different lifetimes. Put raw payloads behind an explicit, audited, time-boxed grant while leaving metrics open. Retain payloads for days or weeks and metadata for much longer, since trend analysis only needs the metadata. Sample payload capture rather than taking everything, while always capturing errors, guardrail triggers, and negative-feedback sessions. And make sure a user deletion request actually reaches the trace store and anything derived from it.
Build it in code: Agentic AI Course · Fundamentals: Core Concepts