Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 23 min read

Hallucination, Grounding, and Citations - Complete Deep Dive

Prerequisites: How LLMs Actually Work, RAG End to End, Evals Used in: LLM-as-Judge, Prompt Injection and the OWASP LLM Top 10, ChatGPT


What is Hallucination?

A hallucination is output that is fluent, confident, well-formed, and not true. It is not gibberish and not an error message. It is the model’s normal behaviour producing a claim with no basis in its training data or the context you supplied, delivered in exactly the same tone as a correct answer.

Real-world analogy: A student who has not read the book but writes a beautiful essay about it. They know the genre, the vocabulary, the shape a good answer takes, and the confident register examiners reward. Everything about the essay is right except its relationship to the book. Crucially, the student is not lying β€” they are doing the task they were actually trained for, which was β€œproduce text that reads like a good essay,” and they did it well.

That last sentence is the whole engineering problem, and it reframes what your job is. You are not hunting a defect. You are managing a known failure mode of the component you chose, the way you manage packet loss on a network or a cache miss on a hot key: with detection, bounds, fallbacks, and a budget.


Why It Is Intrinsic, Not a Bug

A language model is trained to predict plausible continuations of text. That objective is satisfied by a fluent wrong answer just as well as by a fluent right one. Nothing in the training signal distinguishes β€œthis claim is true” from β€œthis claim is the sort of thing that appears in text like this” β€” and for most of the corpus those coincide closely enough that the model becomes extremely good at the second while never learning the first as a separate skill.

Several consequences follow directly, and they explain most of what you will observe in production:

So the goal is not elimination. It is to make the failure detectable, bounded, visible to the user, and measured β€” and those are four different pieces of engineering.


The Types Worth Separating

Lumping every wrong answer into β€œhallucination” is why teams thrash. These five have different detection methods and different fixes, and knowing which one you have tells you what to build.

Fabricated facts. A claim with no basis anywhere β€” not in the retrieved context, not in the corpus, not in reality. Most common when the question genuinely cannot be answered from what was provided and nothing permitted the model to say so.

Fabricated citations and URLs. A reference to a document, section, page, or link that does not exist, or that exists but was never retrieved for this request. The most embarrassing type, because it looks like rigour β€” and the easiest to catch, because it is mechanically checkable.

Unsupported inference. Every individual fact appears in the context, but the conclusion goes beyond what the context supports: one example generalised into a rule, correlation stated as cause, a gap between two documents bridged by invention. This is the type that survives naive grounding checks, because a keyword overlap test passes while the logic does not.

Conflation of similar entities. Two people with the same surname, two product tiers, two policy versions, two API endpoints with near-identical names get merged into one answer that is individually sourced and collectively wrong. Usually a retrieval precision problem wearing a generation costume.

Outdated knowledge. A superseded policy, deprecated parameter, or old price recalled from pretraining, or retrieved from a stale index. The answer was true once, which makes it unusually convincing to reviewers who remember it being true.

Type, detection, mitigation

Type How it shows up Detection method Mitigation
Fabricated facts Confident specifics with no source; often appears when the corpus does not cover the question Verify each claim against retrieved text; faithfulness judge; cross-check structured claims against a system of record Retrieval grounding, forced citation, explicit permission to abstain, relevance threshold before generating
Fabricated citations and URLs Cites a chunk id, page, or link that was never retrieved or does not resolve Mechanical - resolve every identifier against the retrieved set and the chunk store, and reject anything unresolvable Restrict citations to an allowlist of retrieved chunk ids, validate server-side before rendering, never render an unresolved link
Unsupported inference Each fact is present, the conclusion is not Span-level entailment check per claim; a judge rubric with an explicit no-overreach criterion Claim-level attribution, constrained extraction, instructions to answer only what the context states and to flag gaps
Entity conflation Facts from two similar entities blended into one answer Check that all retrieved chunks refer to the same entity; audit retrieval precision on name-collision cases Metadata filters on entity id, reranking, and asking a clarifying question when candidates are genuinely ambiguous
Outdated knowledge A superseded policy, version, or price stated as current Compare the answer against the effective-dated source; monitor index lag; assert on version and date fields Recency filters at retrieval, effective dates stamped on every chunk with instructions to prefer the newest, propagating deletions, a published freshness SLA

The Mitigation Ladder

Weakest to strongest. Each rung costs more and buys more. Climb only as far as your stakes require, and know what each rung still leaves open.

Rung 1 - Better prompting with explicit permission to abstain

Tell the model, in the system prompt, that β€œI could not find this in the provided documents” is a correct and preferred answer. This sounds trivial and is disproportionately effective, because the default framing of a prompt implies the context is sufficient and an answer is expected. Also ask it to state what is missing rather than just refusing.

Still misses: everything, eventually. A prompt is a request, not a control. It reduces the rate and cannot bound it.

Rung 2 - Retrieval grounding

Put the actual source material in the context so the model is doing reading comprehension instead of recall. This is the single largest reduction available, and it changes the failure mode from β€œinvented from weights” to β€œmisread the page” β€” a much more tractable problem. See RAG End to End.

Still misses: the model can ignore the context and answer from pretraining anyway, and retrieving the wrong passages grounds it in the wrong thing. Grounding without a relevance threshold just gives confident answers a plausible-looking backdrop.

Rung 3 - Forced citation with span-level attribution

Require the answer to attach a source marker to each claim, referencing a specific chunk and ideally a specific span within it. Two effects: it makes the answer auditable by a human, and demanding attribution per claim discourages claims that have no source.

Still misses: unverified citations. A model will attach a marker to a chunk that does not contain the claim, which produces the appearance of grounding. Rung 3 alone can make a system less safe by making unsupported answers look sourced.

Rung 4 - Programmatic verification of every citation

Check mechanically that each cited identifier was actually retrieved, and that the cited span actually supports the attached claim. A failed check blocks, strips, or flags β€” it does not warn in a log nobody reads. This is the rung that turns citations from decoration into a control, and it is covered in detail below.

Still misses: correctly-cited claims drawn from a source that is itself wrong, and unsupported inference that stitches together two genuinely cited facts.

Rung 5 - Constrained extraction

For the highest-risk fields, do not let the model paraphrase at all. It may only select and quote spans from the provided context, with the span offsets returned as structured output that you validate against the source text. If the returned quote is not a substring of the context, the response is invalid and never reaches the user. Paraphrase is where meaning drifts; removing paraphrase removes the drift. This pairs naturally with Structured Outputs.

Still misses: fluency and usefulness. Quote-only output is stiff and cannot synthesise across sources, so reserve it for the fields where exactness dominates readability.

Rung 6 - Human review

For output where a wrong answer is expensive or irreversible β€” clinical, legal, financial, safety-relevant, or anything published under your name β€” a person approves before it goes out. The AI drafts and cites; the human decides. Design the review surface so verification is fast: the claim and its cited span side by side, with anything that failed an automated check pre-flagged.

Still misses: throughput, and reviewer attention. Humans rubber-stamp when the queue is long and the last two hundred drafts were fine, so pre-flagging the suspicious subset is what keeps the rung real.


Citation Verification

This is the part most teams get wrong, so it gets its own section. A model will happily cite a document that does not contain the claim. Not occasionally as a glitch β€” routinely, as a natural consequence of generating plausible text. The citation marker is generated by the same process that generated the claim, so it carries no independent evidence.

A citation that has not been mechanically checked is a stronger failure mode than no citation at all, because it converts an unsupported claim into an apparently-sourced one and switches off the reader’s scepticism.

flowchart LR
    Q[User question] --> R[Retrieve and assemble context]
    R --> N[Nothing clears the relevance threshold]
    N --> AB[Abstain - state that the documents do not cover this]
    R --> G[Generate answer with claim level citation markers]
    G --> V1[Resolve every marker against the retrieved chunk ids]
    V1 --> V2[Check the cited span supports the attached claim]
    V2 --> P[Pass - render the answer with deep links]
    V2 --> FL[Fail - strip the claim or flag it or block the response]
    V1 --> FL
    FL --> AB

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class Q client
    class R,G service
    class V1,V2 edge
    class N,FL async
    class P client
    class AB data

Layer one is free and catches the worst cases. Every marker the model emits must resolve to a chunk id that was in the context for this request. Compare against the allowlist of retrieved ids and reject anything else. No model call, a few lines of code, and it eliminates the entire class of invented document references and hallucinated links. If your system renders a URL the model produced rather than one your retriever produced, fix that before anything else on this page.

Layer two checks support, not just existence. Resolving to a real chunk proves the chunk was retrieved, not that it says what the answer claims. Options, cheapest first: exact or near-exact span matching when the claim quotes the source; lexical overlap of key entities, numbers, and dates between claim and cited span, which catches gross mismatches; an entailment or natural-language-inference model scoring whether the span supports the claim; and a judge with a faithfulness rubric for the residual cases. Numbers and dates deserve special handling β€” they are high-consequence, easy to check exactly, and a favourite site of drift.

A failed check must have teeth. Decide the policy per surface and write it down. Strip the unsupported claim and return the rest. Flag it visibly to the user as unverified. Block the response and abstain. Route it to human review. All four are defensible; logging a warning and rendering the answer anyway is not, because it is indistinguishable from having no verification.

Make citations resolvable end to end. Keep, per chunk, a stable id, the source URI, and enough location information to deep-link β€” page, heading, or character offsets. A citation the reader cannot open in one click will not be checked by anyone, which means your grounding story rests entirely on the automated layer.

Instrument the verification rate. The share of responses with at least one failed citation check is one of the highest-signal metrics in the whole system. It moves when retrieval degrades, when a prompt changes, and when a model is swapped, often before any aggregate quality score notices.


Faithfulness Versus Correctness

These get conflated constantly, and they are genuinely different properties.

Faithfulness asks: is every claim supported by the provided context? It is a closed-book question, answerable with the context in hand, and it is what citation verification and faithfulness judges measure.

Correctness asks: is the claim true? That requires knowing the world, or at least trusting the source.

Β  Context is correct Context is wrong
Answer is faithful to context Correct and faithful - the target Faithful but wrong - the model did its job, your corpus is the bug
Answer is unfaithful to context Unfaithful and usually wrong - classic hallucination Unfaithful and accidentally right - undetectable luck, do not count on it

The top-right cell is the one worth internalising. An answer can be perfectly faithful to a retrieved document that is itself outdated, superseded, or simply incorrect. Every grounding technique on this page will pass it, because grounding measures the relationship between answer and source, not between source and reality.

That makes corpus quality a hallucination control. Stale runbooks, superseded policies, contradictory duplicates of the same document, and draft pages that were never deleted all become authoritative-looking hallucinations with a valid citation attached. Practical consequences: retire documents rather than leaving them indexed, stamp effective dates and prefer the newest when two chunks conflict, prefer a single canonical source per topic over five near-duplicates, and propagate deletions promptly. When you report a faithfulness metric, say plainly that it is faithfulness β€” not accuracy.


Weak Confidence Signals

Two signals get proposed as hallucination detectors. Both are weak, both are useful for triage, and neither is a truth score.

Token log-probabilities. Low average or minimum token probability across a span correlates loosely with shaky output. The caveat is fundamental: log-probabilities measure how predictable a token was given the model’s distribution, not whether it was true. A confidently-memorised falsehood scores high; a correct but unusually-phrased answer scores low. Not every provider exposes them, and they are not comparable across models. Use them to route the bottom slice of a distribution to verification, not to label anything as false.

Self-consistency across samples. Ask the same question several times at non-zero temperature and compare the answers. Stable factual content across samples suggests the claim is anchored; answers that name a different date each time are being reconstructed. This is one of the better cheap detectors available, especially for short factual claims, and it costs a multiple of your inference bill per checked response, so sample rather than applying it everywhere. It also fails in a specific way worth knowing: a consistently-memorised error is consistently wrong, so agreement is evidence of anchoring rather than of truth.

The caveat that governs both: fluency correlates poorly with accuracy, and models are frequently confident and wrong. Never surface a model-produced confidence claim β€” β€œI am highly confident” β€” as if it were calibrated. It is generated text about confidence, not a measurement of it. Verification against a source beats every intrinsic signal.


Designing the UX So Uncertainty Is Visible

Hiding uncertainty does not remove risk, it silently transfers the risk to the user. Interface decisions are hallucination controls.


Measuring a Hallucination Rate

Until it is a number with a budget, hallucination is an anecdote that surfaces in escalations. Make it a metric.

Pick the unit. Response-level (β€œdid this response contain at least one unsupported claim”) is easier to label and matches user experience. Claim-level (share of claims that fail verification) is more sensitive and better for diagnosis. Pick one, define it precisely, and do not silently switch.

Build the labelled set. Sample real production traffic, weighted toward the query types where the corpus is thin, and label with a judge validated against human labels, audited by a human on a slice. An unvalidated judge here produces a hallucination rate that is itself a hallucination.

Split the causes. Report unsupported-claim rate, failed-citation rate, and abstention rate separately. Together they are diagnostic: failed citations rising with a flat unsupported-claim rate points at attribution; a collapsing abstention rate means the system has started bluffing where it used to decline.

Set a budget and gate on it. An explicit target, a threshold that fails the build, and a rollback trigger in production. This is error-budget thinking applied to a non-deterministic component, and it is the same release discipline described in Deployment and Reliability.

Track it continuously, not per release. The corpus, the traffic, and the model all drift. Sampled asynchronous grading on live traffic, with per-request traces you can replay, is what turns a one-off audit into a monitored metric β€” see Tracing and Observability.


Bad to Good to Great

Bad - tell the model not to make things up

A system prompt line saying β€œdo not hallucinate” or β€œonly state facts,” and nothing else.

The model has no mechanism to comply. It cannot distinguish a fact it knows from a continuation it generated, so the instruction is a request for a capability it does not have. Nothing is detected, nothing is bounded, nothing is measured, and the first time anyone notices is when a user reports a confidently wrong answer.

Good - retrieval grounding with citations displayed

Retrieve relevant passages, instruct the model to answer only from them, render the citations it emits.

A large genuine improvement, and where most shipped RAG systems stop. The gap is specific and serious: the citations are unverified, so the system produces answers that look sourced whether or not they are, and the display of a citation actively suppresses the reader’s scepticism. Retrieval failure also has no separate handling, so when nothing relevant comes back the model still answers, now with irrelevant context to draw on.

Great - a verified grounding pipeline with a measured rate

  1. Relevance threshold before generating, with a real abstain path when nothing clears it.
  2. Claim-level citation markers restricted to an allowlist of chunk ids retrieved for this request.
  3. Mechanical resolution of every marker, server-side, with unresolvable markers never rendered.
  4. Support checking of the cited span β€” exact match and numeric checks where possible, entailment or a validated judge for the rest.
  5. A written failure policy per surface: strip, flag, block, or escalate. Never log-and-ship.
  6. Constrained extraction for high-risk fields, validated as substrings of the source.
  7. Uncertainty visible in the UI, with sources deep-linked and abstention designed as a legitimate answer.
  8. Corpus hygiene treated as a control β€” effective dates, retirement, deduplication, propagated deletions.
  9. A hallucination rate with a budget, split by cause, gated in CI and monitored on sampled live traffic.
  10. Human review on the high-stakes slice, with automated flags pre-sorting the queue.

One more control belongs in the picture: a retrieved document containing instructions is an attack surface, not just a fact source, and a successful injection produces output that is confidently wrong in a way that looks exactly like hallucination. Treat retrieved text as untrusted data β€” see Prompt Injection and the OWASP LLM Top 10.


When to Use

βœ… Invest in grounding and verification when:

❌ Do not over-engineer when:


Common Interview Questions

Q1: How do you stop an LLM from hallucinating?

You do not, and framing it as elimination is the wrong starting point. The training objective rewards plausible continuation, and a fluent wrong answer satisfies it as well as a fluent right one, so there is no internal truth state to switch on. What you do is manage it: ground answers in retrieved sources so the task becomes reading comprehension rather than recall, require claim-level citations, verify those citations mechanically against the text that was actually retrieved, give the model a permitted way to abstain and a relevance threshold that triggers it, expose uncertainty in the interface instead of hiding it, and track a hallucination rate against a budget. The goal is detectable, bounded, visible, and measured - not zero.

Q2: Your model cites a document that does not contain the claim. How do you catch that?

Two layers, and the first is free. Every citation marker the model emits must resolve to a chunk id that was in the context for this specific request - compare it against the allowlist of retrieved ids and reject anything else. That alone eliminates invented documents and hallucinated URLs, and it needs no model call. The second layer checks support rather than existence: exact span matching where the answer quotes, entity and number overlap between the claim and the cited span, then an entailment model or a validated judge for the residual cases. The part people skip is the consequence - a failed check has to strip, flag, block, or escalate, because an unverified citation is worse than none, since it makes an unsupported claim look sourced and switches off the reader’s scepticism.

Q3: What is the difference between faithfulness and correctness, and why does it matter?

Faithfulness asks whether every claim is supported by the provided context. Correctness asks whether the claim is true. They come apart in a specific and common way: an answer can be perfectly faithful to a retrieved document that is itself outdated or wrong, and every grounding technique will pass it, because grounding measures the answer against the source and never the source against reality. That has two practical consequences. Corpus quality becomes a hallucination control, so retiring stale documents, stamping effective dates, deduplicating near-identical pages, and propagating deletions all matter as much as prompt work. And when you report the metric, call it faithfulness rather than accuracy, because claiming the second from a measurement of the first is how a team convinces itself a broken knowledge base is a working system.

Q4: Can you use the model’s own confidence to detect hallucination?

Only as a weak triage signal, never as a verdict. Token log-probabilities measure how predictable a token was under the model’s distribution, not whether it was true, so a confidently-memorised falsehood scores high and a correct but oddly-phrased answer scores low; they also are not comparable across models and are not always exposed. Self-consistency across several samples is somewhat better - an answer that names a different date each run is being reconstructed rather than recalled - but a consistently-memorised error is consistently wrong, so agreement is evidence of anchoring, not truth. And a model’s stated confidence in prose is generated text about confidence, not a measurement of it. Use these to select which responses get verified; verify against an actual source to decide.

Q5: How would you measure a hallucination rate in production?

Define the unit first - response-level is easier to label and closer to user experience, claim-level is more sensitive - then hold it fixed. Sample real traffic, weighted toward query types where the corpus is thin, and label with a judge whose agreement against human labels you have actually measured, with a human auditing a slice, because an unvalidated judge yields a rate you cannot defend. Report the causes separately: unsupported-claim rate, failed-citation rate, and abstention rate. Read them together, since failed citations rising while unsupported claims stay flat points at attribution, and a falling abstention rate means the system started bluffing where it used to decline. Then give it a budget, fail the build on a regression, set a rollback trigger, and keep grading sampled live traffic so drift in the corpus or the model shows up as a trend rather than an escalation.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access