Hallucination, Grounding, and Citations - Complete Deep Dive
Prerequisites: How LLMs Actually Work, RAG End to End, Evals Used in: LLM-as-Judge, Prompt Injection and the OWASP LLM Top 10, ChatGPT
What is Hallucination?
A hallucination is output that is fluent, confident, well-formed, and not true. It is not gibberish and not an error message. It is the modelβs normal behaviour producing a claim with no basis in its training data or the context you supplied, delivered in exactly the same tone as a correct answer.
Real-world analogy: A student who has not read the book but writes a beautiful essay about it. They know the genre, the vocabulary, the shape a good answer takes, and the confident register examiners reward. Everything about the essay is right except its relationship to the book. Crucially, the student is not lying β they are doing the task they were actually trained for, which was βproduce text that reads like a good essay,β and they did it well.
That last sentence is the whole engineering problem, and it reframes what your job is. You are not hunting a defect. You are managing a known failure mode of the component you chose, the way you manage packet loss on a network or a cache miss on a hot key: with detection, bounds, fallbacks, and a budget.
Why It Is Intrinsic, Not a Bug
A language model is trained to predict plausible continuations of text. That objective is satisfied by a fluent wrong answer just as well as by a fluent right one. Nothing in the training signal distinguishes βthis claim is trueβ from βthis claim is the sort of thing that appears in text like thisβ β and for most of the corpus those coincide closely enough that the model becomes extremely good at the second while never learning the first as a separate skill.
Several consequences follow directly, and they explain most of what you will observe in production:
- There is no internal βI do not knowβ state to consult. The model produces a distribution over next tokens for every prompt, including ones it has no basis to answer. Abstaining is a learned behaviour layered on top, not a native capability, which is why it has to be explicitly permitted and rewarded.
- Fluency and accuracy are separate axes. Confident phrasing is a stylistic property of the output, not evidence about its content. Human intuition reads the two as correlated, which is precisely why hallucinations get through review.
- Specificity is generated, not retrieved. Plausible dates, version numbers, section references, and names are exactly the sort of thing that appears in authoritative text, so the model supplies them. A fabricated citation is not a different phenomenon from a fabricated fact; it is the same mechanism applied to a bibliography.
- Rarer facts are shakier. Claims well-attested across the corpus are reproduced reliably. Claims appearing once, or not at all, get reconstructed from the shape of adjacent knowledge.
- Post-training reduces the rate, it does not remove the mechanism. Alignment work makes models more willing to hedge and decline. It does not give them a truth oracle.
So the goal is not elimination. It is to make the failure detectable, bounded, visible to the user, and measured β and those are four different pieces of engineering.
The Types Worth Separating
Lumping every wrong answer into βhallucinationβ is why teams thrash. These five have different detection methods and different fixes, and knowing which one you have tells you what to build.
Fabricated facts. A claim with no basis anywhere β not in the retrieved context, not in the corpus, not in reality. Most common when the question genuinely cannot be answered from what was provided and nothing permitted the model to say so.
Fabricated citations and URLs. A reference to a document, section, page, or link that does not exist, or that exists but was never retrieved for this request. The most embarrassing type, because it looks like rigour β and the easiest to catch, because it is mechanically checkable.
Unsupported inference. Every individual fact appears in the context, but the conclusion goes beyond what the context supports: one example generalised into a rule, correlation stated as cause, a gap between two documents bridged by invention. This is the type that survives naive grounding checks, because a keyword overlap test passes while the logic does not.
Conflation of similar entities. Two people with the same surname, two product tiers, two policy versions, two API endpoints with near-identical names get merged into one answer that is individually sourced and collectively wrong. Usually a retrieval precision problem wearing a generation costume.
Outdated knowledge. A superseded policy, deprecated parameter, or old price recalled from pretraining, or retrieved from a stale index. The answer was true once, which makes it unusually convincing to reviewers who remember it being true.
Type, detection, mitigation
| Type | How it shows up | Detection method | Mitigation |
|---|---|---|---|
| Fabricated facts | Confident specifics with no source; often appears when the corpus does not cover the question | Verify each claim against retrieved text; faithfulness judge; cross-check structured claims against a system of record | Retrieval grounding, forced citation, explicit permission to abstain, relevance threshold before generating |
| Fabricated citations and URLs | Cites a chunk id, page, or link that was never retrieved or does not resolve | Mechanical - resolve every identifier against the retrieved set and the chunk store, and reject anything unresolvable | Restrict citations to an allowlist of retrieved chunk ids, validate server-side before rendering, never render an unresolved link |
| Unsupported inference | Each fact is present, the conclusion is not | Span-level entailment check per claim; a judge rubric with an explicit no-overreach criterion | Claim-level attribution, constrained extraction, instructions to answer only what the context states and to flag gaps |
| Entity conflation | Facts from two similar entities blended into one answer | Check that all retrieved chunks refer to the same entity; audit retrieval precision on name-collision cases | Metadata filters on entity id, reranking, and asking a clarifying question when candidates are genuinely ambiguous |
| Outdated knowledge | A superseded policy, version, or price stated as current | Compare the answer against the effective-dated source; monitor index lag; assert on version and date fields | Recency filters at retrieval, effective dates stamped on every chunk with instructions to prefer the newest, propagating deletions, a published freshness SLA |
The Mitigation Ladder
Weakest to strongest. Each rung costs more and buys more. Climb only as far as your stakes require, and know what each rung still leaves open.
Rung 1 - Better prompting with explicit permission to abstain
Tell the model, in the system prompt, that βI could not find this in the provided documentsβ is a correct and preferred answer. This sounds trivial and is disproportionately effective, because the default framing of a prompt implies the context is sufficient and an answer is expected. Also ask it to state what is missing rather than just refusing.
Still misses: everything, eventually. A prompt is a request, not a control. It reduces the rate and cannot bound it.
Rung 2 - Retrieval grounding
Put the actual source material in the context so the model is doing reading comprehension instead of recall. This is the single largest reduction available, and it changes the failure mode from βinvented from weightsβ to βmisread the pageβ β a much more tractable problem. See RAG End to End.
Still misses: the model can ignore the context and answer from pretraining anyway, and retrieving the wrong passages grounds it in the wrong thing. Grounding without a relevance threshold just gives confident answers a plausible-looking backdrop.
Rung 3 - Forced citation with span-level attribution
Require the answer to attach a source marker to each claim, referencing a specific chunk and ideally a specific span within it. Two effects: it makes the answer auditable by a human, and demanding attribution per claim discourages claims that have no source.
Still misses: unverified citations. A model will attach a marker to a chunk that does not contain the claim, which produces the appearance of grounding. Rung 3 alone can make a system less safe by making unsupported answers look sourced.
Rung 4 - Programmatic verification of every citation
Check mechanically that each cited identifier was actually retrieved, and that the cited span actually supports the attached claim. A failed check blocks, strips, or flags β it does not warn in a log nobody reads. This is the rung that turns citations from decoration into a control, and it is covered in detail below.
Still misses: correctly-cited claims drawn from a source that is itself wrong, and unsupported inference that stitches together two genuinely cited facts.
Rung 5 - Constrained extraction
For the highest-risk fields, do not let the model paraphrase at all. It may only select and quote spans from the provided context, with the span offsets returned as structured output that you validate against the source text. If the returned quote is not a substring of the context, the response is invalid and never reaches the user. Paraphrase is where meaning drifts; removing paraphrase removes the drift. This pairs naturally with Structured Outputs.
Still misses: fluency and usefulness. Quote-only output is stiff and cannot synthesise across sources, so reserve it for the fields where exactness dominates readability.
Rung 6 - Human review
For output where a wrong answer is expensive or irreversible β clinical, legal, financial, safety-relevant, or anything published under your name β a person approves before it goes out. The AI drafts and cites; the human decides. Design the review surface so verification is fast: the claim and its cited span side by side, with anything that failed an automated check pre-flagged.
Still misses: throughput, and reviewer attention. Humans rubber-stamp when the queue is long and the last two hundred drafts were fine, so pre-flagging the suspicious subset is what keeps the rung real.
Citation Verification
This is the part most teams get wrong, so it gets its own section. A model will happily cite a document that does not contain the claim. Not occasionally as a glitch β routinely, as a natural consequence of generating plausible text. The citation marker is generated by the same process that generated the claim, so it carries no independent evidence.
A citation that has not been mechanically checked is a stronger failure mode than no citation at all, because it converts an unsupported claim into an apparently-sourced one and switches off the readerβs scepticism.
flowchart LR
Q[User question] --> R[Retrieve and assemble context]
R --> N[Nothing clears the relevance threshold]
N --> AB[Abstain - state that the documents do not cover this]
R --> G[Generate answer with claim level citation markers]
G --> V1[Resolve every marker against the retrieved chunk ids]
V1 --> V2[Check the cited span supports the attached claim]
V2 --> P[Pass - render the answer with deep links]
V2 --> FL[Fail - strip the claim or flag it or block the response]
V1 --> FL
FL --> AB
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class Q client
class R,G service
class V1,V2 edge
class N,FL async
class P client
class AB data
Layer one is free and catches the worst cases. Every marker the model emits must resolve to a chunk id that was in the context for this request. Compare against the allowlist of retrieved ids and reject anything else. No model call, a few lines of code, and it eliminates the entire class of invented document references and hallucinated links. If your system renders a URL the model produced rather than one your retriever produced, fix that before anything else on this page.
Layer two checks support, not just existence. Resolving to a real chunk proves the chunk was retrieved, not that it says what the answer claims. Options, cheapest first: exact or near-exact span matching when the claim quotes the source; lexical overlap of key entities, numbers, and dates between claim and cited span, which catches gross mismatches; an entailment or natural-language-inference model scoring whether the span supports the claim; and a judge with a faithfulness rubric for the residual cases. Numbers and dates deserve special handling β they are high-consequence, easy to check exactly, and a favourite site of drift.
A failed check must have teeth. Decide the policy per surface and write it down. Strip the unsupported claim and return the rest. Flag it visibly to the user as unverified. Block the response and abstain. Route it to human review. All four are defensible; logging a warning and rendering the answer anyway is not, because it is indistinguishable from having no verification.
Make citations resolvable end to end. Keep, per chunk, a stable id, the source URI, and enough location information to deep-link β page, heading, or character offsets. A citation the reader cannot open in one click will not be checked by anyone, which means your grounding story rests entirely on the automated layer.
Instrument the verification rate. The share of responses with at least one failed citation check is one of the highest-signal metrics in the whole system. It moves when retrieval degrades, when a prompt changes, and when a model is swapped, often before any aggregate quality score notices.
Faithfulness Versus Correctness
These get conflated constantly, and they are genuinely different properties.
Faithfulness asks: is every claim supported by the provided context? It is a closed-book question, answerable with the context in hand, and it is what citation verification and faithfulness judges measure.
Correctness asks: is the claim true? That requires knowing the world, or at least trusting the source.
| Β | Context is correct | Context is wrong |
|---|---|---|
| Answer is faithful to context | Correct and faithful - the target | Faithful but wrong - the model did its job, your corpus is the bug |
| Answer is unfaithful to context | Unfaithful and usually wrong - classic hallucination | Unfaithful and accidentally right - undetectable luck, do not count on it |
The top-right cell is the one worth internalising. An answer can be perfectly faithful to a retrieved document that is itself outdated, superseded, or simply incorrect. Every grounding technique on this page will pass it, because grounding measures the relationship between answer and source, not between source and reality.
That makes corpus quality a hallucination control. Stale runbooks, superseded policies, contradictory duplicates of the same document, and draft pages that were never deleted all become authoritative-looking hallucinations with a valid citation attached. Practical consequences: retire documents rather than leaving them indexed, stamp effective dates and prefer the newest when two chunks conflict, prefer a single canonical source per topic over five near-duplicates, and propagate deletions promptly. When you report a faithfulness metric, say plainly that it is faithfulness β not accuracy.
Weak Confidence Signals
Two signals get proposed as hallucination detectors. Both are weak, both are useful for triage, and neither is a truth score.
Token log-probabilities. Low average or minimum token probability across a span correlates loosely with shaky output. The caveat is fundamental: log-probabilities measure how predictable a token was given the modelβs distribution, not whether it was true. A confidently-memorised falsehood scores high; a correct but unusually-phrased answer scores low. Not every provider exposes them, and they are not comparable across models. Use them to route the bottom slice of a distribution to verification, not to label anything as false.
Self-consistency across samples. Ask the same question several times at non-zero temperature and compare the answers. Stable factual content across samples suggests the claim is anchored; answers that name a different date each time are being reconstructed. This is one of the better cheap detectors available, especially for short factual claims, and it costs a multiple of your inference bill per checked response, so sample rather than applying it everywhere. It also fails in a specific way worth knowing: a consistently-memorised error is consistently wrong, so agreement is evidence of anchoring rather than of truth.
The caveat that governs both: fluency correlates poorly with accuracy, and models are frequently confident and wrong. Never surface a model-produced confidence claim β βI am highly confidentβ β as if it were calibrated. It is generated text about confidence, not a measurement of it. Verification against a source beats every intrinsic signal.
Designing the UX So Uncertainty Is Visible
Hiding uncertainty does not remove risk, it silently transfers the risk to the user. Interface decisions are hallucination controls.
- Show sources inline and make them open the exact span. One click to the sentence in the source document. The distance between a claim and its evidence determines whether anyone ever checks it.
- Distinguish grounded from ungrounded text. If part of the answer is synthesis beyond the sources, mark it. Uniform presentation of verified and unverified claims is a design decision to obscure the difference.
- Make βnot found in your documentsβ a real, well-designed answer. If abstaining looks like a product failure, your team will tune the system to stop doing it β and a bluffing system is worse than an honest gap. Name what was searched and suggest a refinement.
- Surface failed verification rather than quietly dropping it. βThis claim could not be verified against a sourceβ is information the user can act on.
- Avoid certainty theatre. Confident phrasing, a definitive tone, and a spinner that implies deliberation all raise trust without raising accuracy.
- Keep an escape hatch to the source of truth or a human, and treat the rate at which people take it as a quality signal β see Evals.
- Match the affordance to the stakes. A draft the user will edit tolerates far more uncertainty than a number they will act on directly. For actions rather than text, require confirmation before anything irreversible.
Measuring a Hallucination Rate
Until it is a number with a budget, hallucination is an anecdote that surfaces in escalations. Make it a metric.
Pick the unit. Response-level (βdid this response contain at least one unsupported claimβ) is easier to label and matches user experience. Claim-level (share of claims that fail verification) is more sensitive and better for diagnosis. Pick one, define it precisely, and do not silently switch.
Build the labelled set. Sample real production traffic, weighted toward the query types where the corpus is thin, and label with a judge validated against human labels, audited by a human on a slice. An unvalidated judge here produces a hallucination rate that is itself a hallucination.
Split the causes. Report unsupported-claim rate, failed-citation rate, and abstention rate separately. Together they are diagnostic: failed citations rising with a flat unsupported-claim rate points at attribution; a collapsing abstention rate means the system has started bluffing where it used to decline.
Set a budget and gate on it. An explicit target, a threshold that fails the build, and a rollback trigger in production. This is error-budget thinking applied to a non-deterministic component, and it is the same release discipline described in Deployment and Reliability.
Track it continuously, not per release. The corpus, the traffic, and the model all drift. Sampled asynchronous grading on live traffic, with per-request traces you can replay, is what turns a one-off audit into a monitored metric β see Tracing and Observability.
Bad to Good to Great
Bad - tell the model not to make things up
A system prompt line saying βdo not hallucinateβ or βonly state facts,β and nothing else.
The model has no mechanism to comply. It cannot distinguish a fact it knows from a continuation it generated, so the instruction is a request for a capability it does not have. Nothing is detected, nothing is bounded, nothing is measured, and the first time anyone notices is when a user reports a confidently wrong answer.
Good - retrieval grounding with citations displayed
Retrieve relevant passages, instruct the model to answer only from them, render the citations it emits.
A large genuine improvement, and where most shipped RAG systems stop. The gap is specific and serious: the citations are unverified, so the system produces answers that look sourced whether or not they are, and the display of a citation actively suppresses the readerβs scepticism. Retrieval failure also has no separate handling, so when nothing relevant comes back the model still answers, now with irrelevant context to draw on.
Great - a verified grounding pipeline with a measured rate
- Relevance threshold before generating, with a real abstain path when nothing clears it.
- Claim-level citation markers restricted to an allowlist of chunk ids retrieved for this request.
- Mechanical resolution of every marker, server-side, with unresolvable markers never rendered.
- Support checking of the cited span β exact match and numeric checks where possible, entailment or a validated judge for the rest.
- A written failure policy per surface: strip, flag, block, or escalate. Never log-and-ship.
- Constrained extraction for high-risk fields, validated as substrings of the source.
- Uncertainty visible in the UI, with sources deep-linked and abstention designed as a legitimate answer.
- Corpus hygiene treated as a control β effective dates, retirement, deduplication, propagated deletions.
- A hallucination rate with a budget, split by cause, gated in CI and monitored on sampled live traffic.
- Human review on the high-stakes slice, with automated flags pre-sorting the queue.
One more control belongs in the picture: a retrieved document containing instructions is an attack surface, not just a fact source, and a successful injection produces output that is confidently wrong in a way that looks exactly like hallucination. Treat retrieved text as untrusted data β see Prompt Injection and the OWASP LLM Top 10.
When to Use
β Invest in grounding and verification when:
- Users will act on the output without independently checking it
- The answer must come from a specific corpus rather than general knowledge
- A wrong answer is expensive, hard to reverse, or a compliance matter
- The output is published, sent to a customer, or attributed to your organisation
- Claims involve numbers, dates, policy conditions, or named entities
- You need to be able to answer βwhere did this come fromβ about any response
β Do not over-engineer when:
- The user is the immediate verifier β brainstorming, drafting, and ideation where they read and edit everything
- The task needs no external facts at all, such as reformatting or translating text the user supplied
- Creative generation is the point, where invention is the feature
- The output is already checked by something deterministic downstream, like code that must compile and pass tests
- You have not measured your current rate, in which case measure before adding machinery so you can tell whether it helped
Common Interview Questions
Q1: How do you stop an LLM from hallucinating?
You do not, and framing it as elimination is the wrong starting point. The training objective rewards plausible continuation, and a fluent wrong answer satisfies it as well as a fluent right one, so there is no internal truth state to switch on. What you do is manage it: ground answers in retrieved sources so the task becomes reading comprehension rather than recall, require claim-level citations, verify those citations mechanically against the text that was actually retrieved, give the model a permitted way to abstain and a relevance threshold that triggers it, expose uncertainty in the interface instead of hiding it, and track a hallucination rate against a budget. The goal is detectable, bounded, visible, and measured - not zero.
Q2: Your model cites a document that does not contain the claim. How do you catch that?
Two layers, and the first is free. Every citation marker the model emits must resolve to a chunk id that was in the context for this specific request - compare it against the allowlist of retrieved ids and reject anything else. That alone eliminates invented documents and hallucinated URLs, and it needs no model call. The second layer checks support rather than existence: exact span matching where the answer quotes, entity and number overlap between the claim and the cited span, then an entailment model or a validated judge for the residual cases. The part people skip is the consequence - a failed check has to strip, flag, block, or escalate, because an unverified citation is worse than none, since it makes an unsupported claim look sourced and switches off the readerβs scepticism.
Q3: What is the difference between faithfulness and correctness, and why does it matter?
Faithfulness asks whether every claim is supported by the provided context. Correctness asks whether the claim is true. They come apart in a specific and common way: an answer can be perfectly faithful to a retrieved document that is itself outdated or wrong, and every grounding technique will pass it, because grounding measures the answer against the source and never the source against reality. That has two practical consequences. Corpus quality becomes a hallucination control, so retiring stale documents, stamping effective dates, deduplicating near-identical pages, and propagating deletions all matter as much as prompt work. And when you report the metric, call it faithfulness rather than accuracy, because claiming the second from a measurement of the first is how a team convinces itself a broken knowledge base is a working system.
Q4: Can you use the modelβs own confidence to detect hallucination?
Only as a weak triage signal, never as a verdict. Token log-probabilities measure how predictable a token was under the modelβs distribution, not whether it was true, so a confidently-memorised falsehood scores high and a correct but oddly-phrased answer scores low; they also are not comparable across models and are not always exposed. Self-consistency across several samples is somewhat better - an answer that names a different date each run is being reconstructed rather than recalled - but a consistently-memorised error is consistently wrong, so agreement is evidence of anchoring, not truth. And a modelβs stated confidence in prose is generated text about confidence, not a measurement of it. Use these to select which responses get verified; verify against an actual source to decide.
Q5: How would you measure a hallucination rate in production?
Define the unit first - response-level is easier to label and closer to user experience, claim-level is more sensitive - then hold it fixed. Sample real traffic, weighted toward query types where the corpus is thin, and label with a judge whose agreement against human labels you have actually measured, with a human auditing a slice, because an unvalidated judge yields a rate you cannot defend. Report the causes separately: unsupported-claim rate, failed-citation rate, and abstention rate. Read them together, since failed citations rising while unsupported claims stay flat points at attribution, and a falling abstention rate means the system started bluffing where it used to decline. Then give it a budget, fail the build on a regression, set a rollback trigger, and keep grading sampled live traffic so drift in the corpus or the model shows up as a trend rather than an escalation.
Build it in code: Agentic AI Course · Fundamentals: Core Concepts