Document Parsing and Chunking - Complete Deep Dive
Prerequisites: Embeddings, Vector Databases, Tokens and Cost Math Used in: RAG End to End, Hybrid Search and Reranking, Multimodal
What is Ingestion, and Why Does It Decide Everything?
Retrieval quality is usually lost at ingestion, not at search. This is the least glamorous stage of RAG and the one where most projects actually fail. Teams spend weeks tuning ef_search, swapping embedding models, and adding rerankers, when the real problem is that the answer was severed across two chunks at ingest and no retriever on earth can put it back.
Ingestion is the offline pipeline that turns a raw file into indexable units: detect the format, extract the text in correct reading order, reconstruct the documentβs structure, split it into chunks, attach metadata, and embed. Everything downstream inherits whatever this stage produces.
Real-world analogy: a library with a world-class catalogue and a brilliant reference librarian is still useless if a careless clerk tore every book into random 40-page bundles and reshelved them with no titles. The search desk is not the problem. The shelving was.
Parsing Reality: A PDF Has No Structure
This is the fact that surprises backend engineers most. PDF is a page-description format. It says βdraw glyph T at coordinates x, y in this font.β It does not say βthis is a headingβ, βthis is a table cellβ, βthis paragraph continues from the previous pageβ, or even βread this column first.β Any structure you get from a PDF was inferred by your extractor from glyph positions, and inference fails in specific, repeatable ways:
- Multi-column layouts read out of order. A naive extractor walks glyphs roughly left to right, top to bottom, splicing line 1 of column A with line 1 of column B. The output is grammatically plausible and semantically scrambled β the worst kind of corruption because nothing crashes.
- Tables flatten into token soup. Row and column relationships live entirely in coordinates. Lose them and
Region Q1 Q2 North 412 380 South 290 355is the embedding of nothing in particular, and an LLM reading it will confidently misattribute numbers. - Scanned documents have no text layer at all. Extraction returns empty strings or a handful of stray characters. You need OCR, and OCR introduces its own error rate on tables, handwriting, stamps, and poor scans.
- Headers and footers pollute every chunk. Page numbers, confidentiality notices, and running titles get interleaved mid-sentence at every page boundary. At scale the same footer appears in thousands of chunks, which also wrecks deduplication and crowds result lists.
- Line-level artefacts. Words hyphenated across line breaks, ligatures extracted as odd codepoints, soft hyphens, and non-breaking spaces all silently change tokenization.
Detect this early with a blunt but effective check: sample extracted documents and eyeball the raw text. A short, mechanical review of thirty real documents finds more retrieval bugs than a week of parameter tuning.
| Document type | Parsing hazard | Approach |
|---|---|---|
| Text-layer PDF, single column | Hyphenation, ligatures, running headers and footers | Standard text extraction; strip repeated lines that appear on most pages |
| Text-layer PDF, multi-column | Reading order scrambles columns together | Layout-aware extraction that segments regions before reading text |
| Scanned or photographed PDF | No text layer at all | Detect empty extraction and route to OCR; track OCR confidence as metadata |
| PDF with tables | Cell relationships lost; numbers reattributed | Table detection, then serialize each table to Markdown or CSV and keep it whole |
| HTML pages | Navigation, ads, cookie banners, and footers dominate | Main-content extraction, then split on real heading tags |
| Word documents | Tracked changes and comments extracted as body text | Accept or drop revisions explicitly; use the heading styles as the hierarchy |
| Slide decks | Fragmentary bullets; speaker notes hold the actual content | Chunk per slide, merge notes with slide body, keep the deck title |
| Spreadsheets | No prose; meaning lives in headers and formulas | Emit row-wise records with column names inlined; never split away the header |
| Source code | Fixed-size splits cut functions in half | Syntax-aware splitting on function and class boundaries; prepend the file path |
| Email and chat logs | Quoted reply chains duplicate content endlessly | Strip quoted history; chunk per thread with participants as metadata |
Chunking: Bad to Good to Great
Bad - fixed-size character splitting
Cut every 1,000 characters. It is four lines of code and it is why your retrieval is bad.
Fixed-size splitting is blind to language. It severs sentences mid-clause, splits a definition from the term it defines, and cuts tables in half. Worse, it produces chunks with no topical coherence: a chunk spanning the tail of one section and the head of the next embeds to the average of two unrelated ideas, which is close to nothing. Use it only as the baseline you improve on.
Good - recursive splitting on structural separators
Try to split on the largest structural boundary available and only descend when a piece is still too big: paragraph breaks first, then single newlines, then sentence boundaries, then β as a last resort β characters. Add a small overlap between adjacent chunks so a sentence straddling a boundary appears in full on at least one side.
This is a genuine improvement and the right default when you have no structural signal to work with. Its limit is that it respects typographic boundaries, not the documentβs meaning. A recursive splitter still has no idea that the chunk it just produced sits under βSection 7.3 Termination for Cause.β
Great - structure-aware chunking with heading context
Reconstruct the documentβs real hierarchy first β headings, sections, list structure, table boundaries β then chunk within leaf sections and prepend the heading path to every chunk.
Chunk as stored and embedded:
Source: vendor-msa-2026.pdf
Section: 7. Termination > 7.3 Termination for Cause
---
Either party may terminate this Agreement immediately upon written notice
if the other party commits a material breach and fails to cure such breach
within thirty days of receiving notice of it.
Three things improve at once. The chunk is now interpretable in isolation, so a human or an LLM reading it out of context knows what βeither partyβ refers to. The embedding carries the sectionβs topic, so a query about termination matches even when the body text never repeats the word. And the cited answer can name its source and section, which is what makes grounding checkable.
The cost is real: structure extraction is format-specific work, and there is no universal parser. Budget for it, and fall back to recursive splitting when structure detection fails rather than pretending it succeeded.
Chunk Size, Overlap, and the Small-to-Big Pattern
Chunk size is a genuine tradeoff with no universally right answer:
| Β | Small chunks | Large chunks |
|---|---|---|
| Embedding precision | Sharp β the vector is about one idea | Diluted β an average of several ideas matches nothing strongly |
| Context for the generator | Starved; the answer may be true but unusable | Rich; the model can see the surrounding reasoning |
| Context budget | Many chunks fit in the prompt | Few chunks fit, and cost per query rises |
| Failure mode | Answer severed across a boundary | Relevant sentence buried in mostly irrelevant text |
Since retrieval wants small and generation wants large, stop trying to satisfy both with one unit.
Overlap is the cheap partial fix β repeat a slice of the previous chunk at the start of the next so boundary-straddling sentences survive somewhere intact. A modest fraction of chunk size is the usual convention rather than a tuned optimum. It is not free: it inflates index size and cost proportionally, and it guarantees near-duplicate results in your top-k, so you need deduplication before assembling the prompt or you will spend context budget on the same sentence three times.
Parent-document retrieval, or small-to-big, is the better fix. Embed and search over small precise chunks, but return the larger unit that contains them.
flowchart LR
A[Section text] --> B[Split into small child chunks]
B --> C[Embed child chunks only]
C --> D[Vector index of child chunks]
A --> E[Parent store keyed by parent id]
D -->|match at query time| F[Look up parent id]
E --> F
F --> G[Send parent text to the generator]
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class A client
class B,C service
class D,E data
class F edge
class G async
Retrieval precision comes from the child; generation context comes from the parent. Deduplicate parents when several children from the same section match, which they will.
Contextual enrichment goes one step further: prepend the document title and section path, expand acronyms, and optionally have a cheap model write one situating sentence per chunk at ingest time. All of it is paid once offline and repaid on every query β a favourable trade whenever the corpus is read more often than it is written.
Metadata, Permissions, Tables, and Deduplication
Attach at ingest, not later: source URI, document ID, stable chunk ID, section path, page or character span, created and modified timestamps, content hash, document type, language, and access-control labels.
That last one is not optional and not retrofittable. If a chunk does not carry the identity of who may read it, you cannot enforce permissions at query time β the vector index has no idea which tenant, group, or clearance a chunk belongs to, and an LLM will happily quote a document the user was never entitled to see. Capture the labels during ingestion, keep them denormalized on the chunk so filtering needs no join, and remember that permission predicates are exactly the selective filters that break naive post-filtering, so verify your engineβs filtered-search behaviour. See Authentication and Authorization for the model and Vector Databases for the filtering mechanics. Adding permissions after the fact means a full re-index.
Tables and code need their own paths. Keep a table whole rather than splitting it, serialize it to Markdown or CSV so structure survives as text, and repeat the header row if it truly must be split. A useful pattern is to embed a short natural-language summary of the table while returning the raw table to the generator β the summary matches queries, the raw rows carry the numbers. For code, split on syntactic boundaries so a function is never cut mid-body, and prepend the file path and relevant imports so an isolated snippet is interpretable.
Deduplication runs in two passes. Exact duplicates go by content hash β trivial, and there are always more than you expect once the same PDF exists in three shared drives. Near-duplicates need either a similarity threshold measured on your corpus or a shingling technique like MinHash. The case that matters most is boilerplate: a legal footer repeated across 4,000 documents will produce thousands of near-identical chunks that crowd genuine answers out of your top-k. Strip repeated blocks at parse time and dedupe what survives.
The Full Ingestion Pipeline
flowchart LR
A[Raw source document] --> B[Detect type and route]
B --> C[Extract text with layout and reading order]
C --> D[OCR fallback when no text layer exists]
C --> E[Strip headers and footers and boilerplate]
D --> E
E --> F[Rebuild hierarchy from headings and tables]
F --> G[Structure aware split with heading context]
G --> H[Deduplicate by content hash and similarity]
H --> I[Attach metadata and access control labels]
I --> J[Embed each chunk with the pinned model]
J --> K[Vector index plus parent document store]
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class A client
class B edge
class C,D,E,F service
class G,H,I async
class J edge
class K data
Incremental re-indexing is what makes this pipeline operable rather than a one-off script. Source documents change, and re-embedding everything on every edit is wasteful and slow.
- Hash each source document at fetch time. Unchanged hash means skip the document entirely.
- Derive stable chunk IDs from document ID plus section path plus chunk content hash. Edit one paragraph and only that chunkβs ID changes; the rest keep their existing vectors.
- Diff the new chunk ID set against the indexed set. Insert what is new, delete what disappeared, leave the rest untouched.
- Propagate deletes. This is the step teams skip, and the symptom is severe: the system confidently cites text that no longer exists in the source, with a citation that looks perfectly legitimate.
- Watch churn against your indexβs tolerance for it. Heavy delete-and-reinsert traffic degrades graph indexes through tombstone accumulation, so plan for periodic compaction or rebuild-and-swap.
Two operational notes. Keep the raw extracted text, not just the chunks β every re-chunking experiment otherwise starts by re-parsing and re-OCRing the corpus, which is the slowest and most expensive part. And record the parser version alongside the model version, because a parser upgrade changes chunk boundaries and therefore chunk IDs, which is a re-index just as surely as an embedding model change is.
When to Use
β Invest in structure-aware ingestion when:
- Source documents have real hierarchy that carries meaning β contracts, standards, manuals, policies, API docs
- Answers depend on tables, figures, or numbers where attribution errors are unacceptable
- The corpus contains scanned or multi-column documents, where naive extraction is silently wrong
- Citations must name a section, because grounding checks need a verifiable span
- Access control is per-document or per-section, so labels must ride along with every chunk
β A simple recursive splitter is enough when:
- The corpus is already clean, flat prose β Markdown docs, support articles, transcripts
- Documents are short enough to embed whole, making chunking largely moot
- You are prototyping to find out whether retrieval helps at all; upgrade ingestion once it does
- You have no eval set yet, so you could not tell whether the fancier pipeline actually helped
Common Interview Questions
Q1: Your RAG system retrieves plausible but unhelpful chunks. Where do you look first?
Ingestion, before anything else. Print the raw text of the chunks actually returned and read them. The usual findings are structural, not algorithmic: a multi-column PDF extracted in scrambled reading order, a table flattened into unattributable numbers, running headers spliced mid-sentence, or an answer severed across a chunk boundary so no single chunk contains it. None of these are fixable by tuning the index or swapping the embedding model, and all of them look like retrieval problems from the outside. Only after the chunks read as coherent, self-describing text does it make sense to tune recall, add a reranker, or change models.
Q2: How do you choose chunk size?
Recognize that the question is malformed as asked, because retrieval and generation want opposite things. Small chunks embed precisely but starve the generator of context; large chunks give rich context but dilute the embedding into an average of several ideas and burn context budget. Rather than compromising on one size, use parent-document retrieval: embed and search small child chunks, then return the larger parent section to the generator. That gives precision on the retrieval side and context on the generation side. If you must pick a single size, pick it empirically against an eval set of real questions, and measure whether the answer was actually contained in what you retrieved.
Q3: Why does prepending the heading path to each chunk help so much?
It fixes both interpretability and matching. An isolated chunk saying βeither party may terminate immediately upon written noticeβ does not say which agreement, which parties, or under what conditions β a human cannot use it and neither can an LLM. Prepending
vendor-msa-2026.pdf > 7. Termination > 7.3 Termination for Causemakes the chunk self-describing. It also improves retrieval directly, because the section topic is now inside the embedded text, so a query about termination matches even where the body never repeats the term. And it makes citations specific enough to verify, which is what grounding checks need.
Q4: Where do access-control labels belong in a RAG pipeline?
On the chunk, captured at ingest, denormalized so filtering needs no join. This is a design decision you cannot defer: the vector index has no notion of tenancy or clearance, so if a chunk does not carry who may read it, there is no way to enforce permissions at query time and the generator will quote documents the user was never entitled to see. Retrofitting means a full re-index. Two follow-on points worth raising unprompted β permission predicates are highly selective, which is precisely the case where naive post-filtering silently returns nothing or leaks, so you need pre-filtering or filter-aware search; and when permissions change, the labels on already-indexed chunks have to be updated too.
Q5: A source document is edited. What has to happen?
Not a full re-index, if the pipeline was built for it. Hash each source document so unchanged documents are skipped outright. Derive stable chunk IDs from document ID, section path, and chunk content hash, so editing one paragraph changes exactly one chunk ID and every other chunk keeps its existing vector. Diff the new chunk ID set against what is indexed, then insert additions and delete removals. The critical step is propagating deletes β skip it and the system cites text that no longer exists, with a citation that looks entirely legitimate. Finally, watch the churn rate, because heavy delete-and-reinsert traffic accumulates tombstones in graph indexes and degrades both latency and recall until you compact or rebuild and swap.
Build it in code: Agentic AI Course · Fundamentals: Core Concepts