Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 15 min read

Document Parsing and Chunking - Complete Deep Dive

Stage 3 - Retrieval Lesson 12 of 36

Prerequisites: Embeddings, Vector Databases, Tokens and Cost Math Used in: RAG End to End, Hybrid Search and Reranking, Multimodal


What is Ingestion, and Why Does It Decide Everything?

Retrieval quality is usually lost at ingestion, not at search. This is the least glamorous stage of RAG and the one where most projects actually fail. Teams spend weeks tuning ef_search, swapping embedding models, and adding rerankers, when the real problem is that the answer was severed across two chunks at ingest and no retriever on earth can put it back.

Ingestion is the offline pipeline that turns a raw file into indexable units: detect the format, extract the text in correct reading order, reconstruct the document’s structure, split it into chunks, attach metadata, and embed. Everything downstream inherits whatever this stage produces.

Real-world analogy: a library with a world-class catalogue and a brilliant reference librarian is still useless if a careless clerk tore every book into random 40-page bundles and reshelved them with no titles. The search desk is not the problem. The shelving was.


Parsing Reality: A PDF Has No Structure

This is the fact that surprises backend engineers most. PDF is a page-description format. It says β€œdraw glyph T at coordinates x, y in this font.” It does not say β€œthis is a heading”, β€œthis is a table cell”, β€œthis paragraph continues from the previous page”, or even β€œread this column first.” Any structure you get from a PDF was inferred by your extractor from glyph positions, and inference fails in specific, repeatable ways:

Detect this early with a blunt but effective check: sample extracted documents and eyeball the raw text. A short, mechanical review of thirty real documents finds more retrieval bugs than a week of parameter tuning.

Document type Parsing hazard Approach
Text-layer PDF, single column Hyphenation, ligatures, running headers and footers Standard text extraction; strip repeated lines that appear on most pages
Text-layer PDF, multi-column Reading order scrambles columns together Layout-aware extraction that segments regions before reading text
Scanned or photographed PDF No text layer at all Detect empty extraction and route to OCR; track OCR confidence as metadata
PDF with tables Cell relationships lost; numbers reattributed Table detection, then serialize each table to Markdown or CSV and keep it whole
HTML pages Navigation, ads, cookie banners, and footers dominate Main-content extraction, then split on real heading tags
Word documents Tracked changes and comments extracted as body text Accept or drop revisions explicitly; use the heading styles as the hierarchy
Slide decks Fragmentary bullets; speaker notes hold the actual content Chunk per slide, merge notes with slide body, keep the deck title
Spreadsheets No prose; meaning lives in headers and formulas Emit row-wise records with column names inlined; never split away the header
Source code Fixed-size splits cut functions in half Syntax-aware splitting on function and class boundaries; prepend the file path
Email and chat logs Quoted reply chains duplicate content endlessly Strip quoted history; chunk per thread with participants as metadata

Chunking: Bad to Good to Great

Bad - fixed-size character splitting

Cut every 1,000 characters. It is four lines of code and it is why your retrieval is bad.

Fixed-size splitting is blind to language. It severs sentences mid-clause, splits a definition from the term it defines, and cuts tables in half. Worse, it produces chunks with no topical coherence: a chunk spanning the tail of one section and the head of the next embeds to the average of two unrelated ideas, which is close to nothing. Use it only as the baseline you improve on.

Good - recursive splitting on structural separators

Try to split on the largest structural boundary available and only descend when a piece is still too big: paragraph breaks first, then single newlines, then sentence boundaries, then β€” as a last resort β€” characters. Add a small overlap between adjacent chunks so a sentence straddling a boundary appears in full on at least one side.

This is a genuine improvement and the right default when you have no structural signal to work with. Its limit is that it respects typographic boundaries, not the document’s meaning. A recursive splitter still has no idea that the chunk it just produced sits under β€œSection 7.3 Termination for Cause.”

Great - structure-aware chunking with heading context

Reconstruct the document’s real hierarchy first β€” headings, sections, list structure, table boundaries β€” then chunk within leaf sections and prepend the heading path to every chunk.

Chunk as stored and embedded:

  Source: vendor-msa-2026.pdf
  Section: 7. Termination > 7.3 Termination for Cause
  ---
  Either party may terminate this Agreement immediately upon written notice
  if the other party commits a material breach and fails to cure such breach
  within thirty days of receiving notice of it.

Three things improve at once. The chunk is now interpretable in isolation, so a human or an LLM reading it out of context knows what β€œeither party” refers to. The embedding carries the section’s topic, so a query about termination matches even when the body text never repeats the word. And the cited answer can name its source and section, which is what makes grounding checkable.

The cost is real: structure extraction is format-specific work, and there is no universal parser. Budget for it, and fall back to recursive splitting when structure detection fails rather than pretending it succeeded.


Chunk Size, Overlap, and the Small-to-Big Pattern

Chunk size is a genuine tradeoff with no universally right answer:

Β  Small chunks Large chunks
Embedding precision Sharp β€” the vector is about one idea Diluted β€” an average of several ideas matches nothing strongly
Context for the generator Starved; the answer may be true but unusable Rich; the model can see the surrounding reasoning
Context budget Many chunks fit in the prompt Few chunks fit, and cost per query rises
Failure mode Answer severed across a boundary Relevant sentence buried in mostly irrelevant text

Since retrieval wants small and generation wants large, stop trying to satisfy both with one unit.

Overlap is the cheap partial fix β€” repeat a slice of the previous chunk at the start of the next so boundary-straddling sentences survive somewhere intact. A modest fraction of chunk size is the usual convention rather than a tuned optimum. It is not free: it inflates index size and cost proportionally, and it guarantees near-duplicate results in your top-k, so you need deduplication before assembling the prompt or you will spend context budget on the same sentence three times.

Parent-document retrieval, or small-to-big, is the better fix. Embed and search over small precise chunks, but return the larger unit that contains them.

flowchart LR
    A[Section text] --> B[Split into small child chunks]
    B --> C[Embed child chunks only]
    C --> D[Vector index of child chunks]
    A --> E[Parent store keyed by parent id]
    D -->|match at query time| F[Look up parent id]
    E --> F
    F --> G[Send parent text to the generator]

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class A client
    class B,C service
    class D,E data
    class F edge
    class G async

Retrieval precision comes from the child; generation context comes from the parent. Deduplicate parents when several children from the same section match, which they will.

Contextual enrichment goes one step further: prepend the document title and section path, expand acronyms, and optionally have a cheap model write one situating sentence per chunk at ingest time. All of it is paid once offline and repaid on every query β€” a favourable trade whenever the corpus is read more often than it is written.


Metadata, Permissions, Tables, and Deduplication

Attach at ingest, not later: source URI, document ID, stable chunk ID, section path, page or character span, created and modified timestamps, content hash, document type, language, and access-control labels.

That last one is not optional and not retrofittable. If a chunk does not carry the identity of who may read it, you cannot enforce permissions at query time β€” the vector index has no idea which tenant, group, or clearance a chunk belongs to, and an LLM will happily quote a document the user was never entitled to see. Capture the labels during ingestion, keep them denormalized on the chunk so filtering needs no join, and remember that permission predicates are exactly the selective filters that break naive post-filtering, so verify your engine’s filtered-search behaviour. See Authentication and Authorization for the model and Vector Databases for the filtering mechanics. Adding permissions after the fact means a full re-index.

Tables and code need their own paths. Keep a table whole rather than splitting it, serialize it to Markdown or CSV so structure survives as text, and repeat the header row if it truly must be split. A useful pattern is to embed a short natural-language summary of the table while returning the raw table to the generator β€” the summary matches queries, the raw rows carry the numbers. For code, split on syntactic boundaries so a function is never cut mid-body, and prepend the file path and relevant imports so an isolated snippet is interpretable.

Deduplication runs in two passes. Exact duplicates go by content hash β€” trivial, and there are always more than you expect once the same PDF exists in three shared drives. Near-duplicates need either a similarity threshold measured on your corpus or a shingling technique like MinHash. The case that matters most is boilerplate: a legal footer repeated across 4,000 documents will produce thousands of near-identical chunks that crowd genuine answers out of your top-k. Strip repeated blocks at parse time and dedupe what survives.


The Full Ingestion Pipeline

flowchart LR
    A[Raw source document] --> B[Detect type and route]
    B --> C[Extract text with layout and reading order]
    C --> D[OCR fallback when no text layer exists]
    C --> E[Strip headers and footers and boilerplate]
    D --> E
    E --> F[Rebuild hierarchy from headings and tables]
    F --> G[Structure aware split with heading context]
    G --> H[Deduplicate by content hash and similarity]
    H --> I[Attach metadata and access control labels]
    I --> J[Embed each chunk with the pinned model]
    J --> K[Vector index plus parent document store]

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class A client
    class B edge
    class C,D,E,F service
    class G,H,I async
    class J edge
    class K data

Incremental re-indexing is what makes this pipeline operable rather than a one-off script. Source documents change, and re-embedding everything on every edit is wasteful and slow.

  1. Hash each source document at fetch time. Unchanged hash means skip the document entirely.
  2. Derive stable chunk IDs from document ID plus section path plus chunk content hash. Edit one paragraph and only that chunk’s ID changes; the rest keep their existing vectors.
  3. Diff the new chunk ID set against the indexed set. Insert what is new, delete what disappeared, leave the rest untouched.
  4. Propagate deletes. This is the step teams skip, and the symptom is severe: the system confidently cites text that no longer exists in the source, with a citation that looks perfectly legitimate.
  5. Watch churn against your index’s tolerance for it. Heavy delete-and-reinsert traffic degrades graph indexes through tombstone accumulation, so plan for periodic compaction or rebuild-and-swap.

Two operational notes. Keep the raw extracted text, not just the chunks β€” every re-chunking experiment otherwise starts by re-parsing and re-OCRing the corpus, which is the slowest and most expensive part. And record the parser version alongside the model version, because a parser upgrade changes chunk boundaries and therefore chunk IDs, which is a re-index just as surely as an embedding model change is.


When to Use

βœ… Invest in structure-aware ingestion when:

❌ A simple recursive splitter is enough when:


Common Interview Questions

Q1: Your RAG system retrieves plausible but unhelpful chunks. Where do you look first?

Ingestion, before anything else. Print the raw text of the chunks actually returned and read them. The usual findings are structural, not algorithmic: a multi-column PDF extracted in scrambled reading order, a table flattened into unattributable numbers, running headers spliced mid-sentence, or an answer severed across a chunk boundary so no single chunk contains it. None of these are fixable by tuning the index or swapping the embedding model, and all of them look like retrieval problems from the outside. Only after the chunks read as coherent, self-describing text does it make sense to tune recall, add a reranker, or change models.

Q2: How do you choose chunk size?

Recognize that the question is malformed as asked, because retrieval and generation want opposite things. Small chunks embed precisely but starve the generator of context; large chunks give rich context but dilute the embedding into an average of several ideas and burn context budget. Rather than compromising on one size, use parent-document retrieval: embed and search small child chunks, then return the larger parent section to the generator. That gives precision on the retrieval side and context on the generation side. If you must pick a single size, pick it empirically against an eval set of real questions, and measure whether the answer was actually contained in what you retrieved.

Q3: Why does prepending the heading path to each chunk help so much?

It fixes both interpretability and matching. An isolated chunk saying β€œeither party may terminate immediately upon written notice” does not say which agreement, which parties, or under what conditions β€” a human cannot use it and neither can an LLM. Prepending vendor-msa-2026.pdf > 7. Termination > 7.3 Termination for Cause makes the chunk self-describing. It also improves retrieval directly, because the section topic is now inside the embedded text, so a query about termination matches even where the body never repeats the term. And it makes citations specific enough to verify, which is what grounding checks need.

Q4: Where do access-control labels belong in a RAG pipeline?

On the chunk, captured at ingest, denormalized so filtering needs no join. This is a design decision you cannot defer: the vector index has no notion of tenancy or clearance, so if a chunk does not carry who may read it, there is no way to enforce permissions at query time and the generator will quote documents the user was never entitled to see. Retrofitting means a full re-index. Two follow-on points worth raising unprompted β€” permission predicates are highly selective, which is precisely the case where naive post-filtering silently returns nothing or leaks, so you need pre-filtering or filter-aware search; and when permissions change, the labels on already-indexed chunks have to be updated too.

Q5: A source document is edited. What has to happen?

Not a full re-index, if the pipeline was built for it. Hash each source document so unchanged documents are skipped outright. Derive stable chunk IDs from document ID, section path, and chunk content hash, so editing one paragraph changes exactly one chunk ID and every other chunk keeps its existing vector. Diff the new chunk ID set against what is indexed, then insert additions and delete removals. The critical step is propagating deletes β€” skip it and the system cites text that no longer exists, with a citation that looks entirely legitimate. Finally, watch the churn rate, because heavy delete-and-reinsert traffic accumulates tombstones in graph indexes and degrades both latency and recall until you compact or rebuild and swap.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access