Multimodal - Complete Deep Dive
Prerequisites: Document Parsing and Chunking, Tokens and Cost Math, Structured Outputs Used in: RAG End to End, Evals, Hallucination and Grounding
What is a Multimodal Model?
A multimodal model accepts more than one kind of input β text plus images, sometimes audio, sometimes video β and reasons across them in a single context. Practically, it means you can hand a model a photograph of an invoice and ask βwhat is the total and when is it dueβ instead of building a parser, an OCR step, and a field-matching heuristic first.
The mechanism matters because it sets the cost. Non-text inputs are converted into tokens in the same context window as your text. An image becomes a sequence of embedded patches; audio is usually transcribed or encoded into a token-like sequence. From the modelβs point of view there is one stream. From your point of view that means the practical fact most teams miss: an image is not free, and it is not cheap. It consumes context and it consumes budget, in proportion to its resolution and how many of them you send.
Real-world analogy: Describing a photograph over the phone versus emailing it. The description is compact, searchable, and quotable β and it silently discards everything the describer did not think was important, including the layout. The photo preserves all of it and costs far more to send. Almost every design decision on this page is a choice between those two, made per document type rather than once for the whole system.
An Image Is Not Free - It Becomes Tokens
Providers convert an image into some number of tokens according to a published formula that depends on its dimensions, and larger images are tiled into more patches. Look up the current formula and pricing in your providerβs own documentation, and verify against the usage figures your API responses return. Do not trust a token-per-image number you read anywhere, including here β there is no such constant, it differs by provider and by model, and it changes.
What does transfer are the three structural consequences:
- Resolution is a dial you control. Downscaling before upload cuts tokens directly. The right resolution is the lowest one at which the task still works, found by measurement: run your eval set at several resolutions and look at where accuracy falls off. For reading dense tables or small print that floor is high. For classifying whether a photo contains a receipt at all, it is low.
- Page count multiplies everything. Sending a 40-page PDF as 40 images is 40 times the image cost and a large slice of your context window, before the model reads a word. Triage which pages are relevant first β often with a cheap text pass β and send only those.
- Context is a shared budget. Images compete with retrieved chunks, few-shot examples, and conversation history for the same window. A multimodal request is where context budgeting stops being theoretical, and the accounting method is the same one in Tokens and Cost Math.
The following is an ASSUMPTION-based illustration, not a measurement. Suppose a page image costs an ASSUMED 1,000 tokens at full resolution and an ASSUMED 300 at half. A 10-page document is then 10,000 tokens versus 3,000 β before any prompt or retrieved context. If triage narrows it to the 2 pages that matter, it is 2,000 versus 600. The numbers are invented; the ratio between βsend everything at full resolutionβ and βtriage then downscaleβ is the point, and it is usually an order of magnitude.
The Patterns That Actually Ship
Demos show a model describing a photo. Production uses multimodal for a narrower set of jobs where it clearly beats the alternative.
1. Document understanding where a text parser fails. This is the highest-value pattern by a distance. Document Parsing and Chunking catalogues the hazards: a PDF is a page-description format with no real structure, multi-column layouts read out of order into plausible nonsense, tables flatten into token soup that loses which number belongs to which row, and scanned pages have no text layer at all. A vision model looks at the rendered page and sees the layout the way a person does β so it keeps column order, reads a table as a table, and handles a scan without a separate OCR stage. When extraction quality from the text path is the bottleneck, this is the fix.
2. Chart and diagram interpretation. A chartβs meaning is entirely in its geometry, so text extraction returns axis labels and nothing else. A vision model can describe the trend and read labelled values. It cannot read values the chart does not label, and it will happily estimate one anyway β treat any precise number it produces from a chart as unverified.
3. Screenshot and UI understanding. Reading application state from a screenshot powers support tooling (βwhat error is this user seeingβ), QA triage, and visual regression description. It is also the perception layer for agents that operate a UI, which puts an LLMβs judgement in charge of clicks β review the injection surface in AI Security before building that, because text rendered in an image is still instructions the model reads.
4. Image classification and moderation. Often the least glamorous and most reliable use. A general multimodal model is a strong zero-shot classifier, which makes it excellent for bootstrapping a category taxonomy before you have labels. At volume, a purpose-trained classifier is usually cheaper and faster for the same accuracy, so treat the general model as the starting point and the path to a specialised one.
5. Audio transcription then text processing. The dominant audio pattern is not βan audio-native model reasons about the callβ β it is transcribe first, then run your existing text pipeline on the transcript. Everything you already built for summarisation, extraction, retrieval, and evals keeps working, and the transcript is a durable, searchable, reviewable artifact.
Convert to Text Early or Keep the Modality
This is the main pipeline decision, and it is per document type, not global.
| Β | Convert to text early | Keep the original modality |
|---|---|---|
| Cost | Low - text tokens only, and the conversion happens once | High - image tokens on every request that touches the page |
| Composability | High - chunking, embedding, retrieval, and evals all work unchanged | Low - each stage needs multimodal support |
| Caching and reuse | Convert once, reuse forever | Re-sent per request unless you cache the extraction |
| What survives | The words, and whatever structure your parser inferred | Layout, spatial relationships, stamps, signatures, handwriting, visual emphasis |
| Auditability | The extracted text is reviewable and diffable | Requires a human to look at the page again |
| Best for | Prose documents, anything feeding a RAG index, high query volume | Forms, statements, complex tables, scans, anything where layout carries meaning |
The pragmatic answer for most systems is convert early but keep the original, which means the vision model is part of ingestion rather than part of query serving. Render the page, have the vision model produce structured text or fields once, index that, and store the original page in object storage so a human or a re-extraction can always go back to it. Query-time requests then hit ordinary text retrieval, which is cheap, fast, and already evaluated β see RAG End to End.
The Document Pipeline
flowchart TD
UP[Document uploaded] --> ST[Store the original in object storage]
ST --> TRI[Triage each page - is there a usable text layer]
TRI --> TXT[Text parser path - extract and chunk]
TRI --> VIS[Vision path - render the page to an image]
VIS --> VLM[Vision model reads the page into a strict schema]
TXT --> VER[Field level verification and confidence check]
VLM --> VER
VER --> OK[Verified - persist and index]
VER --> HITL[Low confidence or failed check - human review queue]
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class UP client
class ST data
class TRI edge
class TXT,VLM service
class VIS async
class VER async
class OK data
class HITL client
Five things this shape gets right.
- The original is stored first. Every later stage is then re-runnable when you improve the pipeline, and an auditor can always see what the model actually saw.
- Triage routes per page, not per document. A contract with a scanned signature page and 30 clean text pages should use both paths. Routing the whole document to vision because one page needs it multiplies cost for no benefit.
- The vision model fills a schema. Asking for free prose and parsing it later throws away your strongest defence. Constrain the output as described in Structured Outputs, with an explicit
unreadablevalue available per field so the model has a way to abstain instead of guessing. - Verification sits before persistence, not after. Once a wrong number is in the index it is indistinguishable from a right one. Check it while you still have the page.
- Low confidence routes to a person. A queue that catches a small percentage of pages is cheap. A wrong total on an invoice is not.
Multimodal Embeddings and Cross-Modal Retrieval
Some embedding models place images and text in a shared vector space, so a text query can retrieve images and an image query can retrieve text or similar images. That enables genuinely useful things: searching a product catalogue by description, finding the diagram that matches a paragraph, reverse image lookup over your own corpus.
The caveat is the part to remember: a shared embedding space is weaker than a same-modality one. Cross-modal similarity is trained on paired data, so it captures the kinds of correspondence that appeared in that pairing β broad subject matter, obvious visual attributes β and it captures fine distinctions poorly. Text-to-text retrieval over a good caption frequently beats text-to-image retrieval over the raw image, because a caption is dense, specific, and searchable by everything you already have.
So the practical pattern is usually caption, then index the caption. Have a vision model produce a rich description and any structured attributes at ingestion, store them alongside the image reference, and run your normal hybrid retrieval over the text. Reach for true cross-modal embeddings when the query genuinely is an image, or when captioning the whole corpus is prohibitive. Either way, evaluate it as retrieval β recall at k against a labelled set, exactly as in Embeddings.
Audio Has Its Own Rules
Transcription looks solved until you run it on your own recordings.
- Quality varies by accent, domain vocabulary, and noise, and it varies unevenly. A model strong on clear speech in one accent can degrade sharply on another, on a poor phone line, or on overlapping speakers. Product names, drug names, part numbers, and internal jargon are transcribed as the nearest common word. Where the vocabulary is closed and known, supplying it as a hint or biasing list is the single highest-return fix.
- Diarization β who spoke β is a separate capability from transcription, and it is the harder one. Most downstream tasks need it, because βthe customer agreed to the chargeβ and βthe agent agreed to the chargeβ are different facts extracted from the same sentence.
- Timestamps are what make a transcript auditable. Word or segment level offsets let a reviewer jump to the audio behind a claim, which is the audio equivalent of a citation and the same grounding argument as Hallucination and Grounding.
- Real-time and batch are different products. Streaming transcription must emit partial results and revise them as more audio arrives, so early words can change; batch sees the whole file and is more accurate. Pick per use case, and never compare their accuracy as if they were the same task.
- Downstream errors inherit upstream ones. A summary built on a transcript that misheard a number is confidently wrong with no trace of where it went wrong. Keep transcription and comprehension as separately evaluated stages.
The Failure Modes You Must Plan For
Reading errors on dense content. Tightly packed tables, small fonts, low-contrast scans, and handwriting all raise the error rate. The errors concentrate exactly where they hurt most β digits, decimal places, and which column a value belongs to.
Overconfident reading of low-quality input. Faced with a blurry scan, the model does not say βI cannot read this.β It produces a plausible value in fluent prose. Confidence in the output is uncorrelated with whether the reading was correct.
Hallucinated values in numeric fields. The worst version. A chart with no labelled value still gets a number; a faded total still gets a total. The model is completing a pattern, not measuring pixels.
And the fact that ties them together: a vision model can misread a number with no signal that it did. There is no parse error, no exception, no low-confidence flag by default. A text parser that fails usually fails loudly β empty output, garbled characters, a crash. A vision model fails quietly and fluently, which is worse, because a silent wrong answer propagates into your index, your reports, and your decisions.
Therefore, as a rule rather than a suggestion: field-level verification is mandatory for anything financial, legal, medical, or otherwise consequential. Concretely β cross-check arithmetic that the document itself constrains, such as line items summing to a stated total; validate formats and plausible ranges for dates, currencies, and identifiers; reconcile against a system of record where one exists; require the model to return unreadable rather than guess, and measure how often it does; run a second independent extraction on high-value fields and route disagreements to a human; and sample-audit output against the stored original forever, not just at launch.
Modality Scorecard
| Modality | Best-fit task | Main failure mode |
|---|---|---|
| Photo or natural image | Classification, moderation, attribute extraction | Confident description of detail that is not present |
| Scanned or photographed page | Reading a page with no text layer | Silent character and digit misreads on poor scans |
| Complex-layout page | Field extraction from forms and statements | Values attached to the wrong label or column |
| Chart or diagram | Describing a trend and reading labelled values | Precise numbers invented where the chart labels none |
| Screenshot | UI state understanding and support triage | Text in the image treated as instructions to follow |
| Speech audio | Transcription feeding a normal text pipeline | Domain vocabulary and accents silently mistranscribed |
| Video | Rare - keyframe sampling plus the audio transcript | Cost, plus whatever happened between sampled frames |
Cost, Latency, and Evaluation
Multimodal requests are materially more expensive and slower than text-only, for structural reasons rather than incidental ones: more tokens per request, and heavier prefill over those tokens. Both effects scale with resolution and page count. That pushes the same conclusions from three directions β do the expensive work once at ingestion rather than per query, triage aggressively so the model only sees pages that matter, downscale to the lowest resolution that passes your evals, and cache extraction results keyed on a content hash of the file so a re-upload of the same document is free.
Evaluation needs modality-specific ground truth, which is the part teams underestimate. A text eval set does not test whether a table was read correctly. You need labelled pages with the correct field values, labelled audio with the correct transcript and speaker turns, and labelled images with the correct class β and creating that requires a person looking at the original. Build it small and real: 30 genuinely difficult pages from your own corpus, including the bad scans, beats 500 clean synthetic ones. Score per field rather than per document, because a 95 percent document score that is wrong on the total every time is a failed system. The discipline is the same as Evals, applied to a new input type.
Bad to Good to Great
Bad - send every page to a vision model and trust the output
No triage, full resolution, free-text output parsed with regular expressions, no verification. It works in the demo. In production it costs an order of magnitude more than necessary, and every silent misread lands in your database indistinguishable from a correct value.
Good - route scans to vision, use a schema, spot-check the output
Text layer where there is one, vision where there is not, structured output, and someone reviews a sample each week. This is a working system. What is missing: per-page rather than per-document routing, resolution tuned by measurement, arithmetic and range checks on the fields, an abstention path, and a modality-specific eval set that would catch a regression before users do.
Great - triage per page, extract once into a schema, verify before persisting
- Original stored first in object storage, so every stage is re-runnable and auditable.
- Per-page triage between the text path and the vision path, with a cheap check deciding.
- Resolution and page selection tuned by eval, not by default settings.
- Strict schema output with an explicit unreadable value, so abstention is available and measurable.
- Field-level verification before persistence β arithmetic, format, range, reconciliation β with low-confidence pages routed to human review.
- Extraction cached on a content hash, so the expensive work happens once per document.
- A modality-specific eval set of real difficult pages, scored per field, gating changes to the pipeline.
- Ongoing sample audits against the stored original, because silent misreads never announce themselves.
When to Use
β Reach for multimodal when:
- A text parser is mangling your documents and extraction quality is the bottleneck
- Layout carries meaning β forms, statements, invoices, complex tables, scans
- The input is genuinely not text: photos, screenshots, charts, recorded calls
- You need a zero-shot classifier to bootstrap a taxonomy before you have labels
- A human currently reads the page to get the answer, and the answer is a small set of fields
β Stay text-only when:
- The text layer is clean and complete, which makes vision pure added cost and risk
- Volume is high and the task is narrow enough that a purpose-trained classifier wins
- You cannot build a verification step, and the fields are consequential
- You have no modality-specific ground truth yet, so you would be shipping unmeasured
- The real requirement is exact values from a machine-readable source you could query instead
Common Interview Questions
Q1: Why is an image more expensive than it looks?
Because it becomes tokens in the same context window as the text, and the count scales with resolution and with how many images you send. Providers publish the formula and it differs per model, so the number has to be looked up rather than assumed. Two consequences follow. Resolution is a dial: the right setting is the lowest at which your eval set still passes, which is high for dense tables and low for coarse classification. And page count multiplies everything, so sending a whole PDF as page images burns budget and context before the model reads anything β triage to the relevant pages first, usually with a cheap text pass.
Q2: When do you use a vision model instead of a text parser for documents?
When the structure a text parser has to infer is the thing that keeps breaking. A PDF is a page-description format, so multi-column layouts read out of order into plausible nonsense, tables flatten and lose which number belongs to which row, and scans have no text layer at all. A vision model sees the rendered layout, so it handles those natively. I decide per page rather than per document β a clean text layer goes down the cheap path, and only pages that fail triage go to vision. And I do the extraction once at ingestion into a schema, not per query, because per-query vision calls are expensive, slow, and uncacheable.
Q3: What is the most dangerous failure mode of a vision model, and how do you handle it?
A silent misread. A text parser that fails usually fails loudly with empty or garbled output; a vision model returns a fluent, plausible, wrong value with no error and no low-confidence signal. A faded total still gets a total, an unlabelled chart still gets a number. So verification cannot be optional for consequential fields. I make the schema allow an explicit unreadable value so abstention is possible and measurable, check the arithmetic the document itself constrains such as line items summing to a stated total, validate formats and ranges, reconcile against a system of record where one exists, run a second independent extraction on high-value fields and escalate disagreements to a human, and keep sample-auditing against the stored original in perpetuity.
Q4: How would you build search over a catalogue of images?
Default to captioning at ingestion and indexing the text. A vision model produces a rich description and structured attributes once, those go into the normal hybrid retrieval path, and the image reference sits alongside in object storage. This wins because text retrieval is cheap, well understood, already evaluated, and dense captions capture the specific detail that matters. True multimodal embeddings put images and text in a shared space, which is genuinely useful when the query itself is an image or when captioning the corpus is too expensive β but a shared space is weaker at fine distinctions than a same-modality one, since it only learned the correspondences present in its training pairs. Either way I measure it as retrieval, with recall at k against a labelled set.
Q5: How do you evaluate a multimodal pipeline?
With modality-specific ground truth, and scored per field rather than per document. A text eval set cannot tell you whether a table was read correctly, so I need labelled pages with correct field values, labelled audio with correct transcripts and speaker turns, and labelled images with correct classes β all of which requires a person looking at the original. I keep the set small and deliberately nasty: thirty genuinely hard pages from my own corpus, including the bad scans and the dense tables, beats hundreds of clean ones. Per-field scoring matters because a document that is 95 percent right and always wrong on the total is a failed system. And I evaluate the stages separately, since a summary built on a mistranscribed number is confidently wrong with nothing in the output pointing at the transcription.
Build it in code: Agentic AI Course · Fundamentals: Core Concepts