Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 20 min read

Multimodal - Complete Deep Dive

Stage 7 - Production Lesson 34 of 36

Prerequisites: Document Parsing and Chunking, Tokens and Cost Math, Structured Outputs Used in: RAG End to End, Evals, Hallucination and Grounding


What is a Multimodal Model?

A multimodal model accepts more than one kind of input β€” text plus images, sometimes audio, sometimes video β€” and reasons across them in a single context. Practically, it means you can hand a model a photograph of an invoice and ask β€œwhat is the total and when is it due” instead of building a parser, an OCR step, and a field-matching heuristic first.

The mechanism matters because it sets the cost. Non-text inputs are converted into tokens in the same context window as your text. An image becomes a sequence of embedded patches; audio is usually transcribed or encoded into a token-like sequence. From the model’s point of view there is one stream. From your point of view that means the practical fact most teams miss: an image is not free, and it is not cheap. It consumes context and it consumes budget, in proportion to its resolution and how many of them you send.

Real-world analogy: Describing a photograph over the phone versus emailing it. The description is compact, searchable, and quotable β€” and it silently discards everything the describer did not think was important, including the layout. The photo preserves all of it and costs far more to send. Almost every design decision on this page is a choice between those two, made per document type rather than once for the whole system.


An Image Is Not Free - It Becomes Tokens

Providers convert an image into some number of tokens according to a published formula that depends on its dimensions, and larger images are tiled into more patches. Look up the current formula and pricing in your provider’s own documentation, and verify against the usage figures your API responses return. Do not trust a token-per-image number you read anywhere, including here β€” there is no such constant, it differs by provider and by model, and it changes.

What does transfer are the three structural consequences:

The following is an ASSUMPTION-based illustration, not a measurement. Suppose a page image costs an ASSUMED 1,000 tokens at full resolution and an ASSUMED 300 at half. A 10-page document is then 10,000 tokens versus 3,000 β€” before any prompt or retrieved context. If triage narrows it to the 2 pages that matter, it is 2,000 versus 600. The numbers are invented; the ratio between β€œsend everything at full resolution” and β€œtriage then downscale” is the point, and it is usually an order of magnitude.


The Patterns That Actually Ship

Demos show a model describing a photo. Production uses multimodal for a narrower set of jobs where it clearly beats the alternative.

1. Document understanding where a text parser fails. This is the highest-value pattern by a distance. Document Parsing and Chunking catalogues the hazards: a PDF is a page-description format with no real structure, multi-column layouts read out of order into plausible nonsense, tables flatten into token soup that loses which number belongs to which row, and scanned pages have no text layer at all. A vision model looks at the rendered page and sees the layout the way a person does β€” so it keeps column order, reads a table as a table, and handles a scan without a separate OCR stage. When extraction quality from the text path is the bottleneck, this is the fix.

2. Chart and diagram interpretation. A chart’s meaning is entirely in its geometry, so text extraction returns axis labels and nothing else. A vision model can describe the trend and read labelled values. It cannot read values the chart does not label, and it will happily estimate one anyway β€” treat any precise number it produces from a chart as unverified.

3. Screenshot and UI understanding. Reading application state from a screenshot powers support tooling (β€œwhat error is this user seeing”), QA triage, and visual regression description. It is also the perception layer for agents that operate a UI, which puts an LLM’s judgement in charge of clicks β€” review the injection surface in AI Security before building that, because text rendered in an image is still instructions the model reads.

4. Image classification and moderation. Often the least glamorous and most reliable use. A general multimodal model is a strong zero-shot classifier, which makes it excellent for bootstrapping a category taxonomy before you have labels. At volume, a purpose-trained classifier is usually cheaper and faster for the same accuracy, so treat the general model as the starting point and the path to a specialised one.

5. Audio transcription then text processing. The dominant audio pattern is not β€œan audio-native model reasons about the call” β€” it is transcribe first, then run your existing text pipeline on the transcript. Everything you already built for summarisation, extraction, retrieval, and evals keeps working, and the transcript is a durable, searchable, reviewable artifact.


Convert to Text Early or Keep the Modality

This is the main pipeline decision, and it is per document type, not global.

Β  Convert to text early Keep the original modality
Cost Low - text tokens only, and the conversion happens once High - image tokens on every request that touches the page
Composability High - chunking, embedding, retrieval, and evals all work unchanged Low - each stage needs multimodal support
Caching and reuse Convert once, reuse forever Re-sent per request unless you cache the extraction
What survives The words, and whatever structure your parser inferred Layout, spatial relationships, stamps, signatures, handwriting, visual emphasis
Auditability The extracted text is reviewable and diffable Requires a human to look at the page again
Best for Prose documents, anything feeding a RAG index, high query volume Forms, statements, complex tables, scans, anything where layout carries meaning

The pragmatic answer for most systems is convert early but keep the original, which means the vision model is part of ingestion rather than part of query serving. Render the page, have the vision model produce structured text or fields once, index that, and store the original page in object storage so a human or a re-extraction can always go back to it. Query-time requests then hit ordinary text retrieval, which is cheap, fast, and already evaluated β€” see RAG End to End.


The Document Pipeline

flowchart TD
    UP[Document uploaded] --> ST[Store the original in object storage]
    ST --> TRI[Triage each page - is there a usable text layer]
    TRI --> TXT[Text parser path - extract and chunk]
    TRI --> VIS[Vision path - render the page to an image]
    VIS --> VLM[Vision model reads the page into a strict schema]
    TXT --> VER[Field level verification and confidence check]
    VLM --> VER
    VER --> OK[Verified - persist and index]
    VER --> HITL[Low confidence or failed check - human review queue]

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class UP client
    class ST data
    class TRI edge
    class TXT,VLM service
    class VIS async
    class VER async
    class OK data
    class HITL client

Five things this shape gets right.

  1. The original is stored first. Every later stage is then re-runnable when you improve the pipeline, and an auditor can always see what the model actually saw.
  2. Triage routes per page, not per document. A contract with a scanned signature page and 30 clean text pages should use both paths. Routing the whole document to vision because one page needs it multiplies cost for no benefit.
  3. The vision model fills a schema. Asking for free prose and parsing it later throws away your strongest defence. Constrain the output as described in Structured Outputs, with an explicit unreadable value available per field so the model has a way to abstain instead of guessing.
  4. Verification sits before persistence, not after. Once a wrong number is in the index it is indistinguishable from a right one. Check it while you still have the page.
  5. Low confidence routes to a person. A queue that catches a small percentage of pages is cheap. A wrong total on an invoice is not.

Multimodal Embeddings and Cross-Modal Retrieval

Some embedding models place images and text in a shared vector space, so a text query can retrieve images and an image query can retrieve text or similar images. That enables genuinely useful things: searching a product catalogue by description, finding the diagram that matches a paragraph, reverse image lookup over your own corpus.

The caveat is the part to remember: a shared embedding space is weaker than a same-modality one. Cross-modal similarity is trained on paired data, so it captures the kinds of correspondence that appeared in that pairing β€” broad subject matter, obvious visual attributes β€” and it captures fine distinctions poorly. Text-to-text retrieval over a good caption frequently beats text-to-image retrieval over the raw image, because a caption is dense, specific, and searchable by everything you already have.

So the practical pattern is usually caption, then index the caption. Have a vision model produce a rich description and any structured attributes at ingestion, store them alongside the image reference, and run your normal hybrid retrieval over the text. Reach for true cross-modal embeddings when the query genuinely is an image, or when captioning the whole corpus is prohibitive. Either way, evaluate it as retrieval β€” recall at k against a labelled set, exactly as in Embeddings.


Audio Has Its Own Rules

Transcription looks solved until you run it on your own recordings.


The Failure Modes You Must Plan For

Reading errors on dense content. Tightly packed tables, small fonts, low-contrast scans, and handwriting all raise the error rate. The errors concentrate exactly where they hurt most β€” digits, decimal places, and which column a value belongs to.

Overconfident reading of low-quality input. Faced with a blurry scan, the model does not say β€œI cannot read this.” It produces a plausible value in fluent prose. Confidence in the output is uncorrelated with whether the reading was correct.

Hallucinated values in numeric fields. The worst version. A chart with no labelled value still gets a number; a faded total still gets a total. The model is completing a pattern, not measuring pixels.

And the fact that ties them together: a vision model can misread a number with no signal that it did. There is no parse error, no exception, no low-confidence flag by default. A text parser that fails usually fails loudly β€” empty output, garbled characters, a crash. A vision model fails quietly and fluently, which is worse, because a silent wrong answer propagates into your index, your reports, and your decisions.

Therefore, as a rule rather than a suggestion: field-level verification is mandatory for anything financial, legal, medical, or otherwise consequential. Concretely β€” cross-check arithmetic that the document itself constrains, such as line items summing to a stated total; validate formats and plausible ranges for dates, currencies, and identifiers; reconcile against a system of record where one exists; require the model to return unreadable rather than guess, and measure how often it does; run a second independent extraction on high-value fields and route disagreements to a human; and sample-audit output against the stored original forever, not just at launch.


Modality Scorecard

Modality Best-fit task Main failure mode
Photo or natural image Classification, moderation, attribute extraction Confident description of detail that is not present
Scanned or photographed page Reading a page with no text layer Silent character and digit misreads on poor scans
Complex-layout page Field extraction from forms and statements Values attached to the wrong label or column
Chart or diagram Describing a trend and reading labelled values Precise numbers invented where the chart labels none
Screenshot UI state understanding and support triage Text in the image treated as instructions to follow
Speech audio Transcription feeding a normal text pipeline Domain vocabulary and accents silently mistranscribed
Video Rare - keyframe sampling plus the audio transcript Cost, plus whatever happened between sampled frames

Cost, Latency, and Evaluation

Multimodal requests are materially more expensive and slower than text-only, for structural reasons rather than incidental ones: more tokens per request, and heavier prefill over those tokens. Both effects scale with resolution and page count. That pushes the same conclusions from three directions β€” do the expensive work once at ingestion rather than per query, triage aggressively so the model only sees pages that matter, downscale to the lowest resolution that passes your evals, and cache extraction results keyed on a content hash of the file so a re-upload of the same document is free.

Evaluation needs modality-specific ground truth, which is the part teams underestimate. A text eval set does not test whether a table was read correctly. You need labelled pages with the correct field values, labelled audio with the correct transcript and speaker turns, and labelled images with the correct class β€” and creating that requires a person looking at the original. Build it small and real: 30 genuinely difficult pages from your own corpus, including the bad scans, beats 500 clean synthetic ones. Score per field rather than per document, because a 95 percent document score that is wrong on the total every time is a failed system. The discipline is the same as Evals, applied to a new input type.


Bad to Good to Great

Bad - send every page to a vision model and trust the output

No triage, full resolution, free-text output parsed with regular expressions, no verification. It works in the demo. In production it costs an order of magnitude more than necessary, and every silent misread lands in your database indistinguishable from a correct value.

Good - route scans to vision, use a schema, spot-check the output

Text layer where there is one, vision where there is not, structured output, and someone reviews a sample each week. This is a working system. What is missing: per-page rather than per-document routing, resolution tuned by measurement, arithmetic and range checks on the fields, an abstention path, and a modality-specific eval set that would catch a regression before users do.

Great - triage per page, extract once into a schema, verify before persisting

  1. Original stored first in object storage, so every stage is re-runnable and auditable.
  2. Per-page triage between the text path and the vision path, with a cheap check deciding.
  3. Resolution and page selection tuned by eval, not by default settings.
  4. Strict schema output with an explicit unreadable value, so abstention is available and measurable.
  5. Field-level verification before persistence β€” arithmetic, format, range, reconciliation β€” with low-confidence pages routed to human review.
  6. Extraction cached on a content hash, so the expensive work happens once per document.
  7. A modality-specific eval set of real difficult pages, scored per field, gating changes to the pipeline.
  8. Ongoing sample audits against the stored original, because silent misreads never announce themselves.

When to Use

βœ… Reach for multimodal when:

❌ Stay text-only when:


Common Interview Questions

Q1: Why is an image more expensive than it looks?

Because it becomes tokens in the same context window as the text, and the count scales with resolution and with how many images you send. Providers publish the formula and it differs per model, so the number has to be looked up rather than assumed. Two consequences follow. Resolution is a dial: the right setting is the lowest at which your eval set still passes, which is high for dense tables and low for coarse classification. And page count multiplies everything, so sending a whole PDF as page images burns budget and context before the model reads anything β€” triage to the relevant pages first, usually with a cheap text pass.

Q2: When do you use a vision model instead of a text parser for documents?

When the structure a text parser has to infer is the thing that keeps breaking. A PDF is a page-description format, so multi-column layouts read out of order into plausible nonsense, tables flatten and lose which number belongs to which row, and scans have no text layer at all. A vision model sees the rendered layout, so it handles those natively. I decide per page rather than per document β€” a clean text layer goes down the cheap path, and only pages that fail triage go to vision. And I do the extraction once at ingestion into a schema, not per query, because per-query vision calls are expensive, slow, and uncacheable.

Q3: What is the most dangerous failure mode of a vision model, and how do you handle it?

A silent misread. A text parser that fails usually fails loudly with empty or garbled output; a vision model returns a fluent, plausible, wrong value with no error and no low-confidence signal. A faded total still gets a total, an unlabelled chart still gets a number. So verification cannot be optional for consequential fields. I make the schema allow an explicit unreadable value so abstention is possible and measurable, check the arithmetic the document itself constrains such as line items summing to a stated total, validate formats and ranges, reconcile against a system of record where one exists, run a second independent extraction on high-value fields and escalate disagreements to a human, and keep sample-auditing against the stored original in perpetuity.

Q4: How would you build search over a catalogue of images?

Default to captioning at ingestion and indexing the text. A vision model produces a rich description and structured attributes once, those go into the normal hybrid retrieval path, and the image reference sits alongside in object storage. This wins because text retrieval is cheap, well understood, already evaluated, and dense captions capture the specific detail that matters. True multimodal embeddings put images and text in a shared space, which is genuinely useful when the query itself is an image or when captioning the corpus is too expensive β€” but a shared space is weaker at fine distinctions than a same-modality one, since it only learned the correspondences present in its training pairs. Either way I measure it as retrieval, with recall at k against a labelled set.

Q5: How do you evaluate a multimodal pipeline?

With modality-specific ground truth, and scored per field rather than per document. A text eval set cannot tell you whether a table was read correctly, so I need labelled pages with correct field values, labelled audio with correct transcripts and speaker turns, and labelled images with correct classes β€” all of which requires a person looking at the original. I keep the set small and deliberately nasty: thirty genuinely hard pages from my own corpus, including the bad scans and the dense tables, beats hundreds of clean ones. Per-field scoring matters because a document that is 95 percent right and always wrong on the total is a failed system. And I evaluate the stages separately, since a summary built on a mistranscribed number is confidently wrong with nothing in the output pointing at the transcription.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access