Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 16 min read

Embeddings and Semantic Similarity - Complete Deep Dive

Stage 3 - Retrieval Lesson 10 of 36

Prerequisites: The Math You Actually Need, How LLMs Actually Work, Model Selection and Routing Used in: Vector Databases, RAG End to End, Hybrid Search and Reranking Build it: Lesson 13 - Embeddings and Cosine Similarity implements this as runnable, tested code you can execute offline.


What is an Embedding?

An embedding is a learned vector β€” a fixed-length list of floats β€” that a model assigns to a piece of text, such that geometric closeness in that vector space approximates semantic relatedness. Two chunks about refund policies land near each other. A chunk about refund policies and one about GPU pricing land far apart. You never wrote a rule saying so; the model learned it from data.

Real-world analogy: an embedding model draws a map. On a city map, two hotels printed close together really are close in the city, so you can answer β€œwhat is near me” with arithmetic on coordinates instead of reading every address. An embedding model draws that map for meaning. The catch is the one that matters for your system: the model drew the map, and you did not choose the legend.


Similar Means Whatever the Training Objective Made It Mean

Most embedding models are trained so that text appearing in related contexts ends up nearby. That objective rewards relatedness, not equivalence, and certainly not truth. The consequences are not academic:

This is why embedding search cheerfully returns a document about the opposite of your query and reports high confidence. The vector is not wrong β€” it is answering the question it was trained on (β€œis this the same topic?”) rather than the one you asked (β€œdoes this say what I need?”).

Mitigations exist and you should know them by name: a cross-encoder reranker reads the query and the candidate together and can detect the negation, lexical matching pins exact identifiers, and a grounding check makes the generator cite the span it relied on. See Hybrid Search and Reranking. Do not try to fix this by picking a better embedding model. It is a property of the objective, not a defect of one vendor.


Dense vs Sparse Representations

Β  Dense embeddings Sparse representations
Shape Few hundred to few thousand dims, nearly all non-zero Vocabulary-sized, almost entirely zero
Produced by A neural encoder Term statistics such as BM25, or a learned sparse model
Matches on Meaning and paraphrase Exact tokens and their weights
Strength "how do I get my money back" finds a refund policy that never says β€œmoney back” Finds ERR_TLS_2049 and product SKUs exactly
Weakness Blind to rare literal strings it never learned Blind to synonyms and paraphrase
Interpretability Effectively none You can read which terms fired

Neither is a superset of the other, which is the entire argument for hybrid retrieval. Dense handles the vocabulary mismatch between how users ask and how documents are written; sparse handles the long tail of exact strings that matter enormously in technical and legal corpora.


Dimensionality and the Storage-Recall Tradeoff

Dimension count is a capacity knob. More dimensions give the model more room to encode fine distinctions, up to the point where the extra room is spent on noise. Cost, however, grows with no such ceiling: memory, index build time, and query latency all scale roughly linearly in dimensions, and memory is the binding constraint in practice (see the sizing math in Vector Databases).

Two levers reduce the bill without re-training anything:

Matryoshka-style truncatable embeddings. Some models are trained so the prefix of the vector is itself a usable embedding β€” the most important information is packed into the leading dimensions. You can slice a 1024-dim vector down to 256 and keep most of the retrieval quality, at exactly one quarter of the storage. This only works for models explicitly trained that way. Truncating an ordinary embedding destroys it.

Quantization. Store each dimension as int8 instead of float32 for a 4x cut, or product-quantize for far more. This is an index-level decision, covered in the next page.

The honest procedure for picking dimensions: build a small labelled eval set from real queries, measure recall at your chosen k for full, truncated, and quantized variants, and take the cheapest one that clears your bar. Guessing is how teams end up paying for 3072 dimensions to serve a FAQ.


Normalization and Why Cosine Is the Standard

Cosine similarity measures the angle between two vectors and ignores their length. That is the desirable property, because vector magnitude in a text encoder mostly tracks incidental things β€” length, token frequency, formatting β€” rather than meaning.

q  = [1,  2, 2]     the query
d1 = [2,  2, 1]     a genuinely related document
d2 = [2, -1, 0]     an unrelated document
d3 = [2,  4, 4]     exactly q's direction but twice as long

magnitudes
  |q|  = sqrt(1 + 4 + 4)   = sqrt(9)  = 3
  |d1| = sqrt(4 + 4 + 1)   = sqrt(9)  = 3
  |d2| = sqrt(4 + 1 + 0)   = sqrt(5)  = 2.236
  |d3| = sqrt(4 + 16 + 16) = sqrt(36) = 6

dot products
  q . d1 = 1*2 + 2*2  + 2*1 = 2 + 4 + 2 = 8
  q . d2 = 1*2 + 2*-1 + 2*0 = 2 - 2 + 0 = 0
  q . d3 = 1*2 + 2*4  + 2*4 = 2 + 8 + 8 = 18

cosine = dot / (|a| * |b|)
  cos(q, d1) = 8  / (3 * 3)     = 8/9   = 0.889
  cos(q, d2) = 0  / (3 * 2.236) = 0.000     orthogonal
  cos(q, d3) = 18 / (3 * 6)     = 1.000     identical direction

Read the ranking twice. By raw dot product, d3 scores 18 and d1 scores 8 β€” but d3 is literally 2 * q, carrying no more information than a unit-length copy would. Unnormalized inner product ranking is partly a contest over vector length, and length usually correlates with nothing you care about.

After L2 normalization every vector has length 1, and three convenient facts fall out:

One calibration note: in real text embeddings, cosine scores are compressed into a narrow positive band. Unrelated text often scores 0.6-0.7 and opposite-meaning text scores above 0.9. A raw score of 0.85 is not β€œ85% confident” and cos = -1 essentially never occurs. Never hard-code an absolute similarity threshold without measuring its distribution on your own corpus, and re-measure it after any model change.


The two ends of a retrieval query are not the same kind of text. A query is short, imperative, and often ungrammatical. A corpus chunk is long, declarative, and prose. This is asymmetric search.

Β  Symmetric Asymmetric
Both sides look like Same shape β€” sentence vs sentence Short query vs long passage
Typical task Deduplication, clustering, near-duplicate detection, STS Question answering, document retrieval, RAG
Trained with Sentence-pair similarity objectives Query-passage objectives, often with hard negatives
Failure if you pick wrong Queries retrieve other queries-shaped text and prefer short chunks Paraphrase detection gets sloppy

Models built for asymmetric retrieval frequently expect instruction prefixes β€” a distinct marker on queries versus documents, such as a query: and passage: convention. These are part of the model contract, not decoration. Omitting them, or applying the query prefix at index time, silently costs recall without erroring. Read the model card, apply the prefixes consistently, and record which prefix scheme the index was built with.


One Model, One Index

Query vectors and corpus vectors must come from the same model at the same version with the same preprocessing. Embedding spaces from different models are unrelated coordinate systems; comparing across them produces confident nonsense rather than an error.

flowchart LR
    subgraph IndexTime["Index Time - offline and batched"]
        A[Source documents] --> B[Parse and chunk]
        B --> C[Embedding model at pinned version]
        C --> D[Vector index plus metadata]
    end

    subgraph QueryTime["Query Time - online and latency bound"]
        E[User query] --> F[Embedding model at pinned version]
        F --> G[ANN search over the index]
        G --> H[Top k chunks for the generator]
    end

    D --> G
    C -.->|same weights and same prefix scheme| F

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class E client
    class A,B async
    class C,F service
    class D data
    class G,H edge

So changing the embedding model is a data migration, not a config change. Every vector in the corpus must be recomputed. Budget for it properly:

  1. Pin the model identifier and dimension count in the index metadata, and log it on every query. This turns a mystery outage into a one-line diff.
  2. Re-embed the full corpus into a new index rather than mutating the live one. For a large corpus this is a real batch job with real cost.
  3. Evaluate old versus new on the same eval set before cutting over. A newer model is not automatically better on your corpus.
  4. Cut over atomically behind an alias, keep the old index warm, and be able to flip back. Treat it like any other risky release β€” the deployment reliability playbook applies directly.

An upgrade that looks free because the API is a drop-in replacement is the most expensive mistake in this section.


Domain Mismatch, Chunk Granularity, and Languages

Domain mismatch. A general-purpose model learned the public web’s sense of relatedness. Hand it dense domain jargon β€” ICD codes, statutory citations, chemical names, internal service codenames β€” and terms that are worlds apart to a specialist collapse into one neighbourhood, because the model never saw enough of them to separate them. Symptoms: recall that looks acceptable on generic questions and falls apart on the ones experts actually ask. Options, cheapest first: keep a lexical channel so exact codes still match, add a reranker, expand acronyms at ingest, and only then consider a domain-adapted or fine-tuned encoder.

Chunk-level vs document-level. One vector per document is an average of everything the document says, so a long document’s vector is a blur that matches nothing sharply β€” and even when it wins, you have to feed the whole document to the generator. One vector per chunk is sharp and cheap to feed, but a chunk stripped of its surrounding context can be uninterpretable on its own. The practical answer is chunk-level embedding plus enough retained context to make each chunk self-describing, which is exactly the job of Document Parsing and Chunking.

Multilingual behaviour. Multilingual models place translations of the same sentence near each other, so a Hindi query can retrieve an English passage. Two caveats: quality is uneven across languages, usually tracking training volume, and cross-lingual scores are not comparable to same-language scores, so a single global threshold will quietly favour one language. If cross-lingual retrieval is a real requirement, build a per-language eval slice.


Failure Modes

Failure mode Symptom Fix
Model mismatch between index and query Results look random; scores cluster in a narrow band Log model ID and dimension per query; assert they match the index metadata
Missing or wrong instruction prefix Quiet recall loss; short chunks win too often Apply the model card’s query and passage prefixes; record the scheme in index metadata
Unnormalized vectors with inner product Long or verbose chunks dominate every result list L2 normalize at index and query time
Hard-coded similarity threshold Everything passes, or nothing does, after a model change Calibrate on a labelled set; store the threshold with the model version
Negation and antonym confusion Retrieved chunk states the opposite of the answer Cross-encoder rerank; require the generator to cite a span
Exact identifiers not found SKUs, error codes, and ticket IDs never retrieve Add a sparse or lexical channel and fuse the results
Domain jargon collapse Generic queries fine, expert queries fail Lexical channel, acronym expansion, domain-adapted encoder
Chunks too small or context-free Retrieved text is on-topic but unusable Prepend heading context; use the parent-document pattern
Dimensions chosen by vibes Memory bill dwarfs the value delivered Measure recall at k for truncated and quantized variants; pick the cheapest that clears the bar
Stale index after source edits Confidently cited answers quote deleted text Incremental re-index keyed on content hash

When to Use

βœ… Use embeddings when:

❌ Don’t use embeddings alone when:


Common Interview Questions

Q1: Why cosine similarity instead of Euclidean distance?

Cosine ignores vector magnitude, and magnitude in a text encoder mostly reflects incidental properties like length rather than meaning. The subtler point: once vectors are L2 normalized, cosine, inner product, and Euclidean distance all produce the same ranking, since ||a - b||Β² = 2 - 2 * cos(a, b) is monotone in cosine. So the metric choice is largely moot on normalized vectors, and the real question is whether normalization happens consistently at index and query time. If it does not, inner product ranking becomes a contest over vector length.

Q2: Embedding search returned a document stating the exact opposite of the query. Why, and how do you fix it?

Because embedding models are trained for relatedness, not equivalence. A sentence and its negation share topic, entities, and vocabulary, so they sit very close together β€” often above 0.9 cosine. This is a property of the training objective, not a bug in one vendor’s model, so swapping models will not fix it. The fixes are layered: rerank the top candidates with a cross-encoder that reads query and passage jointly and can see the negation, keep a lexical channel for exact terms, and require the generator to cite the span it used so a grounding check can catch the contradiction.

Q3: What actually happens if you swap the embedding model on a live system?

Retrieval quality collapses, and usually without any error being raised. The new model’s output lives in a different coordinate system than the indexed vectors, so similarity scores become meaningless β€” and if dimensions match, nothing even complains. Swapping models requires re-embedding the entire corpus into a new index, evaluating old versus new on the same eval set, then cutting over behind an alias with the old index kept warm for rollback. Pin the model ID and dimension in index metadata and log them per query, so a mismatch is a one-line diff instead of a week of debugging.

Q4: How do you pick the embedding dimension?

Empirically, not by reputation. Higher dimensions buy capacity for fine distinctions but cost memory, build time, and query latency roughly linearly, and memory is what dominates the bill. Build a labelled eval set from real queries, measure recall at your target k across candidate models and dimensions, and take the cheapest configuration that clears your quality bar. If the model supports Matryoshka-style truncation, test the truncated prefix too β€” a quarter of the dimensions often retains most of the quality, which is a direct 4x storage saving. Note that truncation only works on models explicitly trained for it.

Q5: What is asymmetric search and why does it change your model choice?

In retrieval the two sides differ in shape: a short, often ungrammatical query against a long declarative passage. Models trained on sentence-pair similarity assume both sides look alike, so they underperform here β€” queries tend to retrieve query-shaped text and short chunks win too easily. Models trained on query-passage pairs with hard negatives handle the asymmetry, and many of them require distinct instruction prefixes for queries versus documents. Those prefixes are part of the contract: getting them wrong costs recall silently, so record the scheme in index metadata alongside the model version.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access