Embeddings and Semantic Similarity - Complete Deep Dive
Prerequisites: The Math You Actually Need, How LLMs Actually Work, Model Selection and Routing Used in: Vector Databases, RAG End to End, Hybrid Search and Reranking Build it: Lesson 13 - Embeddings and Cosine Similarity implements this as runnable, tested code you can execute offline.
What is an Embedding?
An embedding is a learned vector β a fixed-length list of floats β that a model assigns to a piece of text, such that geometric closeness in that vector space approximates semantic relatedness. Two chunks about refund policies land near each other. A chunk about refund policies and one about GPU pricing land far apart. You never wrote a rule saying so; the model learned it from data.
Real-world analogy: an embedding model draws a map. On a city map, two hotels printed close together really are close in the city, so you can answer βwhat is near meβ with arithmetic on coordinates instead of reading every address. An embedding model draws that map for meaning. The catch is the one that matters for your system: the model drew the map, and you did not choose the legend.
Similar Means Whatever the Training Objective Made It Mean
Most embedding models are trained so that text appearing in related contexts ends up nearby. That objective rewards relatedness, not equivalence, and certainly not truth. The consequences are not academic:
"the contract does not auto-renew"and"the contract auto-renews"are about the same entity, the same clause, the same vocabulary. Their embeddings are nearly identical. One is the negation of the other.- Antonyms co-occur constantly β hot and cold, approve and reject, increase and decrease β so they sit close together.
- Numbers and identifiers barely move the vector.
invoice 4471andinvoice 7814are near-duplicates in embedding space. - Swapping an entity is a small geometric move but a total semantic one.
ACME's termination rightsvsGlobex's termination rights.
This is why embedding search cheerfully returns a document about the opposite of your query and reports high confidence. The vector is not wrong β it is answering the question it was trained on (βis this the same topic?β) rather than the one you asked (βdoes this say what I need?β).
Mitigations exist and you should know them by name: a cross-encoder reranker reads the query and the candidate together and can detect the negation, lexical matching pins exact identifiers, and a grounding check makes the generator cite the span it relied on. See Hybrid Search and Reranking. Do not try to fix this by picking a better embedding model. It is a property of the objective, not a defect of one vendor.
Dense vs Sparse Representations
| Β | Dense embeddings | Sparse representations |
|---|---|---|
| Shape | Few hundred to few thousand dims, nearly all non-zero | Vocabulary-sized, almost entirely zero |
| Produced by | A neural encoder | Term statistics such as BM25, or a learned sparse model |
| Matches on | Meaning and paraphrase | Exact tokens and their weights |
| Strength | "how do I get my money back" finds a refund policy that never says βmoney backβ |
Finds ERR_TLS_2049 and product SKUs exactly |
| Weakness | Blind to rare literal strings it never learned | Blind to synonyms and paraphrase |
| Interpretability | Effectively none | You can read which terms fired |
Neither is a superset of the other, which is the entire argument for hybrid retrieval. Dense handles the vocabulary mismatch between how users ask and how documents are written; sparse handles the long tail of exact strings that matter enormously in technical and legal corpora.
Dimensionality and the Storage-Recall Tradeoff
Dimension count is a capacity knob. More dimensions give the model more room to encode fine distinctions, up to the point where the extra room is spent on noise. Cost, however, grows with no such ceiling: memory, index build time, and query latency all scale roughly linearly in dimensions, and memory is the binding constraint in practice (see the sizing math in Vector Databases).
Two levers reduce the bill without re-training anything:
Matryoshka-style truncatable embeddings. Some models are trained so the prefix of the vector is itself a usable embedding β the most important information is packed into the leading dimensions. You can slice a 1024-dim vector down to 256 and keep most of the retrieval quality, at exactly one quarter of the storage. This only works for models explicitly trained that way. Truncating an ordinary embedding destroys it.
Quantization. Store each dimension as int8 instead of float32 for a 4x cut, or product-quantize for far more. This is an index-level decision, covered in the next page.
The honest procedure for picking dimensions: build a small labelled eval set from real queries, measure recall at your chosen k for full, truncated, and quantized variants, and take the cheapest one that clears your bar. Guessing is how teams end up paying for 3072 dimensions to serve a FAQ.
Normalization and Why Cosine Is the Standard
Cosine similarity measures the angle between two vectors and ignores their length. That is the desirable property, because vector magnitude in a text encoder mostly tracks incidental things β length, token frequency, formatting β rather than meaning.
q = [1, 2, 2] the query
d1 = [2, 2, 1] a genuinely related document
d2 = [2, -1, 0] an unrelated document
d3 = [2, 4, 4] exactly q's direction but twice as long
magnitudes
|q| = sqrt(1 + 4 + 4) = sqrt(9) = 3
|d1| = sqrt(4 + 4 + 1) = sqrt(9) = 3
|d2| = sqrt(4 + 1 + 0) = sqrt(5) = 2.236
|d3| = sqrt(4 + 16 + 16) = sqrt(36) = 6
dot products
q . d1 = 1*2 + 2*2 + 2*1 = 2 + 4 + 2 = 8
q . d2 = 1*2 + 2*-1 + 2*0 = 2 - 2 + 0 = 0
q . d3 = 1*2 + 2*4 + 2*4 = 2 + 8 + 8 = 18
cosine = dot / (|a| * |b|)
cos(q, d1) = 8 / (3 * 3) = 8/9 = 0.889
cos(q, d2) = 0 / (3 * 2.236) = 0.000 orthogonal
cos(q, d3) = 18 / (3 * 6) = 1.000 identical direction
Read the ranking twice. By raw dot product, d3 scores 18 and d1 scores 8 β but d3 is literally 2 * q, carrying no more information than a unit-length copy would. Unnormalized inner product ranking is partly a contest over vector length, and length usually correlates with nothing you care about.
After L2 normalization every vector has length 1, and three convenient facts fall out:
- Cosine and inner product become the same function, so an index configured for inner product gives cosine ranking for free.
- Euclidean distance becomes a monotone transform of cosine:
||a - b||Β² = 2 - 2 * cos(a, b). Check the endpoints βcos = 1gives distance 0,cos = 0givessqrt(2),cos = -1gives 2. - Therefore cosine, inner product, and L2 produce identical rankings on normalized vectors. The metric argument people have in design reviews is usually moot; the real question is whether normalization happened at all.
One calibration note: in real text embeddings, cosine scores are compressed into a narrow positive band. Unrelated text often scores 0.6-0.7 and opposite-meaning text scores above 0.9. A raw score of 0.85 is not β85% confidentβ and cos = -1 essentially never occurs. Never hard-code an absolute similarity threshold without measuring its distribution on your own corpus, and re-measure it after any model change.
Symmetric vs Asymmetric Search
The two ends of a retrieval query are not the same kind of text. A query is short, imperative, and often ungrammatical. A corpus chunk is long, declarative, and prose. This is asymmetric search.
| Β | Symmetric | Asymmetric |
|---|---|---|
| Both sides look like | Same shape β sentence vs sentence | Short query vs long passage |
| Typical task | Deduplication, clustering, near-duplicate detection, STS | Question answering, document retrieval, RAG |
| Trained with | Sentence-pair similarity objectives | Query-passage objectives, often with hard negatives |
| Failure if you pick wrong | Queries retrieve other queries-shaped text and prefer short chunks | Paraphrase detection gets sloppy |
Models built for asymmetric retrieval frequently expect instruction prefixes β a distinct marker on queries versus documents, such as a query: and passage: convention. These are part of the model contract, not decoration. Omitting them, or applying the query prefix at index time, silently costs recall without erroring. Read the model card, apply the prefixes consistently, and record which prefix scheme the index was built with.
One Model, One Index
Query vectors and corpus vectors must come from the same model at the same version with the same preprocessing. Embedding spaces from different models are unrelated coordinate systems; comparing across them produces confident nonsense rather than an error.
flowchart LR
subgraph IndexTime["Index Time - offline and batched"]
A[Source documents] --> B[Parse and chunk]
B --> C[Embedding model at pinned version]
C --> D[Vector index plus metadata]
end
subgraph QueryTime["Query Time - online and latency bound"]
E[User query] --> F[Embedding model at pinned version]
F --> G[ANN search over the index]
G --> H[Top k chunks for the generator]
end
D --> G
C -.->|same weights and same prefix scheme| F
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class E client
class A,B async
class C,F service
class D data
class G,H edge
So changing the embedding model is a data migration, not a config change. Every vector in the corpus must be recomputed. Budget for it properly:
- Pin the model identifier and dimension count in the index metadata, and log it on every query. This turns a mystery outage into a one-line diff.
- Re-embed the full corpus into a new index rather than mutating the live one. For a large corpus this is a real batch job with real cost.
- Evaluate old versus new on the same eval set before cutting over. A newer model is not automatically better on your corpus.
- Cut over atomically behind an alias, keep the old index warm, and be able to flip back. Treat it like any other risky release β the deployment reliability playbook applies directly.
An upgrade that looks free because the API is a drop-in replacement is the most expensive mistake in this section.
Domain Mismatch, Chunk Granularity, and Languages
Domain mismatch. A general-purpose model learned the public webβs sense of relatedness. Hand it dense domain jargon β ICD codes, statutory citations, chemical names, internal service codenames β and terms that are worlds apart to a specialist collapse into one neighbourhood, because the model never saw enough of them to separate them. Symptoms: recall that looks acceptable on generic questions and falls apart on the ones experts actually ask. Options, cheapest first: keep a lexical channel so exact codes still match, add a reranker, expand acronyms at ingest, and only then consider a domain-adapted or fine-tuned encoder.
Chunk-level vs document-level. One vector per document is an average of everything the document says, so a long documentβs vector is a blur that matches nothing sharply β and even when it wins, you have to feed the whole document to the generator. One vector per chunk is sharp and cheap to feed, but a chunk stripped of its surrounding context can be uninterpretable on its own. The practical answer is chunk-level embedding plus enough retained context to make each chunk self-describing, which is exactly the job of Document Parsing and Chunking.
Multilingual behaviour. Multilingual models place translations of the same sentence near each other, so a Hindi query can retrieve an English passage. Two caveats: quality is uneven across languages, usually tracking training volume, and cross-lingual scores are not comparable to same-language scores, so a single global threshold will quietly favour one language. If cross-lingual retrieval is a real requirement, build a per-language eval slice.
Failure Modes
| Failure mode | Symptom | Fix |
|---|---|---|
| Model mismatch between index and query | Results look random; scores cluster in a narrow band | Log model ID and dimension per query; assert they match the index metadata |
| Missing or wrong instruction prefix | Quiet recall loss; short chunks win too often | Apply the model cardβs query and passage prefixes; record the scheme in index metadata |
| Unnormalized vectors with inner product | Long or verbose chunks dominate every result list | L2 normalize at index and query time |
| Hard-coded similarity threshold | Everything passes, or nothing does, after a model change | Calibrate on a labelled set; store the threshold with the model version |
| Negation and antonym confusion | Retrieved chunk states the opposite of the answer | Cross-encoder rerank; require the generator to cite a span |
| Exact identifiers not found | SKUs, error codes, and ticket IDs never retrieve | Add a sparse or lexical channel and fuse the results |
| Domain jargon collapse | Generic queries fine, expert queries fail | Lexical channel, acronym expansion, domain-adapted encoder |
| Chunks too small or context-free | Retrieved text is on-topic but unusable | Prepend heading context; use the parent-document pattern |
| Dimensions chosen by vibes | Memory bill dwarfs the value delivered | Measure recall at k for truncated and quantized variants; pick the cheapest that clears the bar |
| Stale index after source edits | Confidently cited answers quote deleted text | Incremental re-index keyed on content hash |
When to Use
β Use embeddings when:
- Users phrase things differently from your documents and keyword search misses the paraphrase
- You need semantic grouping β clustering, deduplication, near-duplicate detection, topic routing
- You are building retrieval for a generator and need a cheap, scalable first-stage candidate fetch
- You want classification without labelled data by comparing against a few labelled exemplars
β Donβt use embeddings alone when:
- The query is an exact identifier β error code, SKU, invoice number, commit SHA
- The answer requires aggregation or computation over records; that is a SQL query, not a similarity search
- Correctness hinges on negation, dates, quantities, or entity precision, without a reranker or verification step behind it
- The corpus is small enough that keyword search plus a filter already satisfies users β measure before adding an index
Common Interview Questions
Q1: Why cosine similarity instead of Euclidean distance?
Cosine ignores vector magnitude, and magnitude in a text encoder mostly reflects incidental properties like length rather than meaning. The subtler point: once vectors are L2 normalized, cosine, inner product, and Euclidean distance all produce the same ranking, since
||a - b||Β² = 2 - 2 * cos(a, b)is monotone in cosine. So the metric choice is largely moot on normalized vectors, and the real question is whether normalization happens consistently at index and query time. If it does not, inner product ranking becomes a contest over vector length.
Q2: Embedding search returned a document stating the exact opposite of the query. Why, and how do you fix it?
Because embedding models are trained for relatedness, not equivalence. A sentence and its negation share topic, entities, and vocabulary, so they sit very close together β often above 0.9 cosine. This is a property of the training objective, not a bug in one vendorβs model, so swapping models will not fix it. The fixes are layered: rerank the top candidates with a cross-encoder that reads query and passage jointly and can see the negation, keep a lexical channel for exact terms, and require the generator to cite the span it used so a grounding check can catch the contradiction.
Q3: What actually happens if you swap the embedding model on a live system?
Retrieval quality collapses, and usually without any error being raised. The new modelβs output lives in a different coordinate system than the indexed vectors, so similarity scores become meaningless β and if dimensions match, nothing even complains. Swapping models requires re-embedding the entire corpus into a new index, evaluating old versus new on the same eval set, then cutting over behind an alias with the old index kept warm for rollback. Pin the model ID and dimension in index metadata and log them per query, so a mismatch is a one-line diff instead of a week of debugging.
Q4: How do you pick the embedding dimension?
Empirically, not by reputation. Higher dimensions buy capacity for fine distinctions but cost memory, build time, and query latency roughly linearly, and memory is what dominates the bill. Build a labelled eval set from real queries, measure recall at your target k across candidate models and dimensions, and take the cheapest configuration that clears your quality bar. If the model supports Matryoshka-style truncation, test the truncated prefix too β a quarter of the dimensions often retains most of the quality, which is a direct 4x storage saving. Note that truncation only works on models explicitly trained for it.
Q5: What is asymmetric search and why does it change your model choice?
In retrieval the two sides differ in shape: a short, often ungrammatical query against a long declarative passage. Models trained on sentence-pair similarity assume both sides look alike, so they underperform here β queries tend to retrieve query-shaped text and short chunks win too easily. Models trained on query-passage pairs with hard negatives handle the asymmetry, and many of them require distinct instruction prefixes for queries versus documents. Those prefixes are part of the contract: getting them wrong costs recall silently, so record the scheme in index metadata alongside the model version.
Build it in code: Agentic AI Course · Fundamentals: Core Concepts