Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 15 min read

Lesson 13 - Embeddings and Cosine Similarity

Part 3 - Give It Knowledge Lesson 13 of 24

Code: agentic-course/agentic/retrieval.py Tests: agentic-course/tests/test_retrieval.py Run it: python3 -m unittest tests.test_retrieval -v Concept: Embeddings and Semantic Similarity covers the theory and the interview framing, without code.


What you will build


The idea

An embedding turns text into a fixed-length list of numbers, positioned so that texts the model considers similar land near each other. That last clause is the whole lesson, and it is the clause everyone skips.

β€œSimilar” means whatever the training objective made it mean. Most embedding models are trained on co-occurrence: text that shows up in the same contexts gets placed nearby. Co-occurrence captures relatedness, not equivalence. β€œRefunds are permitted” and β€œrefunds are not permitted” co-occur constantly β€” same vocabulary, same documents, same paragraphs. So they embed close together, and an embedding search for your refund policy can confidently hand back the document stating the exact opposite of what you asked. This is not a bug in the index. It is the objective working as designed.

Hold onto that. Everything else in this lesson is arithmetic; this part is judgement.


The four functions

The entire dense-retrieval stack rests on these. From retrieval.py:

def dot(a: Iterable[float], b: Iterable[float]) -> float:
    return sum(x * y for x, y in zip(a, b))


def norm(a: Iterable[float]) -> float:
    return math.sqrt(sum(x * x for x in a))


def normalize(v: tuple[float, ...]) -> tuple[float, ...]:
    n = norm(v)
    return tuple(x / n for x in v) if n else v


def cosine(a: tuple[float, ...], b: tuple[float, ...]) -> float:
    da, db = norm(a), norm(b)
    if da == 0 or db == 0:
        return 0.0
    return dot(a, b) / (da * db)

Four lines of real maths. dot measures alignment, norm measures length, and cosine divides the first by the product of the second two β€” which cancels length out and leaves pure direction. The zero-vector guard is not decoration: an empty string embeds to all zeros, and test_cosine_handles_zero_vector_without_dividing_by_zero pins the behaviour so a blank chunk returns 0.0 instead of crashing retrieval.


A worked example you can check by hand

Take the query vector q = [1, 2, 2]. Its norm is sqrt(1 + 4 + 4) = sqrt(9) = 3. Now score it against three corpus vectors.

vector dot with q its norm cosine Euclidean distance from q
[4, 8, 8] 4+16+16 = 36 sqrt(144) = 12 36 / (3Β·12) = 1.000 sqrt(9+36+36) = 9.000
[2, 3, 6] 2+6+12 = 20 sqrt(49) = 7 20 / (3Β·7) = 0.952 sqrt(1+1+16) = 4.243
[-2, 1, 0] -2+2+0 = 0 sqrt(5) = 2.236 0 / (3Β·2.236) = 0.000 sqrt(9+1+4) = 3.742

[4, 8, 8] is exactly q scaled by 4 β€” same direction, different length β€” so cosine is exactly 1.0. [-2, 1, 0] is perpendicular, so the dot product is zero and cosine is zero. Those two rows are test_cosine_of_identical_direction_is_one and test_orthogonal_vectors_score_zero; the middle row is test_cosine_worked_example, asserting 20 / 21.

Run it against the real functions:

from agentic.retrieval import cosine, dot, norm

q = (1, 2, 2)
for v in [(4, 8, 8), (2, 3, 6), (-2, 1, 0)]:
    euclid = norm(tuple(a - b for a, b in zip(q, v)))
    print(v, "dot", dot(q, v), "norm", round(norm(v), 3),
          "cos", round(cosine(q, v), 4), "euclid", round(euclid, 3))

The ranking inversion

Read that last column again. By cosine the ranking is [4,8,8] then [2,3,6] then [-2,1,0] β€” perfect match first, perpendicular junk last. By Euclidean distance it is exactly inverted: the perpendicular vector is closest at 3.742, and the perfect directional match is furthest at 9.000.

Euclidean is not confused. It is answering a different question: how far apart are these points, not do they point the same way. [4, 8, 8] is on the same ray as the query but four times as long, so as points they sit far apart. In an embedding space, magnitude is mostly artifact β€” document length, token count, how the encoder pooled β€” while direction carries the meaning. A one-sentence and a one-page document on the same topic should score alike. That is the whole argument for cosine.

The nuance nobody mentions

If your vectors are already unit length, cosine and Euclidean produce the same ranking. Expand the squared distance between two unit vectors:

|a - b|^2  =  |a|^2 + |b|^2 - 2(a . b)
           =  1 + 1 - 2 cos
           =  2 - 2 cos

Distance is a strictly decreasing function of cosine, so sorting by one sorts by the other. Check it: cosine 1.000, 0.952, 0.000 gives unit distances 0.000, 0.309, 1.414. Same order.

So the real reason to default to cosine is not that it ranks better on normalized vectors β€” it ranks identically. It is that cosine is magnitude-invariant whether or not normalization happened. HashingEmbedder.embed returns normalize(...), so its output is unit length and test_embeddings_are_unit_length proves it. But the moment you swap in another embedder, or store a vector you built by averaging, or concatenate two embeddings, that invariant quietly dies. Cosine survives it. Euclidean does not.


The embedder in this repo is not semantic

Be clear about what you are running:

class HashingEmbedder:
    def __init__(self, dim: int = 64) -> None:
        self.dim = dim

    def embed(self, text: str) -> tuple[float, ...]:
        vec = [0.0] * self.dim
        for token in tokenize(text):
            slot = _stable_hash(token) % self.dim
            vec[slot] += 1.0
        return normalize(tuple(vec))

That is a bag-of-words hash. Each token is hashed into one of dim slots and counted. It has learned nothing. It cannot tell that β€œcheap” and β€œinexpensive” are related, because nothing ever taught it a relationship β€” score those two strings and you get 0.0. It exists so the mechanics of this course run offline with no API key and no pip install: dimensionality, normalization, cosine, indexing, filtering, fusion, and metrics are all exercised end to end. Every technique transfers. The semantic quality does not, because there is none.

One detail worth copying into your own code. The hash is hand-rolled FNV-1a:

def _stable_hash(s: str) -> int:
    h = 2166136261
    for ch in s.encode():
        h = ((h ^ ch) * 16777619) & 0xFFFFFFFF
    return h

Python’s built-in hash() for strings is salted per process, so the same token lands in a different slot on every run and your embeddings stop being reproducible β€” index today, query tomorrow, get nothing. test_embeddings_are_stable_across_calls names the reason in a comment and asserts e.embed("refund policy") == e.embed("refund policy").


Three operational rules

One model, both sides, always. A vector only has meaning relative to the model that produced it. Embed your corpus with model A and your queries with model B and you get numbers that look fine, rank plausibly, and are noise. Which means changing embedding model is a full re-index β€” a migration with dual writes, a backfill, and a cutover, not a config flip. Budget for it before you pick a model, because the cost of switching scales with your corpus, not with your enthusiasm.

A raw score means nothing on its own. In high dimensions, distances concentrate: the gap between the nearest and furthest neighbour shrinks as dimensionality grows, so almost everything sits in a narrow band and that band moves with your model, your corpus, and your query phrasing. A score of 0.72 is not β€œ72% relevant”. It is only interpretable relative to other scores in the same result set β€” is this hit well clear of the next three, or are the top ten all within 0.01? A hardcoded global cutoff like if score > 0.8 is a bug waiting for a model upgrade. VectorStore.search takes min_score=0.0, which drops only genuinely zero-overlap hits; treat anything stricter as needing calibration against labelled data you actually collected.

Give every chunk its address. This one costs nothing:

def with_context(self) -> str:
    prefix = " > ".join(p for p in (self.source, self.section) if p)
    return f"[{prefix}]\n{self.text}" if prefix else self.text

An isolated chunk is often uninterpretable. β€œIt must be renewed within 30 days” β€” what must? The heading answered that, and chunking threw it away. Prepending the path fixes both halves of the problem at once: the embedding now carries topical signal from the source and section, and the model reading the retrieved text can see where it came from. VectorStore.add embeds c.with_context(), not c.text, on purpose.

from agentic.retrieval import Chunk

c = Chunk(id="refund", text="Refunds are accepted within 30 days of delivery.",
          source="policy", section="Refunds")
print(c.with_context())
# [policy > Refunds]
# Refunds are accepted within 30 days of delivery.

Exercise

Write cosine similarity yourself, check it against the three worked vectors by hand, then confirm your numbers match the library exactly.

Success criterion: your function agrees with agentic.retrieval.cosine to nine decimal places on all three pairs, and you can state from memory why [4, 8, 8] scores 1.0 while [-2, 1, 0] scores 0.0.

Worked solution ```python import math from agentic.retrieval import cosine, norm def my_cosine(a, b): d = sum(x * y for x, y in zip(a, b)) na = math.sqrt(sum(x * x for x in a)) nb = math.sqrt(sum(x * x for x in b)) return 0.0 if na == 0 or nb == 0 else d / (na * nb) q = (1, 2, 2) expected = {(4, 8, 8): 1.0, (2, 3, 6): 20 / 21, (-2, 1, 0): 0.0} for v, want in expected.items(): mine, theirs = my_cosine(q, v), cosine(q, v) euclid = norm(tuple(x - y for x, y in zip(q, v))) assert abs(mine - want) < 1e-9, (v, mine, want) assert abs(mine - theirs) < 1e-9, (v, mine, theirs) print(f"{str(v):<12} cosine {mine:.6f} euclid {euclid:.3f}") print("\nranked by cosine :", sorted(expected, key=lambda v: -cosine(q, v))) print("ranked by euclid :", sorted(expected, key=lambda v: norm(tuple(x - y for x, y in zip(q, v))))) ``` Output: ```text (4, 8, 8) cosine 1.000000 euclid 9.000 (2, 3, 6) cosine 0.952381 euclid 4.243 (-2, 1, 0) cosine 0.000000 euclid 3.742 ranked by cosine : [(4, 8, 8), (2, 3, 6), (-2, 1, 0)] ranked by euclid : [(-2, 1, 0), (2, 3, 6), (4, 8, 8)] ``` The two rankings are exact reverses. `[4, 8, 8]` is `q` scaled by 4, so the angle between them is zero and cosine is `1.0` β€” but as points they are 9 units apart, which is why Euclidean puts the best match last.

What broke when I wrote this

The first HashingEmbedder used Python’s hash(). Every test passed. Embeddings computed inside a single test run agreed with each other, so nothing looked wrong β€” the hash is only re-salted between processes. The failure mode was persistence: embed a corpus in one run, query it in the next, get garbage. That bug does not show up in a test suite that builds its index in setUp, which is precisely why it is dangerous. Hence the hand-rolled FNV and the explicit stability test.


Checkpoint

Why can an embedding search return a document about the opposite of your query?

Because embedding models are trained on co-occurrence, which captures relatedness rather than equivalence. β€œRefunds are permitted” and β€œrefunds are not permitted” share vocabulary and appear in the same contexts, so they embed close together. Negation is a known weak spot of dense retrieval and a reason to pair it with lexical search.

Cosine and Euclidean give the same ranking on unit vectors. So why default to cosine?

Because squared distance between unit vectors equals 2 - 2Β·cosine, a strictly decreasing function, so the orders match exactly. The reason to prefer cosine is that it is magnitude-invariant whether or not the vectors were normalized β€” and the moment you average, concatenate, or swap embedders, the unit-length assumption silently breaks.

Your team upgrades the embedding model. What work does that create?

A full re-index of the corpus. Query and corpus vectors are only comparable when produced by the same model, so this is a migration: dual writes, a backfill, a cutover, and a rollback plan. Never a config flip.

Why is if score > 0.8: keep a bug?

Distances concentrate in high dimensions, so absolute scores are not calibrated and shift with the model, corpus, and query phrasing. A score is only meaningful relative to the other scores in the same result set. Use relative gaps, or calibrate a threshold against labelled data and re-calibrate on every model change.

What does HashingEmbedder teach you, and what does it not?

It teaches the mechanics β€” dimensionality, normalization, cosine, indexing, filtering, fusion, metrics β€” offline and with no dependencies. It teaches nothing about semantics, because it is a bag-of-words FNV hash that has learned nothing. β€œCheap” and β€œinexpensive” score 0.0.


Theory and interview framing: Become an AI Engineer

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access