The Math You Actually Need - Complete Deep Dive
Prerequisites: The AI Engineer Role, How LLMs Actually Work Used in: Embeddings, Vector Databases, Hybrid Search and Reranking, Latency Engineering
What is βthe math you actually needβ?
It is a short, finite list. Six ideas, all of which fit on this page, none of which require calculus. Product-layer AI work uses maths the way backend work uses maths: you need to read the numbers your tools produce and know what they mean, not derive the tools.
Real-world analogy: you use TLS every day without knowing the proofs behind elliptic curves. But you do need to understand certificate chains, expiry, and what a hostname mismatch means - because those are the things that break, and the error message assumes you know them. Same deal here. Cosine similarity is your certificate chain: you will see the number constantly, and you need to know what a 0.71 means and when it is lying to you.
The βI need a maths degree firstβ blocker is the single most common reason competent engineers never start. It is wrong, and it costs people years. Here is the honest list.
The Short List
| Idea | Why it is on the list | Depth you need |
|---|---|---|
| Vectors | Embeddings are vectors. Everything retrieval touches is a vector | What a coordinate list is. Addition and scaling |
| Dot product | The single operation underneath similarity search | Multiply pairwise, add up. One line of code |
| Cosine similarity | The default ranking metric in every vector store | The formula, plus why magnitude is excluded |
| Basic probability | The model emits a distribution. Sampling, retries, and routing are all probability | Distributions, independent events, expected value |
| Logarithms | Log-probs, perplexity, and cross-entropy are all logged for concrete reasons | What a log does to a product. Why negative log-probs are positive |
| Percentiles | Latency work is entirely percentile work | p50 versus p95 versus p99, and why you cannot average them |
That is it. If you know these six, no part of the product-layer curriculum will be blocked by maths.
1. Vectors and Embeddings
A vector is an ordered list of numbers. [1, 2, 2] is a three-dimensional vector. An embedding is a vector produced by a model from some input - text, an image, an audio clip - such that inputs with similar meaning land near each other in that space. Production text embeddings commonly have a few hundred to a few thousand dimensions.
You need two operations and one property:
- Scaling - multiply every component by a constant.
2 * [1, 2, 2]is[2, 4, 4]. Note this does not change the direction, which turns out to be the whole point below. - Length, or L2 norm - Pythagoras extended to any number of dimensions. For
[1, 2, 2]the norm issqrt(1 + 4 + 4), which issqrt(9), which is exactly3. - Normalizing - divide a vector by its norm to get a unit vector, length 1, same direction.
[1, 2, 2]normalizes to[0.333, 0.667, 0.667].
That is the entire linear algebra prerequisite for retrieval work.
2. Dot Product
Multiply corresponding components, then sum. For a = [1, 2, 2] and b = [2, 3, 6]:
(1 x 2) + (2 x 3) + (2 x 6) = 2 + 6 + 12 = 20
Two facts make this the workhorse:
- It is large when the two vectors point in a similar direction, near zero when they are unrelated, and negative when they point opposingly.
- It is one fused multiply-add per dimension, which modern hardware does absurdly fast. A vector search over a million candidates is a million dot products, and that is why vector search at scale is feasible at all.
The catch: the dot product also grows with the length of the vectors. [4, 8, 8] and [1, 2, 2] point in exactly the same direction, but the dot product with a query is four times larger for the longer one. That is the problem cosine similarity solves.
3. Cosine Similarity - A Worked Example
Cosine similarity divides the dot product by both norms, which cancels out magnitude and leaves only direction. It ranges from -1 to 1, where 1 means identical direction and 0 means unrelated.
cos(a, b) = dot(a, b) / (norm(a) x norm(b))
Take a query vector q = [1, 2, 2], with norm(q) = 3, and three candidate documents. Every number below is verifiable by hand.
| Doc | Vector | dot with q | norm | Cosine | Euclidean distance to q |
|---|---|---|---|---|---|
| d1 | [4, 8, 8] | 4 + 16 + 16 = 36 | sqrt(144) = 12 | 36 / (3 x 12) = 1.000 | sqrt(9 + 36 + 36) = 9.000 |
| d2 | [2, 3, 6] | 2 + 6 + 12 = 20 | sqrt(49) = 7 | 20 / (3 x 7) = 0.952 | sqrt(1 + 1 + 16) = 4.243 |
| d3 | [-2, 1, 0] | -2 + 2 + 0 = 0 | sqrt(5) = 2.236 | 0 / (3 x 2.236) = 0.000 | sqrt(9 + 1 + 4) = 3.742 |
Read the two ranking columns against each other, because this is the actual lesson:
- By cosine: d1 then d2 then d3. Correct.
d1isqscaled by four, so it is a perfect direction match;d3is orthogonal toq, so it is unrelated. - By Euclidean distance: d3 then d2 then d1. Backwards. It ranks the completely unrelated document first and the perfect match last, purely because
d1is a long vector and raw distance punishes that.
That is the intuition for why cosine and not Euclidean: magnitude in an embedding is mostly an artifact, and direction carries the meaning. Two documents about caching should rank as similar whether one is a sentence or a page.
The nuance worth knowing
If every vector is already normalized to unit length, cosine and Euclidean give the same ranking. The algebra: for unit vectors, distance^2 = 2 - 2 x cosine, so distance is a strictly decreasing function of cosine. Check it against the table by normalizing first - d1 normalizes to exactly qβs unit vector, giving cosine 1 and distance sqrt(2 - 2) which is 0; d2 gives cosine 0.952 and distance sqrt(0.095) which is 0.309; d3 gives cosine 0 and distance sqrt(2) which is 1.414. Same order, now correct.
So the real reason to default to cosine is not that Euclidean is wrong in principle. It is that cosine is magnitude-invariant whether or not the vectors happened to be normalized, so it cannot be silently broken by an embedding model, a chunking change, or a migration that skips normalization. It is the safe default. This is also why vector stores offer cosine, dot product, and L2 as separate options - on normalized vectors the first two are the same computation.
import math
def dot(a, b):
return sum(x * y for x, y in zip(a, b))
def norm(a):
return math.sqrt(dot(a, a))
def cosine(a, b):
return dot(a, b) / (norm(a) * norm(b))
q = [1, 2, 2]
for name, d in [("d1", [4, 8, 8]), ("d2", [2, 3, 6]), ("d3", [-2, 1, 0])]:
print(name, round(cosine(q, d), 3))
# d1 1.0
# d2 0.952
# d3 0.0
Write this once, confirm it matches your vector storeβs scores, and you never have to wonder what the number means again.
Where it sits in a real pipeline
flowchart LR
A[User query text] --> B[Embedding model]
B --> C[Query vector]
C --> D[Vector index<br>approximate nearest neighbour]
D --> E[Cosine scores<br>one per candidate chunk]
E --> F[Top k chunks by score]
F --> G[Prompt context]
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class A,G client
class B,F service
class C,E data
class D edge
4. What High Dimensionality Does to Intuition
Your geometric intuition was built in three dimensions and it does not survive the trip to several hundred. Three consequences that matter operationally:
- Distances concentrate. As dimensions grow, the distance between the nearest and farthest points in a random set shrinks toward each other proportionally. βNearestβ becomes a weaker signal than you expect, which is why raw similarity scores are poor absolute quality measures. A cosine of 0.71 is not inherently good - it is only meaningful relative to the other scores in the same result set. Never hardcode a global score threshold without measuring it on your own data.
- Random directions are almost always near-orthogonal. Pick two random high-dimensional vectors and their cosine will be close to 0. So a meaningful similarity is genuinely informative - the space is mostly empty, and unrelated things do not accidentally look related. This is also why a few hundred dimensions can encode an enormous number of distinguishable concepts.
- Exact nearest-neighbour search stops being tractable. The tree structures that make low-dimensional spatial search fast - see Geospatial - degrade to full scans in high dimensions. That is the entire reason approximate nearest neighbour indexes exist, and why you trade a little recall for a lot of speed. Covered in Vector Databases.
You do not need to prove any of this. You need to stop being surprised by it.
5. Probability You Actually Use
Three pieces, all of which show up weekly.
Distributions and sampling. The modelβs output layer is a probability distribution over the vocabulary, and the sampler draws from it. That is the whole mechanism behind temperature and top-p, covered in How LLMs Actually Work. Knowing βthese numbers are non-negative and sum to 1β is enough to read it.
Independent events multiply. This is the most useful single fact in agent design. If one step succeeds 95 percent of the time and a chain has five independent steps, end-to-end success is 0.95^5, which is about 0.774. Your 95 percent component became a 77 percent feature. Push per-step reliability to 99 percent and 0.99^5 is about 0.951. That arithmetic is the argument for fewer steps, for validation at each hop, and for retries with backoff - see Retry and Backoff and Agent Architectures.
Expected value. The right way to reason about cost when you route between models. If a cheap model handles 70 percent of traffic and an expensive one handles 30 percent, and the expensive one costs ten times more per request, your expected cost per request is 0.7 x 1 + 0.3 x 10 = 3.7 cheap-units instead of 10 - a large reduction for a routing rule. The ratios here are illustrative, not measured; substitute your own. Same tool for expected token spend and expected retry cost. See Model Selection and Routing and Back of the Envelope.
6. Logarithms - Why Log-Probs and Perplexity Are Logged
A log turns multiplication into addition and turns tiny numbers into manageable ones. Both properties are the reason this field logs everything.
Underflow. The probability of a sequence is the product of its per-token probabilities. Take 500 tokens each at probability 0.1: the product is 10^-500, and float64 bottoms out around 10^-308. The number becomes exactly zero and you have lost the information. In log space it is 500 x ln(0.1), which is about -1151.3 - a perfectly ordinary float. So APIs expose log-probs, not probs.
Additivity. Because log(a x b) = log(a) + log(b), sequence scoring becomes a sum. Cheaper, numerically stable, and it composes.
Perplexity. Take the mean negative log-probability per token and exponentiate it. If a model assigns 0.5 to every token, the mean negative log-prob is -ln(0.5) which is 0.693, and e^0.693 is exactly 2. Read that as βthe model was effectively choosing between 2 equally likely options at each step.β Uniform 0.25 gives perplexity 4. It is an effective branching factor, which is why lower is better and why it is worth almost nothing for comparing models on your task - it measures next-token uncertainty, not task success. That is what evals are for, in Evals.
Engineering use. Average log-prob over a response is a weak but real confidence signal. It will not tell you an answer is wrong, but unusually low scores correlate with the model being out of its depth, which makes it a usable input for routing borderline responses to a verifier or to human review. Treat it as a triage signal, never as a correctness guarantee - see Hallucination and Grounding.
7. Percentiles
Model calls are slow and their latency distribution has a long right tail, which makes averages actively misleading. A p50 of 800 milliseconds alongside a p99 of 12 seconds is a completely normal shape for an LLM endpoint, and the mean tells you nothing useful about either.
What you need: p50 is the typical experience, p95 and p99 are the experience that generates support tickets, and percentiles do not average. You cannot take the p95 of three replicas and average them to get the system p95 - you need the merged distribution. The full treatment, including why tail latency dominates user-perceived performance, is in Performance Metrics.
Two things specific to LLM systems:
- Measure time to first token separately from total time. They have different causes and different fixes - prefill cost versus output length - and a single end-to-end number hides both. See Latency Engineering.
- Tails compound in chains. Reusing the multiplication rule from above: in a five-step agent where each step independently has a 5 percent chance of landing above its own p95, the probability that at least one step hits its tail is
1 - 0.95^5, about 23 percent. Nearly a quarter of runs feel slow even though every component is individually within budget. That is the case for timeouts, fallbacks, and Circuit Breakers.
Math Idea to Engineering Task
| Math idea | Where it shows up | What you do with it |
|---|---|---|
| Vectors and L2 norm | Embedding storage and migrations | Verify dimensions match and vectors are normalized before indexing |
| Dot product | Similarity scoring in the index | Understand why a million-candidate scan is cheap |
| Cosine similarity | Retrieval ranking | Rank candidates, debug bad results, avoid hardcoded global thresholds |
| Distance concentration | Retrieval tuning | Treat scores as relative within a result set, not absolute quality |
| Distributions and sampling | Temperature and top-p configuration | Pick near-zero for extraction, higher for generation |
| Independent events multiply | Agent and pipeline reliability | Budget per-step reliability backwards from an end-to-end target |
| Expected value | Cost forecasting and model routing | Compute expected spend per request before committing to a router |
| Logarithms | Log-probs and perplexity | Avoid underflow, read provider metrics correctly |
| Average log-prob | Confidence signals and guardrails | Route low-confidence responses to verification or review |
| Percentiles | Latency SLOs | Set p95 targets for time to first token and end to end separately |
What You Can Defer or Skip
Genuinely skippable for product-layer work. Revisit only if you move to the model layer - training, fine-tuning research, inference kernels.
| Topic | Verdict | When it would actually matter |
|---|---|---|
| Matrix calculus | Skip | Deriving a custom training objective or a new layer |
| Backpropagation derivations | Skip | Implementing an architecture from scratch. You need what training does, not the gradient chain |
| Eigendecomposition and SVD | Defer | Dimensionality reduction and embedding analysis - and even then you call a library. Knowing that PCA reduces dimensions is enough |
| Measure theory and formal probability | Skip | Theoretical ML research. Nothing operational depends on it |
| Convex optimization proofs | Skip | Optimizer design. You will pick an optimizer from a config file |
| Information theory beyond intuition | Defer | Knowing cross-entropy loss penalizes confident wrong predictions is the whole operational takeaway |
| Attention mechanics as equations | Defer | The intuition in How LLMs Actually Work covers every decision you will make |
Honest caveat: βskipβ means skip for this role now. If you later want the model layer, these become the real curriculum. Deferring is not the same as never.
Bad to Good to Great - Learning This Material
Bad: start at chapter 1 of a linear algebra textbook. Three weeks of vector spaces and determinants, no code written, no motivation for any of it, and you quit before reaching the one page that was relevant. This is the most common failure pattern and it is a scheduling mistake, not an aptitude one.
Good: learn each idea on demand. Hit cosine similarity while building retrieval, read about it for twenty minutes, move on. The motivation is already there because you have a broken result set in front of you. Much faster, and it sticks.
Great: implement each one once in a few lines, then verify against the library. Write the cosine function above, run it on [1, 2, 2] and [2, 3, 6], confirm you get 0.952, then check that your vector store returns the same score for the same pair. That single exercise converts cosine similarity from a phrase you nod at into a number you can debug. Repeat for log-probs and percentiles. Total investment is a few hours, and it permanently removes maths as a blocker.
When to Use
β Invest in the maths on this page when:
- Retrieval results look wrong and you need to tell a ranking bug from a chunking bug
- You are setting latency SLOs or an error budget for an LLM endpoint
- You are designing a multi-step agent and need an end-to-end reliability estimate
- You are forecasting cost across a routing strategy before building it
β Do not stop to study maths when:
- You have not written an eval yet - measurement beats theory, every time
- The blocker is plumbing. Most early failures are parsing, chunking, and API handling, not mathematics
- You are tempted to study before starting. The list above is short enough to learn alongside building
- The topic is on the skip table and you are at the product layer
Common Interview Questions
Q1: Why cosine similarity instead of Euclidean distance for embeddings?
Because magnitude in an embedding is mostly an artifact and direction carries the meaning. Cosine divides out both norms, so a one-sentence document and a one-page document about the same topic score similarly. Euclidean punishes long vectors: with query
[1, 2, 2], the document[4, 8, 8]is a perfect direction match with cosine 1.0 but sits 9.0 away, while an unrelated orthogonal vector sits only 3.7 away - so Euclidean ranks the unrelated one first. Worth adding the nuance: on unit-normalized vectors the two give identical rankings, since squared distance equals2 - 2 x cosine. Cosine is the safer default precisely because it does not depend on normalization having happened.
Q2: How much maths do you actually need for this role?
Six things: vectors and norms, dot product, cosine similarity, basic probability including expected value, logarithms for log-probs and perplexity, and percentiles for latency work. All of it is arithmetic-level and none requires calculus. I do not need matrix calculus, backprop derivations, or measure theory at the product layer - those belong to the model layer. The genuinely hard parts of this job are evals and reliability engineering, not mathematics, and treating maths as a prerequisite gate is what stops most capable engineers from starting.
Q3: A retrieval result has a cosine score of 0.71. Is that good?
Unanswerable in isolation, and that is the correct answer. Scores are only meaningful relative to the other candidates in the same result set and the same embedding model, because in high dimensions distances concentrate and absolute values shift with the model and the corpus. So 0.71 could be the best match available or could be noise. I would judge it by the score gap to the next candidates and by whether the retrieved chunk actually contains the answer, and I would set any threshold empirically on a labelled set rather than hardcoding a global cutoff that breaks the day the embedding model changes.
Q4: Why do LLM APIs return log-probabilities instead of probabilities?
Two reasons, both practical. Underflow: a sequence probability is the product of per-token probabilities, so 500 tokens at 0.1 each gives
10^-500, which is zero in float64 - in log space that is about-1151, which is fine. And additivity:log(a x b) = log(a) + log(b), so scoring a sequence becomes a sum, which is cheaper and numerically stable. Operationally I use average log-prob as a weak confidence signal for routing uncertain responses to verification, never as a correctness guarantee.
Q5: Every step in your agent is 95 percent reliable. What is the end-to-end number?
For five independent steps it is
0.95^5, about 77 percent - so a component that looks solid produces a feature that fails roughly one run in four. That multiplication is the core argument for keeping chains short, validating output at each hop rather than only at the end, and making steps idempotent so retries are safe. Going to 99 percent per step gets you to about 95 percent end to end. The same arithmetic applies to latency tails: if each of five steps independently has a 5 percent chance of exceeding its p95, roughly 23 percent of runs hit at least one slow step, which is why per-step timeouts and fallbacks matter more than the average looks like it should justify.
Build it in code: Agentic AI Course · Fundamentals: Core Concepts