Similarity Is a Decision
Derive cosine, dot product, and Euclidean and Manhattan distance as invariance choices, prove that L2 normalization makes three of them rank identically, then run a metric sweep on RELATE where the expected effect almost vanishes — and use that negative result to show that how much metric choice matters is itself a property of the representation.
Part I — A Vector Is Not Meaning
Two vectors, four answers
Here are three document vectors, kept to three dimensions so every number is checkable:
x = ( 2.0, 0.0, 0.0 ) a short doc, one strong topic
y = ( 6.0, 0.1, 0.0 ) a long doc, same topic, much more of it
z = ( 0.0, 2.0, 0.0 ) a short doc, a different topic
Before computing anything, look at the geometry. x and y point in almost exactly the same direction but have very different lengths (norm 2 versus about 6). x and z have the same length but point ninety degrees apart. So “is x more like y or z?” comes down to a prior question: which kind of difference should count — a difference in direction, or a difference in size?
Pick your answer, then compute all four common comparison rules:
x ↔ y x ↔ z verdict
dot product (higher = closer) 12.0 0.0 y, overwhelmingly
cosine (higher = closer) ~1.00 0.0 y, near-perfectly aligned
Euclidean L2 (lower = closer) 4.00 2.83 z is closer
Manhattan L1 (lower = closer) 4.1 4.0 z is closer, barely
Nothing about the documents changed between rows. The comparison rule changed, and with it the answer.
Each rule disagrees for a specific reason:
- Cosine divides out both vectors’ lengths, so it sees only that
xandylie on the same axis. Size is invisible to it. Verdict:y. - Dot product equals
‖x‖‖y‖cos θ— alignment times both magnitudes.yis aligned withxand large, so it scores enormously higher. Verdict:y. - Euclidean distance measures straight-line displacement, and it squares each coordinate gap. The
x-coordinate gap toyis 4, which squares to 16 and dominates; the gaps tozare 2 and 2, squaring to 4 and 4. Sozlands much closer. Verdict:z. - Manhattan distance sums the same coordinate gaps without squaring. Now the 4-unit gap to
yand the two 2-unit gaps toznearly cancel out (4.1 versus 4.0), andzwins only by a hair. Verdict:z, weakly.
“Similar” is not a property that two texts have. It is the output of one comparison rule applied to one representation. The rule encodes an assumption about which differences you are willing to ignore.
What each rule refuses to care about
Read the four rules as invariance choices — statements about which transformations leave the comparison unchanged and which differences remain visible.
Dot product. x · y = Σ xᵢyᵢ = ‖x‖‖y‖cos θ. A similarity: larger means closer. Direction and both magnitudes contribute to the scalar, so positive rescaling can change the score and the ranking. Its assumption is not that magnitude is meaningful, but that magnitude is allowed to matter. Whether that helps depends entirely on whether the representation’s norms carry signal you want or nuisance variation you do not. That is an empirical property of the representation, not a general fact about embeddings.
Cosine similarity. cos θ = (x · y) / (‖x‖ ‖y‖). Also a similarity, bounded in [−1, 1]. Its defining invariance is positive rescaling: cos(ax, by) = cos(x, y) for positive a, b. Magnitude is deliberately removed from the comparison, leaving angular alignment. Computing cosine on raw vectors and on separately L2-normalized copies therefore returns the same number up to floating-point error — always, for any nonzero vectors. Remember that; it matters when we read the experiment.
Euclidean (L2) distance. ‖x − y‖ = √Σ(xᵢ − yᵢ)². A distance: smaller means closer. It measures straight-line displacement. It is invariant to translating both vectors by the same offset, but not to independently rescaling them; large coordinate gaps contribute quadratically before the square root.
Manhattan (L1) distance. ‖x − y‖₁ = Σ|xᵢ − yᵢ|. A distance. Like L2 it is invariant to a common translation, but it aggregates coordinate gaps linearly rather than quadratically, so the same total displacement can be ranked differently.
| Rule | Formula | Kind | Key invariance | What remains rank-relevant |
|---|---|---|---|---|
| Dot product | Σ xᵢyᵢ | similarity ↑ | none to independent positive rescaling | direction and both magnitudes |
| Cosine | (x · y) / (‖x‖‖y‖) | similarity ↑ | independent positive rescaling | direction |
| Euclidean (L2) | √Σ(xᵢ − yᵢ)² | distance ↓ | common translation | straight-line displacement |
| Manhattan (L1) | Σ|xᵢ − yᵢ| | distance ↓ | common translation | coordinate-wise absolute displacement |
One algebraic fact worth memorizing
Put every vector on the unit sphere — x̂ = x / ‖x‖ — and three of these rules collapse into one. For unit vectors:
x̂ · ŷ = cos θ (dot product becomes cosine)
‖x̂ − ŷ‖² = 2 − 2 cos θ (Euclidean distance becomes a decreasing function of cosine)
So on L2-normalized vectors, cosine similarity, dot-product similarity, and Euclidean distance produce the same ranking of any candidate set (ties and numerical edge cases aside), because each is a monotonic transformation of the others. This is algebra. It holds for every dataset and every model, before you measure anything.
Manhattan distance does not join this equivalence. Two unit vectors with the same cosine to a query can sit at different L1 distances from it, depending on how their difference is spread across coordinates: a difference concentrated in one coordinate gives a smaller L1 sum than the same-size difference spread across many. Normalizing removes magnitude as a source of disagreement; it does not make every readout equivalent.
Normalization is a lossy transformation, chosen on purpose
L2 normalization sends every nonzero vector to x / ‖x‖, a point on the unit sphere. It deletes one thing — how long the vector was — and preserves another — which direction it pointed. That is not bookkeeping. It changes the hypothesis the downstream rule is even able to express.
Before normalization, two vectors can differ for two reasons: direction and length. After normalization, only direction survives into a cosine/dot/L2 comparison. If the representation encoded something useful in norm — document length, token count, a confidence-like quantity for this model — normalization throws that away. If norm was nuisance variation, normalization is exactly the cleanup you wanted. Neither outcome is automatic; it depends on the representation, and the honest move is to check rather than assume.
This connects straight back to Chapter 3’s information path. Preprocessing, representation, pooling, readout — normalization is a stage on that path, and like the others it can erase a distinction that no later stage recovers.
One more consideration, scoped carefully. Some embedding families — particularly contrastive sentence encoders — are trained with L2-normalized embeddings and a temperature-scaled cosine objective, so their geometry is shaped for angular comparison. Others train with dot-product objectives, ranking losses, classification heads, distillation, or combinations. There is no single rule. The principle is: if the model was trained under one comparison geometry and you serve queries under another, that is a mismatch to validate, not a detail to wave through. Whether it costs measurable quality is a question for an experiment.
Predict the sweep
We are about to rank the RELATE pool under all four rules, raw and normalized, on all-mpnet-base-v2. Reason it through first.
If this encoder’s outputs behave as near-unit-length direction vectors, then normalizing changes little, and on the (near-)unit sphere cosine, dot, and Euclidean must rank identically — algebra guarantees it. Manhattan could differ slightly. Net prediction: a nearly flat sweep.
If instead the outputs carry substantial rank-relevant magnitude variation, the raw dot-product and raw Euclidean conditions should pull away from cosine, and normalizing should visibly change their rankings.
The sweep is a test of which regime this representation is in.
Demonstration: the metric sweep on RELATE
MEASURED on RELATE v0.1, Wave 1 row 1.3 — artifact
experiments/embeddings-from-first-principles/wave1/artifacts/metric-sweep.json. Modelall-mpnet-base-v2; embeddings and query vectors requested un-normalized, with L2-normalized copies formed for the normalized conditions; scored by nDCG@10 over the full pool and over the hard-negative subset.
Spread across all eight conditions: 0.0005 nDCG@10 overall, 0.0013 on hard negatives.
The “metric choice moves retrieval by several points” story did not reproduce. Read what the table does and does not say.
What is established. The eight conditions deliver the same measured retrieval quality at four-decimal resolution, except Manhattan, which is higher by 0.0005 overall and 0.0013 on hard negatives.
What is not established. Equal nDCG@10 is not equal score arrays, not equal rankings, and not equal raw vectors. Aggregate nDCG@10 can hide score differences (two rules can produce very different numbers and the same top-10 order) and can hide small ranking differences (a reshuffle below rank 10, or a swap that does not change the discounted gain at this precision). The artifact measures the aggregate, so that is what we can claim.
Which equalities are algebra and which are measurement.
cosine raw = cosine normalizedis guaranteed. Cosine divides by norms; normalizing first cannot change it. This tells us nothing about whether the vectors were already unit length.cosine = dot = Euclideanwithin the normalized conditions is guaranteed by the unit-sphere identities above.dot raw = dot normalizedandEuclidean raw = Euclidean normalizedare not guaranteed. They are the measured result. The most economical reading: for this encoder and this task, the outputs carry little magnitude variation that matters to top-10 ranking, so dividing it out barely moves nDCG.Manhattansitting 0.0005–0.0013 above the rest is also measurement — real and reproducible, but within the resolution and scope of one benchmark. It is not a “winner.” L1’s linear aggregation happened to order a few hard-negative pairs slightly better.
MEASURED (scope:
all-mpnet-base-v2, RELATE v0.1, nDCG@10): within the tested conditions, changing the comparison rule barely changed retrieval quality. At this evaluation resolution, magnitude contributed little that changed the top-10 task metric. That is weaker — and safer — than claiming the raw vectors were already unit length: the artifact measures retrieval quality, not the norm distribution.
Three kinds of equality
Keep these separate whenever you compare two comparison rules:
- Score equality — the numbers match.
- Ranking equality — the ordered candidate lists match.
- Task-metric equality — the aggregate score (here nDCG@10) matches.
Score equality implies ranking equality implies task-metric equality; none of the reverse implications hold. The algebra gives us score-level and ranking-level equivalence for cosine/dot/L2 on the unit sphere. The artifact gives us task-metric equality across the eight conditions. Using one as evidence for another is the easiest mistake to make with a table like this.
Why the choice barely mattered here
Three levels, three different verdicts:
- Toy level. The four rules disagreed dramatically — cosine and dot said
y, the distances saidz. Metric choice looked decisive. - Algebraic level. L2 normalization collapses cosine, dot, and Euclidean into one ranking. Three of the four “choices” stop being choices the moment you normalize.
- Benchmark level. On this representation and this task, even the surviving freedom — Manhattan versus the rest, raw versus normalized — moved nDCG@10 by less than 0.002.
So the useful conclusion is not “metric choice does not matter.” It is: how much metric choice matters is a property of the representation, the preprocessing, and the task — not a constant. If norms carry rank-relevant information, raw dot product and cosine are genuinely different readouts. If normalization or the learned geometry makes norm variation irrelevant to the task, those choices can collapse toward the same ranking. The metric is therefore downstream of a more important question: what distinctions did the representation make available for the readout to use?
There is a subtler point in the hard-negative column, which barely moved. It is tempting to conclude that if cosine cannot separate a distinction, the representation did not capture it. That is too strong. Distinguish two cases:
- Absent or collapsed. The distinction is not in the vectors — nothing in training rewarded encoding it, and no readout can recover what is not there.
- Present but inaccessible to this readout. The information is in the vectors, but not along the directions a generic cosine compares. A learned, relation-specific projection can sometimes expose it (Chapter 15).
A flat metric sweep is consistent with both. It weakens the hypothesis that swapping among these generic metrics is the missing fix; it does not show that the distinction is absent from the representation. The later RELATE experiments make this separation operational by evaluating generic geometry and relation-specific readouts through the same preference direction: the question becomes not merely “which distance?”, but which readout exposes the relation the task actually cares about?
What this chapter establishes and what it does not
Establishes: the four rules and their invariance assumptions; the algebraic fact that L2 normalization makes cosine, dot, and Euclidean rank-equivalent while Manhattan stays separate; that normalization is a deliberate, lossy information-path decision; and that on all-mpnet-base-v2 / RELATE v0.1, the metric sweep moved nDCG@10 by ≤0.0013 across eight conditions (Wave 1 row 1.3).
Does not establish: a single best metric (it depends on how the model was trained and whether its norms carry signal); that metric choice is generally inert (it is not, for magnitude-carrying representations); that identical nDCG means identical scores, rankings, or vectors; or that a flat sweep proves a distinction is absent rather than merely invisible to cosine.
Lab 4: metric sweep, and what the aggregate hides
MEASURED — artifact
experiments/embeddings-from-first-principles/wave1/artifacts/metric-sweep.json(all-mpnet-base-v2). REPRODUCIBLE —run_wave1.py 1.3. (Distances are negated so every condition ranks “higher is closer”.)
Question. How much of your retrieval score is the comparison rule, and how much is the representation underneath it?
What we ran. One set of embeddings, ranked under 8 conditions — cosine / dot / Euclidean / Manhattan, each raw and L2-normalized — scored by nDCG@10 over the full pool and the hard-negative subset:
| Condition | nDCG@10 (all) | nDCG@10 (hard-neg) |
|---|---|---|
| cosine (raw or normalized) | 0.9518 | 0.9577 |
| dot (raw or normalized) | 0.9518 | 0.9577 |
| Euclidean (raw or normalized) | 0.9518 | 0.9577 |
| Manhattan (raw or normalized) | 0.9523 | 0.9590 |
Spread across all eight: 0.0005 nDCG@10 overall, 0.0013 on hard negatives.
Interpretation. Establishes: for this encoder, “raw” and “normalized” produce the same aggregate retrieval quality, and on the unit sphere cosine, dot, and Euclidean must rank identically. The “several points” effect did not appear and the claim is rescoped to magnitude-carrying representations. Does not establish: that metric choice never matters, or that the eight conditions produced identical scores or rankings — only identical nDCG@10 at four decimals.
Try it yourself
These are reader experiments; no artifact backs them yet.
- Measure the norms. Histogram
‖v‖over the RELATE embeddings. If they cluster tightly near one value, you have found a plausible mechanism for why raw dot product and cosine behave similarly. Then verify it at the ranking level: tight norms alone do not prove identical top-10 lists or identical nDCG.- Look at rankings, not just nDCG. For each query, compare the top-10 lists under raw dot and under cosine. Count how many queries have any difference. Aggregate nDCG can be identical while a handful of lists are reshuffled.
- Hunt for disagreement. Find the queries, if any, where raw dot and cosine put a different item at rank 1. Read those pairs — what do they share?
- Change the representation. Repeat the whole sweep on TF-IDF vectors or a raw-output encoder — something with real norm variation. Watch the raw dot-product condition separate from cosine.
- Separate the three effects. For any two conditions, report: do the scores differ, do the rankings differ, does the task metric differ? They are three questions, and only the third is in the artifact.
Companion component: the similarity spec
The Observatory never stores a similarity number without the rule that produced it.
A stored score without its comparison rule loses its meaning. 0.83 is not a fact about two texts; it is a fact about two vectors under one readout with one normalization policy.
And the spec is necessary provenance, not sufficient for cross-space comparison. Two cosine scores from two models are still not directly comparable: different encoders, normalization regimes, query/document encoder pairs, corpora, and score distributions all shift what a given number means. Even within one space, a threshold belongs to the distribution that produced it — task, domain, and negative set included. Calibrate within each scoped setting (Chapter 14); the spec tells you what quantity is being calibrated.
Failure modes
Each is a mistake, why it is tempting, and the check.
- Train/serve geometry mismatch. Tempting because every rule returns one scalar, so they feel interchangeable. Check: find the comparison geometry the model was trained under; if you serve a different one, run a controlled ranking evaluation before trusting the scores or any threshold built on them.
- Comparing raw scores across spaces. Tempting because
0.83looks like the same quantity everywhere. Check: calibrate within each space; never assume a cosine of 0.83 from model A means what a cosine of 0.83 from model B means. - Assuming cosine removes every nuisance. Tempting because it visibly fixes the length-bias toy. Check: cosine removes magnitude from the comparison and nothing else — anisotropy, hubness, frequency and lexical bias, and corpus structure all survive it (Chapters 6–8).
- Metric-shopping a representation failure. Tempting because trying four formulas is cheap and one might be higher. Check: a flat generic-metric sweep does not tell you whether the distinction is absent from the representation or merely inaccessible to those readouts. Test that question separately. If the signal is absent, no readout can recover it; if it is present but off-axis for the generic metric, a learned relation-specific readout may expose it (Chapter 15).
What this chapter established
- Dot, cosine, L2, and L1 are invariance choices: each names a difference it agrees to ignore. On L2-normalized vectors the first three rank identically — algebra, not measurement — and Manhattan does not join them.
- Normalization is a deliberate lossy step on the information path: it deletes magnitude, keeps direction, and changes what the readout is able to express.
- Identical aggregate nDCG is not identical scores, identical rankings, or identical vectors. Those are three separate claims, and the sweep established only the third.
- How much metric choice matters is a property of the representation, not a constant — which is why the similarity spec never stores a score without its metric, normalization, and training-geometry context.
Next
Part I has treated the vector as a whole and asked how to compare two of them. Part II goes inside a single vector. The next chapter picks one coordinate — dimension 173 — and asks what it means on its own. The answer forces a question the rest of Part II works on: are individual coordinates stable, nameable features, or does the meaning live in directions and combinations of coordinates that no single axis captures?