← Embeddings From First Principles

Similarity Is a Decision

Derive cosine, dot product, and Euclidean and Manhattan distance as invariance choices, prove that L2 normalization makes three of them rank identically, then run a metric sweep on RELATE where the expected effect almost vanishes — and use that negative result to show that how much metric choice matters is itself a property of the representation.

Part I — A Vector Is Not Meaning

Two vectors, four answers

Here are three document vectors, kept to three dimensions so every number is checkable:

x = ( 2.0, 0.0, 0.0 )      a short doc, one strong topic
y = ( 6.0, 0.1, 0.0 )      a long doc, same topic, much more of it
z = ( 0.0, 2.0, 0.0 )      a short doc, a different topic

Before computing anything, look at the geometry. x and y point in almost exactly the same direction but have very different lengths (norm 2 versus about 6). x and z have the same length but point ninety degrees apart. So “is x more like y or z?” comes down to a prior question: which kind of difference should count — a difference in direction, or a difference in size?

Pick your answer, then compute all four common comparison rules:

                         x ↔ y        x ↔ z      verdict
dot product  (higher = closer)   12.0          0.0       y, overwhelmingly
cosine       (higher = closer)   ~1.00         0.0       y, near-perfectly aligned
Euclidean L2 (lower  = closer)    4.00         2.83      z is closer
Manhattan L1 (lower  = closer)    4.1          4.0       z is closer, barely

Nothing about the documents changed between rows. The comparison rule changed, and with it the answer.

Each rule disagrees for a specific reason:

  • Cosine divides out both vectors’ lengths, so it sees only that x and y lie on the same axis. Size is invisible to it. Verdict: y.
  • Dot product equals ‖x‖‖y‖cos θ — alignment times both magnitudes. y is aligned with x and large, so it scores enormously higher. Verdict: y.
  • Euclidean distance measures straight-line displacement, and it squares each coordinate gap. The x-coordinate gap to y is 4, which squares to 16 and dominates; the gaps to z are 2 and 2, squaring to 4 and 4. So z lands much closer. Verdict: z.
  • Manhattan distance sums the same coordinate gaps without squaring. Now the 4-unit gap to y and the two 2-unit gaps to z nearly cancel out (4.1 versus 4.0), and z wins only by a hair. Verdict: z, weakly.

“Similar” is not a property that two texts have. It is the output of one comparison rule applied to one representation. The rule encodes an assumption about which differences you are willing to ignore.

What each rule refuses to care about

Read the four rules as invariance choices — statements about which transformations leave the comparison unchanged and which differences remain visible.

Dot product. x · y = Σ xᵢyᵢ = ‖x‖‖y‖cos θ. A similarity: larger means closer. Direction and both magnitudes contribute to the scalar, so positive rescaling can change the score and the ranking. Its assumption is not that magnitude is meaningful, but that magnitude is allowed to matter. Whether that helps depends entirely on whether the representation’s norms carry signal you want or nuisance variation you do not. That is an empirical property of the representation, not a general fact about embeddings.

Cosine similarity. cos θ = (x · y) / (‖x‖ ‖y‖). Also a similarity, bounded in [−1, 1]. Its defining invariance is positive rescaling: cos(ax, by) = cos(x, y) for positive a, b. Magnitude is deliberately removed from the comparison, leaving angular alignment. Computing cosine on raw vectors and on separately L2-normalized copies therefore returns the same number up to floating-point error — always, for any nonzero vectors. Remember that; it matters when we read the experiment.

Euclidean (L2) distance. ‖x − y‖ = √Σ(xᵢ − yᵢ)². A distance: smaller means closer. It measures straight-line displacement. It is invariant to translating both vectors by the same offset, but not to independently rescaling them; large coordinate gaps contribute quadratically before the square root.

Manhattan (L1) distance. ‖x − y‖₁ = Σ|xᵢ − yᵢ|. A distance. Like L2 it is invariant to a common translation, but it aggregates coordinate gaps linearly rather than quadratically, so the same total displacement can be ranked differently.

RuleFormulaKindKey invarianceWhat remains rank-relevant
Dot productΣ xᵢyᵢsimilarity ↑none to independent positive rescalingdirection and both magnitudes
Cosine(x · y) / (‖x‖‖y‖)similarity ↑independent positive rescalingdirection
Euclidean (L2)√Σ(xᵢ − yᵢ)²distance ↓common translationstraight-line displacement
Manhattan (L1)Σ|xᵢ − yᵢ|distance ↓common translationcoordinate-wise absolute displacement

One algebraic fact worth memorizing

Put every vector on the unit sphere — x̂ = x / ‖x‖ — and three of these rules collapse into one. For unit vectors:

x̂ · ŷ = cos θ                    (dot product becomes cosine)
‖x̂ − ŷ‖² = 2 − 2 cos θ           (Euclidean distance becomes a decreasing function of cosine)

So on L2-normalized vectors, cosine similarity, dot-product similarity, and Euclidean distance produce the same ranking of any candidate set (ties and numerical edge cases aside), because each is a monotonic transformation of the others. This is algebra. It holds for every dataset and every model, before you measure anything.

Manhattan distance does not join this equivalence. Two unit vectors with the same cosine to a query can sit at different L1 distances from it, depending on how their difference is spread across coordinates: a difference concentrated in one coordinate gives a smaller L1 sum than the same-size difference spread across many. Normalizing removes magnitude as a source of disagreement; it does not make every readout equivalent.

Normalization is a lossy transformation, chosen on purpose

L2 normalization sends every nonzero vector to x / ‖x‖, a point on the unit sphere. It deletes one thing — how long the vector was — and preserves another — which direction it pointed. That is not bookkeeping. It changes the hypothesis the downstream rule is even able to express.

Before normalization, two vectors can differ for two reasons: direction and length. After normalization, only direction survives into a cosine/dot/L2 comparison. If the representation encoded something useful in norm — document length, token count, a confidence-like quantity for this model — normalization throws that away. If norm was nuisance variation, normalization is exactly the cleanup you wanted. Neither outcome is automatic; it depends on the representation, and the honest move is to check rather than assume.

This connects straight back to Chapter 3’s information path. Preprocessing, representation, pooling, readout — normalization is a stage on that path, and like the others it can erase a distinction that no later stage recovers.

One more consideration, scoped carefully. Some embedding families — particularly contrastive sentence encoders — are trained with L2-normalized embeddings and a temperature-scaled cosine objective, so their geometry is shaped for angular comparison. Others train with dot-product objectives, ranking losses, classification heads, distillation, or combinations. There is no single rule. The principle is: if the model was trained under one comparison geometry and you serve queries under another, that is a mismatch to validate, not a detail to wave through. Whether it costs measurable quality is a question for an experiment.

Predict the sweep

We are about to rank the RELATE pool under all four rules, raw and normalized, on all-mpnet-base-v2. Reason it through first.

If this encoder’s outputs behave as near-unit-length direction vectors, then normalizing changes little, and on the (near-)unit sphere cosine, dot, and Euclidean must rank identically — algebra guarantees it. Manhattan could differ slightly. Net prediction: a nearly flat sweep.

If instead the outputs carry substantial rank-relevant magnitude variation, the raw dot-product and raw Euclidean conditions should pull away from cosine, and normalizing should visibly change their rankings.

The sweep is a test of which regime this representation is in.

Demonstration: the metric sweep on RELATE

MEASURED on RELATE v0.1, Wave 1 row 1.3 — artifact experiments/embeddings-from-first-principles/wave1/artifacts/metric-sweep.json. Model all-mpnet-base-v2; embeddings and query vectors requested un-normalized, with L2-normalized copies formed for the normalized conditions; scored by nDCG@10 over the full pool and over the hard-negative subset.

Spread across all eight conditions: 0.0005 nDCG@10 overall, 0.0013 on hard negatives.

The “metric choice moves retrieval by several points” story did not reproduce. Read what the table does and does not say.

What is established. The eight conditions deliver the same measured retrieval quality at four-decimal resolution, except Manhattan, which is higher by 0.0005 overall and 0.0013 on hard negatives.

What is not established. Equal nDCG@10 is not equal score arrays, not equal rankings, and not equal raw vectors. Aggregate nDCG@10 can hide score differences (two rules can produce very different numbers and the same top-10 order) and can hide small ranking differences (a reshuffle below rank 10, or a swap that does not change the discounted gain at this precision). The artifact measures the aggregate, so that is what we can claim.

Which equalities are algebra and which are measurement.

  • cosine raw = cosine normalized is guaranteed. Cosine divides by norms; normalizing first cannot change it. This tells us nothing about whether the vectors were already unit length.
  • cosine = dot = Euclidean within the normalized conditions is guaranteed by the unit-sphere identities above.
  • dot raw = dot normalized and Euclidean raw = Euclidean normalized are not guaranteed. They are the measured result. The most economical reading: for this encoder and this task, the outputs carry little magnitude variation that matters to top-10 ranking, so dividing it out barely moves nDCG.
  • Manhattan sitting 0.0005–0.0013 above the rest is also measurement — real and reproducible, but within the resolution and scope of one benchmark. It is not a “winner.” L1’s linear aggregation happened to order a few hard-negative pairs slightly better.

MEASURED (scope: all-mpnet-base-v2, RELATE v0.1, nDCG@10): within the tested conditions, changing the comparison rule barely changed retrieval quality. At this evaluation resolution, magnitude contributed little that changed the top-10 task metric. That is weaker — and safer — than claiming the raw vectors were already unit length: the artifact measures retrieval quality, not the norm distribution.

Three kinds of equality

Keep these separate whenever you compare two comparison rules:

  1. Score equality — the numbers match.
  2. Ranking equality — the ordered candidate lists match.
  3. Task-metric equality — the aggregate score (here nDCG@10) matches.

Score equality implies ranking equality implies task-metric equality; none of the reverse implications hold. The algebra gives us score-level and ranking-level equivalence for cosine/dot/L2 on the unit sphere. The artifact gives us task-metric equality across the eight conditions. Using one as evidence for another is the easiest mistake to make with a table like this.

Why the choice barely mattered here

Three levels, three different verdicts:

  • Toy level. The four rules disagreed dramatically — cosine and dot said y, the distances said z. Metric choice looked decisive.
  • Algebraic level. L2 normalization collapses cosine, dot, and Euclidean into one ranking. Three of the four “choices” stop being choices the moment you normalize.
  • Benchmark level. On this representation and this task, even the surviving freedom — Manhattan versus the rest, raw versus normalized — moved nDCG@10 by less than 0.002.

So the useful conclusion is not “metric choice does not matter.” It is: how much metric choice matters is a property of the representation, the preprocessing, and the task — not a constant. If norms carry rank-relevant information, raw dot product and cosine are genuinely different readouts. If normalization or the learned geometry makes norm variation irrelevant to the task, those choices can collapse toward the same ranking. The metric is therefore downstream of a more important question: what distinctions did the representation make available for the readout to use?

There is a subtler point in the hard-negative column, which barely moved. It is tempting to conclude that if cosine cannot separate a distinction, the representation did not capture it. That is too strong. Distinguish two cases:

  • Absent or collapsed. The distinction is not in the vectors — nothing in training rewarded encoding it, and no readout can recover what is not there.
  • Present but inaccessible to this readout. The information is in the vectors, but not along the directions a generic cosine compares. A learned, relation-specific projection can sometimes expose it (Chapter 15).

A flat metric sweep is consistent with both. It weakens the hypothesis that swapping among these generic metrics is the missing fix; it does not show that the distinction is absent from the representation. The later RELATE experiments make this separation operational by evaluating generic geometry and relation-specific readouts through the same preference direction: the question becomes not merely “which distance?”, but which readout exposes the relation the task actually cares about?

What this chapter establishes and what it does not

Establishes: the four rules and their invariance assumptions; the algebraic fact that L2 normalization makes cosine, dot, and Euclidean rank-equivalent while Manhattan stays separate; that normalization is a deliberate, lossy information-path decision; and that on all-mpnet-base-v2 / RELATE v0.1, the metric sweep moved nDCG@10 by ≤0.0013 across eight conditions (Wave 1 row 1.3).

Does not establish: a single best metric (it depends on how the model was trained and whether its norms carry signal); that metric choice is generally inert (it is not, for magnitude-carrying representations); that identical nDCG means identical scores, rankings, or vectors; or that a flat sweep proves a distinction is absent rather than merely invisible to cosine.

Lab 4: metric sweep, and what the aggregate hides

MEASURED — artifact experiments/embeddings-from-first-principles/wave1/artifacts/metric-sweep.json (all-mpnet-base-v2). REPRODUCIBLE — run_wave1.py 1.3. (Distances are negated so every condition ranks “higher is closer”.)

Question. How much of your retrieval score is the comparison rule, and how much is the representation underneath it?

What we ran. One set of embeddings, ranked under 8 conditions — cosine / dot / Euclidean / Manhattan, each raw and L2-normalized — scored by nDCG@10 over the full pool and the hard-negative subset:

ConditionnDCG@10 (all)nDCG@10 (hard-neg)
cosine (raw or normalized)0.95180.9577
dot (raw or normalized)0.95180.9577
Euclidean (raw or normalized)0.95180.9577
Manhattan (raw or normalized)0.95230.9590

Spread across all eight: 0.0005 nDCG@10 overall, 0.0013 on hard negatives.

Interpretation. Establishes: for this encoder, “raw” and “normalized” produce the same aggregate retrieval quality, and on the unit sphere cosine, dot, and Euclidean must rank identically. The “several points” effect did not appear and the claim is rescoped to magnitude-carrying representations. Does not establish: that metric choice never matters, or that the eight conditions produced identical scores or rankings — only identical nDCG@10 at four decimals.

Try it yourself

These are reader experiments; no artifact backs them yet.

  1. Measure the norms. Histogram ‖v‖ over the RELATE embeddings. If they cluster tightly near one value, you have found a plausible mechanism for why raw dot product and cosine behave similarly. Then verify it at the ranking level: tight norms alone do not prove identical top-10 lists or identical nDCG.
  2. Look at rankings, not just nDCG. For each query, compare the top-10 lists under raw dot and under cosine. Count how many queries have any difference. Aggregate nDCG can be identical while a handful of lists are reshuffled.
  3. Hunt for disagreement. Find the queries, if any, where raw dot and cosine put a different item at rank 1. Read those pairs — what do they share?
  4. Change the representation. Repeat the whole sweep on TF-IDF vectors or a raw-output encoder — something with real norm variation. Watch the raw dot-product condition separate from cosine.
  5. Separate the three effects. For any two conditions, report: do the scores differ, do the rankings differ, does the task metric differ? They are three questions, and only the third is in the artifact.

Companion component: the similarity spec

The Observatory never stores a similarity number without the rule that produced it.

A stored score without its comparison rule loses its meaning. 0.83 is not a fact about two texts; it is a fact about two vectors under one readout with one normalization policy.

And the spec is necessary provenance, not sufficient for cross-space comparison. Two cosine scores from two models are still not directly comparable: different encoders, normalization regimes, query/document encoder pairs, corpora, and score distributions all shift what a given number means. Even within one space, a threshold belongs to the distribution that produced it — task, domain, and negative set included. Calibrate within each scoped setting (Chapter 14); the spec tells you what quantity is being calibrated.

Failure modes

Each is a mistake, why it is tempting, and the check.

  • Train/serve geometry mismatch. Tempting because every rule returns one scalar, so they feel interchangeable. Check: find the comparison geometry the model was trained under; if you serve a different one, run a controlled ranking evaluation before trusting the scores or any threshold built on them.
  • Comparing raw scores across spaces. Tempting because 0.83 looks like the same quantity everywhere. Check: calibrate within each space; never assume a cosine of 0.83 from model A means what a cosine of 0.83 from model B means.
  • Assuming cosine removes every nuisance. Tempting because it visibly fixes the length-bias toy. Check: cosine removes magnitude from the comparison and nothing else — anisotropy, hubness, frequency and lexical bias, and corpus structure all survive it (Chapters 6–8).
  • Metric-shopping a representation failure. Tempting because trying four formulas is cheap and one might be higher. Check: a flat generic-metric sweep does not tell you whether the distinction is absent from the representation or merely inaccessible to those readouts. Test that question separately. If the signal is absent, no readout can recover it; if it is present but off-axis for the generic metric, a learned relation-specific readout may expose it (Chapter 15).

What this chapter established

  • Dot, cosine, L2, and L1 are invariance choices: each names a difference it agrees to ignore. On L2-normalized vectors the first three rank identically — algebra, not measurement — and Manhattan does not join them.
  • Normalization is a deliberate lossy step on the information path: it deletes magnitude, keeps direction, and changes what the readout is able to express.
  • Identical aggregate nDCG is not identical scores, identical rankings, or identical vectors. Those are three separate claims, and the sweep established only the third.
  • How much metric choice matters is a property of the representation, not a constant — which is why the similarity spec never stores a score without its metric, normalization, and training-geometry context.

Next

Part I has treated the vector as a whole and asked how to compare two of them. Part II goes inside a single vector. The next chapter picks one coordinate — dimension 173 — and asks what it means on its own. The answer forces a question the rest of Part II works on: are individual coordinates stable, nameable features, or does the meaning live in directions and combinations of coordinates that no single axis captures?