← Embeddings From First Principles

Dimensions Do Not Mean What You Think

Ask what dimension 173 means and find that the question is not basis-invariant. Rotate a whole embedding space by an orthogonal matrix: every coordinate changes, and cosine, distance, and retrieval do not. Separate basis dependence from superposition, and anisotropy from both.

Part II — Inside the Space

What does dimension 173 mean?

Pull the 173rd coordinate of every vector in your corpus and sort. You get a list of documents ordered by something. Once in a while a coordinate is weakly readable — “this one runs high for questions” — but usually the sorted list has no nameable theme. Dimension 173 is not “formality,” not “sentiment,” not “is about sports.”

And yet the vectors work. Similarity search returns sensible results, clusters fall out, nearest neighbors are reasonable. The usable structure is there. It is just not sitting on the individual axes.

If no single coordinate carries a nameable meaning, where is the structure, and what does that tell us about how to treat the numbers?

A coordinate is a description relative to a basis

Start in two dimensions, where everything is visible.

x = (1.0, 0.0)
y = (0.8, 0.6)

Rotate both vectors by 45 degrees. Their coordinates become:

x' = (0.71, 0.71)
y' = (0.14, 0.99)

Every number changed. Now check the relationships:

                  before      after
‖x‖, ‖y‖          1.0, 1.0    1.0, 1.0        unchanged
x · y             0.80        0.80            unchanged
cos(x, y)         0.80        0.80            unchanged
angle             36.9°       36.9°           unchanged
‖x − y‖  (L2)     0.63        0.63            unchanged
‖x − y‖₁ (L1)     0.80        0.85            changed

A basis is the set of reference directions you measure coordinates against. A coordinate is the coefficient of a vector along one of those directions. You can read the transformation above in two equivalent ways: actively rotate every vector while keeping the coordinate axes fixed, or passively re-express the same geometry in a rotated basis. For the quantities we care about here, the distinction does not matter. Inner products, norms, cosine, angles, and Euclidean distance are unchanged under a joint orthogonal transformation. A quantity that reads coordinates one at a time, such as L1 distance, is not generally invariant.

The lesson that survives to high dimensions is narrower and stronger: useful information in an embedding need not live on any individual raw coordinate. It can live in directions that combine many coordinates. A retrieval system based only on cosine, dot product, or L2 geometry cannot tell which orthogonal basis it has been given.

The rotation experiment

Take the embedding matrix X (n items by d dimensions). Draw a random orthogonal matrix R (d by d) — one with R Rᵀ = Rᵀ R = I. “Orthogonal” covers rotations and also reflections and axis permutations. Form X' = X R, so each row transforms as x' = x R.

The algebra is one line. For any two rows,

⟨x', y'⟩ = (x R)(y R)ᵀ = x R Rᵀ yᵀ = x yᵀ = ⟨x, y⟩.

Inner products are unchanged. Everything built from them follows: ‖x'‖ = ‖x‖, cos(x', y') = cos(x, y), the angle between any two vectors, and ‖x' − y'‖ = ‖x − y‖. So any procedure that reads only those quantities — cosine retrieval, dot-product retrieval, L2 retrieval, k-nearest-neighbors under one of those, PCA’s variance spectrum — receives equivalent information after the same R is applied to every relevant vector, documents and queries alike.

This is a proof, not a prediction. The rotation is guaranteed to preserve the geometry. What an experiment adds is a check that the actual retrieval implementation — with its indexing, its query encoder, its top-k truncation — respects the guarantee rather than breaking it through some coordinate-sensitive shortcut.

Intrinsic to the geometry, or an artifact of the basis

Preserved under a joint orthogonal rotationChanges under a joint orthogonal rotation
dot product, cosine, vector norm, anglesany individual coordinate’s value
Euclidean (L2) distanceper-coordinate variance and its ranking
k-NN lists computed from cosine / dot / L2axis-aligned thresholds and coordinate pruning
retrieval metrics (Recall, MRR, nDCG) computed from those rankingsL1 / Manhattan distance and any k-NN built on it
singular values / PCA variance spectrum“dimension 173 means X”
whether the vector cloud is directionally lopsided (anisotropy)which axis the lopsidedness happens to align with

The precise claim is not “the basis does not matter.” Some real operations are coordinate-sensitive — L1 distance, per-coordinate quantization, axis-aligned pruning, coordinate-level interventions. The claim is narrower and stronger: the basis is not identifiable from rotation-invariant behavior. If all you use is cosine, dot, and L2 geometry, an orthogonally rotated basis is observationally equivalent for those operations, and the model’s native output basis is just one choice among infinitely many that describe the same geometry.

What rotation does and does not say about “dimension 173 means sentiment”

That sentence can mean three different things.

  1. An intrinsic geometric claim: the embedding geometry contains an axis for sentiment, and it is coordinate 173. A generic orthogonal change of basis is a decisive test of the coordinate-specific part of that claim: the semantic direction, if it exists, rotates with the space, while coordinate 173 refers to a different basis direction. Under a dense random rotation it becomes a mixture of many old coordinates. The geometry may contain a sentiment-related direction; the rotation-invariant geometry does not privilege the label “coordinate 173.”
  2. A native-basis descriptive claim: in this model’s emitted basis, coordinate 173 correlates with sentiment. This can be empirically true. It is a fact about one particular basis, and it does not transfer to a rotated copy of the same space — but as a statement about the model’s actual output, it is checkable and can hold.
  3. A mechanistic claim: the model implements sentiment by writing it to coordinate 173. This needs causal evidence — interventions, ablations — well beyond a correlation.

Rotation does not prove that no native coordinate can ever correlate with something interpretable. It proves that such a correlation is a property of the chosen basis, not of the rotation-invariant geometry. Keep the three claims apart and most confusion about “interpretable dimensions” dissolves.

Demonstration: rotate RELATE, and watch a coordinate correlation move

MEASURED on RELATE v0.1, Wave 2 row 2.1 — artifact experiments/embeddings-from-first-principles/wave2/artifacts/rotation-invariance.json. Model all-mpnet-base-v2, L2-normalized embeddings. Document and query vectors are rotated by the same random orthogonal matrix (Q from the QR factorization of a Gaussian matrix); retrieval is scored under three independent such rotations.

nDCG@10, no rotation                 0.951773
nDCG@10, rotation 1 / 2 / 3          0.951773 / 0.951773 / 0.951773
maximum absolute change              4.4 × 10⁻⁷

Retrieval is rotation-invariant to numerical precision, exactly as the algebra requires. Every coordinate of every vector was replaced, and the ranking the retrieval system produced did not move.

Now the coordinate-level view. Before rotation, take the highest-variance coordinate of the native space and correlate it with sentence character length:

corr(highest-variance native coordinate, text length)      before:  +0.293
corr(highest-variance coordinate after one rotation, ...)   after:   −0.38

Read this carefully, because the obvious summary is wrong. The rotation did not destroy any information that was present in the embedding matrix. An orthogonal transformation is invertible: given the rotated vectors and R, you can reconstruct the originals exactly. So whatever information about text length was accessible from the full original representation remains accessible from the full rotated representation. What changed is its alignment with the raw axes: a different coordinate is now the highest-variance one, and its correlation with length has a different magnitude (0.38 versus 0.29) and even a different sign. The single-coordinate readout changed; the underlying representation did not lose information.

MEASURED: on this model and corpus, retrieval quality is invariant under a joint orthogonal rotation (Δ ≤ 4.4 × 10⁻⁷ nDCG@10), while the correlation between the top coordinate and a surface property is not — it changed from +0.29 to −0.38 under a single rotation. A coordinate-level interpretation is a statement about the basis, and the basis is not something the retrieval geometry pins down.

Basis dependence is enough; superposition is a separate claim

It is tempting to explain messy axes by reaching for superposition — the idea, from mechanistic interpretability research, that a model under capacity pressure represents more features than it has dimensions by placing them along non-orthogonal directions and tolerating interference. Resist the reach, at least here.

Everything above requires no superposition. Basis dependence is pure linear algebra. Even a representation with exactly d perfectly independent, orthogonal feature directions can be rotated so that not one of them lands on a coordinate axis; the coordinates become mixtures, and no single one is nameable. That is the entire content of the rotation result, and it holds for the cleanest possible representation.

Superposition is a stronger and different phenomenon. In the usual picture, a representation carries more features than can all occupy mutually orthogonal directions, so some features share representational dimensions. Sparsity can make that packing workable because many features are not active at the same time; when overlapping features are jointly active, interference can appear. A learned direction or unit may therefore respond to multiple otherwise unrelated features. None of that follows merely from basis dependence.

Superposition, if present, is one more reason not to expect a tidy one-axis-per-concept dictionary. But this chapter’s experiments do not measure it. The RELATE results show basis dependence and rotation invariance; they say nothing about how many features a sentence encoder packs into its space or whether its feature directions interfere. Sparse autoencoders and dictionary-learning methods are attempts to find a different, often overcomplete feature dictionary in which activations become sparser and sometimes more readable — not a procedure that “unrotates” a space into its one true semantic basis, and not a source of ground-truth concept labels.

Anisotropy: a property of the cloud, not the axes

Rotation invariance is about how the same geometry looks under different rulers. Anisotropy is a different question entirely: within that geometry, is the cloud of vectors spread evenly over directions, or does it lean?

MEASURED on RELATE v0.1, Wave 1 row 1.4 — artifact experiments/embeddings-from-first-principles/wave1/artifacts/anisotropy.json. Mean cosine over 5,000 random item pairs, per model, on L2-normalized embeddings.

model                  mean random    std    pairs with     effective rank   top singular
                       cosine                 cosine > 0     / nominal dim    value share
all-MiniLM-L6-v2        0.058          0.119    65.6%          259 / 384        0.070
all-mpnet-base-v2       0.075          0.124    72.6%          387 / 768        0.076
mxbai-embed-large-v1    0.337          0.099   100.0%          425 / 1024       0.087
bge-large-en-v1.5       0.403          0.107   100.0%          434 / 1024       0.090
bge-small-en-v1.5       0.451          0.094   100.0%          271 / 384        0.099

For independently sampled unit vectors, the expected dot product is ‖E[x]‖²; because cosine equals dot product on unit vectors, the random-pair mean is therefore related to the squared length of the population mean direction. In a finite corpus, with self-pairs excluded and only 5,000 sampled pairs, the measured mean is an empirical estimate rather than an exact identity. A large positive value is evidence of a shared directional component in the distribution, but it does not mean every vector carries the same component to the same degree.

For bge and mxbai on this corpus, the measured random-pair mean is 0.34 to 0.45 and every sampled pair has positive cosine — strong evidence of substantial common directional structure here. For MiniLM and mpnet, the baseline is much closer to zero (0.06 to 0.08). That does not make them isotropic: isotropy is a property of the full directional distribution. “Closer to zero than bge on this corpus” is the claim these measurements support; “isotropic” is not.

The effective rank here is an entropy-based summary of how broadly variance is spread across the singular spectrum — one number, not the count of semantic concepts and not an intrinsic dimension in the strong sense. Chapter 7 takes dimensionality apart properly. For now: on these five model-corpus combinations, variance was spread across a few hundred directions regardless of whether the nominal dimension was 384 or 1,024.

Anisotropy survives rotation. Rotate the cloud and its lean rotates with it, but the cloud is exactly as lopsided as before, because every pairwise angle is preserved. Rotation can scramble the axes without making a lopsided cloud round. Anisotropy is part of the geometry, so it is intrinsic; only the axis it happens to align with is basis-dependent.

Reading an absolute cosine value

The mean random-pair cosine is an empirical background level for one model on one corpus under one embedding protocol. It is not the mathematical origin of cosine space — cosine zero already means something specific, orthogonality — and it is not a floor. The standard deviations above are 0.09 to 0.12, so plenty of random pairs score below the mean; nothing prevents a pair from scoring near zero or negative even when the average is 0.45.

What it does tell you is that an absolute cosine has no fixed interpretation across spaces. A cosine of 0.6 sits very differently relative to a background distribution centered near 0.06 than to one centered near 0.45. But the mean alone is not enough to call either case “strong,” “weak,” or “noise”: the spread, tails, and relevant-pair distribution matter too. For example, bge-small here has a random-pair mean of about 0.451 with a standard deviation of about 0.094, so 0.6 is roughly 1.6 standard deviations above that sampled background mean; that standardized offset is informative, but it is not a relevance probability and does not by itself choose a threshold.

The right move is to read a score against the distributions it lives in: background pairs and, ideally, known-relevant and known-irrelevant pairs for the actual task. There is no universal correction — centering, percentiles, standardized scores, or a held-out threshold answer different questions. Chapter 8 works through corrections; the Chapter 5 point is only that an absolute cosine needs the context of its space’s score distribution before it earns a semantic interpretation.

What this chapter establishes and what it does not

Establishes: a coordinate is a basis-relative description; inner products, norms, cosine, angles, and L2 distance are provably invariant under a joint orthogonal rotation, and on RELATE / all-mpnet-base-v2 retrieval nDCG@10 was invariant to 4.4 × 10⁻⁷ under three such rotations; a coordinate-property correlation is basis-dependent and was seen to move (+0.29 to −0.38) under one rotation; and the five measured encoders have distinct, non-zero random-pair cosine baselines (0.06 to 0.45) on this corpus.

Does not establish: that raw coordinates are never empirically interpretable in a model’s native basis (claim 2 above can hold); that these encoders use superposition, or how many features they carry; that MiniLM and mpnet are isotropic; that any of the measured baselines are universal constants rather than measurements on RELATE v0.1 with L2-normalized outputs; or that L1-based or coordinate-sensitive retrieval is rotation-invariant — it is not.

Lab 5: the rotation test, and the background distribution

MEASURED — artifacts experiments/embeddings-from-first-principles/wave2/artifacts/rotation-invariance.json and experiments/embeddings-from-first-principles/wave1/artifacts/anisotropy.json. REPRODUCIBLE — run_wave2.py 2.1.

Question. Which quantities survive a change of basis, and which are artifacts of it?

What we ran. On all-mpnet-base-v2 (identified by nDCG@10 0.951773), we scored RELATE retrieval, then rotated every document and query vector by the same random orthogonal matrix and re-scored — three independent rotations. Separately, we correlated the highest-variance coordinate with text length, before and after one rotation, and measured the mean random-pair cosine for five models.

QuantityBeforeAfter rotationKind
nDCG@100.9517730.951773 (max Δ = 4.4 × 10⁻⁷ over 3 rotations)invariant
corr(top-variance coordinate, text length)+0.293−0.38 (one rotation)basis-dependent
mean random-pair cosine0.06–0.08 (MiniLM, mpnet) … 0.34–0.45 (bge, mxbai)unchangedinvariant

Interpretation. Establishes: retrieval built on cosine/dot/L2 is rotation-invariant to numerical precision; a single-coordinate interpretation is not stable under rotation; and, for a fixed embedded corpus, the random-pair cosine distribution is itself rotation-invariant because all pairwise cosines are preserved. The measured baseline varies across the five model–corpus combinations. It is not a context-free constant of a model. Does not establish: what any coordinate “means” — the correlation moved rather than vanished, which is the point.

Try it yourself

Reader experiments; no artifact backs the specifics.

  1. Compute baseline pairwise cosine, dot product, L2, and L1 for a sample of items.
  2. Draw a random orthogonal R (QR of a Gaussian matrix). Rotate documents and queries by the same R.
  3. Verify that cosine, dot, and L2 are unchanged to numerical precision. Then compare L1 across many pairs: unlike the other three it is not invariant, so at least some values will generally move even though a particular pair could coincide by accident. Confirm that cosine retrieval and its nDCG remain unchanged.
  4. Sort items by one chosen coordinate before and after the rotation. Try to name that coordinate both times. Measure its correlation with a surface property such as length. Watch the coordinate-level interpretation shift while retrieval does not.
  5. Estimate the random-pair cosine distribution for your space: mean, standard deviation, and a few quantiles. Take a cosine you would otherwise have called “high” and locate it in that distribution. Does it still look unusual?

Companion component: the basis-invariance guard

Before the Observatory interprets any quantity, it records whether that quantity is intrinsic to the geometry in use or fragile to an equivalent change of basis.

basis_invariance:
  invariant_under_joint_orthogonal_rotation:
    - dot_product
    - cosine
    - l2_distance
    - vector_norm
    - angles
    - knn_and_retrieval_metrics_when_using_an_invariant_metric
    - singular_value_spectrum
  basis_dependent:
    - raw_coordinate_value
    - per_coordinate_variance
    - axis_aligned_threshold
    - coordinate_pruning
    - l1_distance_and_knn_built_on_it

background_geometry:
  random_pair_cosine:
    mean:      <float>
    std:       <float>
    quantiles: <p05, p50, p95>
  top_singular_value_share: <float>
  effective_rank:           <float>
  embedding_normalization:  <none | l2 | ...>

Any report that leans on a basis_dependent quantity — a named dimension, an axis-aligned threshold, a coordinate-pruned index — is flagged for review. The claim may be perfectly valid for the model’s native output basis, but it does not automatically transport to an orthogonally equivalent representation. The report must therefore state the basis/representation version and justify why a coordinate-sensitive operation is intended.

Failure modes

Each is a mistake, why it is tempting, and the check.

  • “Dimension 42 is the sentiment dimension.” Tempting because a sorted coordinate sometimes looks thematic. Check: decide which claim you are making. If you mean that coordinate 42 is an intrinsic axis of the rotation-invariant geometry, a generic change of basis breaks that coordinate-specific claim. If you mean only that coordinate 42 correlates with sentiment in this model’s native output basis, say exactly that and test the correlation there.
  • Axis-aligned pruning by coordinate variance. Tempting because “drop the low-variance dimensions” sounds like harmless compression. Check: the per-coordinate variance ranking is basis-dependent — a rotation reshuffles it. PCA supplies data-defined directions whose principal subspace rotates with the data, so PCA truncation avoids this particular axis-choice defect; it still must be evaluated on the downstream task before low-variance components are declared expendable (Chapter 7).
  • Reading a raw cosine without its background. Tempting because 0.6 feels like a strong score. Check: locate it in the background distribution and, ideally, compare it with known-relevant and known-irrelevant score distributions. A different background mean changes the context of 0.6; the mean by itself does not tell you whether the score is signal or noise.
  • Treating a sparse-autoencoder feature as ground truth. Tempting because a discovered direction with a clean-looking activation set feels like a fact. Check: treat the feature as a hypothesis about a learned readout. Validate its semantic consistency and, when making a causal claim, use interventions rather than labels from activation examples alone.

What this chapter established

  • A coordinate is a description relative to a basis. The model emits one native basis, but infinitely many orthogonally related bases preserve the cosine/dot/L2 geometry retrieval actually uses — so that geometry cannot identify which basis it was handed.
  • A coordinate’s correlation with an interpretable property is therefore basis-dependent: it moved from +0.293 to −0.38 under a single invertible rotation, while retrieval did not move at all. The single-coordinate readout changed; the representation lost nothing.
  • Basis dependence is elementary linear algebra and does not imply superposition, which is a stronger, separate phenomenon this chapter does not measure.
  • The random-pair mean is an empirical background statistic — not the origin of cosine, not a floor. An absolute cosine needs its score distribution before it earns a task-level interpretation, and the basis-invariance guard records which quantities survive a change of basis before the Observatory interprets any of them.

Next

A joint rotation leaves the global geometry — every pairwise angle and distance — exactly intact. That raises the question the next chapter takes up: even when the global geometry is well behaved, is the local geometry equally well behaved everywhere in the space? The next chapter looks directly at neighborhoods — who is near whom, how dense the space is in different places, whether some points show up in far more neighbor lists than they should — and finds the local picture is less uniform than the global one.