Meaning Becomes Geometry
Build a tiny embedding space by hand and turn semantic questions into geometric ones — distance, direction, angle, magnitude, neighborhood. Then use a measured projection experiment to mark exactly where that translation can be trusted and where it only looks trustworthy.
Part I — A Vector Is Not Meaning
A space you can draw
Give six words two coordinates each, by hand:
x = monarchy / institutional power y = gendered association (-1 … +1)
king ( 0.9, 0.6)
queen ( 0.9, -0.6)
duke ( 0.7, 0.5)
duchess ( 0.7, -0.5)
apple (-0.7, 0.1)
orange (-0.7, -0.1)
Now several semantic questions have a geometric form:
- Which words form a group? →
appleandorangesit almost on top of each other, about 0.2 apart, and far from the four titles. The titles form a looser group of their own nearx ≈ 0.8. Grouping shows up as density. - What distinguishes
kingfromqueen? → the vectorking − queen ≈ (0, 1.2). It points almost purely alongy. The difference between them, in this space, is one attribute. (duke − duchess ≈ (0, 1.0)points the same way, but only because we placed the points that way; a learned space owes us no such consistency.) - Is
appleoriented likeking? → the angle between them is obtuse and their cosine is negative. Fruit and titles point into different half-planes. Orientation separates the two categories cleanly, even though, as we are about to see, it is noisier within a category. - Are
kingandqueenrelated? → obviously: they are the two monarchs. But their Euclidean distance is 1.2, whilekingsits 0.22 fromdukeand about 1.12 fromduchess. Rankking’s neighbors by distance andqueencomes third, behinddukeandduchess. The gender axis — which a search for “royal titles” need not care about at all — has quietly reordered the neighborhood.
That last point is the one to carry forward.
Related is not the same as nearby. “These two are related” is a semantic judgment. “These two vectors are close” is a geometric fact about one space under one metric. The two line up often enough to be useful and not reliably enough to skip checking. Whenever a space has a strong axis that is irrelevant to your question, distance will still rank items by that axis.
Even so, the promise is real: every question above did take a geometric form. State it plainly, but state it carefully.
Once information has been represented as vectors, semantic hypotheses can be posed as geometric questions.
Note what that sentence does not say. It does not say the geometric answer is automatically the right semantic answer. The rest of this chapter is about the gap between posing a question geometrically and trusting what the geometry replies.
When is a geometric answer a trustworthy answer to the semantic question?
| Semantic question | Geometric form | Primitive | What makes the geometric answer trustworthy — or not |
|---|---|---|---|
| Are X and Y related? | small separation | distance | A raw distance means little alone. Read it against the distribution of distances in the same space; depending on the data, the metric, and normalization, that distribution can be wide or tightly concentrated (Chapter 8). |
| What distinguishes X from Y? | the direction X − Y | direction | A direction that encodes an attribute in one region need not encode it elsewhere. “This is a semantic axis” is a claim to test across the space, not a property you get for free. |
| Is X oriented like Y? | small angle / high cosine | angle | Cosine ignores magnitude by construction. Its relationship to the dot product and to Euclidean distance depends on whether the vectors are unit length (see below). |
| Which items form a group? | a dense neighborhood | neighborhood, cluster | A cluster is a fact about density. Whether it corresponds to a topic or category is a further claim, checked by sampling members and non-members. Distinctions the training objective never had to make can also collapse two real groups toward one point. |
The geometric primitives
Five quantities do almost all the work in this book.
- Coordinates. The
dnumbers. Individually close to meaningless (Chapter 5); together, a position. - Distance. How far apart two points are. For the hand-built space in this chapter we use Euclidean distance
‖x − y‖. It is one choice among several; Chapter 4 treats the choice of metric as a decision with consequences, not a default. - Direction.
x − y, read as a vector. The hope behind analogy arithmetic is that a given semantic change corresponds to a consistent direction across the space. That is sometimes true locally and rarely exact; treat each instance as testable. - Angle. The angle between
xandy, measured through cosine similarity. Sensitive to orientation, indifferent to length. - Magnitude.
‖x‖. Cosine deliberately discards it. Whether magnitude carries useful signal in a given space — and what that signal is — depends on the model, its objective, its pooling, and its normalization. It is something to measure for a specific representation, not a general property of embeddings.
Two derived ideas:
- Neighborhood. The set of points near a given point. In practice, “what does this vector mean?” is answered as “what is it near, in the space we actually query?”
- Cluster. A region of higher density. A geometric structure first; a semantic category only once that has been checked.
Do not confuse these. Cosine similarity already divides by both vector norms, so L2-normalizing the two vectors first does not change their cosine. What normalization changes is how cosine relates to the other primitives. For unit vectors the dot product equals the cosine, and Euclidean distance becomes a strictly decreasing function of it (
‖x − y‖² = 2 − 2·cos), so all three rank pairs the same way. For vectors of differing lengths they can disagree, and which one you reported starts to matter. Chapter 4 takes this apart properly.
Where the translation leaks
The hand-built space was built to make the geometry legible. Learned spaces differ in specific ways, and each one weakens a correspondence from the table above.
The axes are not given. We named x “power” and y “gender.” A learned model hands you d unlabeled numbers, and the directions that carry an attribute are usually diagonal combinations of coordinates rather than the coordinates themselves (Chapter 5).
A local direction is not a global axis. king − queen and duke − duchess point the same way here because we placed them that way. In a learned space, a direction that separates one pair by an attribute in one region does not automatically separate another pair by the same attribute elsewhere. A successful analogy is evidence for that relation in that part of that space, not proof of a universal semantic axis. Test a direction wherever you intend to rely on it, and report where it holds and where it breaks.
A distance needs its distribution. A Euclidean distance of 1.2 told us nothing until we compared it to the other distances in the same six-point space. The same is true at scale, with an added complication: for some distributions and metrics, most pairwise distances fall into a narrow band, so the gap between “near” and “typical” is small and easily lost in noise (Chapter 8). Calibrate against the space’s own distances before reading anything into one of them.
Distinctions the training signal does not reward can become geometrically weak. If separating two cases does little to improve the training objective, the resulting space may place them very close together. Chapter 1 showed exactly that failure shape for negation and role reversal: changing the assertion could leave cosine similarity almost unchanged. Geometric closeness can therefore mean “these express nearly the same claim” or “this readout barely responds to the distinction between them.” The geometry alone does not tell you which.
Demonstration: RELATE in 2D
MEASURED on RELATE v0.1 — the first 250 items by sorted id,
bge-small-en-v1.5(384-d), L2-normalized, centered, then PCA to two dimensions. Artifactexperiments/embeddings-from-first-principles/wave1/artifacts/pca-projection.json.
Take 250 RELATE items, embed them, and reduce the 384-dimensional vectors to two dimensions with PCA so they can be plotted:
variance retained by the top 2 principal components 15.3%
mean fraction of each item's 10 full-space nearest 20.3%
neighbors still among its 10 nearest in 2D
|Pearson r| of PCA axis 1 with sentence character length 0.64
Read the three numbers separately, because each says something different.
The 15.3% is the share of the embedding’s total variance that the two plotted axes account for. The remaining ~85% is variation among the vectors that a 2D picture cannot show at all.
The 20.3% is a neighborhood-preservation figure, and its exact form matters. For each of the 250 items, take its ten nearest neighbors in the full space and its ten nearest in the 2D projection, measure how much those two sets overlap, then average over all items. The mean overlap is about one fifth. Put differently, averaged across the 250 items, roughly four out of five full-space top-10 neighbor identities are absent from the projected top 10. That does not say every item loses exactly 80% of its neighbors; it says the projection preserves little of the neighborhood structure on average — exactly the structure a reader is tempted to infer by eye.
The 0.64 is the absolute correlation between PC 1 — the highest-variance axis of this PCA projection — and a plain surface feature of the text: how many characters the sentence has. Sentence length may be irrelevant to whatever semantic question brought you to the plot. That is the point. A visually important direction can be strongly associated with a property you never chose to inspect, and nothing in the picture itself tells you whether that property is the one you care about.
MEASURED: this 2D PCA projection keeps 15.3% of the centered embedding variance. Across the 250 items, the mean overlap between each item’s full-space and projected top-10 neighbor sets is 20.3%. PC 1 has |r| = 0.64 with sentence character count. The picture may look orderly even when most top-10 neighbor identities, averaged over items, differ from those in the source space.
Projection methods preserve different aspects of a space and discard or distort others. This chapter has measured PCA here; the lab below also measures t-SNE. UMAP is not measured in this chapter. The exact distortion therefore depends on the method and settings, but the engineering rule is broader than any one projector:
A visualization of an embedding space is another transformation applied on top of a learned transformation. Treat the plot as evidence about the projection, and check any claim meant for the original space in the original space.
More briefly: discover in the picture, verify in the space. A plot is a good place to notice a pattern, form a hunch, spot an outlier, or communicate a structure worth investigating. By itself, it is not confirmation of a structural claim about the source embedding space.
What this chapter establishes and what it does not
Establishes: semantic hypotheses can be posed as geometric questions (relatedness and distance, contrast and direction, orientation and angle, grouping and density); the five primitives and two derived notions; four specific ways the semantic-to-geometric correspondence weakens in learned spaces; and, from a measured projection, that a 2D layout can keep a persuasive-looking structure while, on average, only about one fifth of each item’s ten nearest neighbors survive.
Does not establish: that analogy arithmetic succeeds or fails in general; that any one projection method is best or worst; that clusters seen in a plot are categories; or that the specific numbers (15.3%, 20.3%, 0.64) hold for other corpora, encoders, or projection settings. Those are measured later, per space.
Lab 2: build one, then break one
MEASURED — artifacts
experiments/embeddings-from-first-principles/wave1/artifacts/lab02-break-one.json(12-word demonstration,bge-small-en-v1.5, seed 42) and.../pca-projection.json(250-item RELATE projection). REPRODUCIBLE —venv/Scripts/python experiments/embeddings-from-first-principles/wave1/lab02_break_one.py.
Question. When a real embedding space is flattened to 2D, what survives — and what does the picture add that was never there?
Part A — build a space you fully understand. Hand-place 12 words in 2D in four tight groups: animals (cat, dog, horse), vehicles (car, truck, bus), colors (red, blue, green), actions (run, jump, drive). Measured check: the largest distance within a group is 0.412, the smallest distance between groups is 4.601. Because we chose the coordinates and the grouping rule, we know why the geometry has this structure. The distances do not discover the categories; they faithfully report a distinction we deliberately built into the space.
Part B — break a space you did not build. Embed the same 12 words with bge-small-en-v1.5, rank neighbors by cosine similarity in the full 384-dimensional space, then repeat after reducing to 2D with PCA and with t-SNE (perplexity 4, seed 42). The experiment centers each 2D projection before computing its cosine neighbor ranks:
| Word pair | Full-space rank | PCA rank | t-SNE rank |
|---|---|---|---|
| cat ↔ dog | 1 | 1 | 1 |
| car ↔ truck | 2 | 1 | 3 |
| red ↔ blue | 1 | 1 | 1 |
| cat ↔ car | 3 | 10 | 7 |
| dog ↔ bus | 11 | 11 | 9 |
| red ↔ cat | 5 | 5 | 5 |
The rank is directional even though the table uses ↔ as a compact label: for a row a ↔ b, the script asks where b appears in a’s neighbor list, with rank 1 meaning b is a’s nearest neighbor. Thus cat ↔ dog at rank 1 means dog is the nearest neighbor of cat in all three representations. More strikingly, cat ↔ car is rank 3 in the full embedding space but falls to rank 10 after PCA and rank 7 after t-SNE. The claim is not that cat and car are mutually third-nearest; it is the specific, reproducible cat→car rank recorded by the experiment.
Nothing about the source embedding changed between those columns. The 384 numbers for cat and car are identical throughout; only the derived 2D representation changed. Across all 12 query words, the experiment asks what fraction of each word’s full-space top-3 cosine neighbors remain in its projected top 3, then averages that fraction. The result is 0.78 for PCA and 0.81 for t-SNE. At this toy scale, much of the local structure survives — which is exactly why an individual failure can be easy to miss. The two projectors also disagree about some of the ranks they alter, so changing projector does not remove the need to check the source space.
Now scale up. The 250-item RELATE projection from earlier keeps 15.3% of the variance and only 20.3% of the 10-nearest-neighbor structure, and its leading axis correlates |r| = 0.64 with sentence length. The 12-word case is small enough to hold in your head; the RELATE case is what the same effect looks like once the space is too large to check by eye.
(UMAP was not run for this lab — the dependency is not vendored. It is another nonlinear projection whose output would need the same full-space check; this lab does not measure it.)
Interpretation. Establishes: a 2D plot is evidence about a projection of a representation, and full-space neighbor ranks can differ sharply from what the plot shows — trust the rank, not the layout. Does not establish: that one projector is generally better; here PCA and t-SNE distort different pairs.
Try it yourself
Open
lab02_break_one.pyand edit the three settings at the top:WORDS(the list of terms),MODEL(try"sentence-transformers/all-mpnet-base-v2"in place ofbge-small), andSEED. Re-run and compare the full-space ranks with the PCA and t-SNE ranks. Then add a word with two common senses, such asbat, and see which of its neighbors each projection keeps. Which adjacencies survive every projection you try?
Companion component: the geometry probe
The projection experiment showed that a picture can rearrange neighborhoods while still looking convincing. The response an Observatory should make is to compute the geometric facts from the source space first, and keep them, so any plot can be checked against them rather than trusted on its own.
geometry_probe(space, items) -> {
cosine_similarity: pairwise cosine similarity matrix
euclidean_distance: pairwise Euclidean distance matrix
nearest_neighbors: for each item, the top-k neighbor identities and their ranks,
computed in the FULL space, under a stated metric
direction_samples: (a - b) for labeled contrast pairs
value_distributions: histograms of the pairwise cosine and distance values, so any
single similarity or distance can be read against the whole space
}
Two rules travel with it:
- Report cosine as a similarity and Euclidean as a distance. Do not call a cosine value a “distance” unless you mean
1 − cosineand say so. - Any 2D plot shown to a reader carries a note stating which of its visible adjacencies still hold among the full-space nearest neighbors.
Failure modes
Each of these is a specific mistake, the reason it is tempting, and the check that catches it.
- Reading the plot as the space. Tempting because a clean layout looks like an explanation and well-separated blobs feel like discovered categories. Check: before believing a separation or an adjacency you see, look up those items’ nearest neighbors in the full space with
geometry_probe. - Assuming a global semantic axis. Tempting because one clean analogy generalizes in the mind into a rule. Check: apply the direction in several regions of the space and report where it holds and where it does not.
- Reading a raw distance. Tempting because a single number feels objective. Check: place it against the distribution of pairwise distances in the same space before drawing a conclusion.
- Confusing magnitude with importance. Tempting because “longer vector, stronger signal” is intuitive. Check: measure whether
‖x‖carries task-relevant information in this space. If magnitude is not part of the signal you intend to use, choose a readout that removes it, such as cosine; if it is, preserve and evaluate it rather than assuming either choice is universally correct. - Naming a cluster. Tempting because a shape on a plot invites a label. Check: sample items inside and outside the cluster and confirm the boundary tracks the property you are attributing to it.
What this chapter established
- Semantic hypotheses can be posed geometrically — relatedness as distance, contrast as direction, orientation as angle, grouping as density — using five primitives (coordinates, distance, direction, angle, magnitude) and two derived notions (neighborhood, cluster). Posing the question geometrically is not the same as trusting the answer.
- “Related” and “nearby” are different claims: an axis irrelevant to your question can still reorder a neighborhood.
- A visualization is a transformation applied to a transformation — discover in the picture, verify in the space — and the geometry probe exists to preserve the full-space facts any plot has to be checked against.
Next
We now have ways to interrogate a space — its distances, directions, angles, neighborhoods — and evidence that both the geometry and the pictures we draw of it can mislead. One question has gone untouched the whole chapter: where did this geometry come from? Every space so far was handed to us, either placed by hand or produced by a model we treated as given. The next chapter builds one from scratch, starting from nothing but word counts in a six-sentence corpus, and watches embedding structure emerge as learned compression rather than coordinates handed down by a model.