Change the Model, Change the Universe
Two encoders can produce vectors of the same length without producing vectors in the same coordinate frame. Learn why equal dimension is syntax, not compatibility, and how to measure structural and decision agreement between independently trained spaces without ever comparing their coordinates directly.
Part V — Embedding Spaces Are Not Universal
The same sentence, three universes
Embed one sentence with three models:
model A (384-d): [ ... ]
model B (768-d): [ ... ]
model C (768-d): [ ... ]
A and B differ in length, so nobody expects to compare them coordinate-wise. But B and C are both 768-dimensional. It is tempting to line up their vectors and read something off the alignment directly — to compute cos(B_sentence, C_sentence) and interpret whatever number comes back.
Resist that temptation, and notice exactly why. Independent training does not, by itself, create a correspondence between coordinate 17 of B and coordinate 17 of C. The models may differ in data, objective, architecture, initialization, seed, or several of these — but the decisive fact is simpler: no alignment contract has established that their bases correspond. A raw cross-model cosine between same-length outputs is numerically computable, but it has no authorized semantic interpretation under that missing correspondence. A value near 0.02 would not prove the spaces are unrelated. A value near 0.99 would not prove they are compatible. Both readings assume the one thing that was never established: that B’s basis and C’s basis line up. Without that, the number is not evidence about compatibility — it is a coordinate-dependent calculation with no declared cross-space semantics.
When two models embed the same text, what — if anything — do their outputs have in common, and how do you measure it without assuming their coordinates already correspond?
Why a rotation makes the point precisely
Chapter 5 already proved the fact this chapter needs, and it is worth reusing rather than reinventing a cross-model demonstration that has no artifact behind it. Take one embedding matrix X and apply a single orthogonal rotation:
X' = X R, RᵀR = I
Chapter 5 showed that this joint rotation preserves every quantity retrieval depends on: norms, dot products within the space, cosine similarity within the space, Euclidean distances, and nearest-neighbor behavior under those metrics. X and X' can drive exactly the same retrieval system — same rankings, same nearest neighbors, same everything a downstream consumer would notice — while their individual coordinate axes differ completely. Coordinate 12 of X and coordinate 12 of X' are unrelated numbers describing the same underlying geometry from two different bases.
Now compare X and X' coordinate-by-coordinate — x against x' for the same item. That comparison has no invariant meaning: it depends entirely on which rotation R happened to be applied, and a different R would produce a different, equally uninterpretable answer, without changing anything about the space’s actual behavior.
This is the whole argument, stated plainly:
If a single orthogonal rotation can change every coordinate value without changing any internal retrieval behavior, then two vectors having the same length cannot, by itself, authorize a coordinate-wise comparison between them.
The rotation argument is not a literal model of what happened between two independently trained encoders, and it does not prove that their geometries must be unrelated. In fact, row 3.1 will shortly show substantial structural agreement between some independently trained model pairs. Its role is narrower and stronger: even in the friendliest possible case — exactly the same internal geometry expressed through a different orthogonal basis — equal dimensionality still does not authorize coordinate-wise comparison. Independent training gives us no reason to assume a simpler correspondence than that unless an explicit alignment constraint or measured map establishes one.
Naming this precisely
The current chapter’s opening reflex is to say the two spaces have no shared origin, no shared axes, no shared scale. That is close, but sloppier than the book can afford at this point. Two 768-dimensional model outputs are both, mathematically, elements of ℝ^768 — they share the same zero vector, the same coordinate count, the same ambient space in the linear-algebra sense. And this book’s own Wave 3 harness deliberately standardizes vector norm before comparison — every embedding is L2-normalized before a row 3.1 measurement runs — so “no shared scale” is not even descriptive of how the comparison below is actually computed.
What is actually missing is not a shared origin or a shared scale. It is a shared basis — an authorized correspondence between B’s coordinate axes and C’s coordinate axes. State the claim at the level it is actually true:
There is no justified correspondence between two independently learned models’ coordinate axes, or between their individual coordinate values, unless the spaces were trained or explicitly mapped under a stated alignment constraint.
Prefer no shared coordinate semantics, no authorized basis correspondence, or independently learned coordinate frames over claims about origins, axes, or scale as literal geometric objects. The two arrays live in the same mathematical space in the trivial sense that any two same-length real vectors do; that triviality is exactly why dimension alone was never going to settle anything.
Three layers of cross-space comparison
If coordinate values cannot be compared directly, the question is not “are these spaces comparable” but “at what level.” Three layers, from least to most demanding:
Layer 1 — coordinate correspondence. Can coordinate j in space A be treated as coordinate j in space B? For two independently trained encoders, the default answer is no authorization — and, as the rotation argument shows, equal dimension does not change that answer. This layer is not measured in this chapter; it is the layer this chapter rules out by default.
Layer 2 — structural agreement. Do the two spaces arrange the same items similarly, relationally, without assuming their coordinates already correspond? This is where neighborhood overlap and CKA live: both compare relational or global structure computed over a shared item population, never coordinate j against coordinate j.
Layer 3 — decision agreement. Do the two models make the same task-relevant choices? Exact top-1 retrieval agreement and hard-negative decision agreement live here: not “is the geometry similar” but “did the two systems produce the same measured result or verdict?” Agreement at this layer is still distinct from correctness, a boundary the demonstration below makes explicit.
These layers are not redundant restatements of each other, and the experiment below exists specifically to show that a pair can score high on Layer 2 while leaving real disagreement at Layer 3. Holding them apart is the intellectual work of this chapter — the same discipline Chapter 12 applied to policy versus execution, Chapter 13 to ranking versus score meaning, and Chapter 15 to a scalar versus richer ranking evidence, now applied one level up: across spaces instead of within one.
What can be compared without shared coordinates
Layer 2 and Layer 3 give concrete, computable methods, and it is worth being honest up front about which of them this chapter actually measures and which remain conceptual tools for later use:
- Neighborhood overlap. For each item, take its top-
kneighbors in space A and its top-kneighbors in space B, and measure how much those neighbor sets agree. Measured below, row 3.1. - Linear CKA (Centered Kernel Alignment). A scalar comparing the global linear structure of two centered representation matrices, usable even when the two matrices have different numbers of columns. Measured below, row 3.1.
- Top-1 retrieval agreement. Do the two models return the exact same top-ranked item for the same query? Measured below, row 3.1.
- Hard-negative decision agreement. Do the two models make the same binary call on a specific positive-versus-hard-negative comparison? Measured below, row 3.1.
- Mutual k-NN consistency. The fraction of item pairs that are mutual top-
kneighbors in both spaces — a stricter, symmetric variant of neighborhood overlap. Conceptual here — not computed by row 3.1. - RSA (Representational Similarity Analysis). Correlate the two full pairwise-distance matrices directly, rather than reducing to a single Gram-based scalar as CKA does. Conceptual here — not computed by row 3.1.
- Procrustes residual. The best-orthogonal-alignment error between two representation matrices, giving a direct measure of how far a rotation-only map falls short (Part VI). Conceptual here — not computed by row 3.1.
- Rank correlation of scores. Spearman-style agreement between the two models’ full score orderings, rather than a single overlap or agreement fraction. Conceptual here — not computed by row 3.1.
This mirrors the distinction Chapter 15 drew between the conceptual diagnostic toolbox and the six features Wave 1 row 1.12 actually measured. The list above is a toolbox; the table below reports what one specific experiment, row 3.1, actually ran. Keep those two facts separate as you read the rest of this chapter — this is exactly the conflation the rest of the chapter now works to avoid.
Demonstration: how much do three model pairs agree?
MEASURED on RELATE v0.1, Wave 3 row 3.1 — artifact
experiments/embeddings-from-first-principles/wave3/artifacts/space-comparison.json.
What the implementation actually computes, exactly, for each of the three model pairs:
10-NN overlap. Every corpus item embedding is L2-normalized in each model independently. Item-item cosine similarity is computed within each space; each item’s own entry is excluded (fill_diagonal(S, -2) before ranking). Each item’s top-10 neighbors are taken in space A and in space B. For item i:
overlap_i = |N_A(i) ∩ N_B(i)| / 10
averaged over every corpus item. This is a mean top-10 neighbor overlap fraction — intersection size divided by the fixed constant 10. It is not Jaccard similarity, which would divide by |N_A(i) ∪ N_B(i)| instead; the artifact’s field name (neighborhood_overlap_at10) and this chapter’s prose both use the intersection-over-10 definition throughout.
Linear CKA. Both item representation matrices are centered — X ← X − mean(X), Y ← Y − mean(Y) — then:
CKA(X, Y) = ‖XᵀY‖²_F / ( ‖XᵀX‖_F · ‖YᵀY‖_F )
computed on the shared item rows in both matrices, regardless of whether the two matrices have the same number of columns — which is exactly why CKA can compare a 384-dimensional space to a 1024-dimensional one. In this feature-space form, XᵀX and YᵀY are centered self-products and XᵀY is the centered cross-product; linear CKA has an equivalent formulation in terms of sample Gram matrices. The resulting quantity is invariant to orthogonal transformations and isotropic scaling applied to either representation. It is not invariant to an arbitrary invertible linear transformation; a shear or non-uniform coordinate rescaling generally changes the value. Do not read a high CKA as “the spaces agree up to any linear transform” — that overstates what the quantity is insensitive to.
Top-1 retrieval agreement. For each release query, each model independently scores the full corpus and identifies its own top-ranked item:
top_A = argmax score under model A
top_B = argmax score under model B
agree = (top_A == top_B)
averaged over queries. The field name in the artifact and used throughout this chapter is top-1 retrieval agreement — exact same top-ranked item, in both models, for the same query. This is not top-10 overlap, not Recall@10, and should never be shortened to “retrieval agree@10.”
Hard-negative decision agreement. For each query with at least one grade-3 positive, take the first grade-3 positive, p0 — the same selected-reference-positive convention used in Chapters 10, 11, and 15, not “the correct answer” in some broader sense. For each of that query’s listed hard negatives h, each model independently computes a Boolean:
decision_A = score_A(q, p0) > score_A(q, h)
decision_B = score_B(q, p0) > score_B(q, h)
agree = (decision_A == decision_B)
averaged over every (query, hard negative) comparison. This measures whether the two models make the same Boolean call on whether the selected reference positive outranks a specific hard negative — nothing about whether that call is correct. A tie under the strict > comparison resolves to False for both terms. Two models agree either because both say positive > negative, or because both say positive ≤ negative — the artifact cannot distinguish those two very different situations, and this chapter will not pretend it can.
The measured table:
model pair 10-NN overlap linear CKA top-1 retrieval agree hard-neg decision agree
bge-large vs mxbai-large 0.879 0.987 0.862 0.961
minilm-l6 vs mpnet-base 0.714 0.814 0.747 0.943
mpnet-base vs bge-large 0.705 0.846 0.673 0.911
with pair metadata recorded alongside the measurement, as reported by the harness:
bge-large vs mxbai-large— same dimension (1024), both described by the harness as retrieval-tuned, different creators.minilm-l6 vs mpnet-base— different dimensions (384 vs 768), same broad Sentence-Transformers family.mpnet-base vs bge-large— different dimensions (768 vs 1024), different family, both described as retrieval-tuned.
This is a descriptive measurement on one frozen, controlled corpus — RELATE v0.1’s item population and release queries. No mapping is fit, no train/test split is used, and neighborhood overlap and CKA are computed over the full corpus, not a held-out sample. Do not describe these values as held-out generalization, learned compatibility, or bridge quality; those are different experiments, some of which appear in Part VI.
Reading the result without inventing a cause
Rounded to two significant figures, bge-large and mxbai-large score roughly 0.88 neighbor overlap, 0.99 CKA, 0.86 top-1 agreement, and 0.96 hard-negative decision agreement — the highest measured value on all four properties of the three pairs. mpnet-base and bge-large score lowest on three of the four: 0.71 overlap, 0.67 top-1 agreement, 0.91 hard-negative agreement, though not the lowest CKA. minilm-l6 and mpnet-base sit between them on most measures.
It is tempting, staring at that pattern, to reach for the metadata column and conclude that equal dimension, or shared training family, or shared tuning objective, explains the ordering. Resist it. Only three pairs were measured, and every candidate explanation is confounded with every other one:
bge-large/mxbai-largeshare dimension and both are retrieval-tuned and have different creators.minilm-l6/mpnet-basediffer in dimension but share a training family.mpnet-base/bge-largediffer in dimension and family, while both are retrieval-tuned.
No pair provides a controlled intervention that changes one candidate cause while holding the others fixed. With three measured pairs and at least six plausible explanatory factors — dimension, architecture, training objective, training corpus, family lineage, and creator — no causal attribution is available from row 3.1. The defensible statement is:
Among these three pairs,
bge-large/mxbai-largeshows the highest agreement on all four measured properties. Row 3.1 does not identify why.
Their shared metadata — same 1024-dimensional output, both retrieval-tuned — is worth recording, because it is a reasonable place to look next. It is a hypothesis generator, not a finding. Pair metadata can suggest why two spaces might agree; it does not, by itself, explain why row 3.1 measured what it measured. That restraint is not a hedge — it is the same discipline Chapter 11 applied to margin, Chapter 14 applied to calibration, and Chapter 15 applied to the NLI decrement: measure first, and do not let a plausible story stand in for an experiment that was not run.
The same restraint applies to the reverse claim. Only one of the three measured pairs shares a dimension (bge-large/mxbai-large, both at 1024), and that pair happens to have the strongest agreement of the three — so row 3.1 cannot support “equal dimension predicts nothing,” despite that being the book’s governing intuition from the rotation argument. What the measured set actually supports is narrower and still useful: matching dimension is not evidence of coordinate correspondence — the rotation argument already established that on first principles, independent of any experiment — and structural compatibility has to be measured independently of vector width, which is exactly what this experiment does for three specific pairs.
The mismatch that matters most
The single most important reading in this table is not which pair wins. It is what happens within one pair when you compare its own four numbers to each other.
Take mpnet-base vs bge-large: linear CKA of 0.846 alongside a top-1 retrieval agreement of 0.673. The two representation matrices show fairly strong global linear structural agreement under this measurement — and yet the two models return a different top-ranked item on roughly a third of the evaluated queries. Structural agreement and decision agreement are not the same quantity, and one does not predict the other with any precision this experiment establishes. A model pair can look close by one measurement and diverge meaningfully by another, on the same corpus, at the same moment.
This is not a contradiction to explain away, but the reason is not that top-1 cosine retrieval is coordinate-sensitive. If one space — queries and corpus vectors together — is rotated orthogonally, its cosine ranking is preserved just as Chapter 5 established. The distinction is instead one of scope and aggregation: linear CKA summarizes global representational structure across the shared item sample, while top-1 agreement asks a local, query-specific argmax question. A high global structural score does not mathematically force two independently learned spaces to choose the same nearest winner for every query. High CKA and top-1 disagreement are therefore answers to two different measurements, not evidence that one of them is secretly testing raw coordinate alignment.
State the disagreement exactly, per pair:
bge-large vs mxbai-large: 1 − 0.862 = 0.138 (13.8% of queries, different top-1 item)
minilm-l6 vs mpnet-base: 1 − 0.747 = 0.253 (25.3%)
mpnet-base vs bge-large: 1 − 0.673 = 0.327 (32.7%)
Top-1 disagreement across these three pairs ranges from 13.8% to 32.7%. minilm-l6/mpnet-base is annotated by the harness as a same-family pair, not cross-family, so describing “the cross-family pairs” as a block that disagrees on roughly a third of retrievals is not accurate to what was measured — only mpnet-base/bge-large reaches that figure.
Even bge-large/mxbai-large, the strongest pair by every row 3.1 measure, disagrees on 13.8% of top-1 retrievals on this RELATE v0.1 workload despite a CKA of 0.987 — roughly one in seven evaluated queries returns a different top item under the two models. That is material evidence of decision divergence on this controlled probe, not an estimate that a production workload will also change one query in seven. Production impact has to be measured under the application’s own Chapter 13 evaluation contract. Structural agreement does not authorize an assumption of decision agreement, even at the highest value measured here.
Agreement is not correctness — twice over
Two separate places in this experiment can be misread as measuring quality when they measure only agreement, and both deserve to be stated with equal weight.
Top-1 retrieval agreement. If both models return the exact same wrong item for a query, that counts as agreement. 0.862 for bge-large/mxbai-large does not mean both models are correct 86.2% of the time — it means they returned the same top-ranked item on 86.2% of evaluated queries, correct or not. Row 3.1 does not check either model’s output against a relevance judgment; it checks the two models against each other.
Hard-negative decision agreement. This one is subtler, because it looks like a quality metric — “agreement” sitting next to “hard negative” invites the reader to assume it is measuring how often the positive wins. It is not. A True agreement occurs whenever both models reach the same Boolean verdict, and there are two ways to reach the same verdict: both correctly rank the positive above the hard negative, or both incorrectly fail to. The artifact stores only the agreement fraction, not which of those two cases produced it. 0.911–0.961 across the three pairs is therefore not hard-negative accuracy, not hard-negative Recall, and not evidence that “every model ranked the correct answer on top” — it is evidence only that the two models’ verdicts, right or wrong, coincide most of the time.
This is one of the clearest instances in the whole book of a distinction worth keeping permanently in view: agreement between two systems and correctness of either system are independent facts, and an experiment that measures one says nothing about the other unless it is explicitly designed to.
Because of that, this chapter does not adopt the explanation that near-restatement queries let every model rank the positive first — row 3.1 contains no query-style breakdown, no correctness check against either model’s individual hard-negative performance, and no causal decomposition of any kind. The defensible statement is exactly this narrow:
The three model pairs make the same binary reference-positive-versus-hard-negative decision on 91.1%–96.1% of the measured comparisons on RELATE v0.1. That high agreement does not indicate whether both models were right or both were wrong on those comparisons, and row 3.1 does not isolate why the agreement is high.
A later chapter’s fitted-bridge result (Chapter 21, row 3.7) shows a polarity distinction actually inverting under a specific mapping — a real and striking finding in its own right. It is not, however, an explanation of this measurement. Row 3.1 fits no map and evaluates no bridge; reaching forward to borrow Chapter 21’s result as a cause for a pattern observed here would misuse a finding that has not yet been earned in the book’s own sequence, and would risk pre-teaching an arc Part VI is built to deliver on its own terms. The honest forward pointer is narrower: later chapters will show that preserving coarse structure under a fitted map does not guarantee preservation of every relation-level distinction. That is as far as this chapter goes.
Compatibility is a vector of properties, not a scalar
Look at the two pairs at the extremes of the measured table side by side:
bge-large vs mxbai-large mpnet-base vs bge-large
CKA 0.987 CKA 0.846
N10 overlap 0.879 N10 overlap 0.705
top-1 agree 0.862 top-1 agree 0.673
hard-neg agree 0.961 hard-neg agree 0.911
Within each row, the four numbers are high but not identical to each other, and the gap between the two pairs is not uniform across measurements either — CKA differs by about 0.14 between the two pairs, top-1 agreement by about 0.19. A single label — compatible or incompatible, structurally close or unrelated coordinates — would flatten four genuinely different measurements into one bit and throw away exactly the information that makes this table useful.
Compatibility is property-specific. Two spaces can agree strongly on global structure and still disagree meaningfully on individual decisions.
That is not a hedge to soften an inconvenient result — it is the chapter’s central finding, and it is the reason the next chapter needs to separate identity from compatibility from usability rather than collapsing them into one score. A space-comparison report that returns a single verdict is answering a question nobody should be asking; the honest report returns a small set of measurements, each scoped to what it actually tested.
What structural agreement does not authorize
Suppose a pair reached CKA = 0.99 and near-perfect neighborhood overlap. That measurement, by itself, does not authorize:
- averaging the two models’ raw vectors;
- querying model A’s index with a raw model B query vector;
- reusing model A’s calibration threshold in model B without revalidation;
- interpreting “dimension 17” as meaning the same thing across the two models;
- concatenating the two spaces’ coordinates as though their axes correspond.
The reason traces straight back to the rotation argument: linear CKA can remain unchanged when a representation is rotated, even though the raw coordinate values have changed and the rotation itself would be required to align them coordinate-by-coordinate. High CKA therefore says that certain structural relationships agree while being intentionally insensitive to the basis correspondence needed for direct coordinate operations. It neither identifies that correspondence nor proves that no alignment is needed.
High structural agreement tells you the spaces organize shared items similarly. It does not tell you their coordinates are interchangeable.
The two cases behind “you cannot average vectors from two models” are worth separating, because they fail for different reasons. When dimensions differ, the operation is mechanically impossible — there is no such thing as adding a 384-vector to a 1024-vector. When dimensions match, the operation is numerically computable — nothing stops the arithmetic — but it remains semantically unjustified without an explicit shared-coordinate contract or a validated bridge (Part VI). The precise version of the practical rule:
Direct cross-space coordinate operations are unauthorized by default, even when matching dimensionality makes them numerically possible.
This is deliberately not a claim that such operations are permanently impossible. A validated bridge, built and measured under its own contract, is exactly what later chapters construct to make specific cross-space operations meaningful for specific properties. This chapter establishes the default; it does not foreclose the exception.
Threshold reuse deserves its own line, because it fails for a reason that has nothing to do with coordinates at all. A calibration threshold, per Chapter 14, belongs to a calibration contract fit on one model’s score distribution. Even a pair with very high CKA can have different score scales, different positive/negative separations, and a different EER threshold entirely — CKA says nothing about a score distribution’s shape. Do not reuse one model’s similarity threshold in another model without revalidating it under its own calibration contract. That is a distinct failure mode from coordinate incompatibility, not a restatement of it.
Comparing models the way Chapter 13 already taught
When the practical question shifts from “how similar are these two spaces” to “which model should I actually run,” the right comparison is not a raw score at all — it is Chapter 13’s evaluation contract, held fixed across both models:
- the same candidate corpus;
- the same query workload;
- the same relevance definition;
- the same metric and aggregation;
- each model’s own valid representation protocol (Chapter 13’s model-use protocol — prefixes, asymmetric query/document roles, whatever that specific model requires).
Run both models through that fixed contract and compare outcomes — nDCG under the same relevance definition, hard-negative accuracy under the same reference-positive convention — not cos_A = 0.81 against cos_B = 0.77. Two models’ raw similarity scores do not share a scale (Chapter 8, Chapter 14), so a raw-score comparison is not even a fair fight; it is two numbers from two different measuring instruments, compared as though they came from one. An evaluation-contract comparison is a fair fight, because it asks both models the identical question and only then compares what they returned.
What this chapter establishes and what it does not
Establishes: changing the encoder creates a new space-identity boundary at which coordinate correspondence must be re-established rather than assumed; equal vector dimension does not establish that correspondence, and Chapter 5’s rotation argument shows why even two representations with identical internal geometry can have nonmatching coordinates; a raw cross-model cosine is not a valid compatibility test absent an established correspondence; cross-space comparison can instead measure shared structure (neighborhood overlap, linear CKA) and shared decisions (top-1 retrieval agreement, hard-negative decision agreement) over the same item population, without assuming coordinates already correspond; row 3.1 measured exactly these four quantities, with precise semantics for each; bge-large/mxbai-large scored highest of the three measured pairs on all four properties, without row 3.1 establishing why; top-1 disagreement across the three pairs ranges from 13.8% to 32.7%; structural agreement and decision agreement are measurably different quantities, and compatibility is therefore property-specific rather than a single scalar.
Does not establish: that any specific factor — dimension, training family, training objective, tuning target, creator — caused the observed differences between pairs; that shared training regime explains the bge-large/mxbai-large result; that hard-negative decision agreement reflects hard-negative accuracy, or that its high values arise from near-restatement queries; that disagreement concentrates in hard, rare, or difficult regions of the corpus (no slice-conditioned analysis was run); a universal hierarchy of what semantic properties model pairs usually share; bridge quality, threshold transferability, or statistical significance for any of the reported values; or that these three model pairs represent compatibility behavior on a natural, non-synthetic corpus. It establishes the operating principle the rest of Part V and Part VI depend on: equal dimensions do not imply compatible representation, and structural agreement does not imply decision agreement or coordinate interoperability.
Lab 16: reproduce the four measurements, and their exact boundaries
MEASURED — artifact
experiments/embeddings-from-first-principles/wave3/artifacts/space-comparison.json. REPRODUCIBLE —python run_wave3.py 3.1.
Question. Given two independently trained encoders, how much do they agree — structurally and decisionally — and does that agreement depend on their sharing a dimension?
Step 1 — freeze the comparison contract. Record: the RELATE release and corpus hash; the exact item IDs and query IDs used; both models’ identifiers/revisions; each model’s query/document model-use protocol; L2 normalization (applied to every embedding before comparison); k = 10 for neighborhood overlap; and the exact definitions of each of the four metrics below. A comparison number without this contract is incomplete evidence, exactly as an evaluation result was incomplete without Chapter 13’s evaluation contract and a threshold was incomplete without Chapter 14’s calibration contract.
Step 2 — compute mean top-10 neighbor overlap. For each model, independently: L2-normalize every corpus item embedding; compute item-item cosine similarity; exclude each item’s own entry; take each item’s top-10 neighbors. Then for each item, overlap_i = |N_A(i) ∩ N_B(i)| / 10, averaged across all corpus items. Do not call this Jaccard — there is no union term in the denominator.
Step 3 — compute linear CKA. Using the same shared item rows in both representation matrices (dimensions may differ between the two matrices), center each matrix and compute ‖XᵀY‖²_F / (‖XᵀX‖_F ‖YᵀY‖_F). Note in your own record that this is invariant to orthogonal transformation and isotropic scaling, not to arbitrary invertible linear maps.
Step 4 — compute exact top-1 retrieval agreement. For each query, each model scores the full corpus and returns its own single highest-scoring item; agreement is top_A == top_B, averaged over queries. This measures identical returned item, not correctness of either model.
Step 5 — compute hard-negative decision agreement. For each query with a grade-3 positive, using the first grade-3 positive p0: for each listed hard negative h, each model independently evaluates score(q, p0) > score(q, h); agreement is that the two models’ Booleans match. State explicitly in your notes that both-correct and both-incorrect outcomes count identically as agreement.
Step 6 — reproduce the exact table.
bge-large vs mxbai-large: overlap 0.879 CKA 0.987 top-1 0.862 hard-neg 0.961
minilm-l6 vs mpnet-base: overlap 0.714 CKA 0.814 top-1 0.747 hard-neg 0.943
mpnet-base vs bge-large: overlap 0.705 CKA 0.846 top-1 0.673 hard-neg 0.911
Step 7 (PROPOSED — no artifact backs this) — inspect disagreement examples. For queries where top_A != top_B, inspect query style, relation type, entity, the rank model A’s winner receives under model B, and vice versa. This is the analysis that would be required before claiming disagreement concentrates in any particular kind of query — it has not been run, and no claim about where disagreement concentrates should be made until it has.
Step 8 (PROPOSED — no artifact backs this) — robustness checks. Repeat neighborhood overlap for k = 5, 10, 20 to check sensitivity to the neighbor-count choice; bootstrap over items or queries to obtain an uncertainty interval around each reported value; compute per-slice agreement by domain or relation type; optionally add RSA or a Procrustes residual as an additional structural measurement. None of these currently exist as results — the artifact reports point estimates only, with no resampling, no confidence interval, and no sensitivity analysis to k. Do not describe the measured values as statistically significant, stable, or different from each other in a statistical sense; describe them only as the point estimates they are.
Try it yourself
Run the space-comparison measurement on your own corpus for two models you might swap between. Record the same four numbers under the same contract, and specifically check whether a pair with high CKA still shows meaningful top-1 disagreement, the way
mpnet-base/bge-largedid here. If it does, treat that as evidence that the swap changes retrieval behavior and run a fresh Chapter 13 evaluation before deployment. A bridge becomes relevant only when you need a scoped cross-space operation or migration path without simply re-embedding into one consistent target space; disagreement by itself does not imply that a bridge is the required deployment mechanism.
Companion component: the space-comparison observation
The report that comes out of this measurement records evidence. It does not, by itself, decide whether an operation is safe — that would collapse exactly the detector/policy, evidence/verdict separation this book has maintained since Chapter 6, through Chapter 14’s calibration record and Chapter 15’s signal bundle.
space_comparison_observation:
id / version:
spaces:
a: space_record_ref
b: space_record_ref
measurement_contract:
corpus_ref:
corpus_hash:
query_set_ref:
item_metric: cosine
normalization: l2
neighbor_k: 10
hard_negative_reference_policy: first_grade3
structural:
neighbor_overlap_at10:
linear_cka:
decision_agreement:
top1_retrieval_agreement:
hard_negative_decision_agreement:
optional: # conceptual methods, explicitly not run here
mutual_knn_consistency: not_measured
rsa: not_measured
procrustes_residual: not_measured
rank_correlation: not_measured
slice_results: not_measured
provenance:
artifact_ref:
code_hash:
If a scoped decision needs to be made from this evidence — for example, “is this pair close enough to attempt a bridge” — that decision belongs in a separate object, built on top of the observation rather than folded into it, the same way Chapter 15’s routing_policy consumed a signal_bundle without living inside it:
compatibility_policy:
scope: # the specific operation being authorized, never "compatibility" in general
observation_ref: space_comparison_observation
required_properties:
thresholds:
decision:
Chapter 17 and Part VI develop that policy layer properly, with the identity/compatibility/usability separation it deserves. The rule this chapter needs is narrower and worth stating on its own:
A comparison observation records which properties agreed, under which contract. It does not authorize any cross-space operation by itself.
Both neighbor_overlap_at10 and linear_cka are properties of two spaces and a shared item population — add or remove corpus items and an item’s top-10 neighbor set can change even though neither model changed at all; change the sampled item rows and the CKA value can shift too. Neither number is a timeless property of a model pair’s names. Store the corpus reference and its hash alongside every comparison value, the same way a margin (Chapter 11) is incomplete without its negative-selection rule and an evaluation result (Chapter 13) is incomplete without its evaluation contract.
The difference between CKA and a bridge
Worth stating plainly before Part VI arrives, so the two questions are never quietly merged: linear CKA asks do these two centered representation matrices have similar global linear (Gram) structure, under this sample. A bridge asks can vectors be mapped from one coordinate system into the other while preserving the specific property I care about. These are different questions with different evidence requirements. A pair can show high CKA and substantial neighborhood overlap — as bge-large/mxbai-large does here — while a fitted bridge between them still fails to preserve some particular relation-level distinction a downstream task depends on. Nothing in this chapter’s measurements predicts that outcome either way; it is Part VI’s question to answer, with its own experiment.
Observatory behavior
When two space records are compared, the Observatory should be able to answer: which exact two spaces are being compared; over which corpus and sample, with which hash; at what k, under what normalization and metric; which structural properties were measured to agree, and by how much; which decisions were measured to agree, and by how much; and which comparison methods were not run. Enforcement then separates two different hazards. Raw coordinate operations — querying one space’s index with another space’s vector, averaging vectors, coordinate-wise combination — require matching space identity or a bridge validated for that operation. Threshold reuse is different: a bridge does not automatically validate a score threshold, because thresholds belong to their own calibration contracts and must be revalidated or recalibrated for the resulting score distribution.
The Observatory should surface the comparison report in that form. It should not collapse it into an inferred rule — it must not, for instance, treat CKA > 0.9 as license to mix vectors, however tempting that single threshold looks sitting next to this chapter’s table. Any such rule is itself a scoped policy, built on top of the observation, validated for a specific operation, exactly as Chapter 14 required a validated calibration contract before a threshold could gate anything. Matching dimension is not compatibility, and neither is a high value on any single structural measurement, however strong.
Failure modes
- “Both are 768-d, so they’re comparable.” Dimension is a count of coordinates, not a coordinate correspondence — Chapter 5’s rotation argument shows why equal width guarantees nothing.
- Reading a raw cross-model cosine as a compatibility measurement. It is numerically computable and semantically unauthorized without an established basis correspondence.
- Calling neighbor overlap “Jaccard.” Row 3.1’s overlap divides by the fixed constant 10, not by the union of the two neighbor sets.
- Reading “Retrieval agree@10” where the artifact measured exact top-1 agreement. These are different quantities; name the one that was measured.
- Treating high CKA as “linearly close” in the sense of a low-error map. CKA is invariant to orthogonal transformation and isotropic scaling, not to arbitrary linear transforms, and says nothing about whether a bridge exists.
- Confusing agreement with correctness. Two models can return the same wrong top-1 item, or make the same incorrect hard-negative call, and both count as agreement.
- Calling hard-negative decision agreement “hard-negative accuracy.” The measured field is a Boolean-match rate between two models’ verdicts, not either model’s correctness against ground truth.
- Explaining a pair’s ranking from confounded metadata. Three pairs, at least four candidate causes (dimension, family, objective, creator) — none independently varied. A shared property is a hypothesis, not an explanation.
- Reusing a threshold across models because their spaces “look close.” A threshold belongs to its own calibration contract (Chapter 14); CKA and neighbor overlap say nothing about score-distribution shape.
- Mixing vectors because a bridge might eventually be built. A possible future map is not a currently validated one; the default for direct cross-space coordinate operations is unauthorized.
- Reporting a comparison value without its corpus and sample. Neighbor overlap and CKA are both computed over a specific item population; the number does not belong to the two model names alone.
- Assuming disagreement concentrates in hard or rare regions. A reasonable hypothesis, and explicitly not one row 3.1 tested.
What this chapter established
- Changing the encoder creates a new space identity that cannot inherit coordinate correspondence by assumption. Equal output dimension answers “how many coordinates,” not “do the coordinates mean the same thing” — Chapter 5’s rotation argument already proves that identical internal geometry can wear entirely different coordinate values.
- A raw cross-model cosine between two unaligned spaces is not a valid compatibility test in either direction: neither a low value nor a high one is interpretable without an established coordinate correspondence.
- Cross-space comparison separates into three layers — coordinate correspondence (unauthorized by default), structural agreement (neighborhood overlap, linear CKA), and decision agreement (top-1 retrieval agreement, hard-negative decision agreement). Row 3.1 measured exactly those four quantities, each with a precise definition its friendly name does not convey, and left RSA, Procrustes residual, mutual k-NN consistency, and rank correlation unmeasured.
- Structural agreement and decision agreement are measurably different quantities on the same pair:
mpnet-base/bge-largeshows CKA0.846alongside top-1 agreement of only0.673. Top-1 disagreement runs from 13.8% to 32.7% across the three pairs — real divergence even at the highest measured structural agreement. Compatibility is a property-specific vector of measurements, not one universal scalar. - High structural agreement authorizes nothing on its own — not vector averaging, not cross-space search, not threshold reuse, not coordinate-level interpretation. The space-comparison observation records evidence bound to its corpus, sample, and contract; a separate compatibility policy is what authorizes an operation.
Next
This chapter has shown that a vector from model A and a vector from model B cannot be treated as though they came from one coordinate frame merely because they share a dimension — and that even the most structurally similar pair measured here still disagrees on a nontrivial fraction of the RELATE retrieval decisions this experiment actually measured. Once space is part of what a vector means, that meaning cannot be left implicit. Every stored vector now needs an answer to a question this chapter has made unavoidable: which space produced me? That is space identity — exact, explicit, and the subject of the next chapter.