Versioning the Space
A vector is not just float[d]; operationally it is float[d] plus the identity of the transformation that produced it. Give that identity an exact, content-derived hash, separate it cleanly from compatibility and usability, and read a deliberately dangerous mixed-index probe that shows why the separation has to be enforced before a quality metric ever gets the chance to reveal the damage.
Part V — Embedding Spaces Are Not Universal
The upgrade that broke search quietly
Illustrative scenario — not a measured incident.
Production runs embedding model v1. A better v2 ships. Someone updates the client library. New documents get v2 vectors; the existing documents are still carrying v1 vectors. Nothing errors. Queries are embedded with v2 and scored directly against a mix of v1 and v2 document vectors.
No exception is raised, because none of this is mechanically invalid — a v2 query and a v1 document, at the same dimension, dot-product into a perfectly ordinary float. What actually breaks is quieter than an error: the index now contains vectors from two declared spaces while every query is produced in only one of them, so a portion of the comparisons the system runs no longer satisfy the representation contract either side assumed. Retrieval keeps returning ten plausible-looking results. Nothing about the response shape announces that some of those ten were scored by an unsupported cross-space comparison. That is exactly the danger: silent retrieval-contract violation, not silent catastrophe.
What identifies an embedding space, when can two vectors be treated as belonging to “the same” space, and what has to happen when that identity changes?
The answer comes in three layers, and keeping them separate is the entire engineering payoff of this chapter. First, an exact name for a space — a space_hash — so that identity cannot be established by a version string that can drift out of sync with what actually produced the bytes. Second, a decision procedure for what to do when that hash changes. Third, a look at what a deliberately unsupported mix of two spaces actually costs on one measured case — not to price migration, but to show how easily an unsupported operation can look fine.
The space identity
A vector is only interpretable relative to the space that produced it. A list of 768 numbers carries no coordinates of its own; its coordinates come from the entire pipeline that produced it — which weights, which tokenizer, how tokens were pooled and normalized, what role-specific instruction was applied, how long inputs were allowed to be, in what precision the arithmetic ran. Chapter 16 established this across different models. This chapter extends the same claim one level further: a version label does not decide space identity. If a v1→v2 release changes the identity material below, it creates a new declared space; if only the human-facing label changes while the declared transformation remains identical, the space_hash need not change. That is exactly why the hash, rather than the version string, becomes the runtime invariant.
Chapter 1 already gave this book its first version of a reusable answer — a descriptive space_record carrying a model name, version string, dimension, normalization, pooling, objective, and a note on what “similar” seemed to mean. That record was, and remains, useful provenance. What it was never built to be is a runtime identity check: a version string is a label a human chose to write down, not a guarantee about which bytes actually ran. This chapter matures that descriptive record into an exact one.
Separate what defines the transformation from what merely describes it:
space_identity:
identity_material:
weights_artifact_hash: <canonical hash of the weights bundle — sharded files included>
tokenizer_artifact_hash: <hash of the tokenizer files>
dimension: <int, output>
pooling: <cls | mean | last | ...>
normalization: <none | l2 | whitened(params_hash) | mean_centered(mean_hash)>
query_role_protocol: <query-side instruction/prefix, verbatim, or none>
document_role_protocol: <document-side instruction/prefix, verbatim, or none>
max_sequence_length: <int>
truncation_policy: <head | tail | middle | none>
precision: <fp32 | fp16 | bf16 | int8 — dtype/quantization shifts vectors>
post_processing: <PCA(matrix_hash) | whitening(params_hash) | none>
implementation_revision: <only when it can change output — e.g. a kernel or pooling-code change>
metadata:
model_name: <display name>
release_tag: <mutable marketing/version label>
source_revision: <upstream commit, for provenance>
objective_note: <what the model was trained to make close — Ch1's note>
space_hash: SHA256(canonical(identity_material))
space_hash is computed only from identity_material — the fields that define the declared transformation. metadata stays attached for provenance and readability, exactly like Chapter 1’s objective_note, but never participates in the hash. That split handles a failure in both directions: if the artifact bytes change while the friendly release label stays the same, the identity hash changes; if only a descriptive version label changes while the declared transformation does not, the identity hash can remain stable. This works only if the identity contract is complete and the runtime resolves its material correctly — a hash cannot detect a transformation dependency that was never included in what it hashes.
A few fields deserve emphasis, because they are the ones teams most often omit from an identity contract:
- Weights and tokenizer artifact hashes, not a release tag. Modern models frequently ship as sharded weight files, custom modeling code, or a bundled tokenizer — “hash the weights file” undersells what identity material can require. A canonical hash of the full weights/tokenizer bundle is the safer framing; only the bytes are identity, however they are packaged.
- Query and document role protocols, recorded together, not forced equal. Chapter 13 already established that a valid model-use protocol may apply different instructions or prefixes to queries and documents by design — some retrieval models are trained specifically so that an asymmetric query/document pair of transformations cooperates within one retrieval contract. This chapter does not ask the two protocols to match; it asks that both be recorded, because changing either one changes the declared retrieval-space contract, whether or not the other side moved.
max_sequence_lengthandtruncation_policy. These silently change the vector of any input longer than the window — one of the most common invisible space changes, because nothing about a short test query will ever expose it.- Precision. Dtype and quantization can shift vectors enough to matter for some downstream operations and not others; recording it as identity material does not presume how much it matters, only that a change here forfeits the assumption of sameness.
What a matching hash means, and does not
space_hash identifies a declared transformation contract — the specific set of identity-material fields, canonicalized and hashed. It is a deterministic metadata invariant that this book’s runtime computes and enforces consistently. It is not a promise about bit-for-bit numerical reproduction: floating-point kernels, hardware, and runtime details can prevent exact byte-for-byte vector reproduction even under an unchanged identity contract, and nothing about space_hash claims otherwise. If bitwise reproducibility across machines matters for some downstream use, that is a separate reproducibility contract to define and test — this chapter’s hash answers a narrower and more useful question.
State plainly what matching and differing hashes each license, because both directions get overread in practice:
Matching space_hash means the two vectors claim the same declared embedding-space identity contract. That is the ordinary condition under which this book treats vectors as belonging to one declared coordinate system for raw within-space operations — similarity computation, nearest-neighbor search, coordinate-level transforms defined for that space. It does not mean the vectors are bit-for-bit identical outputs of two separate runs, and it says nothing about semantic quality, task usefulness, or whether the corpus and queries behind the vectors were correctly constructed. A calibrated threshold is a separate matter: even inside one space, applying it is authorized by its Chapter 14 calibration contract, not by hash equality alone. Identity is a contract about the transformation, not a certificate about every operation performed on its output.
A different space_hash means the declared identity differs — somewhere in the identity material, the two pipelines diverge. From that alone, infer exactly one operational rule: do not assume cross-space coordinate compatibility. It does not prove the two spaces are incompatible for every operation, or for any particular operation; it withdraws that default assumption and requires the relevant compatibility property to be measured, the way Chapter 16 compared distinct spaces. A changed identity field is therefore grounds to assert identity difference, but not to assert operational incompatibility. Two different precisions might produce near-identical vectors on typical inputs; an implementation revision might be behaviorally equivalent for the operation you care about; a changed truncation policy might have zero effect on inputs that never approach the length limit. Identity tells you that the contract changed. Compatibility tells you whether the change mattered for a specific property on a specific workload.
same hash → same declared space identity
(default: safe to treat as one coordinate system for raw ops)
different hash → different declared identity
(default: no assumption of coordinate compatibility — measure it)
Identity, compatibility, and usability now separate into three distinct layers, and this chapter’s job is to keep them from ever answering for one another:
| Layer | Question | Nature | Artifact | Established in |
|---|---|---|---|---|
| Identity | What exact declared transformation produced this vector? | deterministic, configuration-derived | space_hash | this chapter |
| Compatibility | Which properties survive well enough to compare, map, or migrate? | empirical, property- and workload-specific | comparison / preservation measurements | Ch 16, Part VI |
| Usability | May the system perform operation X under scope Y? | scoped application policy | usable_for(scope, operation) | Ch 20 |
Identity tells you whether the declared space changed. Compatibility measures what survived. Policy decides what you may do about it.
A derived representation receives a new identity. PCA projection or truncation, whitening, persisted mean-centering, a learned bridge’s output, a Matryoshka prefix — each of these changes the representation contract, so each gets its own space_identity (with post_processing naming the derivation) and its own space_hash, carrying a link back to where it came from:
derived_from:
parent_space_hash: <the space this was built from>
derivation_artifact_hash: <hash of the transformation/parameters that produced it>
A derived space is not automatically incompatible with its parent for every purpose — a mild whitening transform might preserve neighbor structure well, a lossy PCA truncation might not, and neither fact is known until measured. What a derived identity does mean is that nothing calibrated, evaluated, or validated under the parent’s hash carries over by assumption. A threshold fit on the parent’s score distribution, an evaluation observation scored against the parent’s vectors, a preservation claim measured under the parent’s geometry — none of these silently apply to the derived space. They are revalidated under the derived representation’s own hash, exactly as they would be for any other new identity.
What changes, and what does not, when a space changes
Two artifacts from Chapters 13 and 14 are easy to over-invalidate on a space change, and getting this wrong either wastes work or silently reuses stale evidence.
Chapter 13 separated an evaluation card — the task contract: candidate corpus, query workload, relevance definition, metric, aggregation, model-use protocol — from an evaluation observation — the scores one specific representation produced under that card. A space change does not, by itself, require redefining the evaluation card. The whole point of holding the card fixed across a model upgrade is to make v1 and v2 comparable at all: run the same contract against the new space, and bind the resulting scores to a new observation keyed by the new space_hash. What changes is the observation, produced fresh for each identity; the card can — and usually should — stay exactly as it was, precisely so the comparison is fair.
A calibration record, by contrast, is tied directly to one space’s score distribution — its task-defined positive/negative populations, operating rule, and measured error rates (Chapter 14). A new space_hash therefore invalidates carryover by assumption: the old threshold cannot simply be treated as calibrated for the new identity. But identity change alone does not prove that the numerical operating point moved. Re-run the calibration protocol on the new space; if the old threshold still satisfies the required operating objective, record that validation under the new hash, and if it does not, derive a new operating point.
identity changes
→ evaluation card: usually unchanged (same task contract, for a fair comparison)
→ evaluation observation: rerun, bound to the new space_hash
→ calibration: revalidate under the new space_hash; recalibrate/refit if required
The upgrade decision tree
The trigger for a migration decision is not a version label — this chapter has just spent its opening argument establishing why a label cannot be trusted for that. The trigger is a hash comparison.
flowchart TD
D["new embedding deployment"] --> H["compute / resolve space_hash"]
H --> C{"matches the space_hash already in the index?"}
C -->|yes| S["same declared identity — normal within-space operations continue"]
C -->|no| M["identity changed — a migration decision is required"]
M --> Q1{"want one consistent geometry now?"}
Q1 -->|yes| RE["re-embed: move the corpus into the new space — compute cost + migration window"]
Q1 -->|"not yet"| CO["partitioned coexistence: old and new vectors stay separated by space_hash"]
CO --> OR["query each space through its own valid query representation; retrieve within each space; combine ranked outcomes through a separately validated fusion policy"]
RE --> TH["rerun evaluation observations and recalibrate score-dependent decisions under the new space_hash"]
OR --> TH
M --> BR["bridge: only once Part VI validates a measured translation for the specific operation required"]
A release label can change without the underlying transformation changing at all; a transformation can change without any human updating the label. The hash comparison removes dependence on that mutable label provided the runtime’s identity material is complete and correctly resolved. It is the enforcement point for the declared transformation contract, not a magical detector of dependencies the contract forgot to record.
Read the two live branches as a cost choice, not a correctness debate — neither is “the default correct answer,” each has a distinct cost profile:
Re-embedding buys one coordinate system, one index, and simple calibration and operations, at the price of embedding compute, migration time, and a rebuild window.
Partitioned coexistence avoids that up-front cost by keeping old and new vectors in genuinely separate indexes, each queried by its own valid query representation. If the system needs a combined result, the combination happens at the level of ranked outcomes or task scores — not by placing raw vectors from two identities into one geometry. That combination is itself a policy decision requiring its own evidence: reciprocal-rank fusion, a validated common-utility score, or another list-level method are all conceptually available, and none of them is validated by anything measured in this chapter. Partitioned coexistence is a real, legitimate posture — not a stopgap that “cannot” scale — but the fusion step at its far end is an open design question this chapter does not close.
A bridge — translating a vector produced in one space into the other — is a third mechanism, and Part VI’s entire subject. This chapter does not pre-authorize it; a bridge earns its way into the diagram only once it has been measured for the specific operation it will be used for.
Backward compatible in what sense?
A vendor’s claim that a new model version is “backward compatible” can mean several genuinely different things: the same API surface, the same output dimension, the same tokenizer interface, an intentionally aligned coordinate system, or comparable task performance. Those are not interchangeable claims, and none of them can be assumed from the phrase alone. The useful response is a question, not a rejection: backward compatible in what sense, for what operation?
If the operation in question is direct raw-vector mixing or cross-version search, the phrase is insufficiently specified until it names an explicit coordinate-correspondence guarantee. Chapter 16’s structural measurements — neighborhood overlap and CKA — are useful evidence about the pair, but they are not sufficient authorization for raw coordinate operations, because high structural agreement can coexist with an unknown basis correspondence. Validate the intended cross-space operation directly, or use a bridge whose preservation profile covers that operation. If the claim is instead equivalent task performance, the right instrument is Chapter 13’s evaluation contract, run against both versions and compared as outcomes.
Compatibility, as Chapter 16 already established, is property-specific rather than a single scalar, and that discipline carries forward unchanged here. A measured neighborhood overlap of 0.5 between two spaces is evidence against compatibility for operations that require neighborhood preservation — it is not a universal verdict that the pair is “incompatible,” full stop, and a high overlap is equally not a blanket certificate of compatibility for every operation someone might want to perform. Every compatibility claim in this book names a property, a scope, and a measurement:
compatible_for:
neighborhood_structure: <measured / not measured>
top1_decisions: <measured / not measured>
threshold_transfer: <measured / not measured>
bridge_retrieval: <measured / not measured>
relation_preservation: <measured / not measured>
compatible: true as an unscoped field has no place in this architecture. Chapter 20’s usable_for(scope, operation) completes this pattern with a policy layer; this chapter’s job is only to make sure the evidence beneath that policy is never collapsed into one Boolean before it gets there.
Demonstration: an unsupported mixed-index probe
MEASURED on RELATE v0.1, Wave 3 row 3.2 — artifact
experiments/embeddings-from-first-principles/wave3/artifacts/mixed-index-penalty-curve.json.
The scenario this experiment probes is exactly the opening story’s failure: a target-space query, scored directly against an index that silently mixes target-space and foreign-space document vectors. Row 3.2 runs this as a deliberate probe of an operation Chapter 16 already ruled unauthorized by default, to see what the damage looks like when it happens rather than to validate any migration scheme.
The pair, and why it is a proxy. RELATE v0.1’s locally available models offer only one pair sharing a dimension: bge-large (1024-d) and mxbai-large (1024-d) — different models from different creators, both described by the harness as retrieval-tuned. The two other model pairs in this book’s comparison set differ in dimension and are mechanically unable to share one raw index, so the implementation skips them for this row. The artifact labels the pair bge-large(v1) + mxbai-large(v2), but that labeling is a role assignment for the experiment, not a claim about the models’ history — bge-large and mxbai-large are not two versions of one model, and no genuine v1→v2 pair was available to measure. Chapter 16 measured this same pair’s neighborhood overlap at 0.879 and linear CKA at 0.987 — high structural agreement under those two specific measurements. That is not evidence of coordinate alignment: Chapter 16 was explicit that CKA is invariant to exactly the kind of rotation that would make raw coordinate comparison meaningless, so “structurally similar under Chapter 16’s measurements” is the correct description of this pair, not “aligned.”
The exact construction. Item embeddings from both models are L2-normalized independently. With RNG seed 17, exactly half of the corpus item rows are selected as “legacy.” A mixed index is built by taking the bge-large vector for every legacy row and the mxbai-large vector for every other row. Every query is embedded once, in mxbai-large — the target space — and scored by raw dot product against every row of the mixed index, legacy rows included. That is the unsupported operation itself: a target-space query, compared directly against foreign-space coordinates, with nothing in the scoring step aware that roughly half the candidates were produced under a different declared identity.
What was measured. Ranking quality is scored with nDCG@10 only — the implementation computes and persists nDCG@10 exclusively; no Recall@10 is measured or reported by this row, despite what the artifact’s own top-level metric description string says.
clean, all-target (mxbai-large) index nDCG@10 = 0.9270
naive mixed index (half bge-large, seed 17) nDCG@10 = 0.9193
naive penalty = 0.0077
The attempted “calibrated merge.” The implementation also computes a second, adjusted score, stored under the field calibrated_merge_ndcg10. Read exactly what it does: for every legacy row, it subtracts an offset from that row’s score before ranking, where
offset = mean_i( V_target[i] · V_target[i] ) − mean_i( V_source[i] · V_source[i] )
— the difference between each space’s own mean self-similarity. But both V_target and V_source were L2-normalized before this computation ran, and for a unit-length vector v, v · v ≈ 1 by construction. So offset ≈ 1 − 1 ≈ 0, and subtracting approximately zero from every legacy score leaves the ranking essentially unchanged from the naive mix:
calibrated_merge_ndcg10 = 0.9193 (identical to the naive mixed result)
residual_penalty_after_merge = 0.0077 (identical to the naive penalty)
The field name calibrated_merge_ndcg10 is the artifact’s historical label, worth preserving for provenance — but read literally, it overstates what happened. The attempted offset correction is degenerate under L2 normalization and performs no meaningful calibration. It is not evidence that calibration was unnecessary because the spaces were already well aligned, not evidence that the merge strategy is validated, and not evidence that per-space score calibration is generally ineffective for migration. It is evidence of exactly one narrow thing: this particular offset, applied to unit-normalized vectors, collapses to a no-op. This is precisely the discipline the rest of the book applies everywhere else — read the implementation behind a label, not the label itself, before trusting what it claims to measure.
Reading the result correctly
State only what this one measured case supports:
On one RELATE v0.1 half-corpus partition (seed 17), scoring
mxbai-largequeries directly against a mixed index whose legacy half carries rawbge-largevectors lowered nDCG@10 from0.9270to0.9193— an absolute difference of0.0077. This is one measured cross-model proxy pair, structurally similar under Chapter 16’s measurements, on one random partition of one synthetic corpus.
Everything the current draft was tempted to conclude beyond that sentence goes further than the evidence supports, and each overreach is worth naming so it does not creep back in.
This is not a lower bound. A single observed penalty, on one pair and one partition, establishes nothing about what a more divergent pair would produce. A pair with lower structural agreement could show a larger penalty, a similar one, a smaller one, or a different failure mode altogether — none of that has been measured. Describing 0.0077 as “the minimum silent cost to budget” claims a mathematical property this single data point does not have.
This is not a penalty curve. The artifact is named mixed-index-penalty-curve.json, and that name is aspirational rather than descriptive of its contents — the two dimension-mismatched pairs are skipped entirely, leaving exactly one numeric observation. One point is not a curve, and the filename should not be allowed to imply otherwise. Read the implementation, not the label.
No causal link between overlap and penalty was measured. It is tempting to read 0.879 overlap and 0.0077 penalty side by side and conclude that the small penalty exists because the overlap is high. Row 3.2 varies nothing — it does not hold other factors fixed while moving structural agreement up or down. The hypothesis that mixed-index harm grows as spaces diverge is reasonable and worth stating explicitly as a hypothesis; it is not something this experiment establishes:
HYPOTHESIS — NOT MEASURED BY ROW 3.2. Mixed-index penalty may grow as the two spaces’ structural agreement falls. Testing this would require several pairs (or controlled transformations) spanning a range of measured divergence — row 3.2 supplies exactly one pair.
No all-source baseline exists. The implementation computes a clean all-target run, a naive mixed run, and an adjusted mixed run. It does not compute an all-bge-large baseline, so no such row belongs in any table drawn from this artifact.
The measured value may depend on the specific partition. Seed 17 selects one particular half of the corpus as legacy. Which items happen to land on which side of that split — and in particular, whether relevant items for a given query happen to fall into the legacy half — could plausibly shift the measured penalty. No repeated-seed analysis, distribution, or confidence interval exists; 0.0077 is a point estimate for this one partition, not a stable property of the pair.
The experiment does not isolate why the penalty is small. A modest overall nDCG drop is consistent with several different underlying mechanics — most relevant items might remain on the target side of the split, foreign-space candidates might rarely surface near the top of the mixed ranking, or this specific partition might simply be favorable. Row 3.2 reports the aggregate outcome only; it does not decompose which of these is actually happening, so 0.0077 should not be read as evidence that raw cross-space scores are “almost valid” in general.
What the experiment does establish is narrower and, read correctly, more useful than any of the overreaches above: an unsupported cross-space operation, run on a structurally similar pair, did not collapse retrieval quality or produce an obviously broken aggregate result. It still returned a plausible ranking while violating the representation contract. That is not reassurance — it is the precise danger the opening scenario described. A modest aggregate degradation is compatible with a serious provenance error, so identity should be enforced as a runtime invariant rather than left for a downstream quality metric to diagnose. Row 3.2 does not measure how detectable a 0.0077 change would be in any production monitoring system, so the chapter should not claim that it would or would not be noticed there.
Lab 17: reproduce the mixed-index probe, and its exact boundary
MEASURED — artifact
experiments/embeddings-from-first-principles/wave3/artifacts/mixed-index-penalty-curve.json. REPRODUCIBLE —python run_wave3.py 3.2.
Question. What does an unsupported raw cross-space comparison actually cost, on one measured case — and how much of that answer generalizes?
Step 1 — compute both space identities. Construct the identity-material contract for bge-large and for mxbai-large per this chapter’s schema, and hash each. Record space_hash_A and space_hash_B before running anything, so the rest of the lab is explicit about which two declared identities are being mixed.
Step 2 — build the clean target baseline. L2-normalize all mxbai-large item embeddings. Score every release query, embedded in mxbai-large, against the full clean index. Measured: nDCG@10 = 0.9270.
Step 3 — build the naive mixed index. L2-normalize bge-large item embeddings as well. With RNG seed 17, select exactly half the corpus item rows as legacy. Build the mixed index by taking the bge-large vector for legacy rows and the mxbai-large vector for the rest. Score the same queries, still embedded in mxbai-large, against this mixed index. Measured: nDCG@10 = 0.9193, a naive penalty of 0.0077.
Step 4 — inspect the attempted adjustment. Compute each space’s mean self-similarity — mean(V · V) over all item rows, per space. Because both spaces were L2-normalized in Step 2 and Step 3, both means come out to approximately 1, so their difference (the “calibration” offset) is approximately 0. Apply it anyway to the legacy scores and recompute nDCG@10: measured 0.9193, identical to the naive result. State the conclusion exactly: no useful calibration occurred; the adjustment was a no-op under unit normalization.
Step 5 — state the exact boundary of what this establishes. One pair (a structurally similar cross-model proxy, not a genuine version pair). One partition (seed 17, no repeats). One synthetic corpus. No Recall@10. No all-source baseline. No penalty curve. No lower bound. No validated merge or fusion strategy.
Step 6 (PROPOSED — no artifact backs this) — partition sensitivity. Repeat the half-corpus split over many random seeds; report the resulting penalty’s mean and spread. No result currently exists.
Step 7 (PROPOSED — no artifact backs this) — a genuine version pair. Repeat the entire measurement on real v1 and v2 releases of one model, rather than a same-dimension proxy. Measure the Chapter 16 space comparison for that pair alongside the mixed-index penalty, to see whether a genuinely divergent pair behaves differently from this structurally similar one. No result currently exists.
Step 8 (PROPOSED — no artifact backs this) — inspect foreign-space candidate behavior. For queries where the mixed index disagrees with the clean baseline, measure how often legacy (bge-large) candidates enter the top-k under mxbai-large queries, whether relevant legacy candidates are systematically suppressed, and whether irrelevant legacy candidates spuriously rise. This is the analysis that would isolate why the aggregate penalty is small — it has not been run.
Step 9 (PROPOSED — no artifact backs this) — a real dual-index fusion policy. Query the bge-large index with a bge-large query and the mxbai-large index with an mxbai-large query, separately and validly, then combine the two ranked lists with an explicit, evaluated fusion method — reciprocal-rank fusion, or a validated common-utility score — and compare the fused result against the clean target baseline from Step 2. No result currently exists; this is the corrected version of the “migration mechanism” the naive mix and the degenerate offset both failed to provide.
Try it yourself
Compute
space_hashfor two model versions or two models you might swap between. Run the clean-target and naive-mixed configurations from Steps 2–3 on your own corpus and query workload. Record Chapter 16’s structural measurements alongside the penalty, but do not predict the penalty from them: row 3.2 contains only one pair and establishes no mapping from overlap or CKA to mixed-index harm. Compare your measured result with this chapter’s0.0077only after both experiments are explicit about their pair, partition, workload, and metric.
Companion component: the space registry
The registry is where identity, compatibility evidence, and migration state actually live at runtime — the enforcement layer that makes the three-way separation real rather than aspirational.
space_registry:
spaces:
<space_hash>:
identity: space_identity # Chapter 17
parent_space_hash: <or none, if native>
derivation_ref: <or none, if native>
indexes: [index_ref, ...]
calibration_records: [calibration_record_ref, ...] # Ch14, per space_hash
evaluation_observations: [evaluation_observation_ref, ...] # Ch13, per space_hash
evaluation_cards: # Ch13 — can be shared across a version upgrade
<card_id>: evaluation_card
vectors_require: space_hash # enforced at write time
raw_coordinate_policy:
if query.space_hash == index.space_hash:
allow
else:
require validated_bridge_for(source_hash, target_hash, operation) # Part VI
multi_space_orchestration:
# querying two indexes independently, each in its own valid space, then
# combining ranked outcomes — permitted without a bridge; the fusion step
# is its own policy and needs its own evaluation
allow
bridges:
<source_hash, target_hash, bridge_id>:
preservation_profile_ref: <Part VI>
allowed_operations: [ ... ]
migration:
source_space_hash:
target_space_hash:
mode: <reembed | partitioned | bridged>
coverage: <fraction of corpus in target space>
backfill_status:
cutover_criterion_ref: <a policy decision, not an identity fact>
Two distinctions in this schema do real work and are easy to blur in a simpler version. First, raw_coordinate_policy and multi_space_orchestration are not the same gate. What the registry denies by default is a raw coordinate operation crossing a hash boundary — comparing a vector from one space directly against vectors from another, averaging or concatenating across identities, inserting one space’s vectors into another’s raw index. Querying two properly separated indexes independently, each with its own valid query representation, and combining the two ranked outcomes afterward is not a raw coordinate operation at all — it never puts two identities’ coordinates in the same geometric comparison — and the registry permits it without requiring a bridge. The fusion step that combines the two outcome lists is its own policy decision, and it needs its own evaluation before being trusted, but the orchestration itself is not the thing this chapter forbids.
Second, evaluation cards live outside the per-space bucket, while evaluation observations and calibration records live inside it. A card can be written once and reused across a version upgrade specifically so v1 and v2 can be compared fairly under one fixed contract; the observations and calibration records it produces are bound to whichever space_hash generated them and are never shared across identities.
A migration.cutover_criterion_ref deliberately points to a policy object rather than embedding a number like “cut over at 95% coverage” directly into this schema. What fraction of coverage is sufficient to cut over is an application decision, informed by risk tolerance and the cost of carrying a dual-write window — not a fact the identity layer is positioned to assert.
Embedding Observatory behavior
By the end of this chapter the Observatory should be able to answer: What space_hash produced this vector, and what declared transformation does that hash identify? Is this a derived space, and if so, what is its parent? Does a given index contain vectors from more than one hash? Is a query about to be compared directly against vectors carrying a different hash, and if so, is there a validated bridge for that specific operation, or is this an unsupported comparison in progress? Which calibration record and which evaluation observations belong to this exact space? What migration is currently in progress, and what fraction of the corpus has reached the target space? Is a bridge registered for a given source/target pair and operation, and has the property that operation depends on actually been measured, or only asserted?
The Observatory’s job is to surface this evidence, not to infer a policy from it. It should not decide that a neighborhood overlap of 0.9 makes mixing “safe,” and it should not decide that a measured penalty of 0.008 means “migrate now” or “migrate never.” Those are scoped policy calls belonging to Chapter 20’s usability layer, built on top of what the Observatory reports — never computed by the Observatory itself as an inferred threshold.
The Observatory should catch an identity-contract violation before a quality metric ever has to reveal it. Row 3.2 showed why that ordering matters: an unsupported cross-space comparison changed nDCG@10 from
0.9270to0.9193on one structurally similar proxy pair without producing an obviously catastrophic aggregate result. The experiment says nothing about a production system’s natural metric variance or alert sensitivity; it shows instead that identity is directly observable at write/query time and therefore should not be inferred indirectly from later quality movement.
Failure modes
- Untagged vectors. A vector without a recorded
space_hashcannot be safely routed for any raw operation — there is no identity to check it against. - Trusting a mutable model label as identity. A release tag can move onto new bytes without changing; the underlying transformation can change without the tag moving. Only the identity-material hash tracks the actual pipeline.
- Overinterpreting hash equality. A matching hash means the same declared transformation contract — not bitwise-identical output across machines, and not a claim about task quality.
- Overinterpreting hash inequality. A different hash withdraws the assumption of compatibility; it does not itself prove every property is incompatible. Measure before concluding either way.
- One global
compatibleBoolean. Compatibility is property- and operation-specific — Chapter 16’s lesson, still true here. Name the property. - Mixed raw-vector index. Queries and documents crossing an identity boundary silently, exactly as row 3.2 deliberately probed — no type error, a quietly degraded metric.
- Rebuilding the evaluation card unnecessarily on a model change. The task contract (Chapter 13) can stay fixed across an upgrade specifically to keep v1 and v2 comparable; only the observation needs to be rerun.
- Reusing a calibration threshold across a space change by assumption. A threshold belongs to a Chapter 14 calibration contract; a new
space_hashmeans the old calibration no longer transfers automatically. Revalidate the operating point on the new space, and recalibrate if the required error/cost target is no longer met. - Trusting the
calibrated_merge_ndcg10field by name alone. Its offset collapses to approximately zero under L2 normalization; read the implementation, not the label, before trusting what a field claims to measure. - Calling one mixed-index observation a lower bound or a penalty curve. One pair, one seed, one partition — a single point, not a bound and not a curve.
- Fusing two separately queried indexes by raw score without validating that the scores share a scale. Two calibration thresholds do not, by themselves, put two similarity distributions on one common cardinal scale; rank-level fusion avoids that problem but still needs its own evaluation.
- Assuming every cross-space interaction requires a bridge. Multi-space orchestration — querying each space correctly and combining outcomes afterward — needs no bridge at all; only raw coordinate operations do.
What this chapter established
- A vector is not just
float[d]; operationally it isfloat[d]plus the identity of the transformation that produced it. Chapter 1’s descriptivespace_recordmatures here into an exactspace_identity— hashed identity material, unhashed descriptive metadata, and aspace_hashcomputed only from the former. - A matching
space_hashmeans the same declared identity — not bit-for-bit reproducibility, and never by itself a claim about task quality. A different one supports only the narrow inference do not assume coordinate compatibility, not the stronger claim that the spaces are incompatible. Identity, compatibility, and usability remain three separate layers, and none substitutes for another. - A derived representation — PCA, whitening, a bridge’s output, a Matryoshka prefix — receives its own identity and lineage back to its parent, and nothing calibrated or evaluated under the parent carries over by assumption. Chapter 13’s evaluation card can stay fixed to keep an upgrade comparison fair; the observation must be rerun and the operating point revalidated under the new hash.
- Migration triggers from a hash comparison, not a version label, because a label and the underlying transformation drift independently of each other. Re-embedding and partitioned coexistence are two legitimate postures with different cost profiles; a bridge is a third that Part VI has to earn through measurement.
- Row 3.2’s mixed-index probe cost
0.0077nDCG@10 on one pair and one partition — one observation, not a lower bound and not a curve — and itscalibrated_mergefield is an offset that collapses to zero under L2 normalization. An unsupported cross-space operation that still looks almost fine is precisely why the registry denies raw coordinate operations across a hash boundary by default, while still permitting multi-space orchestration that never mixes coordinates at all.
Next
Identity tells us why two spaces’ vectors cannot simply be mixed. Migration tells us why a team might want another option besides a full re-embed. Row 3.2 showed that the unsupported shortcut — raw cross-space scoring — can look almost fine on one measured case, which is exactly why it cannot be allowed to happen by assumption. Part VI now asks the constructive question this chapter has deliberately left open: can a translation between two spaces be explicitly learned and measured, and what would it have to preserve before the runtime may treat it as authorization for a specific operation?
T(E_A(x)) ≈ E_B(x)
The next chapter starts from that ≈, and from everything this chapter has just established about what “close enough” is not permitted to mean by default.