What Should a Translation Preserve?
A translation objective is incomplete until it names the geometry it is trying to preserve. Hold architecture, data, and seed fixed, sweep the weight on a source pairwise-cosine regularizer, and watch most target-facing Wave 6 metrics decline while cluster ARI refuses to cooperate — a controlled demonstration that stronger pressure toward one invariant can work against properties the destination consumer needs.
Part VI — Crossing Embedding Spaces · Whose geometry counts as success?
“Preserve the geometry” is not a complete instruction
A bridge translates vectors from space A into space B. The natural instinct is to say:
Preserve the source geometry while you translate.
That sounds obviously correct. If two source points are close, keep them close. If two source points are far apart, keep them far apart. Preserve pairwise cosine, distances, neighborhoods.
But Chapter 16 established something uncomfortable: different encoders do not merely rotate the same universe. They can disagree on neighborhoods, density, rank order, hard distinctions, and calibration — real structural agreement can coexist with genuine, measured disagreement about local structure. Chapter 22 then supplied a concrete instance of the consequence: a bridge trained to reconstruct paired target points achieved perfect top-10 counterpart recovery and strong but non-perfect top-1 recovery while reproducing less of the surrounding target neighborhood under the measured agreement and local-order diagnostics. If A and B genuinely organize the same objects differently, an instruction to preserve A’s internal shape can actively work against getting the translated vectors to behave like B.
So the real question is not:
How do we preserve the geometry?
It is:
Which geometry should the translation preserve, and which downstream behavior is the preservation claim supposed to protect?
Two native spaces, one derived representation, and four relational objects
For paired objects x_i, define exactly what Chapter 18 already established:
A_i = E_A(x_i) native source representation
B_i = E_B(x_i) native target representation
T_i = T(A_i) derived bridge-output representation
T_i may share the target model’s dimensionality, and a training objective can push it to behave numerically close to B_i. It remains, by provenance, a derived representation — not native target identity, per Chapters 17 and 18. A bridge record still carries source_space_hash, target_space_hash, and derived_output_space_hash as three separate identities, and nothing this chapter measures collapses that distinction.
That leaves not two relational structures to compare, but at least four, and the difference between them is where most confusion about “preserving geometry” actually lives:
S_A(i,j) = cos(A_i, A_j) source-native internal geometry
S_B(i,j) = cos(B_i, B_j) target-native internal geometry
S_T(i,j) = cos(T_i, T_j) derived-output internal geometry — relationships among translated points
S_TB(i,j) = cos(T_i, B_j) derived-query -> native-target behavior — how a translated point ranks native items
The first three are internal comparisons within one representation. The fourth is a cross-representation comparison, and it matters because it is exactly what several of Chapter 22’s structural-fidelity metrics actually measure: agreement_at_10 and order_preservation_local compare how T_i relates to native target items B_j, not how T_i relates to other translated points T_j. A training objective that pulls S_T toward S_A is answering a different question than one that pulls S_TB toward S_B, and a consumer who cares about one will not automatically be served by a loss aimed at the other. Keeping these four objects straight is the technical backbone of everything that follows.
A translation objective can ask for different equalities among them:
source-isometry: S_T ≈ S_A — internal derived-output geometry matches internal source geometry
point reconstruction: T_i ≈ B_i — each translated item lands where its own paired target sits
target behavior: S_TB ≈ S_B — translated points rank native target items the way native queries would
task fidelity: D(T_i) ≈ desired_task_behavior(x_i) — the derived representation supports the consumer's required decision
These objectives can agree. Illustrated informally, they can also fight each other.
An illustrative conflict
ILLUSTRATIVE — this example is constructed to show the mechanism, not drawn from a persisted artifact.
Suppose source encoder A places, for some object x:
A-space neighborhood of x
1. paraphrase
2. same topic
3. negation
while target encoder B places:
B-space neighborhood of x
1. paraphrase
2. entailment
3. same topic
Negation has dropped out of B’s top three. Now impose a loss insisting the translated output preserve A’s pairwise cosine relationships exactly — negation stays close, because it was close in A. But reproducing B’s behavior over the same candidate set requires moving it further away. Both constraints cannot be satisfied at once, not because of an optimization bug, but because the two encoders genuinely disagree about which items belong near x. This is a conflict between invariants, and it is exactly the kind of disagreement the measured experiment below turns into a controlled intervention rather than a hypothetical.
Four preservation objectives, named precisely
1. Preserve the source shape. A source-geometry loss compares internal pairwise relationships before and after translation:
L_source = mean((S_A(i,j) - S_T(i,j))^2) over sampled pairs (i, j)
This is not “preserve source geometry” in every sense — it is specifically source pairwise-cosine preservation, one particular property among several a source representation carries (local neighbor membership, local order, cluster structure, relation contrasts, and calibration behavior are others, and each would need its own objective or evaluation contract). As a finite-weight regularizer, it penalizes changes in source pairwise angles; it does not guarantee that those angles remain unchanged. That is a well-matched constraint when A and B are believed to represent the same underlying geometry in different coordinates — an orthogonal basis change, for instance. It can be a poor assumption whenever the two encoders’ native neighborhoods genuinely disagree.
2. Reconstruct the target points. With paired data:
L_point ≈ ‖T_i − B_i‖² + (1 − cos(T_i, B_i))
— the actual measured construction below, computed on L2-normalized vectors. This says: put each translated item where B put the same item. It directly supports counterpart recovery. It does not, on its own, guarantee that neighborhoods, local order, or clusters come along for free — Chapter 22’s headline result is the concrete demonstration.
3. Reproduce target-index behavior. Instead of matching S_A, use the target encoder’s ranking behavior as the teacher:
S_B(i,j) = cos(B_i, B_j)
S_TB(i,j) = cos(T_i, B_j)
p_j = softmax(S_B(i,j) / τ)
q_j = softmax(S_TB(i,j) / τ)
L_target_rank = KL(p || q)
This asks the translated point to rank the same native target candidates the way a native query would — it models S_TB, the derived-query-to-native-target relationship, not the internal S_T geometry among translated points. That distinction matters concretely: if the deployed use is searching an existing native B index, S_TB is the relevant behavior. If the deployed use instead replaces an entire B corpus with translated A vectors — so that translated points get compared against other translated points, not native ones — the relevant behavior shifts to S_T ≈ S_B, a different, unmeasured objective this chapter does not build. The consumer’s actual data path decides which of these two is the real target.
4. Preserve the downstream task. Sometimes neither internal geometry nor cross-representation ranking is the real contract. For retrieval, L_task might be a ranking loss against relevance labels; for clustering, a cluster-consistency objective; for a calibrated threshold, operating-point or FAR/FRR preservation. This says: preserve what the consumer actually uses, whatever either native geometry happens to look like.
Demonstration: turning an invariant into an intervention
MEASURED on the Wave 6 cross-space benchmark —
mxbai-embed-large-v1→Qwen3-Embedding-8B, the same frozen 2,336/589 deterministic item-ID split used in Chapter 22. Artifactswave6/artifacts/cross-space-benchmark.jsonandwave6/artifacts/vsp-sweep.json. Supervision is paired throughout — the correct correspondence is supplied, so the experiment isolates what the loss does, not whether correspondence can be discovered.
First, remove a capacity confound. A narrow-hidden MLP could not be trusted as a fair platform for a loss intervention if it were simply too small to fit a good map in the first place. Compare a paired two-hidden-layer GELU network (Linear(d_source, hidden) → GELU → Linear(hidden, hidden) → GELU → Linear(hidden, d_target), output L2-normalized, trained 400 epochs with Adam, lr=1e-3, weight_decay=1e-5, seed 0) at two hidden widths against the paired ridge baseline from Chapter 22:
cosine R@1 agreement@10 local order cluster ARI
paired ridge 0.8837 0.8676 0.7964 0.6987 0.4750
paired MLP, hidden 256 0.8527 0.7742 0.7947 0.6629 0.4142
paired MLP, hidden 1024 0.8959 0.8676 0.8324 0.7561 0.4619
At hidden width 256, the network trails the ridge baseline on every displayed column, although agreement@10 is very close (0.7947 vs 0.7964). At hidden width 1024 — still far below the target’s 4096-d width, but matching the source’s own 1024-d width — it exceeds ridge on cosine, matches the ridge point estimate on Recall@1 (0.8676), and exceeds it on agreement@10 and local order, with a slightly lower cluster ARI. This does not eliminate every possible capacity question, and no uncertainty analysis establishes statistical equivalence between any of these point estimates. What it does establish is narrower and sufficient for this chapter’s purpose: the hidden-1024 network is a reasonable fixed-capacity platform for testing what a loss term does, since it is not obviously capacity-starved relative to the linear baseline on the headline point/neighborhood/order measures.
Now hold everything fixed and change exactly one thing. Same source, same target, same corpus, same split, same architecture, same optimizer, same 400 epochs, same seed 0. Add a source pairwise-cosine vector-space-preservation (VSP) term. Concretely, at every training step, with weight vsp > 0, the implementation samples 4,096 random index pairs (a, b) from the training set and adds
vsp × mean( (cos(T_a, T_b) − cos(A_a, A_b))² )
to the base point-reconstruction loss — an explicit, sampled penalty pulling the derived-output internal geometry S_T toward the source-native internal geometry S_A, exactly the “preserve the source shape” objective named above. Sweep the coefficient vsp over {0, 1, 3, 10, 30, 100}, changing nothing else:
vsp weight cosine R@1 R@10 MRR agree@10 local order triplet cluster ARI
0 0.8959 0.8676 1.0000 0.9262 0.8324 0.7561 0.9267 0.4619
1 0.8929 0.8625 1.0000 0.9239 0.8287 0.7517 0.9262 0.4202
3 0.8793 0.8387 1.0000 0.9060 0.8121 0.7252 0.9225 0.5635
10 0.8581 0.8302 1.0000 0.9002 0.7917 0.6949 0.9037 0.5373
30 0.8029 0.7742 0.9983 0.8667 0.7355 0.5845 0.8861 0.4577
100 0.7033 0.5756 0.9847 0.7341 0.6421 0.3976 0.8615 0.4313
Read the trend precisely, because it is not what a first glance suggests. Six of the eight recorded fields — cosine_to_target, recall@1, mrr, agreement@10, local order, and triplet agreement — decline at every successive weight increment, strictly and without exception, across all six sweep points. recall@10 is non-increasing and stays perfectly saturated at 1.0000 through weight 10, only beginning to fall at weight 30 and 100. That is already seven of eight fields consistent with “more source-cosine pressure, worse target-facing behavior.”
The eighth field breaks the pattern outright. cluster_preservation_ari moves 0.4619 → 0.4202 → 0.5635 → 0.5373 → 0.4577 → 0.4313 — it is genuinely non-monotonic, and its highest value in the entire sweep occurs at weight 3, not at weight 0. Do not describe this sweep as producing uniform, monotonic degradation across every measured target-side property — that claim is false on its face for this one field, and the exception is worth taking exactly as seriously as the six-field trend it interrupts. This is a direct continuation of Chapter 22’s lesson that cluster ARI did not discriminate its high-structural-agreement control from the headline pair either: the VSP sweep reinforces the warning that cluster ARI is not moving in lockstep with local neighborhood and order fidelity in this experiment. No sensitivity analysis — repeated KMeans seeds, a different k, a bootstrap over test items — exists to diagnose why; the honest statement stops at “non-monotonic in this one run,” not “improved by VSP” and not “noisy.”
One epistemic boundary matters more than any single number here. The artifact records the coefficient on the source-cosine term at each sweep point — the pressure applied — not a directly measured, held-out source-fidelity score. Nowhere in this experiment is S_A versus S_T correlation, or source-neighborhood overlap between A_i and T_i, computed and persisted on the held-out test items. So the defensible statement is about pressure and consequence, not about two measured fidelities trading off against each other point for point:
Increasing the coefficient on a source pairwise-cosine preservation term systematically worsens most measured target-facing properties on this pair, under this fixed training recipe. It does not, by itself, demonstrate a measured one-dimensional trade between an achieved source-fidelity score and an achieved target-fidelity score, because achieved source fidelity was never independently measured on held-out data.
Nor does the response scale in any simple proportion to the weight — the jump from weight 30 to 100 costs far more on cosine_to_target (0.8029 → 0.7033) than the jump from 10 to 30 (0.8581 → 0.8029), and recall@10’s response is essentially flat until weight 30 before dropping. Avoid describing the degradation as proportional to the coefficient; it is nonlinear and metric-specific.
Even the memorable extremes are worth stating exactly rather than dramatically. At weight 100, cosine_to_target falls from 0.8959 to 0.7033, recall@1 from 0.8676 to 0.5756, agreement@10 from 0.8324 to 0.6421, and local order from 0.7561 to 0.3976 — substantial degradation on several properties, not a collapse: recall@10 is still 0.9847, meaning the paired target remains in the top ten for the overwhelming majority of held-out objects even at the sweep’s most aggressive weight. This is Chapter 22’s lesson again, now inside a training-objective experiment rather than a static comparison: coarse counterpart recovery can remain remarkably robust while finer point ranking and structural fidelity deteriorate sharply under the same intervention.
A preservation loss can apply increasing pressure toward the wrong invariant — and make the properties the consumer needs worse.
What caused it, and what did not
Because this sweep holds architecture, data, split, optimizer, and seed fixed and varies only the VSP coefficient, it is fair to call the coefficient an experimental intervention and to describe its effect causally, within the narrow scope actually tested:
Within this one fixed Wave 6 training recipe, on this one source/target pair, increasing the coefficient on the source-cosine preservation term causes a systematic deterioration in most measured target-side point, counterpart-recovery, neighborhood, and order metrics.
Two things this experiment does not license, both worth stating plainly because they are easy to reach for. First, the native linear CKA between mxbai and Qwen on this benchmark is 0.8707 — real, measured context confirming the two encoders share substantial global structure without being identical. It is context, not an explanation: only one source/target pair received the VSP sweep, so nothing here establishes that a lower-CKA pair would necessarily pay a larger cost for the same coefficient, or that CKA bounds or predicts the size of any such trade. A future experiment sweeping the same VSP term across several pairs with different native structural agreement, to test whether the cost correlates with CKA, is a real and worthwhile extension — and it remains exactly that, an unrun proposal, not a result this chapter reports.
Second, an earlier chapter’s evidence is sometimes tempting to reach for here and should not be. Chapter 21’s Wave 3 result — the MiniLM → mpnet Procrustes bridge’s paraphrase-minus-negation cosine gap flipping from +0.033 native to −0.107 bridged — compares a bridged relation contrast against the native target’s own relation contrast. It does not measure, and was never claimed to measure, whether that bridge’s training preserved or destroyed something the source representation carried. Using it here as evidence of a source-geometry-versus-target-geometry conflict would misattribute what that number actually compares; it remains useful only as a separate, earlier reminder that pointwise paired fitting can coexist with a target-relation contrast changing sign — a related but distinct lesson, not additional evidence for this chapter’s causal claim.
Also worth restating plainly: this is one training seed and one deterministic data split, for every point on the sweep. No repeated-seed training, no bootstrap over held-out items, and no repeated corpus split exist anywhere in this evidence. Do not describe any of these differences as statistically significant, reliably present, or bounded within a known confidence interval — every number in this section is a point estimate from one controlled run per weight, and a repeated-seed sweep is a real, unmeasured extension worth proposing rather than assuming.
Representability is not discoverability
A paired bridge is told, during training, which object in A corresponds to which in B. An unpaired setting supplies only two point clouds and must recover correspondence as well as learn the map — a genuinely different, harder problem. Wave 6 includes a control that deliberately breaks pairing rather than solving it: the same ridge construction, fit on mxbai → Qwen training data whose target rows have been randomly permuted before fitting, with the correct correspondence restored only for test-time evaluation.
mxbai -> qwen ridge, SHUFFLED training correspondence
recall@1 0.0034
recall@10 0.0102
agreement@10 0.0421
against 0.8676, 1.0000, and 0.7964 for the correctly paired ridge bridge. This is a shuffled-correspondence control, not an “unpaired-chance floor” — that historical phrasing invites exactly the wrong reading. It measures what happens when a supervised estimator that expects true correspondences is deliberately fed a random permutation instead of one. It establishes plainly that the correct paired signal is doing substantial work for this estimator on this benchmark. It does not establish a lower bound on what a genuine unpaired alignment method could achieve — published unpaired translators such as vec2vec and mini-vec2vec (Chapter 18) are designed to learn cross-space alignment without supplied one-to-one pairs, whereas this control deliberately destroys the correspondence expected by a paired ridge estimator and makes no attempt to recover it. Conflating the two would misread a control designed to isolate one question as an answer to a different one:
Can this map family REPRESENT a useful transport, when correspondence is supplied?
-> the paired result shows yes, on this benchmark
Can a training signal DISCOVER that transport without supplied correspondence?
-> a separate, harder question this control does not test
The paired experiments in this chapter establish something narrower and still valuable: this architecture and training recipe can reach a useful mapping on this benchmark when correspondence is supplied. That is evidence the map family is representationally workable here — not a formal statement about its representational ceiling, and not evidence about whether an unpaired objective could discover a comparable map on its own.
Choosing the authority: source, target, or task
There is no universal ranking among source geometry, target geometry, and downstream task behavior. Authority comes from what the consuming system actually treats as invariant — and naming that consumer is the step that has to happen before a loss term gets written down.
Source authority fits when source relationships are themselves the contract. A coordinate migration within one known geometry — a basis change, a precision change, a storage-format change — explicitly wants an isometry. A source-calibrated consumer that remains authoritative treats the target encoder as a computational carrier, not a new semantic standard — but notice the precision this requires: a source-calibrated consumer typically depends on a specific threshold, a score distribution, or a rank order, not automatically the entire pairwise-cosine matrix. If source calibration is the real authority, the right objective is to preserve or evaluate that calibration behavior directly (Chapter 14’s discipline), not to assume L_source’s pairwise-cosine term is the correct proxy for it. And a design goal of round-trip archival transport — wanting A→B→A to recover the original relationships — is a real motivation for a source-preserving objective, but Chapter 21 already established that a round trip is its own end-to-end transformation path; a source-isometry loss does not guarantee round-trip fidelity by itself, and that fidelity still has to be measured directly, the way row 3.9 measured it, rather than assumed from the training objective’s intent.
Target authority fits when translated vectors will be consumed as if they were native target vectors. If the consumer specifically searches an existing native B index, feeds vectors into a B-trained clustering or routing system, or has calibrated behavior against native B scores, the bridge’s obligation shifts from “do not disturb A” to “behave like B where the consumer actually looks” — which may require deliberately changing A’s internal structure. Note again that “target geometry” is not one property either: target-index ranking behavior (S_TB ≈ S_B) and derived-output internal geometry among translated points (S_T ≈ S_B) are different objectives, relevant to different consuming architectures, and naming which one a given consumer needs is part of the same discipline.
Task authority fits when a downstream decision, not either native geometry, is the real contract. Suppose both A and B place a negated claim uncomfortably close to its assertion — faithfully reproducing either geometry preserves the same mistake. If the application genuinely depends on distinguishing polarity, the correct target is neither S_A nor S_B; it is the decision itself. Chapter 21’s supervised structured-negative readout showed that a learned decision function operating entirely within source space could outperform native source cosine on one held-out task — worth citing here for exactly what it demonstrated: a source-space supervised readout, not a cross-space bridge, and its lesson for this chapter is narrow — a downstream task may prefer a learned decision geometry that neither native encoder’s default cosine exposes well, which is a reason to consider task-aware supervision when a task-specific contract exists, not evidence about translation across spaces.
The correct principle is a decision rule, not a fixed hierarchy:
Use the most direct, trustworthy signal for the property the actual consumer relies on. If a task label directly represents that property, it may be the right authority. If it does not — if it is incomplete, narrow, noisy, or misaligned with the actual future use — then geometric or calibration evidence may remain the more trustworthy signal, even where task labels exist. Availability of a label does not, by itself, make it authoritative.
Choose the invariant before choosing the loss
A bridge specification should therefore record not just a fitting method, but a declared training objective and the authority behind it:
preservation_objective:
authority: <source_geometry | target_behavior | downstream_task>
property: <source_pairwise_cosine | target_rank | task_ndcg | calibration | ...>
reference: <native_source | native_target | task_labels>
consumer_requirement_ref: <what this objective is trying to serve — not invented here>
This deliberately does not include a self-declared success_bar — whether a given achieved value is “enough” is Chapter 20’s authorization question, informed by a consumer’s own predeclared requirement, never something a training objective invents on its own. Recording the objective’s authority prevents a specific, common category error: optimizing an elegant, easy-to-write geometric quantity that the downstream consumer never actually uses.
A proposed next objective — not a measured one
The next natural experiment is not a bigger or more adversarial network. It is to target the failure directly. For each paired training object i, compute native target similarities from B_i to a candidate set {B_j}; compute translated similarities from T(A_i) to the same candidates; distill the target ranking or similarity distribution (the L_target_rank construction above); keep a pointwise term so the paired target itself remains identifiable:
L = λ_point L_point
+ λ_rank L_target_rank
+ λ_task L_task # optional, only where a task contract exists
Crucially, no term here asks S_T to equal S_A — source geometry participates only when a consumer’s contract explicitly requires it. This is a proposal, not a measured result. It is not evidence that it will close the gap Chapter 22 identified, and it is not the obvious universally correct choice — it specifically targets S_TB ≈ S_B, the right objective for a consumer searching a native target index, and a different, unbuilt objective would be needed for a consumer whose real requirement is internal derived-output geometry among translated points, or for one whose requirement is task relevance rather than either geometry. Candidate-set selection and the temperature τ both matter to what this loss actually enforces, and neither has been tuned or measured here. The measured result in this chapter is the diagnosis that motivates trying it — not a claim about what it will find.
Failure modes
- Treating “preserve geometry” as self-explanatory. Name the reference (source-native, target-native, or derived-output) and the property (pairwise cosine, local order, cluster structure, calibration, task decision) before writing a loss term.
- Treating a derived bridge output as native target identity.
T_iis a derived representation with its own identity, regardless of how numerically close it sits toB_i. - Treating source pairwise cosine as all of source geometry. It is one measurable property; local order, neighbor membership, clusters, relation contrasts, and calibration are separate properties this chapter’s VSP term does not directly target, even though changing the learned map can affect them indirectly.
- Reading a VSP coefficient as a measured source-fidelity score. The artifact records applied pressure, not an achieved, held-out source-preservation value.
- Claiming every target-side metric degraded monotonically. Six of eight fields did,
recall@10was non-increasing, and cluster ARI was genuinely non-monotonic, peaking at weight3. - Calling the degradation proportional to the VSP weight. The response is nonlinear and differs by metric.
- Attributing the size of the trade to native CKA. Only one pair received the sweep; CKA
0.8707is context, not an established cause or bound. - Reusing the Wave 3 paraphrase-negation reversal as evidence about source geometry. It compares bridged and native-target relation contrasts, never a source-side quantity.
- Calling Chapter 21’s supervised readout a cross-space bridge. It operates entirely within source space; the target encoder’s vectors never participate.
- Calling the shuffled-correspondence control an unpaired-translation floor. It sabotages a supervised estimator’s training labels; it says nothing about methods built to infer correspondence.
- Assuming the paired result proves a formal representational ceiling. It shows this architecture can reach a useful map when correspondence is supplied — not the best possible map, and not a bound on discoverability.
- Assuming task labels automatically outrank geometric or calibration evidence. Authority depends on whether the label directly represents the consumer’s actual requirement, not on whether a label happens to exist.
- Calling a low VSP weight “safe” or “inert.” No consumer requirement or uncertainty analysis was declared to support that judgment; cluster ARI already changes non-monotonically across the low-weight points, so even the direction of every metric’s response is not uniform.
- Letting the training-objective record issue an authorization. It documents what a model was trained to preserve and why;
ALLOW/DENY/usable_forremain Chapter 20’s separate responsibility.
What this chapter establishes and what it does not
Establishes: source-isometry, point reconstruction, target-index behavior, and task fidelity are distinct training objectives that can conflict when the two encoders organize the same objects differently; on the measured mxbai → Qwen benchmark, adding a source pairwise-cosine (VSP) term to an otherwise-good paired point-reconstruction map, at a fixed capacity chosen by controlled comparison against a ridge baseline, produced strict declines across six of eight measured target-side fields as the VSP coefficient rose from 0 to 100, a non-increasing recall@10 that held at 1.0 through weight 10, and a non-monotonic cluster ARI that peaked at weight 3; the size and shape of this response is nonlinear, not proportional to the coefficient; a shuffled-correspondence control shows correct pairing is doing substantial work for this supervised estimator, without bounding genuine unpaired alignment; and a training objective’s authority should come from the actual consuming contract, decided before the loss is written, and recorded alongside the fitted transformation.
Does not establish: that source-geometry preservation is universally harmful — only that this specific, sampled pairwise-cosine term conflicted with most measured target-side properties on this one pair and recipe; a measured trade-off curve between achieved source fidelity and achieved target fidelity, since achieved source fidelity on held-out data was never independently measured; that native CKA causes or bounds the size of this trade; that target-rank distillation will resolve the conflict Chapter 22 identified — it is a proposed next experiment, not a result; that unpaired translation cannot work at this or larger scale; that cluster ARI’s non-monotonic behavior has been explained; or that any universal precedence exists among source, target, and task authority outside the specific consumer contract that names one of them.
Lab 23: turn the invariant into an intervention, and preserve every exception
MEASURED — artifacts
wave6/artifacts/vsp-sweep.jsonandwave6/artifacts/cross-space-benchmark.json. REPRODUCIBLE —python vsp_sweep.py(built onrun_wave6.py’sfit_mlp).
Question. Holding everything else fixed, what does adding pressure to preserve source pairwise cosine actually do to target-facing behavior — and does every measured property respond the same way?
Step 1 — freeze the benchmark and architecture. Source mxbai-embed-large-v1, target Qwen3-Embedding-8B, the same Wave 6 frozen 2,336/589 deterministic item-ID split as Chapter 22. Network: two hidden GELU layers, hidden width 1024, Adam(lr=1e-3, weight_decay=1e-5), 400 epochs, seed 0.
Step 2 — state the exact base point loss. mean‖T_i − B_i‖² + mean(1 − cos(T_i, B_i)), computed on L2-normalized predictions against L2-normalized native target vectors — normalized paired target-point reconstruction, not an unspecified generic reconstruction loss.
Step 3 — state the exact VSP term. At each training step, sample 4,096 random training-index pairs (a, b); compute cos(A_a, A_b) (source-native) and cos(T_a, T_b) (derived-output, from the current network state); add vsp × mean((cos(T_a,T_b) − cos(A_a,A_b))²) to the base loss.
Step 4 — sweep. vsp ∈ {0, 1, 3, 10, 30, 100}, nothing else changed.
Step 5 — reproduce all eight target-side fields at every weight, exactly as tabulated in the Demonstration above — not a hand-picked subset.
Step 6 — classify the observed trends precisely:
strictly decreasing at every step:
cosine_to_target, recall@1, mrr, agreement@10, local order, triplet agreement
non-increasing (flat, then falling):
recall@10 — saturated at 1.0000 through weight 10, then 0.9983, then 0.9847
non-monotonic:
cluster ARI — 0.4619, 0.4202, 0.5635 (peak), 0.5373, 0.4577, 0.4313
Step 7 — state exactly what was not measured. No held-out source pairwise-cosine fidelity (correlation or error between S_A and S_T on test items) is persisted anywhere in this artifact. Do not report or imply an achieved-source-fidelity curve.
Step 8 — interpret narrowly. Increasing source-VSP pressure trades against most measured target-side properties on this pair and training recipe. This is an objective conflict, demonstrated under controlled conditions — not a universal law, and not a one-dimensional fidelity trade-off, since only one side of that trade-off was actually measured.
Step 9 — preserve the ARI exception explicitly in any summary. Never describe this sweep’s result as “every target metric declined.”
Step 10 (PROPOSED — no artifact backs this) — a direct source-fidelity audit. On held-out test items, measure the correlation or error between S_A(i,j) and S_T(i,j), and source-neighborhood overlap between A_i’s native neighbors and T_i’s neighbors among translated points. No such result currently exists.
Step 11 (PROPOSED — no artifact backs this) — the target-rank intervention. Fit the same architecture and training budget with L_point + λ_rank L_target_rank in place of the VSP term, and compare the resulting counterpart-recovery and structural-fidelity profile against both the point-only baseline and the VSP sweep. No such result currently exists.
Step 12 (PROPOSED — no artifact backs this) — uncertainty. Repeat the sweep across several training seeds; repeat the clustering step across several KMeans seeds and k values; bootstrap the held-out test set. No such result currently exists.
Try it yourself
Before looking at the sweep table, write down the one or two metrics your own consumer would actually complain about if they moved — not “geometry” in the abstract, name the field. Then find the lowest VSP weight in the six measured points at which that specific field falls below your own predeclared bar. If your consumer’s requirement instead depends on something this sweep never measured — held-out source fidelity, calibration transfer, a specific relation contrast — say so explicitly rather than reading a proxy off this table. As a further exercise, fit the same fixed-capacity network with
L_point + L_target_rankat matched training budget and report which columns move relative to the point-only baseline; no such result currently exists in this book’s artifacts.
Companion component: the training-objective record
This record documents what a fitted transformation was trained to preserve and why — deliberately distinct from Chapter 21’s preservation observation (what was measured to have survived) and Chapter 20’s authorization trace (whether that measured result is sufficient for a named consumer). None of the three should collapse into either of the others.
bridge_training_record:
bridge_candidate_id:
identities:
source_space_hash:
target_space_hash:
derived_output_space_hash:
supervision:
mode: paired | unpaired
correspondence_ref:
model:
architecture: 2-hidden-layer GELU MLP
hidden_width: 1024
optimizer: adam(lr=1e-3, weight_decay=1e-5)
epochs: 400
seed: 0
objective_authority:
authority: source_geometry | target_behavior | downstream_task
property: source_pairwise_cosine | target_rank | task_ndcg | calibration | ...
consumer_requirement_ref: <the actual downstream need this objective is meant to serve>
loss:
point_reconstruction:
weight: 1.0
reference: native_target
formula: "‖T_i - B_i‖² + (1 - cos(T_i, B_i)), normalized"
source_vsp:
weight: <swept value>
reference: native_source
sampled_pairs_per_step: 4096
formula: "vsp * mean((cos(T_a,T_b) - cos(A_a,A_b))^2)"
evaluation_observation_refs:
counterpart_recovery: <Ch22 observation ref>
structural_fidelity: <Ch22 observation ref>
task_fidelity: <Ch21/Wave 3 observation ref, if applicable>
selection:
policy_ref: <not yet declared for this sweep>
development_metric_ref: <not yet declared — this sweep reports descriptively, no held-out dev split was used for model selection>
provenance:
corpus_hash:
artifact_ref: wave6/artifacts/vsp-sweep.json
code_hash:
Two fields deserve a specific caution. objective_authority.consumer_requirement_ref should point to an actual stated downstream need — never be left to imply that the training objective invented its own justification after the fact. And selection is written honestly for this chapter’s own experiment: the VSP sweep reported here used no separate development split for choosing a “best” weight — it is a descriptive sweep, not a model-selection procedure — so the record says exactly that rather than retroactively claiming a selection discipline that was not followed. A production pipeline choosing among these six candidates would need a genuine train/dev/test split, with the dev set deciding the weight and the test set reserved for one final, unbiased evaluation — a real requirement this chapter’s own artifact does not itself satisfy, and should not pretend to.
Nothing in this record contains ALLOW, DENY, or usable_for. The bridge registry can now answer not only “what did this bridge preserve?” (Chapters 21–22) but “what was it trained to preserve, and under whose authority?” — while “is that enough for this consumer?” remains, unambiguously, Chapter 20’s question.
Embedding Observatory progression
By the end of this chapter, the Observatory should be able to show, for any bridge candidate, not only what survived but what the optimizer was explicitly told to preserve and why:
TRAINING OBJECTIVE
point-target term (weight, reference)
source-VSP term (weight, reference, sampling)
target-rank term (if used)
task term (if used)
OBJECTIVE AUTHORITY
source_geometry | target_behavior | downstream_task
the specific property named
the consumer requirement this objective was meant to serve
MEASURED OUTCOME
counterpart-recovery observation ref (Ch22)
structural-fidelity observation ref (Ch22)
task-fidelity observation ref (Ch21, where applicable)
AUTHORIZATION
a separate Chapter 20 trace — never populated from this record directly
A reader inspecting a bridge candidate should be able to trace, end to end: which property the loss was built to protect, which authority justified that choice, what actually survived under measurement, and — as a fully separate step — whether a named consumer’s policy considers that enough.
Where Part VI lands
Chapter 21 turned preservation into a readable profile and taught that a metric’s name is not its meaning. Chapter 22 split that profile’s headline into counterpart recovery and structural fidelity, and showed the two can diverge sharply on the same translated vectors. This chapter split “preserve the geometry” itself into source, target, and task authorities, and demonstrated — under a controlled sweep that isolated one training-objective coefficient while holding everything else fixed — that increasing pressure toward one invariant can systematically work against properties a different consumer needs, with even the response to that single coefficient disagreeing across metrics. Together, Part VI now separates five questions that should never be collapsed:
identity tells us what representation we have (Ch17)
translation tells us how it was derived (Ch18–19)
authorization separates evidence from permission (Ch20)
measurement tells us what survived and what each metric means (Ch21–22)
objective authority tells us what we should have trained it to preserve (Ch23)
A translation does not succeed because it finds the right point, and it does not succeed because it preserves some geometry. It succeeds when it preserves the property the destination consumer actually relies on — measured against the right reference, and trained toward that reference deliberately rather than by default.
Next
That discipline was built around one transformation: a learned map between two encoders’ spaces. A cross-space bridge is not the only operation that returns a different representation of the same underlying object — compressing a document into a summary does, and applying a targeted semantic edit does too. Part VI asked what a translated vector owes its destination. Part VII asks the structurally identical question one level up, where the stakes are sharper because the output genuinely contains less information than the input: when a representation is made smaller, which relationships and which claims are allowed to disappear? The next chapter, Can a Smaller Representation Preserve a Larger One?, opens exactly where this one closes — name the transformation, name the reference, name the property, then measure whether it actually survived.