Building an Embedding Runtime
The capstone: RELATE 1.0, the executable embedding runtime this book's research forced into being. Exact space identity, measured geometry, hard negatives, calibration, cross-space comparison, bridges, preservation profiles, compression, and semantic operators — composed through one doctrine: identity is exact, compatibility is measured, usability is policy-scoped, and lineage composes while permission does not.
Part VIII — Embeddings Become Infrastructure
Why an embedding runtime became necessary
An embedding system used to look like this:
text → vector → cosine → top-k
Twenty-five chapters ago that was a reasonable sketch. It no longer is, because along the way this book kept asking questions that sketch has no place to answer:
- Which exact pipeline — model, revision, normalization, prefix, truncation — produced this vector? (Chapters 1, 17)
- Can two spaces even be compared, and does comparing them license mixing their vectors? (Chapters 16, 20)
- Does a fitted map between two spaces preserve retrieval, or fine relations, or a calibrated threshold — and are those three different questions? (Chapters 18–21)
- Does a compressed or whitened representation still support the task that mattered, or only the geometry that was easy to check? (Chapters 7, 8, 24)
- Does a semantic edit admit a reusable vector operator at all, and how would we know before trusting one? (Chapter 25)
- What happens after a chain of these — a translation, then a compression, then an edit? Does whatever passed the first hop still hold?
None of these questions has a script-sized answer. Each one produced a companion component — a space_record, a neighborhood_report, a calibration_record, a bridge, a preservation_profile, a compression_record, a transformation_record — because each is a measurement, and a measurement needs somewhere to live that is not the next paragraph of prose. Once there are seven or eight of these, kept as separate scripts, a system built from them accumulates a specific failure mode: nothing enforces that a translated vector still carries its bridge’s scope, or that a compressed one still carries its retention knee, or that anyone checks before mixing two spaces that were never compared. The instruments exist. Nothing makes using them the path of least resistance.
That is the problem an embedding runtime solves — not a smarter model, not a better metric, but a system whose default behavior is to demand the measurement before granting the use.
What would a system look like if every transformation of a representation carried its own identity, and no transformed representation could earn scoped permission until measurement said what it was good for?
RELATE 1.0
That system exists. It is called RELATE, and version 1.0 is what this chapter describes.
RELATE began, outside this book, as a narrower research result: a small ridge projection that recovers relation-specific structure — cyclomatic complexity, nesting depth, call-site count — from frozen CodeBERT embeddings, and orders candidates by that recovered structure instead of by raw cosine. Chapter 15 already brought that result in as external evidence: on the frozen historical assets, Chebyshev distance in the recovered coordinate space ordered structurally closer candidates correctly 73.3% of the time, against 53.2% for raw cosine and 53.3% for raw Euclidean distance — a hard-negative-ordering gain a generic similarity readout could not supply on its own. That result is still preserved, byte-for-byte, as HISTORIC evidence in RELATE’s own test suite; it is not recomputed, because the original CodeBERT assets are not on this machine, and pretending otherwise would violate the same evidence discipline this book has enforced throughout.
But the narrow result was never the whole idea. The idea was the sentence this book keeps sharpening from different directions: geometry is evidence about a representation, not permission to use it. RELATE 1.0 is that sentence, compiled into a runtime. Its core object graph — space identity, native comparison, hard-negative evaluation, calibration, signal bundles, bridges, preservation profiles, compression cartridges, semantic operators, transformation lineage — is not a proposal. It runs. Every number in this chapter that is not explicitly marked otherwise comes from executing real code against a real benchmark suite, and the runtime’s own tests pin its behavior with hashes that fail loudly if it drifts.
This chapter is not a RELATE tutorial dropped into the book. It is the argument of the last twenty-five chapters, read back as architecture — the point where “measure before you trust” stops being only a maxim you apply by hand and becomes a system that refuses to grant scoped permission without measured evidence attached.
The three layers
Everything RELATE does sorts into exactly three layers, and keeping them apart is the discipline the rest of this chapter exists to demonstrate.
IDENTITY. What produced this representation? Exact and hashable. Two vectors either came from the same declared pipeline or they did not, and that question has a yes/no answer computed from a hash, never from inspection or intuition.
COMPATIBILITY. What relationships, structure, or task behavior survive when the representation changes? Empirical and task-dependent. It is never a single scalar, because — as Chapters 16, 20, and 21 measured repeatedly — coarse structure and fine distinctions can move in opposite directions under the same translation.
USABILITY. What is a transformed representation permitted to be used for? Policy-scoped and fail-closed. A scope that was never measured is not silently assumed safe; it comes back unknown, and unknown means no.
Same identity does not prove quality. Different identity does not prohibit comparison. A bridge does not grant usability. Measurement earns scoped permission.
Those three layers are not a metaphor laid over the code after the fact. They are the literal shape of relate.spaces, relate.evaluation, and relate.retrieval/relate.evaluation.preservation — three packages, three responsibilities, and (as the sections below show) a rule against a value from one layer ever silently standing in for a value from another.
Registering exact spaces
Chapter 17’s space_record asked what produced a vector. RELATE’s answer is SpaceIdentity:
from relate.spaces.identity import SpaceIdentity
space = SpaceIdentity(
model="all-mpnet-base-v2",
dimensions=768,
revision="", tokenizer="", prefixes="",
sequence_length=0, truncation="", dtype="float32",
normalize=True, post_processing="",
)
space.space_hash # sha256 over the sorted identity dict, first 16 hex chars
The hash is computed, not assigned — the model, revision, tokenizer, prefixes, sequence length, truncation rule, dtype, normalization flag, and post-processing string all feed a canonical JSON encoding, and space_hash is its digest. Two SpaceIdentity values with identical fields deterministically produce the same registry identity. Changing any identity field changes the canonical payload and therefore the identity being declared; RELATE does not infer sameness from geometry or from matching dimensionality. A SpaceRegistry then maps hashes to identities and refuses to answer for a hash it has never seen:
runtime.spaces.require(some_hash) # RelateError: unknown space_hash
Observatory.attach(embeddings, space=...) is the enforcement point for Observatory-managed representations: it checks the matrix is two-dimensional, checks its width against the declared space, registers the identity, and hands back the matrix. Low-level NumPy evaluators deliberately remain usable on bare arrays — describe_geometry(vectors) can measure geometry without pretending to know where the vectors came from. The stronger boundary appears when an operation claims identity, lineage, cross-space meaning, or scoped usability: there RELATE requires the relevant SpaceIdentity and refuses to manufacture provenance that was never supplied. A caller can always discard provenance on purpose; the runtime’s job is not to pretend it survived.
Derived spaces reuse the same mechanism through derive_space: a PCA cartridge, a bridge’s output, a compressed prefix all get a new hash that records the parent hash and the derivation. Matching space_hash establishes declared identity — not quality, and not compatibility. A different hash does not prohibit anything either; it is simply the fact that triggers every gate in the rest of this chapter.
Measuring a native space
Chapter 2’s geometry_probe and Chapter 7’s dimensionality_report become one function, describe_geometry, returning a GeometryReport:
from relate.evaluation.geometry import describe_geometry
report = describe_geometry(vectors)
report.cosine_mean, report.cosine_std # random-pair background (Ch 5, 8)
report.effective_rank, report.participation_ratio # spectral concentration (Ch 7)
report.intrinsic_dimension.estimate # TwoNN, with its own provenance
Every field carries the caution Chapters 5–8 spent whole chapters earning. intrinsic_dimension is an IntrinsicDimensionEstimate — method, estimate, sample count, parameters — never a bare float, because an intrinsic dimension is an estimator’s opinion, not a fact the space hands over. The pair-sampling policy (PairSamplingSpec) is recorded alongside the cosine statistics — a full census under max_pairs, a seeded uniform sample above it — so that two cosine means measured under different sampling regimes are never compared as if they were the same quantity. describe_geometry diagnoses geometry and stops there: no plots, no PCA coordinates, no UMAP layout. Chapter 2’s warning that a visualization is another transformation is honored by omission — geometry reports geometry; a picture is a separate, later choice a caller can make and is then responsible for.
On RELATE’s own seeded benchmark spaces (benchmarks/geometry/, n = 1,200, nominal d = 256) the shape vocabulary produces exactly the split Chapters 5 and 7 predicted from real encoders: an isotropic space sits near cosine 0.00 with effective rank ≈ 230; a concentrated space sits near 0.49 with effective rank ≈ 96; a genuinely low-rank space (six true dimensions plus noise) reports an effective rank of 6.0 and a TwoNN estimate of 5.3 regardless of its nominal 256 coordinates. These are synthetic fixtures, not a re-run of Chapter 5’s five real encoders — the point is not that the numbers match, it is that the same instrument, pointed at a different space, tells a different and internally consistent story every time. benchmarks/geometry/’s own regression suite asserts something stronger than a report: rotate the space (X @ Q) and the cosine mean, cosine std, and effective-rank ratio drift by less than 1e-9 — Chapter 5’s rotation-invariance proof, now a passing test rather than an argument.
Evaluating retrieval and hard negatives
Chapters 10 and 11 built the vocabulary — win rate against a typed distractor, margin against the hardest negative. RELATE keeps that vocabulary but makes the semantics exact: a HardNegativeCase says an anchor should prefer its positive over a deceptive negative, and evaluate_hard_negatives scores every case with an injected scorer — cosine, Euclidean, a RelationProjection readout, a bridge, anything callable — never assuming which:
from relate.evaluation.hard_negatives import evaluate_hard_negatives, HardNegativeCase
report = evaluate_hard_negatives(cases, vectors, scorer)
report.accuracy, report.mean_margin, report.by_relation
margin = positive_score − negative_score; |margin| ≤ tie_tol is an explicit TIE, not silently folded into a win or a loss, because a near-zero margin is exactly the evidence Chapter 11 built calibration around. The outcome vocabulary is WIN / TIE / LOSS throughout — Chapter 10’s binary “did the distractor win” becomes a three-way, per-case, per-relation, per-group record, with worst_groups surfaced automatically so an aggregate accuracy can never quietly hide the relation that is actually failing.
Two evidence layers stay separate on purpose. The CodeBERT numbers (0.733 relation-readout accuracy vs. 0.532/0.533 for raw cosine/Euclidean) are preserved as HISTORIC reference evidence; relative to the current replay they are also EXTERNAL, because the original assets are absent and the result is not recomputed. expected/historic-reference.json freezes those numbers instead of laundering them into a new measurement. What the current evaluator does reproduce is the observation’s shape, on adversarially mined synthetic negatives: a relation-specific readout scores 0.725, while raw cosine and Euclidean both collapse to 0.100 — harsher than the historic figures because these mirror negatives are deliberately the hardest ones in the pool, not a fixed historic sample. Same structural finding, honestly different numbers, for an honestly different reason.
Calibrating operating points
A score is not an operating policy — Chapter 14’s central claim — and CalibrationRecord refuses the shortcut that claim warns against: there is no is_match(score). There is decide(score), returning ACCEPT, REJECT, or ESCALATE, the middle option Chapter 14 measured as the honest majority outcome on hard negatives:
record.decide(score) # ACCEPT | REJECT | ESCALATE
On RELATE’s seeded calibration fixtures (benchmarks/calibration/, 1,500 scores per side), the same scorer family under two negative distributions produces two entirely different operating pictures — AUC 0.925 / EER 0.159 / ambiguity band 0.126 for separated negatives, against AUC 0.774 / EER 0.296 / ambiguity band 0.478 when the negatives overlap the positives. Calibration belongs to a distribution and a task, not to a model name. In the hard fixture, nearly half the pooled scores fall inside the benchmark’s ambiguity band; the comparison with the book’s differently defined Wave 1 escalation result comes next, with the two quantities kept separate.
A CalibrationRecord is provenance-bound: space_hash, corpus, task, and a CalibrationScope (domain, query type) travel with the threshold. staleness_against(...) does not return a bare boolean; it names which dimensions changed — space_hash, domain, scorer, the negative set’s content hash — so a caller learns why a threshold went stale, not just that it did. Shifting the negative distribution by +0.25 on RELATE’s fixture moves the EER threshold by +0.109, the same order of magnitude as the book’s own ~0.10 cross-domain drift (Chapter 14, rows 1.10–1.11) — measured again, on different synthetic scores, landing in the same neighborhood for the same underlying reason: the operating point is a property of the negatives it was fit against.
Two of these numbers deserve separating rather than blurring together. The hard-regime AUC (0.774) sits close to the book’s own Wave 1 figure (0.75) — plausibly the same phenomenon under different synthetic scores. The size of the safe-decision gap does not match as closely: RELATE’s hard-regime ambiguity band, defined as the fraction of pooled scores with FAR and FRR both under 0.10, comes to 0.478; the book’s Wave 1 escalate band, defined by a different procedure, was 86%. Both say the same qualitative thing — a respectable AUC can still leave a large region where a fixed threshold cannot decide safely — but they are two different metrics on two different score distributions, and stating one figure as a reproduction of the other would be exactly the false-equivalence this book has spent thirteen chapters warning against.
Retrieval, calibration, and signal bundles
Chapter 15’s diagnostic vector — score plus margin plus local geometry plus calibration state — is SignalBundle, and Observatory.inspect_result is the function that composes one after a search has already happened:
bundle = runtime.inspect_result(
query_vector=q, candidate_vector=c, context_vectors=corpus,
candidate_index=i, observation=hard_negative_observation,
calibration=record,
)
bundle.available_signals # which fields are measured, not which are bad
inspect_result does not search and does not decide. It reads the score from an injected scorer, the margin from a supplied HardNegativeObservation, local density and hub in-degree from the existing neighborhood primitives, and the calibration outcome from a supplied CalibrationRecord — pure composition of already-measured evidence, exactly Chapter 15’s “orchestration only” design. available_signals distinguishes unmeasured from measured-and-bad: a None field means nobody computed it; a bad float still means somebody did. External evidence — a cross-encoder, a second encoder, a verifier’s verdict — lives in a separate ExternalSignals field and is never merged into the geometric fields’ numbers, because Chapter 10 already showed why: an unvalidated second model is not automatically a second opinion worth trusting.
On RELATE’s seeded signal-bundle fixture, score alone reaches balanced accuracy 0.677 on a deliberately hard mix; adding the geometric block — margin, density, hubness — lifts that to 0.906; a simulated generic external verifier, right about the easy cases and coin-flip on the hard ones, contributes essentially nothing once stacked on top (0.906 unchanged, within noise). That verifier is explicit fiction, constructed to model a correlated-error pattern rather than to stand in for a measured NLI system — no NLI dependency, no provider call exists anywhere in the runtime. The current fixture therefore supports the narrower claim that adding an untargeted second signal does not automatically improve a strong geometric bundle. The book’s historic 0.76 → 0.90 → 0.897 result (Wave 1 row 1.12) remains separate evidence that showed a similar pattern with a real external model.
RetrievalPolicy.route(bundle) is the consumer, not a field on the bundle itself — routing logic belongs to policy, not to evidence. A missing margin routes straight to verify; a thin margin routes to verify before it routes to accept; a calibration escalation overrides the margin route entirely. Retrieval and verification stay separate function calls, exactly Chapter 12’s boundary: search returns candidates, inspect_result characterizes one, and nothing in between quietly does both.
Comparing spaces without mixing them
Chapter 16 asked whether two models agree. compare_native_spaces answers it without ever authorizing what comes next:
from relate.evaluation.cross_space import compare_native_spaces, require_same_space_for_mixing
report = compare_native_spaces(
source_vectors=A, target_vectors=B,
correspondence=corr, source_space_hash=..., target_space_hash=...,
)
report.geometry.cka, report.neighborhood.mean_overlap, report.counterpart.top1
CorrespondenceSet is the explicit row map between the two matrices — the ids, the source rows, the target rows, and a content hash computed from all three — so row order never silently defines correspondence, a discipline this book’s own cross-space chapters insisted on and RELATE now enforces at the type level: the constructor raises if two ids collide, if a row repeats, if the supplied hash disagrees with the recomputed one. compare_native_spaces returns a SpaceComparisonReport holding geometry (CKA, cosine- and distance-matrix correlation), neighborhood overlap, counterpart recovery as its own separate object, and a hard-negative delta side by side — never collapsed into one number, because Chapter 21’s central finding is that these move independently.
On RELATE’s own seeded native-pair fixture, that independence is exact and visible in a single report: CKA 0.948, ten-nearest-neighbor overlap 0.513, counterpart recall@1 a perfect 1.000, and a hard-negative delta of −0.123 concentrated in exactly the assertion-sensitive relations — negation and temporal mismatch degrade by more than 0.10, topic-related content stays within 0.05 of native. Coarse structure looks strong; the fine distinction the book has been chasing since Chapter 1 fails anyway. A permuted control — the same vectors, rows shuffled — collapses every measure to its chance floor (CKA 0.006, overlap 0.005, counterpart 0.000), confirming the comparison is measuring real structure, not an artifact of the statistic.
Comparing is allowed for distinct hashes. Directly mixing their native vectors is not: require_same_space_for_mixing raises unless the hashes are identical, and the benchmark asserts that denial. A bridge does not switch that guard off or make the native hashes equivalent; it produces candidate vectors under a new derived identity, which must then be measured and judged on its own evidence. Different spaces may be compared; raw vectors from them may not be silently treated as one space.
Bridges
Chapter 20’s directional map is now split along the boundary the later experiments earned. Bridge contains the transformation and its fit provenance — source hash, target hash, a direction string that is "{source}->{target}" by construction (the object cannot represent an undirected map even by accident), a fitted mapping/bias under one affine contract, and a BridgeFitProvenance recording the anchor correspondence hash, anchor count, method, hyperparameters, seed, and the code identity of the fitting routine. What the bridge was shown to preserve is deliberately not stored on the bridge itself; that evidence belongs to the PreservationProfile.
bridge = runtime.fit_bridge(
source, target, source_space=A, target_space=B,
correspondence=train_anchors, method="ridge", params={"alpha": 1.0},
)
candidate = bridge.transform(source) # produces vectors — judges nothing
derived = runtime.bridge_space(A, bridge, B) # a NEW space_hash, never B's own
A bridge can, in the runtime’s own words, “return perfectly shaped garbage” — transform checks dimensions and finiteness and nothing else. Detecting garbage is the measurement spine’s job, deliberately kept out of the producer. And bridge_output_space is where Chapter 17’s lineage rule becomes load-bearing: candidate vectors live in the target’s coordinate system, but stamping them with the target’s own space_hash would collapse coordinate compatibility into representation identity — exactly the confusion Chapter 17 warned against. The derived space instead records all three facts at once — parent, bridge id, target reference — and the target’s native hash is asserted, in the benchmark’s own tests, never to be reused.
Only three fitted families ship as production producers — Procrustes (orthogonal), linear, and ridge — plus deliberately boring controls: the identity no-op, a constant-target-centroid floor, and a seeded random map. An MLP producer was tried and dropped: it is a documented negative result, not a missing feature. On RELATE’s bridge benchmark (train-fit, held-out eval rows), the controls do exactly what controls should — the centroid collapses counterpart recall@1 to 0.002 and ten-NN overlap to 0.013; the seeded random map’s counterpart recall@1 (0.005) is barely better, though its neighborhood overlap (0.386) is considerably higher, because a random linear map at least keeps points spread out even though it aligns none of them correctly. Among the real producers, Procrustes coincides with the identity no-op here (1.000 / 0.596 / −0.122) because the fixture’s shared nuisance backbone dominates and a rigid rotation cannot express the dimension-weakening the true translation needs; linear and ridge relax that rigidity and buy real neighborhood fidelity (0.794 and 0.788 overlap) at a smaller hard-negative cost (−0.035, −0.032). No producer wins every criterion, and which one looks best depends on which criterion you read. Training reconstruction — how well a fitted map recovers its own training anchors — is labeled a FIT DIAGNOSTIC in its own file and is never treated as preservation evidence: train and evaluation correspondences carry distinct, separately asserted content hashes, so excellent training reconstruction cannot quietly stand in for held-out preservation.
Measuring preservation
A fitted map is a claim about coordinates. A PreservationProfile is the claim about what that map is good for, and RELATE keeps its three steps visible rather than collapsing them:
measurement (the evaluators above) → comparison to an explicit authority (ReferenceFrame) → scope-specific verdicts from declared policy.
profile.usable_for("retrieval") # bool — a lookup over an explicit PASS
profile.explain("threshold_transfer") # names exactly which gate failed, and by how much
ReferenceFrame is the callout Chapter 21 deserved and now has: TARGET_NATIVE (did the translation reproduce what the target space itself does — the default for a bridge), SOURCE_NATIVE (did the transformation keep what the source did — the honest question for compression), or TASK_GOLD (a labeled ground truth, when one exists). Every PreservationResult in a profile carries its frame explicitly; nothing infers one silently. “Preserved” is incomplete until you say relative to what.
A PreservationPolicy names a scope and a tuple of Requirements — min_value, max_delta, min_ratio, and so on — each either required or advisory. Judging a policy against measured results produces a ScopeVerdict: PASS if every required gate clears, WARN if only an advisory one misses, FAIL if a required gate misses or was never measured at all — an unmeasured requirement fails closed, not open. usable_for(scope) is a lookup over that verdict, nothing more; an unrecognized scope name returns False because no verdict exists for it, never because the runtime guessed. explain(scope) renders the reasoning a human actually needs: which gate failed, its bound, the observed value, and the reference frame it was judged against.
On RELATE’s preservation benchmark — coarse source translated toward a fine target, the exact direction that exposes what a translation quietly drops — every fitted producer and the no-op alike PASS retrieval and neighborhood-use scopes (counterpart recall@1 1.00, overlap ≈ 0.60) while every one of them FAILs threshold transfer, because the translated score distribution simply does not sit where the target’s calibrated threshold expects it. Retrieval PASS coexisting with threshold-transfer FAIL is not a bug in the benchmark. It is the finding: a bridge can be good enough to find the right neighborhood and still be dangerous behind a fixed accept/reject line. Per-relation visibility makes sure the aggregate accuracy cannot hide this either — negation degrades +0.20 to +0.23 in every producer’s profile (the “failed distinction” made visible, not averaged away), while topic-related content stays essentially flat.
Compressing representations
Chapter 24 asked whether a smaller representation preserves a larger one. RELATE answers per task, under SOURCE_NATIVE authority — did the compression keep what the source itself did, the only question compression can honestly be asked:
cartridge = fit_pca(vectors, output_dimensions=12, source_space_hash=source.space_hash)
compressed_space = runtime.derive_transformation_space(
parent=source, transformation_id=cartridge.transformation_id,
parameters={"method": "pca", "output_dimensions": 12}, dimensions=12,
)
evaluation = runtime.evaluate_compression(cartridge, vectors, vectors, ...)
Three production cartridges exist: PCA, a seeded random projection as its control, and prefix truncation — and prefix truncation carries a field, training_support: unknown, that says explicitly: slicing the first k coordinates of an ordinary model is not Matryoshka truncation unless training provenance says the model was actually trained for nested prefixes. Calling it Matryoshka without that provenance would be exactly the false promise Chapter 24 warned against.
On RELATE’s compression sweep (64/32/16/8 dimensions, three methods, declared per-task policies), the honest middle case is pca-32: it FAILs the retrieval-scoped policy while PASSing the fine-ordering policy — the same cartridge, two different scopes, two different verdicts, which is exactly what a “safe dimension” question should return once the question is scoped instead of global. pca-64 passes retrieval and fails threshold transfer; pca-8 fails everything; the random and prefix controls fail every declared scope at every tested width in this fixture. That is a measured outcome, not a general explanation of either method: random projection is deliberately unstructured, while prefix truncation can be privileged by training in models designed for it — and this benchmark’s training_support is explicitly unknown. Reading the task knees this sweep declares — retrieval needs 64, fine ordering needs 16, no tested width passes threshold transfer — produces the same shape Chapter 7 measured on real RELATE embeddings: the useful statement is “this compression is usable for X under this evidence,” never “32 dimensions is enough.” And the sharpest single number in the sweep is pca-8: it keeps CKA 0.99 — near-perfect coarse geometry — while its fine-ordering policy still fails, negation degrading by −0.10. High coarse preservation does not verify semantic faithfulness. Explained variance is provenance metadata here, never permission.
Testing semantic operators
Chapter 25 asked whether an edit admits a reusable operator, or only a rung 0 (identity) or rung 1 (delta) answer. RELATE runs that exact bake-off, and its vocabulary is worth adopting precisely because it removes an ambiguity the book’s own prose was carrying:
identity_map names the operator, not the texts: identity_map on a transformation means no vector-space map was needed for this evaluation to clear its bar — the embedding barely moved. It is never a claim that the source and target sentences mean the same thing.
NONE_PASS is a scientific result, filed with the same rigor as a PASS, not an implementation failure to apologize for.
from relate.transformations import identity_map, fit_constant_delta
delta = fit_constant_delta(source, target, relation="claim_weakened", source_space_hash=h)
evaluation = runtime.evaluate_operator(delta, source, target, correspondence=corr, reference_space=target_space)
evaluation.profile.verdict_for("operator_fidelity").verdict
evaluate_operator requires the transformation artifact to name its trained relation — an operator, by the canonical vocabulary, is relation-bound, never a free-floating vector. Its authority is TARGET_NATIVE: the transformed content, embedded normally, is the ground truth an operator’s candidates are judged against — you edit the text, embed the edit, and ask whether the operator’s vector-space prediction lands where the real embedded edit landed.
RELATE’s operator benchmark runs this on real frozen relate-doc-0.1.0 transformation pairs — 260 pairs across nine exact classes, real content hashes, asserted-distinct train/eval splits — through a documented deterministic mirror encoder, because no sentence-transformer model runs anywhere inside src/relate. That mirror reproduces the book’s Wave 5 taxonomy structurally, at its own honestly different numbers:
transformation mirror identity cos simplest passing operator
active_to_passive 0.995 identity_map
present_to_past 0.995 identity_map
relation_swap 0.993 identity_map
claim_strengthened 0.995 identity_map
temporal_shift 0.994 identity_map
claim_weakened 0.831 constant_delta
formal_to_informal 0.682 NONE_PASS
verbose_to_concise 0.784 NONE_PASS
statement_to_negation 0.661 NONE_PASS
Five identity_map, one constant_delta, three NONE_PASS — exactly the shape Chapter 25 measured on all-mpnet-base-v2, where the real encoder’s identity cosines ran 0.88–0.98 for the five identity-like classes and 0.67–0.75 for the three that admitted no passing operator. The mirror’s own numbers are not the book’s numbers, and the benchmark’s README says so directly: the geometry here is a documented deterministic synthetic stand-in over the real content pairs, not a re-run of the sentence encoder. What survives across both evaluations is the scoped taxonomy: five classes clear the declared identity_map policy, claim weakening clears constant_delta, and register, length, and negation remain NONE_PASS across the tested ladder. That does not make those relations universal geometric laws, nor does it establish that weakened claims always form one reusable direction in another encoder or corpus. The stronger result is the held-out one: linear/affine train to near 1.0 reconstruction on the small NONE-class sets and still fail evaluation — reported explicitly as a fit diagnostic, never as preservation evidence, the identical discipline Chapter 25 insisted on for training reconstruction anywhere in this book.
One transformation doctrine
Bridges, compression cartridges, and semantic operators look like three different features. They are three producers behind one contract, and naming that contract is the point of this section.
flowchart TD
P["producer — bridge, compression cartridge, or semantic operator"]
P --> T["transform(source) — candidate vectors, nothing judged"]
T --> D["derived space — a NEW space_hash, parent + transformation recorded"]
D --> A["explicit authority — ReferenceFrame: TARGET_NATIVE | SOURCE_NATIVE | TASK_GOLD"]
A --> M["evaluate_transformation — the SAME measurement path for every producer"]
M --> PR["PreservationProfile — results, then scope-specific PASS / WARN / FAIL"]
PR --> U["usable_for(scope) / explain(scope)"]
Observatory.evaluate_transformation is that single measurement path, and the three façades this chapter has walked through — evaluate_bridge_full, evaluate_compression, evaluate_operator — differ only in which ReferenceFrame they default to and which counterpart-recovery behavior makes sense for their shape. None of them contains its own metric code. A bridge, a compression cartridge, and a fitted operator all satisfy the same VectorTransformation protocol — a transformation_id, a source_space_hash, a transform method — which is precisely why they can share one evaluator without the runtime needing to know, in that evaluator, which kind of producer it is judging.
One more distinction the contract makes structural rather than incidental: vector transformations act on vectors already produced (values -> values — a bridge, a PCA cartridge); content transformations act upstream, on text (content -> content' -> embed -> values), and deliberately expose no transform method on vectors at all — their outputs only enter measurement after being embedded normally, the same discipline Chapter 24’s compression-vs-content distinction needed and Chapter 25’s operator evaluation now enforces by construction.
Every producer, whatever it is, ends at the same derived-identity fact: derived_transformation_space constructs a derived identity from the parent and transformation rather than reusing the parent’s identity, even when dimensionality is unchanged. A PCA cartridge that happens to preserve the input width still gets a new space_hash. A transformation creates a new representation identity and an obligation to measure what survived — whether or not a human remembered to ask for the measurement. That obligation is a convention this runtime follows everywhere, consistently, not a rule the type system enforces on a caller who skips evaluate_transformation entirely; RELATE 1.0 makes the path exist for every producer family, and makes it the only measurement path, without claiming to compel its use.
Transformation lineage
Observatory.lineage(space_hash) walks a derived space back through its ancestry — candidate first, parent, its parent’s parent — and returns a tuple of DerivationRecords: space_hash, parent_hash, transformation_id, kind, reference_hash. Nothing in that record carries a verdict.
runtime.lineage(compressed_bridge_output.space_hash)
# (DerivationRecord(kind="transformation", ...),
# DerivationRecord(kind="bridge", ...))
That silence is deliberate, and it is the sharpest sentence this chapter’s doctrine produces:
Transformation lineage composes; preservation permission does not.
Chain a bridge into a compression cartridge — native → bridge → compress — and the lineage walk shows exactly that chain, cleanly, however many hops long. But suppose A → B measurably PASSes the retrieval scope, and B → C also measurably PASSes it. RELATE does not conclude that A → C passes. There is no API that lets a caller ask the composed question by citing the two individual verdicts; the composed candidate is a new derived space with its own hash, and it needs its own evaluate_transformation call against its own explicit authority before any scope verdict exists for it at all. A lineage this clean is not evidence of safety. It is a record of how much fresh measurement is now owed.
The complete Observatory run
Everything above is one runnable path. python -m relate.cli demo walks it end to end — register, compare, bridge, derive, judge, compress, operate, explain — on synthetic spaces built so that a coarse-signal direction and a fine, assertion-sensitive direction are genuinely different in the data, exactly the confound Chapters 16 and 21 spent several chapters teasing apart. This is the actual, unedited output of that command:
SOURCE demo-coarse [a1a5354ed043547f]
TARGET demo-fine [70a68f6a3dae3c1c]
NATIVE CKA 0.86 NN overlap 0.51 counterpart top-1 0.94
BRIDGE ridge a1a5354ed043547f->70a68f6a3dae3c1c anchors 320
DERIVED demo-coarse [11588b564503b381] (native target hash never reused)
VERDICT retrieval PASS
VERDICT threshold_transfer FAIL
WHY FAIL — threshold_transfer / - failed: calibration_transfer/far_increase = 0.075 (requires far_increase max_delta 0.05; frame target_native, reference 0.1)
COMPRESSION
VERDICT compression/retrieval PASS
VERDICT compression/threshold_transfer FAIL
OPERATOR
VERDICT operator/constant_delta PASS (1.000)
VERDICT operator/identity_map FAIL (0.579)
DOCTRINE transform -> derive identity -> measure -> scoped verdict; lineage composes, permission does not.
Read it the way the rest of this book taught you to read a measured table. NATIVE reports coarse agreement (CKA 0.86) is real but far from perfect, and neighborhood overlap (0.51) is meaningfully lower than counterpart recovery (0.94) — the split Chapter 16 predicted, on the first two lines. BRIDGE fits a ridge map on 320 anchor rows; DERIVED shows the candidate space getting its own hash, distinct from both its coarse parent and its fine target, with the runtime’s own comment confirming the target’s native hash was never reused. The two VERDICT lines for the bridge repeat the preservation-profile shape from earlier in this chapter exactly: retrieval PASSes, threshold transfer FAILs, and WHY names the actual failed gate — a false-accept-rate increase of 0.075 against a 0.05 bound, off a native reference FAR of 0.10. COMPRESSION repeats the same split on a PCA cartridge of the fine space. OPERATOR is the section worth sitting with: in this specific synthetic scenario, the constant-shift edit passes its operator-fidelity policy at a perfect 1.000, while the operator built to represent no map at all — identity_map — fails at 0.579, because this demo’s synthetic edit is not geometrically invisible the way five of the nine real RELATE-DOC transformations were. That is not a contradiction of Chapter 25’s finding; it is the same measurement machinery correctly reporting a different fact about a different, deliberately constructed transformation. The final DOCTRINE line is the runtime’s own one-sentence summary of everything sections 4–13 just walked through by hand.
To reproduce it: pip install -e ".[dev]" and python -m relate.cli demo. To reproduce any individual claim in this chapter rather than take the transcript’s word for it, every benchmark family under benchmarks/ ships the same two-command contract —
python benchmarks/<family>/run.py # write expected/
python benchmarks/<family>/run.py --check # verify byte-identical
— which means every number in this chapter is not merely quoted; it is pinned. If the runtime’s behavior ever drifts, the --check invocation fails loudly rather than letting a stale number sit quietly in a README, or in this chapter.
What RELATE does not establish
This is the section the whole chapter has been earning the right to write plainly, and it should read as trustworthy because of what it refuses to claim, not despite it.
RELATE does not establish semantic truth, factual correctness, or that any candidate it returns is right. Retrieval delivers candidates; the runtime has no verification layer, and Chapters 10 and 15 already showed why bolting on a generic one would be its own unvalidated claim.
RELATE does not establish universal cross-model compatibility. Every comparison and every bridge is scoped to the two spaces it was actually measured on; nothing about a good result on one pair licenses an inference about a third space.
RELATE does not establish that a good CKA implies safe mixing, that counterpart recovery implies preserved neighborhoods, or that preserved neighborhoods imply preserved fine semantics. Every native-pair and preservation benchmark in this chapter measured these as separate, independently movable quantities precisely because collapsing them was the mistake Chapters 16 and 21 were written to prevent.
RELATE does not establish that retrieval preservation implies calibration transfer. The preservation benchmark’s own headline result is the opposite: retrieval PASS and threshold-transfer FAIL, in the same profile, for every fitted producer tested.
RELATE does not establish that compression preserving global geometry preserves claims. pca-8 keeping CKA 0.99 while failing fine ordering is the sharpest counterexample the runtime itself produces, and Chapter 24’s claim-conditioned faithfulness question — does a compressed or summarized representation still support a specific claim — stays exactly where the book left it: measured in the Transformation Wave’s artifacts, correctly not promoted to a shipped RELATE evaluator.
RELATE does not establish that a fitted semantic operator proves a linguistic law. An identity_map or constant_delta PASS is a scoped verdict about a tested producer; NONE_PASS is the equally scoped selection outcome when none of the tested producers clears the policy. All of them are conclusions on one corpus, in one embedding space, at one data scale — the operator benchmark’s own README states the small-sample caveat directly, and a passing rank-1 or affine map at a larger scale remains an open, explicitly flagged question, not a claim this runtime makes.
RELATE does not establish that two passing transformations compose into another passing one. Lineage tracks ancestry with no verdict field anywhere in it; the previous section’s negative test is the point, not an oversight.
RELATE does not establish that usable_for(scope) means anything beyond the policy and evidence that produced it. An unrecognized scope fails closed by construction, and a scope that passed under one declared policy is not thereby validated against a different, undeclared one.
Two smaller, honest gaps round this out, named rather than hidden: RELATE has no cost or migration estimator — a real re-embedding decision needs a budget this runtime does not compute — and no ANN or vector-index instrumentation; both are index and deployment infrastructure, correctly out of scope for a runtime about representation measurement. And the “every transformation obligates measurement” doctrine that runs through this whole chapter is, honestly, a convention the runtime follows everywhere consistently — not a rule a type checker enforces against a caller who chooses to skip evaluate_transformation and use a producer’s .transform() output directly. The path to measurement exists, uniformly, for every producer family. Walking it remains a choice.
Failure modes
- Trusting a
space_hashmatch as a quality signal. Tempting because identity is the one exact thing in the system. Check: identity answers “same pipeline,” never “good pipeline” — quality is a downstream, task-scoped question. - Treating a different
space_hashas a reason two spaces cannot be compared. Tempting because the mixing guard makes distinct hashes feel forbidden. Check:compare_native_spacesis built exactly for this; only mixing without a bridge is denied. - Reading
identity_mapas “the texts are identical.” Tempting because the name invites it. Check: it names operator identity — no vector map was needed under this bar — never semantic identity of the content. - Reading
NONE_PASSas a bug to fix. Tempting because every other row in an operator table has a passing entry. Check: it is the scientific result for that transformation at this data scale; report it, do not chase a rung that keeps failing held-out. - Reading counterpart recall@1 as neighborhood or fine-relation preservation. Tempting because a perfect
1.000looks conclusive. Check: the native-pair and preservation benchmarks both hold these as separate results precisely because they diverge. - Reading a PASS on one scope as usability for another. Tempting because “it passed” feels total. Check:
usable_foris per scope; retrieval PASS and threshold-transfer FAIL coexist in nearly every profile this chapter showed. - Assuming lineage implies permission. Tempting because a clean derivation chain looks like a paper trail of approvals. Check: lineage is ancestry only; every hop needs its own fresh
evaluate_transformationcall. - Skipping
evaluate_transformationbecause a producer’s.transform()already “worked.” Tempting because the vectors came out the right shape. Check: a bridge can, in the runtime’s own documentation, “return perfectly shaped garbage” — shape is not preservation.
Companion component: the runtime, end to end
Content / vectors
│
▼
Space Registry SpaceIdentity, space_hash, attach() refuses spaceless vectors
│
▼
Measurement Spine
├─ geometry describe_geometry, PairSamplingSpec, TwoNN with provenance
├─ hard negatives evaluate_hard_negatives, WIN/TIE/LOSS, HardNegativeReport
├─ neighborhoods compare_neighborhoods, counterpart_recovery (kept separate)
├─ calibration CalibrationRecord, decide(), staleness_against()
└─ signal bundles inspect_result, SignalBundle, geometric vs external
│
▼
Cross-space comparison compare_native_spaces, CorrespondenceSet, mixing guard
│
▼
Transformations (one contract, three producer families)
├─ bridges Procrustes / linear / ridge (+ controls), directional
├─ compression PCA / random / prefix (SOURCE_NATIVE authority)
└─ operators identity_map / constant_delta / linear / affine (TARGET_NATIVE authority)
NONE_PASS is a selection outcome, not a producer
│
▼
Derived Space new space_hash, parent + transformation recorded
│
▼
Preservation Profile results → ReferenceFrame → scope-specific PASS/WARN/FAIL
│
▼
usable_for(scope) / explain(scope) the only door usability opens through
Every box above names a real object this chapter has already shown running. Nothing in this diagram is aspirational, and nothing in it should be read as more than what it is: a runtime that carries the book’s measurements as data and refuses to convert a measurement it does not have into a permission it grants anyway.
What this chapter — and the book — established
Retrieval, clustering, deduplication, cross-model migration, compression, and editorial memory all run on the representation layer this book spent twenty-six chapters measuring — and they inherit its limits whether or not a system was built to notice them. The operational limits stopped being vague cautions and became either measurements or explicit boundaries on what the evidence permits us to claim: the objective’s bias toward aboutness over assertion, the arbitrary basis, the effective dimension against the nominal one, the hubs, the near-but-wrong tail, the uncalibrated score, the incompatible second space, the lossy bridge, the forgetful compression, the edit that admits no operator at all.
RELATE 1.0 is what happens when those measurements stop being a chapter’s worth of prose and become a runtime’s default behavior: identity computed from a hash rather than asserted; compatibility read off a measured comparison rather than assumed from matching dimensions; usability granted per scope, from a declared policy, over evidence that either exists or does not — and when it does not, the answer is FAIL, not silence read as consent.
The book began with a vector is not meaning. Twenty-five chapters of instruments led somewhere more precise and more operational than that opening sentence could say on its own:
Geometry is evidence about a representation, not permission to use it.
RELATE turns that principle into three explicit layers and one closing rule that this chapter’s own transformation, section by section, kept proving true of itself:
Identity is exact. Compatibility is measured. Usability is policy-scoped. Transformation lineage composes; permission does not.
A vector is a list of numbers. What you are allowed to conclude from it, and what you are allowed to do to it next, were never the same question — and now, in a system you can install and run, they never quietly collapse into one.