Can a Smaller Representation Preserve a Larger One?
The text got shorter; the embedding did not. Audit a five-layer preservation stack against its actual implementation and discover that whole-document cosine, self-document recovery, and query-conditioned similarity can all stay silent while a compressed summary reverses a relation or changes a number — and that only an external NLI verifier, itself fallible, catches most of it.
Part VII — What Survives Transformation
Two vectors for one document
full_doc → E(full_doc) one vector, dimension d
summary(doc) → E(summary(doc)) one vector, dimension d
Both vectors have exactly the same dimension. What became smaller is the text — the number of words available to encode — not the vector. One embedding was produced from a full document; the other from a much shorter compression of it, passed through the identical embedding pipeline. This is not the dimensionality reduction Chapter 7 covered, and it is not the cross-space translation Part VI built — both the full document and its compression are embedded natively, by the same encoder, into the same space. The question is:
How much of the document’s representational geometry survives when the document’s text is compressed — and which failures can each available measurement actually see?
A high cos(E(full_doc), E(summary)) establishes one specific thing: the compressed text still lands close to the full document under this encoder’s notion of global similarity. It does not, on its own, establish factual entailment, preserved polarity, preserved numbers, preserved relation direction, preserved query behavior, or preserved calibration. This chapter exists because — as the measured demonstration below shows directly — high global proximity and serious local factual failure can coexist on the very same compressed text.
Global similarity tells you that the encoder still sees roughly the same broad object. It does not tell you that every assertion inside the object survived.
The converse deserves equal care. If the compressed embedding drifts — cos(E(full), E(summary)) falls noticeably — that drift could come from benign paraphrase, a register change, shortening itself, a shifted lexical balance, an actually dropped claim, or genuine corruption. Drift is evidence that the representation changed; it is not, by itself, a diagnosis of why. Diagnosing why requires another, more specific measurement — exactly the discipline Chapters 9, 14, 15, and 21 already established for a bare score in other contexts.
Compression, three ways
- Truncation. Keep the first
Ntokens. Cheap, lossy, biased toward the introduction. - Extractive summary. Select the most central sentences. Preserves original wording; may miss synthesis across sentences.
- Abstractive summary. A model rewrites the document into a denser form — key claims, entities, relations. Most compression, most risk of omission or fabrication, potentially the best geometric match if it correctly captures the document’s semantic center.
Terminology. “Cartridge” is used loosely in this chapter for any compact artifact meant to stand in for a document — a short summary, or its single embedding. The word also names a specific published technique (Eyuboglu et al., 2025): a trained KV cache distilled from a corpus by self-study, loaded at inference in place of putting the corpus in context. That is a materially richer object than one embedding vector and is out of scope here; this chapter asks specifically what a single embedding of a compressed text preserves.
Generic measurements one could use, and the ones this chapter actually measures
It is worth separating the wider space of things one could check from what Wave 4 actually persisted, exactly the way Chapter 22 separated generic bridge diagnostics from Wave 6’s measured fields.
Generic preservation measurements one could construct: direct drift (cos(E(full), E(summary))); neighborhood preservation (does the compressed embedding retrieve the same corpus documents as the full one?); query-answering preservation (a full Recall@k or nDCG comparison, full-document index versus compressed-document index, over a labeled query set); directional residual analysis (E(full) − E(summary), projected onto interpretable probes to guess what dropped); an information-retention curve across compression ratios.
What Wave 4 actually measured, on RELATE-DOC v0.1: whole-document cosine drift; a specific, asymmetric top-5 candidate-overlap statistic and a separate own-document top-1 recovery statistic (row 4.1); a five-layer blind-spot detection matrix built from global drift, own-document recovery, query-conditioned score drop, claim-conditioned score drop, and external NLI verification (rows 4.2–4.3); and a standalone whole-document-cosine faithfulness gate calibrated at its equal-error point (row 4.4). It did not measure a full Recall@k query-ranking comparison, and it did not analyze the residual vector E(full) − E(summary) in any way. Both remain genuinely useful ideas — and both are marked PROPOSED below, not reported as findings.
Demonstration, Part A: the retention curve
MEASURED on RELATE-DOC v0.1 — corpus hash
111cc2bb9008557d632061f98a0e4f847b91293e31bcdb2d78285a35de7f7914, embedding modelall-mpnet-base-v2. Artifactexperiments/embeddings-from-first-principles/wave4/artifacts/retention-curve.json.
RELATE-DOC v0.1 is a synthetic, designed corpus, and its documents are short — averaging approximately 78 words each. That scale matters directly to how these results should be read: local corruptions are relatively prominent in a 78-word document, and nothing here measures what happens to a 3,000-word report or a book chapter, where the same kind of localized change would occupy a much smaller share of the whole. Treat every result in this chapter as evidence about this specific, short, controlled corpus — not as a claim about production-scale documents, long reports, or arbitrary summarizers.
All seven measured compression conditions, exactly as persisted:
method cosine to full overlap@5 (asymmetric) own-source-doc top-1
abstractive 25% 0.9703 0.7556 1.0000
abstractive 10% 0.9308 0.7111 1.0000
extractive 50% 0.9103 0.7111 1.0000
extractive 25% 0.8924 0.6933 1.0000
truncation 50% 0.9503 0.7644 1.0000
truncation 25% 0.8881 0.7022 1.0000
truncation 10% 0.8457 0.6756 1.0000
Two things stand out immediately, and one of them needs a careful implementation audit before it can be read at all.
Own-source-document top-1 recovery is perfect across every single measured condition — even truncation to 10% of the original length. This is a direct instance of Chapter 22’s counterpart recovery, one level up: a compressed text retrieving its own full source document, among the full-document corpus, is exactly the compression-domain analogue of a translated vector retrieving its own paired native target. It succeeds completely here, at every compression ratio and every method tested — the cleanest, most robust result in this demonstration.
The overlap@5 field is not a conventional symmetric neighborhood-overlap statistic, and reading it as one overstates what it shows. The implementation builds the native candidate set as each full document’s top-5 nearest other full documents, with the document itself explicitly excluded. It builds the compressed candidate set as the compression’s top-5 nearest full documents — without excluding the compression’s own source document. Because own-document top-1 recovery is 1.0000 in every measured condition, that source document occupies one slot in the compressed candidate list that structurally cannot appear in the native list (which excluded it by construction). This caps the measured statistic’s practical maximum at 4/5 = 0.8 whenever self-recovery holds — so a value like 0.7556 or 0.7644 sits much closer to this implementation’s effective ceiling than a casual “76% of the neighborhood survived” reading would suggest. This is Chapter 21’s lesson again, applied to a new artifact: read the computation, not the friendly field name. A cleaner future metric would exclude the source counterpart from the compressed candidate list before comparing neighborhoods, exactly as Chapter 22’s agreement_at_10 does — that corrected metric is a real, worthwhile extension, and no such corrected value currently exists in this artifact.
No global ranking across methods is supported by seven points measured at unmatched ratios. Truncation at 50% (cosine 0.9503) beats extractive at 50% (0.9103); abstractive at 25% (0.9703) beats truncation at 25% (0.8881). Different compressors clearly preserve different things at different ratios in this small, controlled table — but the ratios are not matched cleanly enough across methods, nor sampled densely enough, to support declaring any one method universally superior. Treat this as a retention curve — evidence that preservation changes with method and ratio — not as evidence of a universal compression “knee.” Where the acceptable operating point actually sits depends on the consumer’s own requirement, the chosen preservation metric, the method, and the corpus; nothing here locates one fixed, corpus-independent knee.
The layered preservation stack, defined from its implementation
A single comparison cannot carry the weight of “did compression preserve the document,” so Wave 4 builds a stack of five increasingly conditioned checks — a layered preservation stack, not a scalar semantic checksum and not a formally nested detector hierarchy where each layer is guaranteed to catch everything the previous layer caught. Each layer answers a genuinely different, precisely scoped question, checked against the actual code rather than its field name.
flowchart TD
L1["L1 global similarity: is full<->compression cosine below the calibration band's lower bound?"] --> L2["L2 self-document recovery: does the compression still retrieve its OWN source full document at rank 1?"]
L2 --> L3["L3 query-conditioned drop: did any known query's cosine to the compression fall more than 0.15 below its cosine to the full doc?"]
L3 --> L4["L4 targeted-claim drop: did the specifically corrupted claim's cosine to the compression fall more than 0.08 below its cosine to the full doc?"]
L4 --> L5["L5 targeted-claim NLI: does an external entailment model judge the compression to still entail the original claim?"]
L1 — global similarity gate. cos(E(full_doc), E(compression)), flagged when it falls below the lower bound of a calibration band derived from a faithful-vs-own-document distribution and a faithful-vs-other-document distribution (the band is [0.3652, 0.8652] on this benchmark — derivation below). This is a deliberately coarse, permissive whole-document check intended to flag only very large representational drift. The faithful-vs-other-document distribution motivates that role, but Wave 4 does not separately report a wrong-document-swap or generic large-topic-drift detection rate, so do not turn the threshold’s design intent into an additional measured result.
L2 — self-document recovery. The historical artifact field is named L2_neighborhood, and that name is misleading enough to correct explicitly. The implementation computes argmax_j cos(E(compression), E(full_doc_j)) ≠ source_document_id — it asks only whether the compression’s single nearest full document, among the entire corpus, is still its own source document. This is not top-k neighborhood overlap in any conventional sense; it is a stricter, single-document counterpart-recovery check, closely related to row 4.1’s own_doc_top1 field but computed as part of the blind-spot detector rather than the retention curve. Call it L2 — source-document recovery (historical field: L2_neighborhood) and never read it as measuring whether the document’s broader neighborhood survived.
L3 — query-conditioned score-drop probe. For every query known to be associated with the document, compute cos(E(query), E(full)) − cos(E(query), E(compression)); flag if any single drop exceeds 0.15. This is not question-answering correctness, not Recall@k, and not proof that a question can no longer be answered from the compressed text — it is a query-conditioned embedding-score-drop detector, and nothing stronger.
L4 — targeted-claim score-drop probe. For the specific atomic claim the corruption targeted, compute cos(E(claim), E(full)) − cos(E(claim), E(compression)); flag if the drop exceeds 0.08. This is not entailment and does not establish that the claim survived — it is a claim-conditioned embedding-score-drop detector, and the demonstration below shows precisely why that distinction matters: a claim can retain high similarity to a compressed text while asserting something false.
L5 — targeted-claim external NLI verification. Using cross-encoder/nli-deberta-v3-base with the compressed text as premise and the original targeted claim as hypothesis, flag if the entailment probability falls below 0.5 or the contradiction probability exceeds 0.4. This is a genuinely different kind of signal from the first four — a separate, assertion-sensitive model, not another view of the same embedding geometry — and it is the layer that actually distinguishes this stack’s coarse checks from its fine one.
Every one of these five checks is source-grounded, not reference-free in any strong sense: L3 depends on the document’s known queries, L4 and L5 depend on the document’s own atomic claims. None of them requires a separately authored human reference summary, but all of them lean on information extracted from the source document itself. “Reference-free faithfulness gate” is a reasonable informal label only if that scope is kept in view.
Demonstration, Part B: the blind-spot matrix
MEASURED on RELATE-DOC v0.1, Wave 4 rows 4.2–4.3 — artifact
experiments/embeddings-from-first-principles/wave4/artifacts/blindspot-matrix.json.
RELATE-DOC gives each document a faithful control summary plus seven controlled corruption families, each targeting one named atomic claim in a specific, localized way: number_dropped, number_changed, relation_reversed, negation_inserted, minority_entity_dropped, temporal_value_shifted, conclusion_changed. The five-layer stack runs against every instance:
corruption n L1 L2 L3 L4 L5
faithful (control) 45 0.000 0.000 0.000 0.000 0.000
number_dropped 35 0.000 0.000 0.057 0.314 1.000
number_changed 35 0.000 0.000 0.000 0.000 0.886
relation_reversed 25 0.000 0.000 0.000 0.000 1.000
negation_inserted 45 0.000 0.000 0.000 0.044 1.000
minority_entity_dropped 15 0.000 0.000 0.000 0.600 1.000
temporal_value_shifted 35 0.000 0.000 0.000 0.114 1.000
conclusion_changed 45 0.000 0.000 0.000 0.400 1.000
Before reading the pattern, one crucial design fact has to sit alongside it: the numeric thresholds for L1, L3, and L4 were selected to produce zero flags on the faithful control. L2 has no tuned threshold — it is the fixed top-1 self-document-recovery test — and it also happens to produce zero faithful-control flags in this run. The zero row is therefore partly a designed condition for the thresholded embedding probes, not an independently surprising discovery. The informative question this matrix actually answers is: with those thresholded probes deliberately set to leave faithful compressions alone, and L2 evaluated by its fixed recovery rule, which corruption families still become visible, and to what degree?
L1 and L2 detect nothing, for every corruption family, at 0.000. Under L1, no corrupted compression falls below the coarse global-drift threshold. Under L2, every corrupted compression still retrieves its own full source document at rank 1. Those are two different reasons for the same 0.000 detection rate; neither measurement is conditioned on the specific claim that was altered.
L3 registers almost nothing — a single nonzero cell, number_dropped at 0.057. A dropped number occasionally shifts a known query’s cosine enough to cross the 0.15 threshold. For every other corruption family, no instance crosses that flag threshold. The experiment does not establish that those query scores were literally unchanged; only that any changes remained below the detector’s cutoff.
L4 is genuinely uneven, not a clean omission-versus-mutation partition. Reading the exact per-class rates: minority_entity_dropped 0.600, conclusion_changed 0.400, number_dropped 0.314, temporal_value_shifted 0.114, negation_inserted 0.044, number_changed 0.000, relation_reversed 0.000. The intuitive story — “L4 catches deletions, misses mutations” — is a reasonable first approximation but not the exact result: temporal_value_shifted and negation_inserted, both mutations rather than deletions, still register small nonzero detection rates, while conclusion_changed (a rewrite, not a clean deletion) registers a substantial 0.400. The defensible statement is narrower than a clean taxonomy: L4 is strongest on omissions and edits that substantially move the targeted claim’s own embedding away from the compressed text, and weak or entirely blind whenever the altered claim remains lexically and semantically close to the original — which is exactly what happens when “Helios acquired Pine” becomes “Pine acquired Helios,” or when a single digit changes in an otherwise identical sentence.
L5 is the only layer with strong detection across every corruption family — but not perfect detection of every instance. Six of the seven families sit at 1.000; number_changed sits at 0.886, meaning this external NLI verifier missed roughly 11.4% of changed-number instances in this run. State the result at exactly this precision: L5 is the only measured layer that detects every corruption class strongly, with per-class detection ranging 0.886–1.000, and zero flags among the 45 faithful controls in this run. That last clause is a real and encouraging result on this sample — it is not proof of a zero false-positive rate in any broader or production setting, and no uncertainty analysis (bootstrap, repeated sampling, a second corpus) exists to extend it beyond this one measured run. Some corruption classes carry small sample sizes worth keeping in mind — minority_entity_dropped has only n = 15 — which is one more reason to read every rate here as a point estimate from one controlled experiment, not a stable population parameter.
The failure is not that embeddings are useless. It is that broad relatedness and assertion-level faithfulness are different properties, and this benchmark’s local corruptions are specifically the kind of change broad relatedness is built to average away.
Omission versus assertion mutation
The blind-spot matrix draws a useful line between two broad kinds of local failure, worth naming because they stress different detectors.
Omission — the compressed text no longer provides a claim the source document made. Dropping a minority entity or a number are the clearest measured examples. These can move the targeted claim’s own embedding away from the compression, which is one reason L4 sometimes catches them.
Assertion mutation — the compressed text still contains a claim-shaped statement but changes a value, polarity, or relation direction. A changed number and an inserted negation can directly contradict the original claim; a relation reversal changes who did what to whom and can make the targeted assertion incompatible without being its literal logical negation. These edits can nevertheless leave the targeted claim’s embedding almost untouched because the corrupted sentence remains structurally and lexically close to the original — “Pine acquired Helios” can sit nearly on top of “Helios acquired Pine” in embedding space even though the role assignment changed.
L4 has useful sensitivity to some omissions and some mutations, but it is weak or completely blind on several assertion mutations in this benchmark. L5’s external, assertion-sensitive model is much more sensitive across those families in this run — and its purpose is precisely to add a different kind of signal, not simply a more finely conditioned embedding comparison.
Two different threshold experiments — do not conflate them
Two numbers in this chapter both look like “the” cosine threshold for faithfulness, and they are not the same experiment, do not share a purpose, and must never be substituted for each other.
Row 4.2’s L1 threshold, 0.3652, is the lower bound of a calibration band derived from two distributions: faithful-summary-to-own-document cosine (mean 0.9703) and faithful-summary-to-other-document cosine (mean 0.1220). It defines a deliberately permissive detector, tuned to flag only very large drift while leaving the faithful control entirely alone — and under that detector, L1’s measured detection rate is 0.000 for every corruption family. This threshold answers: is the compression so far from the original that it might be the wrong document, or catastrophically broken?
Row 4.4’s equal-error threshold, 0.9856, comes from a completely separate calibration experiment: distinguishing 45 faithful summaries from 150 selected topic-preserving corruptions, purely by whole-document cosine, at the threshold where false-accept and false-reject rates meet.
faithful mean cosine 0.9703
topic-preserving-corruption mean cosine 0.9578
equal-error threshold 0.9856
FAR (corruption passes as faithful) 0.3533
FRR (faithful summary rejected) 0.3556
FAR here is the fraction of topic-preserving corruptions that pass the gate as faithful; FRR is the fraction of genuinely faithful summaries the same gate wrongly rejects. At the measured equal-error operating point they are roughly 35.3% and 35.6%. That is poor balanced discrimination for a standalone faithfulness gate. A consumer with highly asymmetric error costs could choose a different threshold, but it would need to report the resulting FAR/FRR tradeoff rather than treating the equal-error point as a universal operating policy.
This is not a contradiction with L1’s 0.000 detection rate above — it is a different experiment answering a different question, at a different threshold, over a different population. L1’s result says a deliberately permissive, faithful-control-calibrated drift gate flags nothing. Row 4.4’s result says the equal-error threshold found by this finite threshold sweep still misclassifies roughly a third of each class. It is not the “best achievable” threshold under every possible application objective; it is the measured point where the two error rates are closest. State the correct, precise conclusion rather than either extreme:
Whole-document cosine has some real distributional signal — the faithful and corruption means genuinely differ,
0.9703versus0.9578— but its balanced standalone discrimination is poor on this benchmark. It is not literally zero information, and any operational gate would still require an application-specific error tradeoff and validation.
Cross-wave connections: a recurring boundary, not one proven mechanism
Three separate experiments, at three separate points in this book, land on a version of the same finding without ever isolating one shared causal mechanism between them:
- Wave 1 (Chapter 11) found that negation and paraphrase can sit at nearly identical cosine similarity under a raw scorer.
- Wave 3 (Chapters 18–21) found that a fitted cross-space bridge can preserve broad structure — a high relation-profile Pearson correlation — while a specific relation contrast, paraphrase minus negation, reverses sign.
- Wave 4, here, finds that whole-document and claim-conditioned embeddings can miss a reversed relation or a changed value inside a compressed document, while every embedding-based layer stays silent.
These are three experiments showing the same recurring boundary, not three independent confirmations of one isolated causal mechanism — no experiment in this book traces why embedding geometry behaves this way across all three settings, only that it repeatedly does. The durable, book-wide statement these three results jointly support is:
Embedding geometry is systematically better at broad relatedness than at assertion-sensitive semantics — role direction, polarity, and exact values in particular.
External faithfulness-evaluation research motivates the same pattern from outside this book’s own artifacts: localized factual errors — a swapped entity, a changed number, a reversed relation, a negated clause — are produced by local edits that move a document-level embedding very little (Kryscinski et al., 2020), and question-conditioned checking (Wang et al., 2020) together with sentence-level NLI aggregation (Laban et al., 2022) recover much of what document-level similarity misses. That literature motivates the layered design; it does not replace Wave 4’s own measured, mpnet-specific numbers, and Wave 4’s result should not be inflated into a claim about embedding models in general — only all-mpnet-base-v2 was measured here, and nothing in this chapter establishes that the blind-spot pattern is model-independent.
A compressed artifact is equivalent-for-an-operation, not equivalent
This is the Part VII counterpart to Chapter 17’s lesson that compatibility is property-specific, and it deserves the same weight here. A compressed artifact is never simply “equivalent to the document.” A summary that preserves self-document recovery perfectly, and whole-document similarity almost perfectly, can still fail a consumer that needs the exact acquisition direction, or the exact figure, or the polarity of a claim.
A compressed artifact is never simply “equivalent to the document.” It is equivalent-for-an-operation, under a stated preservation contract.
A consumer doing topic browsing may tolerate some local losses that a factual-question-answering consumer cannot, while another topic-browsing policy may still require exact entities or dates. Neither judgment follows automatically from the operation’s name or from any single layer above — the consumer has to state the requirement, then check the layer or layers that measure it.
Two further systems consequences carry over directly from earlier chapters and are worth stating plainly rather than assuming.
A compressed index is a different corpus, even under the identical embedding model. Replacing every full document in an index with its compressed counterpart changes the corpus population, per Chapter 17’s discipline: full_corpus_hash ≠ compressed_corpus_hash. Neighborhood distributions, retrieval scores, and any calibrated operating point built against the full-document index do not automatically transfer to the compressed one. The new corpus identity invalidates calibration carryover by assumption: revalidate the old operating point on the compressed corpus, and recalibrate or refit only if it no longer satisfies the consumer’s requirement. The resulting evaluation observation remains bound to the new corpus identity.
Shorter text does not automatically make vector search cheaper. If both the full document and its compression produce one embedding of the same dimension, and the index still stores one vector per document, the vector count and vector width are unchanged — ANN index storage and per-vector similarity cost are not automatically reduced by compression. What compression can plausibly reduce is text storage, embedding-time token cost, downstream context size, and re-reading cost — none of which Wave 4 measured directly. Whether an index built from compressed text is actually cheaper to search depends on whether the system also reduces the number of stored vectors, which is a separate design choice this chapter’s evidence does not speak to.
What this chapter establishes and what it does not
Establishes: content compression before embedding produces two same-dimensional vectors from texts of very different length, and the Part VI preservation vocabulary — name the reference, name the property, measure it — applies directly; a high cos(E(full), E(summary)) establishes coarse embedding proximity, not factual faithfulness, and drift alone is evidence of change without diagnosing its cause; across seven measured compression conditions on RELATE-DOC v0.1, own-source-document top-1 recovery holds perfectly (1.0000) at every ratio and method, including truncation to 10%, while overlap@5 uses an asymmetric candidate construction whose practical ceiling is 0.8, not 1.0, under these conditions; the five-layer preservation stack — global similarity, self-document recovery, query-conditioned drop, claim-conditioned drop, external NLI verification — is a profile of distinct sensors, not a nested hierarchy or a scalar checksum; with L1, L3, and L4 thresholds selected to leave the faithful control at zero flags, and L2 evaluated by its fixed top-1 recovery rule, L1 and L2 detect none of the seven measured corruption families, L3 detects almost none (a single nonzero cell at 0.057), L4 responds unevenly across corruption families (0.000–0.600), and L5 detects every family strongly (0.886–1.000) with zero flags on the faithful control in this run; a separate equal-error calibration experiment produces FAR 35.3% and FRR 35.6%, showing poor balanced standalone discrimination by whole-document cosine; and this pattern echoes, without mechanistically explaining, similar boundaries found in Waves 1 and 3.
Does not establish: that global geometric preservation guarantees factual faithfulness — the measured counterexamples show that it is insufficient, not that geometric preservation and faithfulness are opposites; a corrected, symmetric neighborhood-overlap result (none was computed); measured query-answering correctness, Recall@k, or task-ranking preservation (L3 measures only query-conditioned cosine drop); a universal ranking among truncation, extractive, and abstractive compression; a universal compression “knee” independent of consumer, corpus, and method; that any embedding model besides all-mpnet-base-v2 behaves this way; that L5’s external NLI verifier is a truth oracle, catches every instance, or generalizes beyond this run’s zero-faithful-flag observation; that the equal-error threshold is optimal for every application-specific error tradeoff; that automatic vector-search savings follow from shorter text; or that any measured preservation profile, by itself, authorizes replacing a source document with its compression for any named consumer.
Lab 24: reproduce the retention curve, the blind-spot matrix, and the drift-gate calibration
MEASURED — artifacts
wave4/artifacts/retention-curve.json,blindspot-matrix.json,drift-gate-calibration.json(RELATE-DOC v0.1,all-mpnet-base-v2,cross-encoder/nli-deberta-v3-base). REPRODUCIBLE —python run_wave4.py.
Question. How far can a document be compressed before the compression stops behaving like a faithful stand-in — and which specific failures does each available layer actually catch?
Part A — retention curve. Reproduce all seven compression conditions exactly:
method cosine overlap@5 own-doc top-1
abstractive 25% 0.9703 0.7556 1.0000
abstractive 10% 0.9308 0.7111 1.0000
extractive 50% 0.9103 0.7111 1.0000
extractive 25% 0.8924 0.6933 1.0000
truncation 50% 0.9503 0.7644 1.0000
truncation 25% 0.8881 0.7022 1.0000
truncation 10% 0.8457 0.6756 1.0000
State plainly: own-document top-1 recovery is 1.0 at every condition — the cleanest measured result. State equally plainly: overlap@5 compares an asymmetric candidate construction (native top-5 excludes self; compressed top-5 includes the source document, which occupies a top rank in every condition), so its practical ceiling here is 0.8, and reading 0.76 as “76% of a conventional neighborhood survived” overstates it.
Part B — blind-spot matrix. Reproduce the full seven-family-plus-faithful-control table with all five layers, exactly as persisted above. Classify precisely: L1 and L2 at 0.000 across every family; L3 nonzero only for number_dropped (0.057); L4 ranging 0.000–0.600, strongest on minority_entity_dropped and weakest on number_changed/relation_reversed; L5 ranging 0.886–1.000 with 0.000 on the faithful control. State explicitly: the numeric thresholds for L1, L3, and L4 were chosen to zero out the faithful control; L2 is a fixed top-1 recovery test with no tuned threshold and also yields zero faithful-control flags in this run. The informative question is what remains visible under those stated detector rules, not an independently surprising absence of false positives.
Part C — drift-gate calibration. Reproduce the separate row-4.4 experiment: n_faithful = 45, n_corruptions = 150, faithful mean cosine 0.9703, topic-preserving-corruption mean cosine 0.9578, equal-error threshold 0.9856, FAR 0.3533, FRR 0.3556. State explicitly that this threshold (0.9856) and row 4.2’s L1 threshold (0.3652) come from two different experiments with two different purposes and must never be treated as competing estimates of one quantity.
Part D (PROPOSED — no artifact backs this) — production follow-ups. Repeat across additional embedding models; repeat on longer, natural (non-templated) documents; bootstrap per-corruption detection rates to obtain uncertainty; build a corrected, symmetric neighborhood-overlap metric that excludes the source counterpart from the compressed candidate list; measure actual query-ranking or QA-correctness preservation with a real relevance-labeled query set; evaluate a domain-specific verifier alongside the general-purpose NLI model; evaluate multiple compression generators head-to-head at matched ratios; and define consumer-specific stand-in acceptance bars, none of which exist in the current evidence base.
Try it yourself
Compress a small set of your own documents at several ratios and methods, and reproduce the same three-part structure: a retention curve (global cosine, a corrected symmetric neighborhood-overlap statistic, and own-document recovery), a blind-spot matrix built from corruptions realistic for your domain, and a separately calibrated whole-document drift gate. Before trusting any layer’s zero-detection result, check whether its threshold was tuned against the same faithful control you are now evaluating — a detector calibrated to avoid false positives on one population says nothing about a different one.
Companion component: the compression preservation observation
This artifact records evidence only — it does not decide whether a compressed artifact may replace its source for any consumer. That decision is Chapter 20’s, built on top of this evidence, never emitted from inside it.
compression_preservation_observation:
compression_id:
source:
source_doc_id:
source_doc_hash:
compressed_artifact:
text_hash:
method: truncation | extractive | abstractive
ratio:
generator_ref:
representation:
space_hash:
full_corpus_hash:
compressed_corpus_hash: # a DIFFERENT corpus, even under the same space_hash
evaluation_contract_ref:
measurements:
global_similarity:
cosine:
calibration_band_ref: # row 4.2's band, [0.3652, 0.8652] on this benchmark
source_doc_recovery:
own_doc_top1:
historical_field_name: "L2_neighborhood — misleading; this is self-recovery, not neighborhood overlap"
neighborhood_observation:
metric: overlap_at_5
k: 5
implementation_note: "asymmetric candidate sets — native top-5 excludes self, compressed top-5 does not; practical ceiling 0.8 when self-recovery holds"
value:
query_conditioned:
threshold: 0.15
per_query_score_drop:
claim_conditioned:
threshold: 0.08
per_claim_score_drop:
external_verification:
verifier_id: cross-encoder/nli-deberta-v3-base
per_claim_probabilities:
flag_rule: "flag if entailment_probability < 0.5 OR contradiction_probability > 0.4"
derived_status: supported | contradicted_or_unsupported | not_evaluated
blindspot_profile_ref:
drift_gate_calibration_ref: # row 4.4 — a SEPARATE experiment from the L1 threshold above
uncertainty:
bootstrap: false
repeated_corpora: false
repeated_models: false
provenance:
artifact_refs:
code_hash:
Authorization is a genuinely separate object, referencing this observation rather than being derived automatically from it:
standin_authorization:
consumer:
operation: # e.g. topic browsing, factual QA context, search routing, compliance archive
requirement_ref:
evidence_ref: compression_preservation_observation
decision: ALLOW | CONDITIONAL | DENY | NOT_EVALUATED
No single formula — “L2 clears a bar AND no L4 claim lost AND no L5 contradiction” — applies universally across every consumer. A search-routing consumer may care most about self-document recovery and neighborhood behavior; a factual-QA consumer may require claim coverage and verified entailment; a compliance archive may tolerate almost no semantic change at all; a topic-browsing consumer may tolerate some local factual losses, but only if its declared requirement actually permits them. Each of these is a different standin_authorization entry, over the same underlying compression_preservation_observation, decided under Chapter 20’s requirement-plus-evidence machinery — never a single hard-coded rule baked into the measurement itself.
Failure modes
- Calling textual compression vector compression. The embedding vectors are the same width; only the source text got shorter.
- Calling high whole-document cosine factual faithfulness. It establishes coarse embedding proximity, and this chapter’s own blind-spot matrix is the direct counterexample.
- Treating drift as diagnosis. A moved embedding says the representation changed; it does not say why, or what specifically was lost.
- Calling row 4.1’s
overlap@5a conventional symmetric neighborhood-overlap metric. Its native and compressed candidate sets are built asymmetrically, capping its practical ceiling at0.8here. - Calling row 4.2’s
L2_neighborhoodfield neighborhood preservation. It is own-source-document top-1 recovery — a strict, single-document counterpart check. - Calling L3 question-answering correctness. It is a query-conditioned cosine-drop probe with a fixed
0.15threshold, nothing more. - Calling L4 claim survival. It is a claim-conditioned cosine-drop probe with a fixed
0.08threshold; the experiment itself shows a claim can retain high similarity while asserting the opposite. - Calling the NLI layer truth verification. It is one external model, with its own thresholds and its own measured miss rate (
0.886onnumber_changed). - Claiming L5 caught every corrupted instance. It missed roughly 11.4% of
number_changedcases in this run. - Treating zero faithful-control flags as broad, calibrated evidence of a low production false-positive rate. The numeric thresholds for L1, L3, and L4 were selected to produce that zero; L2 has no tuned threshold and simply recovers the faithful source document top-1 in this run. L5’s zero is an observed result under its stated verifier rule, not a guarantee.
- Conflating row 4.2’s L1 threshold (
0.3652) with row 4.4’s equal-error threshold (0.9856). Two different experiments, two different purposes. - Saying whole-document cosine carries no signal at all. The faithful and corruption means genuinely differ (
0.9703vs0.9578); the problem is poor standalone separability, not zero information. - Claiming every corruption sits at cosine
0.97. No such per-class global-cosine table is persisted; only detection rates are. - Forcing L4’s results into a clean deletion-versus-mutation taxonomy.
temporal_value_shiftedandnegation_insertedare mutations with nonzero detection;conclusion_changedis not a clean deletion and still scores0.400. - Generalizing the blind-spot pattern beyond
all-mpnet-base-v2. Only one embedding model was measured here. - Putting
usable_as_standininside the empirical observation. That is an authorization decision, belonging to a separate Chapter 20 trace referencing this evidence. - Assuming a compressed index automatically saves vector-search cost. Same vector count and dimension means no automatic ANN saving. Shorter text may reduce text storage, embedding-time token cost, or downstream context cost, but Wave 4 did not measure those savings either.
- Treating the seven measured retention points as a universal compression knee. The right operating point is consumer-, metric-, method-, and corpus-specific.
What this chapter established
- Content compression before embedding produces two same-dimensional vectors from texts of very different length: the vector did not shrink, the source text did. High whole-document cosine establishes coarse embedding proximity, not factual faithfulness — and drift is evidence of change, not a diagnosis of its cause.
- The five-layer preservation stack is a profile of distinct sensors — global similarity, self-document recovery, query-conditioned drop, claim-conditioned drop, external verification — not a scalar checksum and not a guaranteed monotonic hierarchy. Each layer’s friendly name overstates what it actually computes.
- With thresholds deliberately set to leave the faithful control at zero flags, the embedding layers see almost nothing. L1 and L2 detect
0.000of every corruption family; L3 reaches0.057on one family alone; L4 is uneven — strongest on a dropped minority entity (0.600), completely blind on changed numbers and reversed relations (0.000). Only the external verifier detects every family (0.886–1.000), and even it is not perfect on every instance. - Whole-document cosine carries real distributional signal and poor standalone discrimination — the calibrated gate runs at FAR
35.3%/ FRR35.6%. The pattern echoes, without mechanistically explaining, the boundary Waves 1 and 3 measured: embedding geometry is systematically better at broad relatedness than at role direction, polarity, and exact values. - A compressed artifact is never simply “equivalent to the document.” It is equivalent-for-an-operation, under a contract naming which layer the consumer actually depends on. A compressed index is a different corpus even under an identical
space_hash, shorter text does not automatically reduce vector-index cost, and the preservation observation stays separate from thestandin_authorizationthat decides replacement.
Next
We compared a document to its own compression, using the difference between two embeddings only as evidence that something changed — never as a diagnosis of what changed or as a reusable object in its own right. The next chapter asks the more dangerous question directly: can the difference between two vectors, E(after) − E(before), be reused as a transformation rule on entirely new content? Chapter 24 treated a residual only as a symptom. Chapter 25, From Deltas to Operators, asks whether that symptom can be turned into a tool.