← Embeddings From First Principles

Retrieval Is a Policy

A retrieval system is not an embedding model. It is a versioned composition: representation, eligible candidates, scoring and fusion, selection, execution strategy, and context assembly. Measure a three-policy ablation on RELATE, discover that its hard-negative metric never observes the reranking stage, and learn the rule that generalizes: an ablation is only valid if the metric passes through the stage you changed.

Part III — Retrieval Is an Experiment

“Just retrieve the relevant documents”

There is no such operation. Chapter 11 closed on a single insight: a score never travels alone — it arrives with a negative-selection rule that is part of the measurement. This chapter widens the frame. A retrieval result never travels alone either. It arrives wrapped in a query construction, a representation, a candidate-eligibility rule, a scoring and fusion policy, a selection rule, an execution strategy, and a context-assembly step that decides what a downstream model actually sees.

A retrieval result is not produced by a model. It is produced by a model inside a policy.

Reusing the vocabulary Chapters 9–11 already built rather than inventing a new one:

    flowchart TD
    Q[query] --> QC["query construction — raw / rewrite / expand / decompose / HyDE"]
    QC --> RE["representation — model, normalization (Parts I–II)"]
    RE --> CE["eligible candidate universe — authorization, filters, corpus/index version"]
    CE --> SC["scoring / fusion policy — metric(s), how multiple scores combine"]
    SC --> SEL["selection policy — candidate depth N, threshold, tie rule, final k"]
    SEL --> RC[retrieved candidates]

    CE -.->|"execution strategy attempts to compute the scoring/selection policy (Chapter 9)"| EX["exact scoring | ANN | compressed | distributed"]
    EX -.-> SC

    RC --> CA["context-assembly policy — dedup, ordering, token budget, citations"]
    CA --> DM["context supplied to the downstream model"]
  

Read this as one concrete policy graph, not the definition of every retrieval system. Several of its choices could be made differently and still be exact: authorization could run before or after ranking; a threshold could apply before or after a second-stage score; a query could be decomposed into several sub-queries whose candidate sets merge before scoring. The ordering and composition of these operations is itself part of the policy, not a detail beneath it.

Four families of decision are worth separating explicitly:

  1. Choices that define the eligible information — which documents could possibly be returned (authorization, filters, corpus and index version).
  2. Choices that define the scoring and ranking objective — the target ordering a perfect, exact computation would produce.
  3. Execution mechanisms that attempt to compute that objective — exact scoring or an approximation of it (Chapter 9’s fidelity-versus-policy distinction, unchanged here).
  4. Post-retrieval decisions about what context downstream actually receives — de-duplication, ordering, budget, citation attachment.

Changing a link changes the policy or its execution, and may change the resulting candidates or context — not automatically will. An intervention that leaves the ranking identical is itself evidence, not a null result to discard (Chapter 6’s discipline again).

One more correction to a phrase that is easy to overreach with: the context handed to a downstream model is not “the system’s memory.” A system may also carry conversation state, persistent memory, tool outputs, or user-supplied context that never passed through this retrieval chain at all. Call what this chapter produces the retrieval-supplied context, or the evidence this retrieval path presents downstream — a precise, bounded object, not a synonym for everything the system knows.

Three time scales

Not every stage in the diagram lives at the same lifecycle. Separating them clarifies what “changing the policy” actually costs, and what rolling it back actually restores.

  • Build-time / corpus-time. Chunking, document preprocessing, the embedding model and version used to produce the corpus vectors, sparse-index construction, ANN graph construction, index version. Changing these usually means rebuilding artifacts, not flipping a runtime flag.
  • Query-time. Query transform and representation, candidate eligibility, the fusion and ranking rule, candidate depth, any second-stage scorer, thresholds, final k, tie rules.
  • Context-time. De-duplication, ordering, token budget, citation attachment, grouping or compression of the retrieved set.

Chunking is the clearest case of a decision that looks like a simple parameter but is not: it changes the corpus representation, which means it changes the dense and sparse indexes built from it. A policy that references “chunk size = 256” is really referencing a specific chunked-corpus artifact and everything built on top of it — and rolling that policy back means restoring that corpus artifact, not just a config value.

Each stage is a hypothesis, not a law

It is tempting to state, for each stage, what it generally does to quality. Resist that — Chapter 10 already showed what happens when a plausible-sounding second stage is trusted without measurement. Better to record, per stage, the decision it changes, the failure it plausibly risks, and what to measure before trusting it:

StageWhat it changesPlausible riskWhat to measure
Chunkingthe unit that gets embedded and indexedtoo large dilutes a match and wastes budget; too small splits an answer across chunks that neither retrieves alonerecall and hard-negative behavior across chunk sizes, on the actual corpus
Query transform (rewrite / expand / decompose / HyDE)the query representation(s) and how their results combinea rewrite can broaden or narrow intent unpredictably; a multi-query aggregation can dilute or amplify a signalbefore/after recall and hard-negative behavior on the same query set
Candidate eligibilitywhich items can be returned at allan authorization bug can leak or hide content; a stale filter can silently shrink the poolaudit that eligible ≠ intended, independent of ranking quality
Scoring / fusion (e.g. dense + sparse)the target ordering itselfa fixed blend weight and mismatched score scales can move quality either directionfull-corpus and hard-negative metrics before and after, on the identical query set
Second-stage scorerre-scores or re-fuses some or all candidatesits training objective and input orientation may not match the distinction you need (Chapter 10)per-distinction, per-query-style evaluation against the exact decision the first stage got wrong
Selection (threshold, candidate depth, final k, tie rule)how much survives, and from how deep a poola depth cutoff can exclude the item a later stage would have promoted; a threshold set by feel rarely transfers (Chapter 14)score-distribution-based calibration, not intuition
Diversity / de-duplicationwhich near-duplicates coexist in the final settoo little wastes slots on repeats; too much can drop a second distinct true answercoverage of distinct correct answers, not just relevance score
Context assemblyordering, budget, citationsmore retrieved text is more tokens and more distractors, not automatically more knowledgedownstream task quality, not retrieval score alone

Reranking deserves its own note, because it is the stage a reader is most tempted to trust on architecture alone. A cross-encoder can jointly attend over a query-candidate pair, which gives it capacity that a bi-encoder cosine lacks. Chapter 10 measured what that capacity actually bought: the same kind of NLI-trained cross-encoder improved one distractor type and worsened four others. Cross-encoder architecture is not a validated reranker. A second-stage scorer is another decision function the policy has adopted, and it earns its place the same way every other stage does — by measurement against the specific distinction it is meant to fix.

Policy, provenance, execution, trace, evaluation

A policy is not the whole story of what ran. Keep five things distinct:

  • Policy — the intended decisions: which stages run, in what order, with what parameters.
  • Provenance — the concrete, versioned artifacts the policy is bound to: corpus version, chunked-corpus identity, embedding-space version, sparse-index version, ANN index build, second-stage model revision, implementation version.
  • Execution — what actually computed the policy for a given query (exact or approximate; Chapter 9).
  • Trace — a record of what happened for one query: the eligible-candidate count, the candidate set after each stage, the scores at each stage, the final selection.
  • Evaluation — whether the resulting behavior was any good, against a stated metric and workload.

A policy plus its provenance is closer to a deployable unit than a policy alone. “The same policy” run against a re-embedded corpus, a rebuilt ANN index, or an updated second-stage model revision is not the same system, even if every parameter value in the policy object is unchanged. And a policy is not exhaustively described by parameters plus the referenced model or space records: observed behavior also depends on the runtime artifacts above and on the workload the queries come from. Two teams can bind an identical policy object to different corpus builds and get different retrieval behavior — the policy did not “fully describe” the outcome; it was one necessary input.

This also repairs a tempting overcorrection. It is not that the model does not matter and only the policy does — the embedding model and space themselves need versioning, evaluation, and rollback exactly as the policy does. Nor does the embedding model simply “set the ceiling”: once a policy introduces an independent signal — BM25, a second encoder, a cross-encoder, metadata, a verification step — that signal can distinguish something the base embedding’s cosine never separated on its own, for better or worse. A model identifier alone is not a retrieval-system version. The deployable, versioned unit is closer to policy + the spaces, scorers, and corpus/index artifacts it references.

Demonstration: same corpus, three policies

MEASURED on RELATE v0.1, Wave 1 row 1.13 — artifact experiments/embeddings-from-first-principles/wave1/artifacts/policy-ablation.json. Model all-mpnet-base-v2 held fixed throughout.

The intuitive prediction: adding a lexical signal, and then a second-stage scorer with real capacity to attend over the pair, should recover quality a plain dense ranking misses. Three policies test that prediction on the same 269 queries and the same 1,173 items.

Policy A — dense. Rank every item by cosine to the query (the harness’s L2-normalized vectors make dot product equal cosine, Chapter 4). Measured: nDCG@10 (all) = 0.9518, nDCG@10 (hard-negative pool) = 0.9577.

Policy B — a specific fusion, not “hybrid search” in general. For each query, every item’s score becomes 0.6 × dense_cosine + 0.4 × (BM25 / max BM25 for this query). Items absent from the query’s BM25 results get zero sparse contribution. This is a fixed linear blend with per-query max normalization — it does not calibrate the two score distributions against each other, and the 0.6/0.4 weighting was not tuned for this evaluation. Measured: nDCG@10 (all) = 0.9255, nDCG@10 (hard-negative pool) = 0.9288. Both metrics are lower than dense alone. Several things could explain that: a genuine lexical-signal mismatch, the untuned weight, the max-only normalization, the mixed query styles, or an interaction among these. The artifact does not isolate which. Resist the tempting shorthand “BM25 injected lexical noise” — that is one hypothesis among several, not a measured finding.

Policy C — hybrid top-50, fused with an NLI entailment score, not a pure rerank. This is where precision matters most. The implementation takes policy B’s ranking, keeps the top 50 (the candidate depth, N = 50 — distinct from the final k = 10 that nDCG@10 evaluates), scores each of those 50 pairs (query, candidate) with the same NLI cross-encoder used in Chapter 10, and computes 0.5 × hybrid_score + 0.5 × entailment_probability. Only those 50 items are reordered by the fused score; items ranked 51st and below keep their hybrid-policy order. So policy C is a 50/50 score fusion of the hybrid score and an NLI entailment probability, applied to the top 50 of the hybrid ranking — not “B plus an NLI rerank” in the sense of letting the NLI model decide the order outright. Measured: nDCG@10 (all) = 0.8560, lower again.

Now the suspicious number. Policy C’s nDCG@10 (hard-negative pool) is reported as 0.9288 — identical to policy B’s, to four decimal places. Before concluding “the NLI stage had zero effect on hard-negative discrimination,” inspect how that number was actually computed.

The implementation scores the graded candidate pool — the positives and hard negatives with a relevance grade for that query — by re-sorting it with a score dictionary sc. For policy C, sc is populated by the hybrid fusion step and is never updated by the top-50 NLI re-fusion; the NLI-fused order lives only in a separate ranked list used for the all-corpus metric. Because the hard-negative metric sorts the graded pool by sc, and sc is identical between policy B and policy C, the hard-negative nDCG for policies B and C is guaranteed to be equal by construction, regardless of what the NLI stage does. The persisted field hardneg_gain_dense_to_rerank = -0.0289 is 0.9288 − 0.9577 — the dense-to-hybrid difference on this metric, arrived at through the “rerank” label. It does not isolate the NLI stage’s effect on hard-negative ranking at all.

policy                                            nDCG@10 (all)   nDCG@10 (hard-negative pool)
A — dense                                            0.9518                0.9577
B — hybrid (0.6 dense / 0.4 max-normalized BM25)     0.9255                0.9288
C — hybrid top-50, fused 0.5×hybrid + 0.5×NLI-ent.   0.8560                0.9288 *

* The hard-negative metric for policy C is computed from the pre-NLI hybrid score dictionary. It does not observe the NLI-fused top-50 ordering and cannot be read as evidence about that stage’s effect on hard-negative discrimination.

MEASURED: on RELATE v0.1 / mpnet-base, this specific 0.6/0.4 max-normalized dense–BM25 fusion lowered nDCG@10 from 0.9518 to 0.9255 (all-corpus) and 0.9577 to 0.9288 (hard-negative pool). Adding a top-50 fusion of that hybrid score with NLI entailment probability lowered all-corpus nDCG@10 further, to 0.8560. The hard-negative metric reported for that third policy is numerically identical to the second policy’s because the implementation computes it from a score dictionary the NLI stage never touches — it is not evidence that the NLI stage left hard-negative discrimination unchanged.

What this establishes. Dense scores 0.9518 on this probe. This specific hybrid fusion scores lower on both reported metrics. This specific hybrid-plus-NLI top-50 fusion scores lower still on the metric it does observe. The hard-negative metric cannot tell us the effect of the NLI stage, because its computation bypasses the stage that changed.

What this does not establish. Why the fusion hurt (score scale, weighting, query distribution, and the NLI orientation are all plausible contributors, none isolated); that hybrid retrieval is worse than dense retrieval in general — this is one fusion rule on one corpus; that this NLI stage’s effect on hard-negative discrimination was zero, positive, or negative — the artifact simply does not measure it; that policy layering inherently helps or inherently hurts; any latency, cost, or throughput comparison — the artifact contains no timing; or that the result would hold with a query population that has a genuine, uncontrolled gap between query and answer rather than RELATE’s controlled construction.

Do not reach for the retrieval literature to paper over this. Papers showing hybrid or reranking gains on other benchmarks do not establish that this fusion, this weighting, or this second-stage orientation would help this corpus, and citing them here would substitute someone else’s measurement for the one this chapter is responsible for. What the negative result actually demonstrates is more useful than a borrowed success story: policy components are hypotheses, and they must be ablated, not assumed helpful — including the ones that sound obviously smarter.

Policy ablation vs. measurement ablation

Row 1.13 is a clean lesson in a distinction worth keeping permanently:

A stage ablation is only evidence about a stage when:

  1. exactly one intentional policy difference separates the two conditions;
  2. the same evaluation workload is used for both;
  3. the metric’s computation actually passes through the changed stage;
  4. both variants have recorded provenance (so “same policy, different artifacts” cannot masquerade as a stage effect);
  5. ideally, cost and latency are measured alongside quality.

Condition 3 is where policy C’s hard-negative comparison fails. The ranking changed; the metric that was supposed to observe it did not. An ablation is only an ablation if the metric passes through the stage being ablated — a rule that generalizes far past retrieval, to any pipeline where a downstream evaluation is wired to an intermediate representation that a later change bypasses.

What this chapter establishes and what it does not

Establishes: retrieval behavior is produced by a composition — representation, eligible candidates, scoring/fusion, selection, execution strategy, and context assembly — not by a model name; stage ordering is itself part of the policy; build-time, query-time, and context-time decisions have different rollback costs; on RELATE v0.1 the measured dense baseline scores 0.9518 / 0.9577 (all / hard-negative), the measured hybrid fusion scores 0.9255 / 0.9288, and the measured hybrid-plus-NLI-top-50 fusion scores 0.8560 on the metric it actually observes; and the persisted hard-negative number for that third policy is not valid evidence about the NLI stage, because the metric’s computation never consumes the reordered top 50.

Does not establish: an optimal policy for any workload; that the embedding model is unimportant or that it alone sets a ceiling; why either added stage reduced the measured metrics; that hybrid retrieval or NLI-based fusion generally hurts; any timing, cost, or latency comparison; or that this result would reproduce on a query population with a genuine, unconstructed gap between query and answer.

Lab 12: run the ablation, then audit whether the metric could have seen it

MEASURED — artifact experiments/embeddings-from-first-principles/wave1/artifacts/policy-ablation.json (mpnet-base, fixed corpus). REPRODUCIBLE — run_wave1.py 1.13.

Question. Which policy stage actually changes quality on your stack — and can your evaluation see it?

Step 1 — freeze the base environment. RELATE version and hash, the 269-query set, the dense space and model revision, the sparse index and version, the relevance definition, and the nDCG implementation.

Step 2 — define policy A exactly. Dense cosine ranking, full corpus. Measured: nDCG@10 (all) = 0.9518, nDCG@10 (hard-negative pool) = 0.9577.

Step 3 — define policy B exactly. 0.6 × dense_cosine + 0.4 × (BM25 / query-max BM25), items missing from the BM25 map scored 0 on the sparse term. Measured: 0.9255 / 0.9288.

Step 4 — define policy C exactly. Hybrid top-50 (candidate depth N = 50), NLI entailment probability on (query, candidate), fused 0.5 × hybrid + 0.5 × entailment, reordering only those 50; ranks 51+ keep their hybrid order. Measured all-corpus: 0.8560.

Step 5 — audit the hard-negative metric before trusting it. Trace which score dictionary the hard-negative pool is sorted by for policy C. It is the pre-NLI hybrid dictionary — the same one used for policy B. The persisted 0.9288 for policy C is therefore not a post-NLI evaluation of hard-negative discrimination; it is arithmetically forced to match policy B on this metric.

Step 6 — design the corrected ablation (PROPOSED — no result exists yet). Recompute the hard-negative metric using the actual post-stage score or rank for each policy — for policy C, the fused top-50 score where the item falls in the top 50, and the unchanged hybrid score otherwise. Then measure, per policy: all-corpus nDCG@10; the corrected hard-negative nDCG@10; per-distractor-type behavior (Chapter 10’s taxonomy); how many graded-relevant items survive into the top-50 candidate depth in the first place; and, separately, p50/p95 latency and per-query cost. None of these numbers are reported here.

Step 7 — change one stage at a time. Resist assembling hybrid + threshold + rerank + diversity + a new chunk size in one pass and asking what helped. Where a factorial design is affordable, run it; otherwise increment one stage, hold everything else fixed, and re-run steps 2–5 for each new condition.

Step 8 — let evaluation produce evidence, and let operational constraints choose the policy. A measured quality delta is an input to a deployment decision, not the decision itself — latency budget, cost, and risk tolerance belong to the application, exactly as Chapter 6 separated a detector from a policy.

Try it yourself

Reproduce Steps 2–5 on your own corpus and query set. Before looking at any table, write down, for each stage, what score dictionary or ranking your evaluation metric actually reads from. If a later stage writes to a different variable than the one your metric consumes, you have found the same class of bug this chapter did — fix the wiring before trusting the ablation, not after.

Companion component: the retrieval policy, with a trace boundary

The reusable artifact is not a flat bag of knobs. It is a versioned composition that references immutable artifacts, and it is paired with a lightweight trace so that a metric’s blind spots — like the one just found — are visible rather than silent.

retrieval_policy:
  id / version:

  corpus:
    corpus_version:
    chunking_artifact_id:        # a build-time reference, not a runtime knob
    dense_index_id:
    sparse_index_id:

  query:
    transform:                   # raw | rewrite | expand | decompose | hyde
    representation_role:

  candidate_policy:
    eligibility:                 # filters
    authorization:                # the correctness/security rule — never optional
    candidate_depth:              # N entering the next stage, distinct from final k

  scoring_policy:
    dense_metric:
    sparse_score:
    fusion_rule:                 # e.g. "0.6 dense + 0.4 max-normalized BM25"

  second_stage:
    scorer_id:                   # NOT listed under "spaces" — a pairwise scorer is not an embedding space
    input_orientation:           # e.g. "(query, candidate)"
    candidate_depth:
    score_fusion:                # e.g. "0.5 hybrid + 0.5 entailment"

  selection_policy:
    threshold:
    final_k:
    tie_break:
    diversity: {method, lambda} | none

  execution:                     # Chapter 9
    method: exact | ann | ...
    build_params:
    search_params:
    fidelity_ref:

  context_assembly_ref:

  evaluation_refs:                # points to eval runs, never inlines a bare number

retrieval_trace(policy_version, query_id):
  eligible_candidate_count:
  candidates_after_scoring:       # ids + scores, or a hash
  candidates_after_second_stage:  # ids + scores, or a hash — MUST differ from the above if the stage changed anything
  final_selected_ids:
  context_token_count:

The design principle is not observability for its own sake. It is that the trace gives every stage an observable before/after boundary, so a question like “did the metric actually see this stage?” has a concrete, checkable answer instead of an assumed one. With this schema the row 1.13 bug becomes visible immediately: candidates_after_second_stage for policy C would show a reordered top 50, while the hard-negative metric’s input would not — a mismatch the trace surfaces rather than hides.

The Observatory should be able to answer, for any reported number: which policy version ran; which corpus, index, space, and scorer versions it referenced; what entered and left each stage; what changed between two policy variants; whether the reported metric actually observed the changed stage; and what quality, cost, and latency evidence belongs to that exact configuration. That is a materially stronger claim than “we use model X.”

Failure modes

  • “We use model X.” Tempting because it is the most visible decision. Check: a model name specifies none of the eligibility, fusion, selection, execution, or context-assembly choices that also determine what a query returns.
  • Calling a score-fusion stage “the reranker.” Tempting because “rerank” is the familiar word. Check: if two scores are blended, record the blend, the weights, and the normalization — policy C is a 50/50 fusion over a top-50 candidate depth, not a pure NLI rerank.
  • Evaluating a stage with a metric wired around it. Tempting because the metric already exists and returns a number. Check: trace which score or ranking the metric actually consumes before trusting an “unchanged” result — this chapter’s own hard-negative column is the cautionary example.
  • Changing several stages at once. Tempting because it is faster to try everything together. Check: attribution requires isolating one intentional difference per comparison.
  • Treating build-time and query-time knobs alike. Tempting because they can look like config values in the same file. Check: chunking and index construction may require rebuilding artifacts, not flipping a flag — rollback needs the artifact, not just the parameter.
  • Leaving stage order implicit. Tempting because the code just runs top to bottom. Check: threshold-before-rerank and threshold-after-rerank are different policies; write the order down.
  • Reporting quality without cost. Tempting because quality is the headline number. Check: a stage can improve ranking while violating a latency or budget constraint — this chapter measured no timing at all, which is itself a gap to close before deploying.
  • Treating retrieved context as memory or truth. Tempting because “the model’s memory” is a vivid phrase. Check: it is selected evidence, assembled under a specific policy, from a specific candidate set — Chapter 10’s distinction between finding a candidate and verifying a claim still applies to whatever this policy hands downstream.

What this chapter established

  • A retrieval system is a composition — representation, eligible candidates, scoring and fusion, selection, execution strategy, context assembly — with its own ordering, not a model plus implicit defaults. Build-time, query-time, and context-time decisions carry different rollback costs; chunking in particular is a corpus-artifact decision wearing a parameter’s clothing.
  • Policy, provenance, execution, trace, and evaluation are five distinct objects. A policy plus its bound artifacts approximates a deployable unit better than a policy alone — and far better than a model name alone.
  • An ablation is only valid evidence about a stage when the metric’s computation actually passes through that stage. The measured hybrid-plus-NLI policy’s hard-negative figure fails exactly that test: the metric never consumed the reordered candidates, so the number is not evidence about the stage it appears to evaluate.
  • The retrieval policy object, paired with a per-query trace, gives every stage an observable before/after boundary — so a metric’s blind spots become visible rather than assumed away.

Next

Application retrieval quality is a joint property of the representation, the corpus, the policy wrapped around them, the execution strategy, and the evaluation definition used to judge the result. Once that policy is made explicit and held fixed — corpus frozen, candidates fixed, fusion and selection rules written down — a cleaner question becomes askable for the first time: with everything around it held still, how good is the representation itself? Part IV takes up exactly that question.