Calibration
A cosine of 0.81 is a geometric score, not an operational verdict. Define a task-specific positive and negative class, build their score distributions, and derive an operating point from AUC, an approximate equal-error threshold, and a quantile-based abstention band — then compare separately fitted operating points across domains.
Part IV — Measuring the Representation
What is 0.81?
Illustrative scenario — invented values for intuition, not a measured result.
A pipeline decides two documents are “duplicates” if their cosine exceeds 0.8. Someone picked 0.8 because it looked reasonable. Imagine the two populations that number is actually sitting between:
distribution of cosine for KNOWN duplicate pairs: mean ~0.79
distribution of cosine for KNOWN non-duplicate pairs: mean ~0.71
If those two populations overlap the way this sketch implies, the threshold sits inside both: a true duplicate at 0.78 is rejected, an unrelated pair at 0.82 is accepted. This chapter’s actual experiment, later, replaces the sketch with measured positive and negative score samples and derived discrimination/operating statistics. The artifact does not persist the full distributions themselves, and its 85.6% figure is the coverage of a specific quantile-derived band rather than a generic measure of distributional overlap.
The correction worth making immediately: cosine 0.81 is not meaningless. Under a stated representation and metric it is a precise geometric fact about two vectors — Chapters 1 through 4 spent considerable effort earning that precision. What 0.81 does not carry on its own is an application decision. It does not know whether your task calls two things “duplicates,” “the same claim,” “relevant,” or “close enough” — and it cannot, because that is not a fact about the vectors. It is a fact about a decision problem someone has to define.
A score is geometry. A threshold is policy. Calibration is the measurement that connects the two.
One more distinction worth fixing before going further, because “calibration” is an overloaded word in machine learning. In some contexts calibration means turning a model’s raw score into a genuine probability — a predicted 0.8 should correspond to the event actually happening about 80% of the time (Platt scaling, isotonic regression, and similar methods live there). That is not this chapter. Here, “calibration” means selecting and validating an operating rule for a raw similarity score — where to accept, where to reject, where to abstain. Nothing in this chapter claims cosine has been turned into a calibrated probability; it remains a geometric score with a measured relationship to two labeled classes.
A similarity score is interpretable only relative to the score distributions of a task-defined positive class and a task-defined negative class. What are those classes, how well does the score separate them, and what operating rule — if any — should the application derive from that evidence?
The calibration contract
Chapter 13 installed the idea that an evaluation score depends on an evaluation contract — query population, candidate universe, relevance definition, metric, aggregation, model-use protocol. A calibrated operating point depends on the same kind of contract, one layer down:
calibration contract
=
space / model version
× scoring rule (metric, normalization)
× task-defined positive class
× task-defined negative class
× sampled population and sampling rule
× domain / slice
× calibration procedure (how the operating point is derived)
× operating objective (EER, target FAR/FRR, loss/cost rule, abstention policy, ...)
A threshold is not a property of a model. It is an artifact derived for a particular decision, under a particular contract — and “0.8435” means nothing portable until every term on the right-hand side is attached to it. Chapter 13 asked which representation orders labeled candidates better. Chapter 14 asks a narrower and more consequential question: given one representation and one scalar score, where — if anywhere — should an application act automatically?
Positive and negative are decision labels, not relatedness labels
It is tempting to read “positive” as “semantically related” and “negative” as “unrelated.” Resist that. In a calibration problem, positive and negative are defined by the decision being automated, and a pair can be strongly related — same topic, same entities, nearly identical wording — while still belonging to the negative class for that decision.
The demonstration below calibrates a same-claim versus different-claim decision. Its classes:
positive (same claim): equivalent, paraphrase
negative (different claim): negation, contradiction, temporal-mismatch, relation-swap
A negation pair is topically about exactly the same thing as its source — often sharing most of its wording — and it is still a negative example here, because the decision this calibration serves is “do these two sentences assert the same claim,” and a negated sentence does not. This is precisely the distinction Chapter 10 installed between relatedness and assertion compatibility, now doing real work: a calibration population built from topically related versus unrelated pairs would be calibrating a different, easier decision than the one an application asking “same claim?” actually needs.
Building the distributions
- Positive set. Pairs known to belong to the positive class for the decision you are calibrating — not merely related in some general sense.
- Negative set, chosen for a stated purpose. Two different, both legitimate, constructions answer different questions. Operational calibration samples a population intended to resemble the cases the decision rule will face in deployment — if most negatives a production system meets are unrelated content, an operational negative set should reflect that. Stress-test calibration deliberately concentrates selected near-misses — for example, topically matched or structured perturbations — to probe a harder, explicitly scoped decision regime. Chapter 11 already showed the same scorer can look comfortably separated against random negatives and much less separated against structured ones; substituting one distribution for the other without saying so produces an operating point for a different competition. The demonstration below is a stress-test construction, and it says so.
- Inspect the distributions, not just their means. A mean alone cannot reveal overlap, tails, or multimodality — the actual shape a decision rule has to live inside. The measured artifact used below stores only summary means and derived statistics, not full histograms or quantile curves; a stronger calibration record would keep a reference to the raw scores or a histogram artifact alongside the summary, and that is flagged as a proposed improvement to the artifact later in this chapter, not something already available.
- Read the region where both classes have mass. That region is where a single threshold cannot cleanly separate the two — the starting intuition for an abstention band, made precise below.
Operating points, not vibes
From the two labeled score sets:
- False acceptance rate (FAR) at threshold
t— the fraction of negative examples scoring≥ t. - False rejection rate (FRR) at
t— the fraction of positive examples scoring< t.
Both are class-conditional rates, and it is worth being exact about what they are not. FAR is not “the probability that an accepted pair is actually wrong,” and FRR is not “the probability that a rejected pair was actually right.” Those quantities depend on how common positives and negatives are in the population the rule actually meets — their prevalence — not only on FAR and FRR. If true negatives outnumber true positives one hundred to one, a FAR of 5% can still mean that most accepted results are false positives, because there were simply so many more negatives available to leak through at that rate. The fixed RELATE calibration view used below contains 350 positive and 343 negative pairs, so the measured sample happens to be close to 1:1. That ratio is a property of this selected calibration view, not an estimate of how often a deployed system will encounter a same-claim pair versus a different-claim one.
- ROC curve — FAR plotted against
1 − FRRacross every threshold. AUC summarizes separability across all thresholds at once — it is threshold-independent, in the sense that no single operating point has to be chosen to compute it. It is not task-independent: it still depends on which pairs were labeled positive and negative, how they were sampled, which scorer produced the numbers, and what preprocessing the vectors went through. A useful intuition for the number itself: AUC is approximately the probability that a randomly drawn positive example scores higher than a randomly drawn negative example, drawn from the same labeled sample, subject to the usual tie-handling convention. AUC asks whether positives tend to outrank negatives. It does not ask where an application should act — that is a separate, later step. - Choose an operating point from a stated target, not a vibe. “FAR must not exceed 1%” for a merge pipeline that cannot afford to fuse distinct records. “FRR must not exceed 5%” for a recall-critical filter. The target should come from the cost of each error in the application it serves.
- An abstention (or “escalation”) band is one explicit policy choice among several, not the only honest design. It is reasonable exactly when the application can tolerate deferring a decision — to a reranker, a second signal (Chapter 15), or a human. An application that must return a binary answer, or that can tolerate a higher FAR or FRR than a three-way rule implies, may reasonably choose differently. Measurement does not choose the policy; it supplies the numbers the policy is chosen from — the same separation Chapter 6 drew between a detector and the action taken on it.
flowchart TD
P["positive set — pairs in the task-defined positive class"] --> H["score distributions, per class"]
N["negative set — operational sample, or a stated stress-test population"] --> H
H --> AUC["AUC — threshold-independent discrimination, not a policy"]
H --> SW["sweep candidate thresholds; measure FAR(t) and FRR(t) at each"]
SW --> T["pick an operating point from a stated FAR/FRR target or cost ratio"]
T --> D{"score vs the chosen policy"}
D -->|"below the reject bound"| REJ[reject]
D -->|"above the accept bound"| ACC[accept]
D -->|"in between, if abstention is the chosen policy"| ESC["escalate — reranker / second signal / human"]
REJ --> RC["bind the operating point to the full calibration contract; revalidate on change"]
ACC --> RC
ESC --> RC
Why a threshold needs revalidation, not blind transfer
A threshold’s validity depends on the calibration contract that produced it. Several things can change that contract: the model or embedding space, the metric or normalization, the positive/negative label definition, the domain, the query or pair population, the sampling scheme, or the input protocol. A change to any of these does not mathematically guarantee the old threshold is now wrong — but it does mean the assumptions under which it was derived no longer obviously hold, so transfer has to be revalidated rather than assumed. A model version bump can reshape a score distribution; it does not automatically do so, and the honest response to any of these changes is the same: rerun the calibration sample under the new conditions and compare the resulting distributions and operating points, rather than assuming a mechanism and skipping the measurement.
Deterministic calibration is not the same as statistically stable calibration
A calibration procedure should be reproducible: the same labeled sample and the same code should produce the same threshold, FAR, FRR, and band, every time. Fix the seed for any sampling step, version the labeled calibration set, and record the derived operating point as data. That answers one question — can I reproduce this operating point? — and it is worth answering well.
It does not answer a different question: would this operating point survive being derived from another, equally representative sample? A deterministic procedure can reproducibly overfit a small or unrepresentative calibration set — the same 350-and-343-pair sample, run twice, will produce the identical threshold 0.8435 both times, and that tells you nothing about how much the threshold would move if the sample had been drawn again. Neither row 1.10 nor row 1.11, below, includes a bootstrap or a held-out validation split; both report the operating point derived from and measured on the same sample. That is an acceptable, clearly-scoped design for a controlled demonstration. It is not the same claim as an independently validated deployment error rate, and the rest of this chapter is careful not to conflate the two.
Demonstration: calibrating a same-claim decision on RELATE
MEASURED on RELATE v0.1, Wave 1 rows 1.10 and 1.11 — artifacts
experiments/embeddings-from-first-principles/wave1/artifacts/calibration.jsonandthreshold-drift.json. Modelall-mpnet-base-v2, raw cosine.
The decision: same claim versus different claim. Positive class — equivalent and paraphrase pairs. Negative class — negation, contradiction, temporal-mismatch, and relation-swap pairs: deliberately selected near-claim confusions rather than a random or operational sample of what a deployed same-claim filter would typically encounter. This is a stress-test calibration by construction. It probes these four specified relation families; it does not establish that they are the hardest possible negatives in RELATE or in deployment.
n_positive = 350 n_negative = 343
positive mean cosine = 0.8742 negative mean cosine = 0.7597
AUC = 0.7474
closest equal-error point on a 200-point threshold sweep:
threshold = 0.8435
FAR = 0.2391
FRR = 0.2457
quantile-derived escalation band (positive 5th percentile to negative 95th percentile):
fraction of pooled scores inside the interval = 0.856
Read each number for exactly what it measures, because each one answers a narrower question than it looks like it does.
The means look encouraging on their own — 0.8742 against 0.7597 is a real, sizeable gap. It does not, by itself, say how separable the two full distributions are; that is what AUC is for.
AUC = 0.7474 means that for a positive example and a negative example drawn at random from this labeled sample, the positive one scores higher about 74.7% of the time. That is a mid-level, not a trivial, separation — substantially better than chance (0.5), well short of clean separation (approaching 1.0) — and it is a statement about ordering across every possible threshold, not about any single operating point. It does not, by itself, tell an application where to draw a line.
The equal-error point is approximate by construction, not by measurement error. The implementation does not solve analytically for the threshold where FAR exactly equals FRR; it sweeps 200 linearly spaced thresholds across the observed score range and keeps the one minimizing |FAR(t) − FRR(t)|. The result — threshold 0.8435, FAR 0.2391, FRR 0.2457 — is the closest point the sweep found, not an exact equal-error solution; the two rates differ by 0.66 percentage points. Call it an approximate equal-error point, and read what it says plainly: if this raw cosine rule is forced to balance the two error types on this stress-test sample, both land around 24%. Moving the threshold can reduce one class-conditional error while increasing the other; this near-EER result therefore does not say that neither FAR nor FRR can be pushed below 24%. It says that raw cosine does not make both error types small at the operating point that balances them.
The escalation band is one specific quantile construction, and its 85.6% is a policy consequence of that construction, not a discovery about the data’s intrinsic ambiguity. It is built from the 5th percentile of the positive scores and the 95th percentile of the negative scores; the reported fraction is the share of all pooled positive-and-negative scores that fall between those two quantile-derived bounds. On this sample that interval captures 85.6% of the pooled scores. That means: if an application adopted this particular 5%-tail construction as its abstention rule, then 85.6% of this stress-test sample would be routed to the escalation region rather than decided automatically by threshold alone. It does not mean 85.6% of cases are intrinsically undecidable, that no other signal or classifier could resolve them, that a human reader could not tell them apart, or that every application must escalate this share of its traffic — a narrower band, a different target, or an entirely different policy would move that number, and Chapter 15 exists specifically to ask whether other signals pulled from the same pair can decide some of what this one scalar, under this one construction, cannot.
Domains complicate the picture further, and in a specific, bounded way — not the way it is easy to first assume. Row 1.11 does not calibrate a threshold on one domain and measure how it performs when applied to another; no transfer experiment was run. What it does is fit an independent approximate equal-error threshold, using the same 200-point sweep, separately within each of RELATE’s five domains, using only that domain’s positive and negative pairs:
domain n_pos n_neg fitted EER threshold
everyday-statements 27 28 0.7631
corporate-events 79 84 0.8134
product-support 36 16 0.8256
geo-civics 187 189 0.8578
biomed-claims 21 26 0.8588
measured spread (max − min): 0.0957
The valid conclusion: the independently fitted equal-error operating point varies across these five domains, from 0.7631 to 0.8588. That is measured, and it is enough to make the chapter’s point without overreaching. What it does not establish is a transfer result — nobody applied everyday-statements’ 0.7631 threshold to biomed-claims and measured what broke. It also does not establish a stable population property: the domain sample sizes range from 16 negative pairs (product-support) to 189 (geo-civics), and no confidence interval or resampling estimate was computed for any of the five thresholds. Smaller strata ordinarily warrant greater concern about sampling variability, but this artifact does not quantify that variability. The spread is measured; its sampling stability is not. And “threshold-drift” — the artifact’s file name — should not be read as evidence of change over time; this is a cross-domain comparison on one frozen release, not a time-series observation, and this chapter reserves “drift” for an actual measured change across time, version, or distribution.
MEASURED: on this stress-test same-claim/different-claim sample, raw cosine achieves AUC
0.7474; its approximate equal-error operating point sits at threshold0.8435with FAR0.2391and FRR0.2457; a 5%-tail quantile construction would route85.6%of the pooled sample to an escalation band rather than an automatic decision; and an equal-error threshold fit independently within each of five RELATE domains ranges from0.7631to0.8588, a spread of0.0957. The artifact reports each domain’s positive and negative sample counts, but it does not report uncertainty or sampling stability for those fitted thresholds.
What this chapter establishes and what it does not
Establishes: a similarity score requires a task-defined positive class and a task-defined negative class before it means anything for a decision, and those classes are decision labels, not synonyms for related and unrelated; FAR and FRR are class-conditional rates that depend on the calibration sample and do not, by themselves, give the probability an accepted or rejected result is correct; AUC measures threshold-independent discrimination and is not itself an operating policy; on this RELATE stress-test sample, AUC is 0.7474, the approximate equal-error operating point is threshold 0.8435 (FAR 0.2391, FRR 0.2457), the 5%-tail quantile band captures 85.6% of pooled scores, and independently fitted per-domain equal-error thresholds range from 0.7631 to 0.8588; and a deterministic calibration procedure guarantees reproducibility, not statistical stability.
Does not establish: a universal or portable threshold; that AUC alone is a sufficient summary for choosing a policy; that the 85.6% band figure measures intrinsic ambiguity rather than one quantile construction’s consequence; that any threshold transferred across domains was actually tested; that the domain-threshold spread is a stable, resampling-confirmed population property; FAR-≤1%-or-FRR-≤5%-style targets (none were computed here); or the outcome of a three-way accept/reject/escalate policy’s own realized FAR and FRR (only the band’s coverage fraction was measured, not the policy’s resulting error rates).
Lab 14: derive an operating point, and name exactly what it is
MEASURED — artifacts
wave1/artifacts/calibration.jsonandwave1/artifacts/threshold-drift.json(mpnet-base). REPRODUCIBLE —python run_wave1.py 1.10 1.11.
Question. Given one representation, one scorer, and one task-defined decision, what can raw cosine alone tell an application to do?
Step 1 — define the calibration contract before computing anything. Space record / model revision, normalization, metric, the exact decision being calibrated (“same claim vs. different claim”), the positive relation classes, the negative relation classes, RELATE release and corpus hash, the sampling method (here: a fixed stress-test view, not an operational sample), domain or slice, and code hash.
Step 2 — build the labeled score sets. Positive: equivalent, paraphrase (n = 350). Negative: negation, contradiction, temporal-mismatch, relation-swap (n = 343). State plainly that this is a stress-test population, not a claim about deployment prevalence.
Step 3 — inspect the distributions, honestly bounded. Measured means: positive 0.8742, negative 0.7597. No standard deviation, quantile curve, or histogram is stored in this artifact; do not invent one. If your own calibration run can afford it, persist quantiles or a histogram reference — flag that as a proposed artifact improvement, not something this run already has.
Step 4 — measure threshold-independent ordering. AUC 0.7474. State what it says (positives tend to outrank negatives, about three times in four) and what it does not (it does not choose where to act).
Step 5 — derive the approximate equal-error point exactly as implemented. Sweep 200 linearly spaced thresholds across the observed score range; keep argmin_t |FAR(t) − FRR(t)|. Measured: threshold 0.8435, FAR 0.2391, FRR 0.2457. Call it an approximate equal-error point — the two rates are close, not identical.
Step 6 — derive the escalation band exactly as implemented. t_reject_candidate = the 5th percentile of positive scores; t_accept_candidate = the 95th percentile of negative scores. Measured: 85.6% of the pooled sample falls between them. State plainly that this is one selective-decision construction among several possible ones, not the unique or natural ambiguity band.
Step 7 — inspect domain-conditioned operating points, with sample sizes attached.
| Domain | n_pos | n_neg | Fitted EER threshold |
|---|---|---|---|
| everyday-statements | 27 | 28 | 0.7631 |
| corporate-events | 79 | 84 | 0.8134 |
| product-support | 36 | 16 | 0.8256 |
| geo-civics | 187 | 189 | 0.8578 |
| biomed-claims | 21 | 26 | 0.8588 |
Measured spread: 0.0957. Label these separately fitted domain-specific thresholds — not a cross-domain transfer measurement.
Step 8 (PROPOSED — no artifact backs this) — test threshold transfer. Fit a threshold on one domain, or on a pooled split, freeze it, apply it unchanged to a different domain or a held-out slice, and report the resulting FAR and FRR there. This is the experiment that would actually measure transfer error; it has not been run for this chapter.
Step 9 (PROPOSED — no artifact backs this) — separate stability from held-out validation. To estimate sampling uncertainty, bootstrap the calibration pairs and report the distribution or interval of the chosen threshold and derived FAR/FRR; do the same within domains if the sample sizes support it. Separately, to validate an operating rule, choose the threshold on a fitting split, freeze it, and measure FAR/FRR on a held-out split. A single held-out split tests transfer to new examples; it does not by itself provide a sampling interval for the fitted threshold. Neither experiment exists in the current artifacts.
Try it yourself
Build your own positive and negative classes for the decision you actually need to automate — not a proxy for “related,” the decision itself. Decide up front whether you are sampling an operational population or deliberately stress-testing with near-misses, and say which. Plot the score distributions, compute AUC, sweep for the operating point your application’s cost ratio actually needs (a target FAR, a target FRR, or an abstention band), and re-run the whole procedure on a second slice — a different domain, a different time window, a different query type — before trusting the threshold anywhere else. Report the observed shift honestly, as a spread you measured, not as a transfer error you assumed.
Companion component: the calibration record
The calibration record extends Chapter 13’s evaluation card one layer down: the evaluation card defines the comparison, the calibration record defines the decision rule derived from one representation’s scores under it.
calibration_record:
id / version:
evaluation_card_ref: <from Ch 13 — the workload this population was drawn from>
space_record_ref: <from Ch 1/17>
decision:
positive_definition: <task-defined classes, not "related">
negative_definition:
population_kind: <operational | stress_test>
score:
metric:
normalization:
population:
release:
corpus_hash:
sample_source:
sampling_rule:
domain_or_slice:
n_positive:
n_negative:
distributions:
positive_summary: <mean, and quantiles/histogram_ref if available>
negative_summary:
raw_scores_or_histogram_ref: <PROPOSED where not yet stored>
discrimination:
auc:
operating_rule:
method: <equal_error_sweep | target_far | target_frr | abstention_band | ...>
threshold_or_band:
target:
measured_on_calibration_sample:
far:
frr:
abstention_fraction:
validation_ref: <held-out or bootstrap result, if one exists — absent here>
provenance:
code_hash:
seed:
The design correction worth stating plainly: the operating point is inseparable from the classes and the population used to derive it. A threshold recorded without decision, population, and population_kind is not reusable evidence — it is a number that happened to work on a sample nobody can reconstruct. And measured_on_calibration_sample is deliberately named to be distinct from validation_ref: this chapter’s own row 1.10 and 1.11 only ever populate the former. Reporting a calibration-sample FAR or FRR as though it were an independently validated deployment error rate would erase exactly the distinction the last two sections spent effort building.
Embedding Observatory behavior
The Observatory should never surface a bare number like threshold = 0.8435 as a property of mpnet-base. It should surface the calibration record that produced it:
same-claim decision
RELATE v0.1 stress-test calibration population (350 positive, 343 negative)
positive = equivalent | paraphrase
negative = negation | contradiction | temporal-mismatch | relation-swap
raw cosine, mpnet-base space
approximate equal-error threshold = 0.8435
calibration-sample FAR = 0.2391, FRR = 0.2457
held-out validation: not measured
And it should warn — not silently refuse, and not declare the number invalid — when the live system’s contract no longer matches the one on record: a different space or model version, a different metric or normalization, a different positive/negative definition, a different population or domain, or a different data release. The correct signal is “requires revalidation,” not an automatic “threshold invalid.” A change to the contract means the assumptions behind the number are no longer confirmed to hold; it does not by itself prove the number is now wrong.
Failure modes
- A bare threshold. Tempting because
similarity > 0.8reads like a specification. Check: no distributions, no FAR/FRR, no calibration contract behind it — it is an uninterpretable number, not a rule. - Calling semantically related pairs “positive.” Tempting because “related” and “positive” feel synonymous. Check: positive means positive for the decision being calibrated — a negation pair can be maximally related and still belong in the negative class.
- Calibrating a near-miss decision on easy negatives. Tempting because easy negatives are simple to sample. Check: an operating point fit against unrelated pairs measures a different, easier competition than the one a near-miss decision will actually face.
- Treating a stress-test distribution as deployment prevalence. Tempting because the calibration sample is the only data on hand. Check: a stress-test sample answers “how bad can it get against hard confusions,” not “how often will this actually happen in production.”
- Confusing FAR with “probability the accepted result is wrong.” Tempting because both sound like an error rate. Check: that conditional probability also depends on class prevalence, which FAR does not encode by itself.
- Calling the quantile band intrinsic ambiguity. Tempting because 85.6% sounds like a discovered fact about the data. Check: it is the coverage of one specific 5th/95th-percentile construction — a different band definition gives a different number.
- Reusing a threshold across a changed contract without validation. Tempting because the old number is easy to keep. Check: model, metric, population, domain, or label definition can all move the score distributions; measure the consequence rather than assuming it.
- Calling separately fitted domain thresholds “transfer failure.” Tempting because the spread looks alarming on its own. Check: transfer failure is a claim about applying one threshold elsewhere and measuring what breaks — that experiment was not run here.
- Trusting determinism as proof of reliability. Tempting because the same run always gives the same number. Check: reproducibility is not the same claim as statistical stability across another representative sample.
- Reporting calibration-sample FAR/FRR as held-out performance. Tempting because it is the only measured number available. Check: selection and validation are different steps; this chapter’s own artifacts only performed the first.
What this chapter established
- A raw similarity score has geometric meaning and no self-contained application-decision meaning. It becomes decision evidence only relative to a stated calibration contract.
- Positive and negative are labels in a decision problem, not stand-ins for semantically related and unrelated — this chapter’s negative class is deliberately, topically close to its positives. FAR and FRR are class-conditional rates, distinct from the probability that an accepted result is actually correct, which also depends on prevalence.
- AUC measures threshold-independent discrimination, not a chosen operating policy. Accept/reject/escalate is one policy an application may adopt, not a mathematical inevitability — and the measured equal-error point is an approximate 200-point sweep result, not an exact solution.
- The measured tail-quantile band captures 85.6% of the pooled sample as a consequence of that specific construction, not as proof that 85.6% of cases are inherently undecidable. The per-domain thresholds were each fitted independently, with no transfer test run and no stability estimate computed.
- Deterministic calibration guarantees reproducibility, not statistical stability. Fitting and validation are different steps, and this chapter’s experiment performed only the first. The calibration record binds a threshold to the exact decision, population, and procedure that produced it.
Next
Even after defining the decision precisely and sampling deliberately difficult near-misses, raw cosine leaves a costly trade-off: at the measured near-equal-error point both class-conditional error rates are about 24%, while one particular 5%-tail abstention rule would route 85.6% of this stress-test sample to the middle region. Calibration makes that trade-off explicit; it does not prove that abstention is mandatory or that the middle cases are intrinsically undecidable. The next chapter asks whether the scalar itself is discarding useful evidence, pulling several additional signals from the same geometry — margin, local density, neighborhood structure, rank — before paying for another model.