You Cannot Optimize What You Cannot Measure
Define the metric, freeze the evaluation protocol, and measure the program — then see why a measurable objective can make the wrong behavior systematically optimizable.
Chapter 7 produced material. This chapter produces evidence.
That makes optimization possible. It does not yet make optimization safe.
no metric
→ no systematic optimization
metric
→ behavior can be optimized
wrong metric
→ the wrong behavior can be optimized systematically
This chapter establishes the surface on which search can operate. Chapter 9 asks what that surface rewards when the metric becomes the target.
flowchart TD
P[frozen program] --> E[evaluate]
C[frozen cases] --> E
M[metric] --> E
E --> PC[per-case results]
PC --> AG[aggregate result]
PC --> FI[failure inspection]
No optimization happens here. Note that the per-case results feed two readings — the aggregate and the failure inspection — and this chapter’s finding comes from the second. The question is narrow:
How well did this frozen program perform under this frozen protocol?
The answer turns out to be less interesting than what we find while computing it. The aggregate number is fine. The per-case breakdown contains a defect that shapes the next nine chapters.
1. A baseline is an identity, not a number
A baseline is not “whatever the code did yesterday.” It is a program identity, a model identity, a dataset identity, a metric identity, and a protocol — and if any of those is missing, the number cannot be compared to anything later.
BASELINE_IDENTITY = {
"program_id": "EditorialRewriteProgram",
"module_strategy": "analyze_predict__rewrite_predict__assess_predict",
"model_provider": "Ollama",
"model_name": "qwen3:latest",
"dataset_id": "book-editorial-corpus",
"split_role": "dev",
"metric_name": "editorial_metric",
"metric_version": "v1",
"optimization_allowed": False,
}
In a real system this record is generated, never typed. The generated version covers program fingerprint, dataset fingerprint, split fingerprint, provider and LM-configuration fingerprints, metric fingerprint, DSPy version, seed, budget, and source revision.
That collection has a name in this book: the evaluation protocol fingerprint. It exists to make one sentence meaningful.
A candidate that reports a higher number under a different protocol has not beaten this baseline. It has run a different experiment.
That sounds pedantic until chapter 12, where an optimizer reports a gain, and the only way to know whether it is real is to confirm that the baseline and the candidate were scored by the same function over the same cases under the same model configuration.
2. The metric: a hard gate and a structural score
The metric has two stages, and the first one is not a score at all.
Stage one is a deterministic hard gate. In the frozen v1 implementation it checks two non-negotiable lexical constraints: every required entity must still be present, and no forbidden term may appear. Its answer is binary.
That is narrower than “is the output usable?” Empty output, sentence count, excessive length, and other shape problems are handled elsewhere in the scoring or validation path. Keeping the exact boundary visible matters because Chapter 9 will attack what the gate does not know how to inspect.
import re
def contains_term(text: str, term: str) -> bool:
if " " in term.strip():
return term.lower() in text.lower()
return re.search(rf"\b{re.escape(term.lower())}\b", text.lower()) is not None
def passes_hard_gate(example, candidate: str) -> bool:
required_ok = all(
contains_term(candidate, entity)
for entity in example.required_entities
)
forbidden_ok = all(
not contains_term(candidate, term)
for term in example.forbidden_terms
)
return required_ok and forbidden_ok
A rewrite that drops a required entity or emits a forbidden term is not treated as slightly worse under this protocol. It scores zero.
That is a policy choice, not a universal property of metrics. The useful distinction is that a gate encodes disqualification, while a weighted score ranks candidates that remain eligible. If you decide that an entity failure is disqualifying, do not let unrelated fluency points buy it back.
Stage two scores what survives the gate, as a weighted sum of three components:
score = 0.35 · changed (did the program edit at all?)
+ 0.25 · scope (did it stay proportionate to the task?)
+ 0.40 · overlap (how close is it to the reference rewrite?)
The weights sum to 1.0, but the component names are more generous than their implementations.
changed is binary: after token normalization, the candidate either differs from the source or it does not.
scope is also binary. It returns zero for an empty candidate or for one longer than max(original_tokens + 4, 1.35 × original_tokens). It does not penalise aggressive shortening unless the candidate becomes empty.
overlap is asymmetric reference-token recall:
|tokens(reference) ∩ tokens(candidate)| / |tokens(reference)|
Extra candidate tokens do not directly reduce it, and word order does not matter.
This is editorial_metric at version v1. It is deterministic, cheap, inspectable, and therefore useful. It is also a deliberately incomplete proxy. Chapter 9 will exploit exactly the asymmetries you can already see in these definitions.
Note one thing now, because it becomes important in section 6. Drop the changed term and you get:
0.25 + 0.40 = 0.65
0.65 is the score for an unchanged candidate when scope passes and reference overlap is perfect. In other words, if the reference itself says leave this sentence exactly as it is, the correct identity output forfeits the entire 0.35 changed component by construction.
That number is going to come up a lot.
3. Running it
DSPy’s evaluator takes a devset and a metric and returns an aggregate plus per-case results.
import dspy
from common.metrics import editorial_metric
def dspy_editorial_metric_v1(example, pred, trace=None) -> float:
return editorial_metric(example, getattr(pred, "rewritten_text", ""), version="v1").score
def evaluate_baseline(program, examples):
evaluator = dspy.Evaluate(
devset=examples,
metric=dspy_editorial_metric_v1,
display_progress=True,
display_table=False,
failure_score=0.0,
)
return evaluator(program)
Two naming decisions are worth explaining, because both were mistakes we had to correct.
The metric is called editorial_metric and takes version as a required argument with no default. Earlier drafts had four names for what were really two functions — editorial_metric, editorial_metric_v1, score_editorial_output, score_editorial_output_v1 — and a default version, which meant a call site could silently change meaning when the default moved. Version is now explicit at every call.
The DSPy-facing wrapper is dspy_editorial_metric_v1, not dspy_editorial_metric. When chapters 10 to 13 hand a metric to an optimizer, the version being optimized against is the single most consequential fact about that run, and it should be legible in the call:
optimizer = dspy.BootstrapFewShot(metric=dspy_editorial_metric_v1)
Nobody reading that line has to guess what is being maximised.
4. What we measured
The frozen EditorialRewriteProgram, on the 11-case development split, under the canonical local Qwen3/Ollama dependency.
On the canonical development case ed-003, the program produced:
“I don’t think we should go there,” Anna said, “because it’s unsafe.”
Score: 0.877, stable to three decimal places across every run we made. Latency around 5.45 seconds.
The development aggregate:
| Run | Dev mean, v1 | Dev mean, v2 |
|---|---|---|
| First | 0.7893 | 0.7893 |
| Second | 0.7992 | 0.7992 |
The baseline is approximately 0.79.
That phrasing is deliberate. Two reported executions of unchanged program state produced development means of 0.7893 and 0.7992, an observed span of about 0.010. Chapter 6 shows why we should not promote that span into a universal ± noise constant: fresh-session drift and within-session order effects are both part of the measurement process.
Reporting 0.7893 as the baseline would therefore give the point estimate more authority than it has.
Compare that with the fixture this replaced, which reported 0.8769230769 — ten decimal places describing one sentence about Anna. The digits were real in the sense that arithmetic produced them. They were not ten decimal places of evidence.
The book carries forward about 0.79, while the machine artifacts retain full precision and the repeated runs retain their exact values. A candidate earns the word improvement only after a paired comparison shows that its delta survives the variability of the evaluation process.
Two deterministic sanity baselines, on ed-003:
| Baseline | Score |
|---|---|
| No edit | 0.6500 |
Mechanical and then → comma rule | 0.6500 |
EditorialRewriteProgram | 0.8769 |
On ed-003, the mechanical rule is a no-op because the sentence contains no and then, so it collapses to the same output as the no-edit baseline.
These are harness checks, not serious competitors. They show that the deterministic metric path executes as expected and that the LM program can score above an unchanged candidate on this particular row. They do not establish that an arbitrary mechanical edit is worth 0.65 or that 0.65 universally means poor quality.
One more result, which we did not expect and should have: in the two reported development-baseline runs, v2 equals v1 case for case. The semantic judge recorded no violated constraints on those baseline outputs.
Recall what v2 changes. It keeps the same hard gate and the same structural score, then applies a semantic factor from the judge; if any stated semantic constraint is violated, the final score is capped at 0.30. So agreement with v1 means the semantic component had nothing to penalise in these particular baseline runs.
That makes v2 useful in a different way from a finer-grained ruler. In the measurements so far, it behaves like a tripwire: ordinary baseline output leaves the structural score untouched, while specific semantic regressions such as ed-025 and later ed-035 make the two metrics separate sharply.
That is an empirical pattern in this experiment, not a guarantee that v2 only diverges after optimization. Any future baseline output that violates a named semantic constraint would trigger it too.
5. The aggregate hides everything interesting
Two programs can both score 0.75:
Program A: mostly good, one invalid output
Program B: always valid, quietly renames entities
The mean cannot distinguish them, and the difference is the only thing that matters. So the runner stores per-case evidence — output, score components, provider and configuration identity, latency, tokens, and whether the result is eligible as model evidence at all.
def unpack_evaluation(result) -> list[dict]:
return [
{
"case_id": example.case_id,
"target_failure": example.target_failure,
"score": float(score),
"prediction": {
"rewritten_text": getattr(prediction, "rewritten_text", None),
"rationale": getattr(prediction, "rationale", None),
"risk": getattr(prediction, "risk", None),
},
"reference": {
"reference_rewrite": example.reference_rewrite,
"required_entities": list(example.required_entities),
"forbidden_terms": list(example.forbidden_terms),
},
}
for example, prediction, score in result.results
]
The target_failure field from chapter 7 now earns its place. Grouping by it turns one number into a diagnosis:
| Target failure | Dev cases | Mean score |
|---|---|---|
semantic_drift | 2 | 0.938 |
entity_change | 2 | 0.825 |
style_failure | 3 | 0.808 |
constraint_violation | 1 | 0.782 |
invalid_output | 1 | 0.650 |
unnecessary_edit | 2 | 0.650 |
The aggregate of about 0.79 hides where the score comes from. This particular run’s grouped table shows two development categories at exactly 0.650.
But 0.650 does not have one universal interpretation. For one category it reflects an output that failed to make useful progress. For the restraint cases below, the same number can be the arithmetic penalty imposed on a correct identity decision.
That ambiguity is itself evidence that the aggregate needs the per-case record beside it.
6. The floor, and what the metric cannot represent
Two categories are pinned to the do-nothing floor, for completely different reasons. One is a program failure. The other is a metric failure, and it is ours.
invalid_output at 0.650 is a program-side failure in this development run. The case is designed to exercise output-shape risk, but the frozen v1 hard gate does not actually test general output shape; it only checks required entities and forbidden terms.
So a candidate can survive the gate and still land at the structural floor. The target_failure label tells us what the case was designed to probe; the component record tells us what the metric actually observed.
unnecessary_edit at 0.650 is different. The two development cases in this category, ed-034 and ed-037, have identity references: the correct reference output is exactly the source sentence.
For any such identity-reference case, a program that correctly leaves the sentence unchanged scores:
changed: 0 (nothing changed — correct)
scope: 1.0 (perfect — correct)
overlap: 1.0 (identical to the reference — correct)
score = 0.35·0 + 0.25·1 + 0.40·1 = 0.65
For identity-reference cases, the metric caps a correct unchanged output at 0.65 by construction. The changed component is 35% of the score and it pays for editing regardless of whether editing was required.
The full corpus contains six cases labelled unnecessary_edit, but that label does not mean all six references are byte-identical to the source. Some ask for restraint or minimal intervention rather than literal identity. The defect is therefore better stated as a property of the metric: whenever no change is the correct answer, v1 has no way to award the missing 0.35.
This is the Chapter 2 problem arriving on schedule. The contract has no explicit semantic state for “no edit needed,” and the metric independently encodes a prior that changing text is always worth points. Those two design choices can compound: an unnecessary edit may outscore a correct decision to leave the sentence alone.
Chapter 18 finds the same tension from another direction when literal reference-blocking collides with cases whose admissible answer can equal the reference. The lesson propagates because the representation, metric, and leakage policy all made assumptions about what a valid edit must look like.
We did not silently repair v1 after seeing this defect. Once Chapters 10 through 13 optimize against a frozen metric, changing its weights or adding a special case would create a different experiment and invalidate direct comparison with the runs already recorded.
That does not mean the defect is comparison-neutral. A candidate that edits an identity-reference case can gain the 0.35 changed credit and therefore outrank a candidate that correctly declines. v2 inherits the same structural blend, so its semantic judge does not fix this particular incentive either.
A proper repair would require a new metric version — ideally one that scores the decision whether to edit explicitly rather than assuming change is always desirable — followed by a new experimental cycle.
For this cycle we freeze the defect, publish it, and keep its direction visible when interpreting later scores.
A metric with a documented defect is more useful than a metric with an undocumented one. But notice what has happened. Before Chapter 9 attacks the metric deliberately, routine inspection of the per-case table has already found a place where the objective encodes the wrong preference.
7. What deterministic checks can catch
Not every failure needs a model to detect it. A surprising number of production text failures are cheap, boring string problems, and catching them before a judge is involved saves cost and removes noise:
helper text leaked into the output
prompt fragments in the response
TODO or placeholder markers
markdown artifacts in plain-text fields
unbalanced quotation marks
terminal punctuation changed
more than one sentence returned
length wildly disproportionate to the original
Every one of those has a deterministic component, so none needs an LM judge merely to detect its surface form.
That does not mean the frozen v1 hard gate already checks all of them; it does not. Its gate is limited to required entities and forbidden terms. Additional deterministic validators can live before scoring or become future hard-gate rules, but adding one changes the evaluation protocol and must be versioned.
The broader rule is: if a property can be checked deterministically and policy says failure is disqualifying, enforce it in software rather than asking a model to rediscover it.
The failure taxonomy for the editorial program is the target_failure vocabulary from chapter 7. Its repository-repair analogue in chapter 20 has the same shape:
patch does not apply
tests still fail
unrelated files changed
validation unavailable
candidate produced by fallback
Evaluation is where these become visible, which is only true if you store enough per-case detail to see them.
What Usually Goes Wrong
| Symptom | Likely cause | How to diagnose it | What to change |
|---|---|---|---|
| The baseline cannot be reproduced | Program, model, data, or protocol identity was not recorded | Compare run manifests and fingerprints | Persist the full evaluation protocol |
| A candidate looks better after the metric changed | Baseline and candidate ran under different protocols | Compare protocol fingerprints | Treat the comparison as invalid until they match |
| Evaluation crashes on one malformed output | The metric assumes well-formed predictions | Re-run with the traceback | Make the metric defensive and assign a failure score |
| The aggregate looks fine and users complain | The metric is too narrow to see the failure | Group per-case scores by expected failure mode | Add components that can see what users notice |
| A whole category sits at one exact value | A metric component or floor may dominate the category | Decompose the value into component weights and inspect the cases | Fix it in a new metric version, or publish the defect and its direction |
| Scores are quoted to many decimal places | Arithmetic precision is being confused with evidential precision | Repeat the same frozen program under the actual execution schedule | Keep full precision in artifacts; use appropriately rounded prose plus the repeated values |
| Optimization starts before a baseline exists | Evaluation and optimization were blurred together | Look for candidate state predating the baseline record | Freeze the baseline first |
Conclusion
We have a frozen baseline: a named program, a named metric with an explicit version, and a fingerprinted protocol. Two reported executions put its 11-case development mean at 0.7893 and 0.7992, so the prose-level baseline is about 0.79, not one sacred four-decimal number. ed-003 remains a useful walkthrough at about 0.877, and v2 agrees with v1 case for case in those reported baseline runs.
Two things from this chapter travel further than the number.
The first is that the aggregate was the least informative thing we computed. Grouping by expected failure mode turned about 0.79 into a diagnosis, and the diagnosis exposed a defect: the two development unnecessary_edit cases sit at exactly 0.650.
That is not a coincidence. On identity-reference cases, an unchanged candidate gets full scope and overlap credit but loses the entire 0.35 changed component. The broader six-case unnecessary_edit category is not uniformly identity-reference, so the correct general claim is structural rather than numerical: v1 systematically under-rewards cases where restraint is the right decision.
The second is what that implies about the whole enterprise. We built v1 from deterministic components and put two explicit disqualifying checks in front of its weighted score. That still leaves properties the gate does not inspect and a changed-term incentive that can reward the wrong behavior.
We found one of those defects by reading a grouped table, not by any clever adversarial technique.
So the obvious question is what else is in there.
We removed the assumption that optimization can begin before measurement, and the assumption that a larger number beats a baseline regardless of the protocol that produced it.
What remains is the metric itself, which so far we have mostly used rather than challenged. The quality comparisons in Chapters 4 and 5 depend on its v1 structural score, and the optimizer chapters will deliberately search against that frozen objective.
We now know one place where the objective is wrong before the optimizer has even touched it. Its only remaining qualification cannot be that we wrote it.
That is not a good enough reason to trust it.
What happens when the measurement stops being an observer and becomes the target of search?