When the Metric Becomes the Target
Attack the objective on purpose, separate scoring from diagnosis, and decide what to do once optimization can exploit an instrument you have proved is incomplete.
Chapter 8 built a metric and found a defect in it by reading a table. Six cases where the correct answer is capped at 0.65, discovered by grouping per-case scores and noticing that a whole category sat on an exact number.
That was luck dressed as diligence. This chapter looks on purpose.
The urgency comes from what happens next. A DSPy optimizer does not want your program to be good. It wants your number to go up, and it will find every route to that outcome, including the ones you did not intend to leave open.
bad metric
+
capable optimizer
=
efficiently optimized mistake
By the end of this chapter we will have found a class of candidate that scores between 0.85 and 0.96 while saying the opposite of what the author wrote, built a second metric that catches it, and validated that second metric rather than trusting it.
Then we have to decide what to do about it, and that decision is the hinge of the book.
1. The metric ladder
Different tasks tolerate different instruments.
| Metric type | Useful when | Weakness |
|---|---|---|
| exact match | There is exactly one correct string | Useless for open-ended rewriting |
| deterministic invariants | Violations are computable | Says nothing about quality |
| reference comparison | The reference captures target behavior | Penalises valid alternatives |
| semantic similarity | Meaning preservation matters | Rewards bland paraphrase |
| LM judge | Human-like judgment is required | Bias, drift, cost, provenance |
| human preference | Real acceptance is the only truth | Slow, sparse, expensive |
| composite | Several failure modes matter at once | Harder to explain, harder to attack |
Editorial rewriting needs a composite, because its failures are not all the same kind. Renaming a character is a hard violation. Producing flat prose is a soft quality judgment. And the dangerous cases are neither — they are candidates that satisfy every implemented check while violating something the metric never encoded.
Those are the ones an optimizer finds.
2. Attacking the metric
An attack suite is a fixture for your metric. Where chapter 7’s corpus tests the program, this one tests the instrument, and the questions are adversarial:
Can a clearly bad candidate score well?
Can a trivial copy exploit the score?
Can verbosity inflate it?
Can lexical closeness beat actual quality?
Can a candidate reverse the meaning and survive?
The suite is deterministic. Every attack candidate is constructed by string manipulation from the case’s own fields, so no model runs, the results are reproducible, and the whole thing costs nothing:
def adversarial_predictions(example):
return {
"copy_original": dspy.Prediction(rewritten_text=example.sentence),
"copy_reference": dspy.Prediction(rewritten_text=example.reference_rewrite),
"reference_plus_junk": dspy.Prediction(
rewritten_text=example.reference_rewrite + " This is clearly improved."
),
"rename_entity": dspy.Prediction(
rewritten_text=str(example.reference_rewrite).replace("Jalen", "Jason")
),
"empty_output": dspy.Prediction(rewritten_text=""),
"over_compressed": dspy.Prediction(rewritten_text=first_clause(example.sentence)),
"lexical_alternative": dspy.Prediction(rewritten_text=paraphrase_far(example)),
"semantic_flip": dspy.Prediction(rewritten_text=reverse_meaning(example)),
}
The results split cleanly into two groups.
The deterministic gate did its job. Required-entity renames failed the hard gate and scored 0.0. Empty outputs failed the hard gate and scored 0.0. The scope component reduced the reference-plus-junk attack below a clean reference, so appending confident-sounding filler does not pay.
The semantic attacks walked straight through.
3. The failure the metric cannot see
Three attacks, one per canonical case. Each takes the reference rewrite — the known-good answer — and reverses its meaning while changing as few words as possible.
ed-001. The scene is tense; Jalen is afraid. The attack changes the fear to calm. Score under v1: 0.960.
ed-002. Mira is tired because she walked all day through rain and mud. The attack negates the causal relationship. Score under v1: 0.927.
ed-003. Anna is warning her brother that a place is unsafe. The attack turns the warning into approval and changes unsafe to safe. Score under v1: 0.846.
Every one passed the hard gate, preserved the required entities, avoided the forbidden terms, returned a single sentence, and stayed proportionate in length. Every one shares most of its vocabulary with the reference, which is exactly why the overlap component — 40% of the score — rewards it.
A candidate can satisfy every property the metric knows how to inspect
while violating the property the task actually requires.
The raw scores are alarming, but compare like with like. Chapter 8’s 0.79 is an eleven-case development mean; the three attack scores above are individual-case scores. They cannot honestly be ranked against one another.
The directly comparable ed-003 result is more useful. The frozen program scores about 0.877 on that case. The meaning-reversing attack scores 0.846. The attack is slightly worse — and still implausibly close to a successful rewrite for a candidate that reverses Anna’s warning.
That is the failure we need: v1 places a deliberately wrong candidate in the same high-score region as acceptable output because the property that failed is absent from the score.
Notice also what the metric had available and did not use. Every row carries semantic_constraints. For ed-003, the fixture records do not change the warning and do not change the speaker. Those requirements were written down in chapter 7, stored with the row, and checked by nothing in v1. The information was present. The instrument had no way to consume it.
4. Building v2
The fix is to consume it. Version 2 keeps everything v1 does and adds a third stage:
1. hard gate deterministic, binary, unchanged
2. structural score 0.35·changed + 0.25·scope + 0.40·overlap, unchanged
3. semantic check an LM judge, one call per constraint ← new
The judge is a DSPy program like any other, with a typed contract:
from typing import Literal
import dspy
class SemanticConstraintJudge(dspy.Signature):
"""Decide whether a one-sentence rewrite still satisfies a single stated
requirement about meaning."""
original_sentence: str = dspy.InputField(
desc="The sentence before editing."
)
candidate_rewrite: str = dspy.InputField(
desc="The proposed one-sentence rewrite."
)
editorial_goal: str = dspy.InputField(
desc="What the edit was asked to achieve."
)
requirement: str = dspy.InputField(
desc="The single meaning requirement to check."
)
verdict: Literal[
"preserved",
"violated",
"unclear",
] = dspy.OutputField(
desc="Whether the rewrite satisfies the requirement."
)
reason_code: Literal[
"meaning_preserved",
"meaning_reversed",
"meaning_weakened",
"meaning_strengthened",
"entity_altered",
"scope_changed",
"register_changed",
"output_shape_invalid",
"cannot_determine",
] = dspy.OutputField(
desc="The primary reason for the verdict."
)
explanation: str = dspy.OutputField(
desc="One sentence explaining the verdict."
)
Three design decisions worth naming.
One call per constraint, not one call per case. A judge asked to evaluate four properties at once will blur them. Asked about one property, it either finds a violation or does not, and the result is attributable.
unclear is a first-class verdict. A judge that must choose between pass and fail will guess, confidently. Parse failures also land in unclear, so a malformed response degrades the score rather than crashing the run.
Every judgement records provenance — model, provider, configuration, signature schema fingerprint, input fingerprint, judge version. Chapter 6 explained why: if the judge changes, the score changed even though the candidate did not.
The components combine like this:
semantic_factor = fraction of constraints judged "preserved"
score = structural × (0.30 + 0.70 × semantic_factor)
if any constraint is "violated": cap the score at 0.30
if the hard gate fails: score is 0
Run ed-001’s semantic flip through it:
structural score: 0.960 (unchanged — it is still lexically excellent)
constraints: 2 of 2 violated
semantic_factor: 0.0
0.960 × (0.30 + 0.70 × 0.0) = 0.960 × 0.30 = 0.288
0.960 becomes 0.288. The other two flips land near 0.29 by the same route.
5. Validating the judge
Here is the step that separates a metric from a hope.
A judge is a program. It has inputs, outputs, failure modes, and no more inherent claim to correctness than the program it judges. Adding an LM judge and trusting it because it is an LM is the same mistake as trusting a chain-of-thought trace because it is reasoning — and it matters more here, because this component decides what counts as success for everything downstream.
So the judge gets its own fixture and its own two gates.
Gate 1 — sensitivity. Does it catch known-bad?
| Attack | Constraints violated | v1 | v2 |
|---|---|---|---|
ed-001 fear → calm | 2 of 2 | 0.960 | 0.288 |
ed-002 causal negation | 2 of 2 | 0.927 | ≈0.29 |
ed-003 warning → approval | 1 of 2 | 0.846 | ≈0.29 |
Three of three caught. Agreement with the known labels: 1.00.
Gate 2 — specificity. Does it reject known-good?
This gate is the one people forget, and skipping it makes Gate 1 meaningless. A judge that returns violated for everything catches all three attacks perfectly and is worthless.
So run it against all 44 reference rewrites — text that is correct by construction.
cases evaluated: 44
constraint judgements: ~90
verdicts of "violated": 0
verdicts of "unclear": 0
parse failures: 0
Zero false positives. The judge accepts every known-good rewrite and rejects every known-bad one. Cost is about 1.3 seconds and 1,800 tokens per case — roughly a doubling of the evaluation budget.
For the record, so the run is auditable rather than merely asserted:
judge model ollama_chat/qwen3:latest (same endpoint as the task program)
temperature 0.0
max_tokens 512
think / cache off / off
seed 7
call granularity one call per named constraint
dataset fingerprint 40ab192b…d4fe46
per judgement model, config, signature-schema fingerprint,
input fingerprint, prompt fingerprint, verdict, reason code
Every number in the two gates above is reproducible from that configuration and the frozen fixture.
What the validation does not establish
The attribution problem is deeper than the reason code. On ed-001, the judge marked both named constraints violated with reason meaning_reversed, although neither constraint literally says to preserve Jalen’s fear. The task-level direction is useful — the judge notices a real semantic regression — but the per-constraint attribution is approximate.
That matters because v2 consumes the verdict, not merely the reason code. If a violated verdict is attached to the wrong requirement, it can still lower the score. The validation therefore supports v2 as a guardrail on the cases we tested; it does not make every per-constraint verdict authoritative.
The closed reason_code vocabulary remains useful for machine-readable records, but aggregation does not make attribution true. Do not build a failure-mode dashboard on these codes until the attribution itself has been validated.
The judge shares the task model. It runs on qwen3:latest, the same weights and endpoint as the program it evaluates — the role collapse chapter 6 warned about. It passed both gates cleanly, which is evidence, not proof.
Two gates on a small suite is a beginning. Three attacks and 44 clean rewrites establishes that the judge is not obviously broken. It does not characterise its behavior on the failure modes nobody thought to construct — and the whole lesson of this chapter is that those exist.
6. What v2 fixes, and what it does not
v2 catches the three canonical semantic flips used to validate it. It does not make the metric correct, and it does not establish that every form of semantic drift is now covered.
It inherits every structural defect of v1. In particular, whenever an unchanged answer is correct, the shared structural blend still withholds the 0.35 changed credit. The full corpus contains six cases labelled unnecessary_edit, but that label is broader than literal identity-reference behavior, so the defect should be stated as a property of the metric rather than as a universal 0.65 cap on all six rows. A judge cannot fix a weighting rule that rewards change by construction.
In the two reported Chapter 8 baseline runs, v2 equals v1 case for case because the judge records no violated constraints on those outputs. In the evidence collected so far, v2 therefore behaves less like a finer ruler and more like a tripwire: ordinary baseline output leaves the structural score untouched, while particular semantic regressions make the two metrics separate sharply.
That is an observed pattern, not a rule that v2 only fires after optimization.
And v2 is expensive, non-deterministic, and dependent on a model. v1 is deterministic and cheap. v2 adds an LM call per constraint and inherits the task model’s blind spots.
So we have two metrics and neither is adequate. v1 is an inexpensive ranking signal with known blind spots. v2 can expose some failures v1 cannot see, but it is costly, model-dependent, and still not a diagnosis of what should change.
7. A score is not a diagnosis
There is a second inadequacy here, and it is easy to miss because it is not about accuracy at all.
ed-035 — a marketing sentence about a water filter, which will matter enormously in Chapters 11 through 14 — can be damaged by deleting one word. Under v1 that costs 0.033. Under v2 it costs 0.700.
Neither number tells you the word was regularly.
That distinction is invisible while you are only asking whether one candidate should rank above another. It becomes decisive when the optimizer must use evaluation not only to select a candidate but to propose the next change.
An evaluator therefore has two independent jobs, and until now we have audited only the first:
| Evaluator output | Question it answers | v1 | v2 |
|---|---|---|---|
| Selection signal | How should this candidate rank? | Available, but incomplete | Available, somewhat broader |
| Search-guidance signal | What failed, in a form that can guide the next proposal? | Absent | Weak — verdicts exist, attribution is not yet reliable |
A good selection signal does not imply a good search-guidance signal. A system can rank candidates correctly enough to choose between them while still being unable to explain which program behavior should change next.
That distinction costs us little with BootstrapFewShot and MIPROv2 because their search procedures can operate on scalar evaluations. It becomes central with GEPA, whose reflection model needs textual failure information to generate the next instruction proposal.
Hold on to the shape of it.
flowchart LR
S[scalar score, e.g. 0.967] --> R[ranking: is this candidate better?]
S -. does not provide .-> D[repair target: what should change next?]
N[named failure, e.g. qualifier:regularly] --> D
0.967 says where a candidate sits in a ranking. Failed constraint: qualifier:regularly identifies a repair target. A reflective optimizer needs both kinds of information: a score to decide whether the child scored better, and a diagnosis rich enough to propose the next child.
8. The decision
We have proved the objective is inadequate — twice over, on accuracy and on diagnosability. What should an engineer do next?
There are two answers and this book is going to use both.
The first is to optimize through it. This is what most projects do, not out of carelessness but because the metric is what you have, the deadline is real, and the number does go up. It is worth seeing precisely what that looks like, because a great many readers of this book are currently living inside it.
So the next four chapters carry one negative-control sequence forward. Chapter 10 freezes the compile boundary. Chapter 11 gives BootstrapFewShot the v1 objective. Chapter 12 gives the same objective to MIPROv2. Chapter 13 gives it to GEPA. In each case, v1 remains the search objective and the semantic check stays outside the main search loop as a second reading of the resulting behavior.
This is a negative control, not a failed attempt. We already know the instrument is incomplete. What we do not yet know is how three different optimization mechanisms will express that incompleteness: demonstration selection, instruction search, and reflection do not fail in the same way merely because they share an objective.
The second is to change the objective. Once you have proved the instrument is wrong, continuing to optimize against it is not rigour, it is sunk cost. The engineering response is to build an objective that can carry the property you actually need.
For this task, that means abandoning reference comparison entirely and asking what is deterministically checkable about a good rewrite:
did the named entity survive?
did the number survive, unchanged?
did the scope-bearing qualifier survive?
did the negation survive?
did the forbidden term stay out?
is it within the length budget?
Every one of those checks can be deterministic, free, instant, and — crucially — nameable. A failure is not merely a lower score; it can be qualifier:regularly, negation:not, or number:30.
That does not make the new objective correct by construction. Determinism removes model variance from the check; it does not prove that the check encodes the right rule, handles every representation, or composes correctly with the others. The replacement objective will have to be attacked just as aggressively as this one.
But it gives us something v1 and v2 do not: a failure channel whose vocabulary is tied directly to the declared constraint. That is the property the second half of the experiment needs.
The new objective runs on the same 44 sentences, with the same program, model, and optimizer mechanisms. The experimental intervention is the objective and the feedback it exposes. Chapter 1 promised two evidence regimes; this is where the first reaches its limit, and why the second half of the book changes what counts as evidence rather than merely trying a stronger optimizer.
What Usually Goes Wrong
| Symptom | Likely cause | How to diagnose it | What to change |
|---|---|---|---|
| Training score rises while outputs get worse | The metric rewards a proxy | Read the highest-scoring outputs, not the lowest | Add a component that sees the real property |
| A semantic reversal scores highly | The property is not represented at all | Construct meaning-reversal attacks from your references | Add a validated semantic check |
| Copying the original scores well | No change incentive in the metric | Run a no-edit baseline | Reward change where change is required |
| Entity renames survive | Similarity outweighs invariants | Test adversarial renames | Make entity preservation a gate, not a component |
| The judge agrees with everything | Gate 2 was never run | Score known-good references and count violations | Measure false positives before trusting verdicts |
| The judge’s reason codes look precise | Attribution was never validated, only direction | Check whether the named constraint is the broken one | Use verdicts for scoring, not reason codes for reporting |
| A reflective optimizer proposes nothing useful | Its score may rank candidates, but its feedback does not identify an actionable failure | Inspect the exact score-and-feedback payload the reflection model received | Keep selection signal and search-guidance signal separate; make failures nameable before blaming reflection |
| The metric crashes mid-optimization | It assumes well-formed predictions | Fuzz malformed and empty outputs | Defensive access, and a defined failure score |
| You keep optimizing a metric you distrust | Sunk cost, or the metric is all you have | Ask what a candidate would have to do to be wrong and still win | Build an objective that can express the property |
Conclusion
We attacked our own metric and it failed in the way that matters most. Three candidates that reversed the author’s meaning while preserving the vocabulary scored 0.960, 0.927, and 0.846 — every one above the frozen program’s baseline of 0.79.
The information needed to catch that had been sitting in the fixture since chapter 7, consumed by nothing. v2 consumes it, through a judge with a typed contract, one call per constraint, unclear as a real verdict, and provenance on every judgement. It reduces those three attacks to roughly 0.29. Then we did the part that is easy to skip: 3 of 3 known-bad caught, 0 false positives across 44 known-good. Both numbers were required; either alone is compatible with a broken judge.
And still neither metric is adequate. v1 supplies a cheap ranking signal while remaining blind to load-bearing semantics. v2 catches the tested semantic failures but is expensive, model-dependent, and weak at attribution. Neither gives reflective search the diagnostic channel we will eventually need.
We removed the assumption that a reference rewrite is the same thing as editorial quality. We removed the more dangerous assumption that passing every implemented check means the intended meaning survived. And we removed one that has been implicit since chapter 8: that evaluation has only one job.
For ordinary candidate selection, a scalar may be enough to rank alternatives. For reflective optimization, evaluation has a second job: it must expose failure in a form capable of guiding the next proposal. Selection signal and search-guidance signal are different engineering objects.
That is the hinge. Across Chapters 10 through 13 we will deliberately keep the inadequate scalar objective and watch three optimizer mechanisms interact with it. Then we will change the objective and the feedback channel rather than continuing to spend search budget on an instrument we already know is wrong.
What does it actually look like when a capable optimizer pursues an objective that does not represent what you want — and what changes when the objective can finally name the thing that failed?