Change One Thing
Hold the corpus, program, model, optimizer configurations and evaluation harness fixed, replace only the objective, and test whether optimization changes direction when failure becomes deterministic and nameable.
The previous three chapters ran three optimizers against a metric we had already proved was blind, and got what chapter 9 predicted: a search that pursued the objective faithfully and damaged the thing the objective could not see.
That result has an obvious interpretation and a correct one, and they are not the same.
The obvious interpretation is that the optimizers underperformed. The correct one requires an experiment, because “the objective was the problem” is a hypothesis and we had not tested it. Everything we knew was consistent with a second explanation — that DSPy’s optimizers do not do much on a small local model regardless of what you point them at.
So this chapter changes exactly one thing.
flowchart LR
F[everything else held fixed] --> E[same experiment, re-run]
M1[editorial_metric_v1] -->|swapped for| M2[deterministic constraint satisfaction]
M2 --> E
Held fixed: the 44 sentences, the program architecture, the model and its quantization and temperature, the optimizer classes and budgets and seeds, the splits and ordering, and the seven-session paired harness. Only the objective changes. That is the whole design. It is not a new task, a bigger model, or a better search. It is the same experiment with the ruler swapped.
What came back was not a clean optimizer victory. It was more informative: the three optimizers separated according to the kind of signal their search procedures could use.
1. Resisting the temptation to fix everything
The strongest instinct here is to improve the experiment while you are in it. The new objective is deterministic and free — no LM calls, instant to compute — so an optimizer could afford far more search than the editorial runs allowed.
We did not raise a single budget.
Bootstrap kept its 0.70 threshold and four demonstrations. MIPRO kept two candidates and four trials. GEPA kept max_metric_calls = 66 and seed 13. Every configuration is copied verbatim from the previous chapter’s runs.
The reason is worth stating because it is easy to get wrong. If the objective changes and the budget changes and the result improves, you have learned nothing about which change did it. A cheap metric is an invitation to run a better experiment later. It is not a licence to run a different one now.
The same applies to the corpus, the split, the model, and the paired harness. Both dataset and split fingerprints match the earlier chapters exactly.
2. An objective that can name its failures
Chapter 9 argued that a metric has two jobs — measuring the right property, and being able to say what went wrong — and that ours did neither well. The replacement is built around both.
Every case gets a frozen set of hard_constraints, derived mechanically from the sentence, goal and context, then reviewed by a human and frozen before anything runs. Six kinds, each binary:
| Constraint | Check |
|---|---|
entity:<X> | token X still present |
number:<N> | numeric value N present and unchanged |
qualifier:<W> | scope-bearing token or phrase W still present |
negation | negation count equivalent to the original |
forbidden:<X> | token or phrase X still absent |
max_words:<N> | no more than N words |
qualifier constraints come from scanning for scope-bearing tokens: regularly, only, may, might, up to, at least, within, usually, often, sometimes, partly, approximately, generally, typically, unless, if. Multi-token phrases are matched as phrases — up to is not two independent checks.
The scoring rule keeps chapter 8’s separation between gates and scores. An entity, number, negation, forbidden or change-rule failure sends the case to zero; partial credit must not reward a candidate for breaking a mandatory part of the contract. Otherwise, qualifier and max_words supply the fractional selection signal. The reported frac result is the mean of these per-case scores, not the proportion of all 315 checks satisfied. That distinction prevents numerous easy forbidden checks from dominating the objective.
Two properties of this objective matter more than its details.
The reference rewrite is not part of scoring or optimization. It is used once, before the freeze, as a diagnostic probe: if a known-good rewrite conflicts with a derived constraint, the guard raises an assertion and requires a human decision grounded in the task contract. It cannot delete or relax the constraint itself. During compilation and evaluation, neither the reference nor a target rewrite enters the metric or its feedback, so no gold answer can leak into search through this objective.
A failure has a name. Not 0.967. qualifier:regularly.
And the canonical test: ed-035, the water-filter sentence the previous chapter’s optimizers damaged by deleting one word, produces the constraint qualifier:regularly. Dropping that word is now an explicit named failure rather than a rounding error.
3. Deterministic is not the same as correct
Here is where the chapter earns its place, because the natural conclusion after chapter 9 is deterministic metrics are the safe kind, and that conclusion is wrong.
Before a single model call, we scored the frozen baseline’s existing outputs under the new objective and inspected the results. The audit found three implementation bugs that falsely zeroed eight cases, plus a fourth protocol bug in the guard used to investigate six constraint-reference conflicts.
Whitespace forbidden terms. Some rows carried forbidden values like "\n" and " ". Stripped, these became the empty string, and the empty string is a substring of everything. Four cases were being scored as permanent failures.
Punctuation-only edits read as unchanged. The change-rule gate compared normalised text and treated punctuation as noise. ed-003’s entire editorial goal is dialogue punctuation — the correct rewrite adds quotation marks and splits a sentence. Three cases were falsely zeroed for making exactly the edit they were asked to make.
A surface-form normalisation miss. entity:200C did not match the equivalent surface form 200°C. One case falsely zeroed.
Six constraint-reference conflicts. Six mechanically derived constraints were not satisfied by the known-good reference rewrite. That did not prove that the constraints were wrong or that the reference was decisive. It proved that the derivation, task contract and answer key disagreed, and that an explicit decision was required before the objective could be frozen.
The first three bugs scored eight of 44 cases wrongly despite involving no model, randomness or judgement call. The fourth failure was different and more dangerous: instead of producing a visibly wrong score, the validation guard silently changed the objective before the contradiction could surface. Determinism buys reproducibility and inspectability. It does not buy correctness. A deterministic metric can be wrong the same way every time, allowing a stable error to look like a stable fact.
A deterministic metric is not automatically a correct metric. It becomes trustworthy because its rules are inspectable and testable.
Twelve unit tests now cover the objective. That is not ceremony — every one of the bugs above is a case those tests would have caught.
The guard that adapted to fit
That fourth failure produced the most important procedural fix.
The first implementation had a guard: when a reference rewrite could not satisfy a derived constraint, the guard removed the constraint. It seemed reasonable. It was silently editing the measuring instrument to agree with the answer key.
That is the same failure as tuning a metric until your program passes, arrived at from the opposite direction and dressed as hygiene.
The replacement never mutates anything. It runs the check, and asserts that every constraint a reference cannot satisfy has an explicit human decision recorded against it — justified from the sentence, goal, context, and task contract, never from the reference.
Six decisions resulted:
| Case | Constraint | Decision | Justification |
|---|---|---|---|
ed-012 | forbidden:It was 1994 | removed | Unanchored substring; also fires inside the acceptable hedge “I think it was 1994” |
ed-012 | qualifier:might | kept | Scope-bearing uncertainty marker; the context says the narrator is deliberately uncertain |
ed-017 | max_words:13→15 | raised | The goal requires “tended to score” for “also scored” — two tokens the constraint did not allow |
ed-021 | max_words:15→17 | raised | Repairing “said…allegedly” is inherently longer |
ed-025 | negation:1 | removed | The goal is plain language; “not later than X” has no negation in its plain form “within X” |
ed-044 | qualifier:if | removed | “if” heads the politeness frame the goal targets; it does not restrict the request |
Final tally across 315 derived constraints: 309 accepted unchanged, one kept after review, three removed, two raised. Every deviation is on the record with a reason that does not mention the reference.
The negation check also had to get smarter rather than simpler. ed-011 rewrites “did not say very much” as “said little” — a lexicalised negative, and a good edit. The check now credits words like little, few, unable, hardly, lacks, fails toward an under-count, directionally: it can close a gap, never open one. ed-011 passes; ed-017, which genuinely adds a negation, still fails. The residual risk — a rewrite lexicalising a negation with a word outside the list — is documented and accepted rather than hidden.
4. Measure the instrument before you run
Chapter 7 introduced the arithmetic of what a fixture can detect. This experiment needed the same check on the objective, and it is the step that would have saved us most.
If the frozen baseline already satisfies almost every constraint, there is nothing for an optimizer to improve, and a flat result would read as the diagnosis was wrong when it actually means the instrument has no room.
So we scored the baseline first, with zero model calls, reusing the seven frozen baseline sessions already on disk. The bracketed values below are the observed session minima and maxima:
| Instrument | Baseline | Governing? |
|---|---|---|
| Fractional-only score | 0.806 [0.775–0.829] | Yes — 19 points of headroom |
| 100%-satisfaction rate | 0.761 [0.730–0.784] | Yes — 24 points |
| Fixable failures per session | ~9 (qualifier 3, negation 1, max_words 5) | Yes |
satisfied / total, all kinds | 0.946 | No — diluted by 137 forbidden checks |
| Zero-factual-violation rate | 0.973 | No — the baseline barely corrupts facts |
The last two rows are the important ones. The obvious aggregate — constraints satisfied over constraints checked — reads 0.946 and looks like there is nothing to do. It is diluted by 137 forbidden checks the baseline passes trivially, because a program that was never going to use a forbidden term gets full marks for not using one.
Choosing the wrong headline number would have made a real 19-point gap look like a 5-point one, and we would have concluded the experiment could not run.
This produces a discipline worth generalising: before running anything, ask which of your candidate aggregates can actually move, and declare that one as governing. An average over checks your program already passes is a measure of how many easy checks you wrote.
The headroom is also narrow and specific, and the report is honest about it. Entities, numbers and forbidden terms are already near-perfect — those constraints act as downside detectors, where an optimizer can only hold or break them. All the room is in qualifiers, negation and length. Which is exactly the ed-035 story.
5. Two arms
Arm 1 — disclosed. The constraint list is added to the program’s inputs. This is a compliance task and a positive control: if optimization cannot improve compliance when the program is told exactly what to satisfy, the harness is broken and nothing downstream can be trusted. It does not isolate the objective, because the program’s contract changed.
Arm 2 — undisclosed. The program sees only sentence, goal and context, exactly as in every previous chapter. The constraints exist only on the evaluation side. The program must infer which tokens are load-bearing.
Arm 2 is the causal experiment, and it reproduces the original situation precisely: nobody told the model that regularly mattered in the previous chapter either.
GEPA’s feedback in both arms is failed-constraint names only — Failed constraint: max_words:12 — never a reference, never a target sentence. Across 149 compile feedback calls in the two arms, zero reference leaks.
6. What happened
Arm 2, the development family, 7 paired sessions. frac is the fractional score; factual violations zero the case.
| Program | frac (mean, sd) | 100% satisfied | Factual hard fails | Paired Δ |
|---|---|---|---|---|
| Baseline | 0.9091, sd 0 | 0.909 | 0/77 | — |
| BootstrapFewShot | 0.9091, sd 0 | 0.909 | 0/77 | +0.0000, sd 0 |
| MIPROv2 (baseline state returned) | 0.8442, sd 0.044 | 0.844 | 0/77 | −0.0649† |
| GEPA | 1.0000, sd 0 | 1.000 | 0/77 | +0.0909, sd 0 |
† MIPRO returned the byte-identical baseline state. The negative value is the paired difference between fresh generations from identical program states under residual decoding nondeterminism, not a compiled-candidate regression. Section 7 separates the state-level null from the output-level variance.
Set against the previous chapter:
| Optimizer | Editorial v1 (paired) | Constraint objective, Arm 2 |
|---|---|---|
| BootstrapFewShot | +0.0003 | +0.0000 |
| MIPROv2 | +0.0199 | selects the default; no candidate |
| GEPA | no accepted child | +0.0909, 7/7 sessions |
Two of the three optimizer outcomes changed direction. MIPRO went from selecting a reproducibly higher-v1 candidate to returning the default state. GEPA went from rejecting every child to selecting one in every session. Bootstrap remained net-flat under both objectives, although the new per-case accounting reveals an exact trade rather than noise.
These are different scales measuring different properties, so do not subtract one column from the other. What is comparable is direction, reproducibility, and what each optimizer actually learned.
7. Reading the three results
MIPRO produced a state-level null, not a regression. It ran 59 compile metric calls, 48 above threshold, and found no instruction that beat the default on the constraint valset, so it returned the baseline program byte-identical. Fresh calls from that identical state still crossed the max_words cliff often enough to produce the −0.0649 paired estimate. The experiment therefore records two separate facts: compilation selected no program change, while generation remained nondeterministic.
Why did its previous gain vanish? Because that gain came from instruction changes scored against a valset containing ed-035 under a lexical metric. The constraint objective does not reward those changes. The search space did not shrink; the reward for searching it did.
Bootstrap’s flat result is a trade, not noise. This is the sharpest mechanical finding in the chapter. On the development family it fixes ed-032 — 13 words down to 10, satisfying max_words — in all seven sessions. And in all seven sessions it returns ed-038 byte-identical to the original, which fails the change rule and scores zero.
Plus seven, minus seven, net exactly zero, reproducibly.
The demonstrations were drawn from easy training cases where a minimal edit sufficed, so they taught conservatism. That is useful on a case needing a shorter sentence and fatal on a case needing an edit at all. The previous chapter called Bootstrap’s flatness noise. It is not noise — it is a specific, reproducible failure mode of demonstration-only search on this program.
GEPA improved, and its mechanism is legible. It selected a child in every session, +7 cases and −0, with the compile valset moving 0.909 → 1.000. Its rewritten analyze instruction includes explicit negation handling, “preserve scope-bearing tokens”, and “stays within a maximum of 15 words”.
It reached that from feedback strings reading Failed constraint: max_words:12.
Same search algorithm, same budget of 66 metric calls, same model, same seed, same corpus. In the previous chapter GEPA rejected all three of its children. Here it accepts one, every time.
The result supports a narrower diagnosis: under the lexical metric, GEPA lacked actionable feedback. A reflective optimizer assumes that a failure carries enough information to suggest a change. Chapter 9 argued that a blended scalar does not identify such a target. In this experiment, replacing that scalar feedback with named failures was sufficient to produce a reproducibly selected instruction change. Because the child was selected and measured on the same development set, this establishes a change in search behaviour, not generalisation.
8. ed-035, and the floor
Two results that close threads running since chapter 2.
Nobody dropped regularly. Every program, both arms, all seven sessions produced one distinct output:
Using the filter regularly helps reduce limescale build-up in your kettle.
qualifier:regularly: PASS, 7/7.
Be precise about what this shows. The frozen baseline always kept regularly — the previous chapter’s regression was introduced by the optimized candidates. So this is a regression-prevention result: under an objective where deleting that word is a named failure, no optimizer deletes it.
editorial v1 : deleting "regularly" costs 0.033, and the optimizer does it
constraint : deleting "regularly" is a named failure, and it doesn't
The restraint floor is gone. All six unnecessary_edit cases score full credit unchanged — 42/42 for baseline, Bootstrap, MIPRO and GEPA in Arm 2. Chapter 2 predicted that a contract which cannot express “leave it alone” would propagate into a metric that penalises restraint; chapter 8 measured it as a hard 0.65 cap on 14% of the corpus. An objective that treats an unchanged output as legitimate on cases where restraint is correct simply does not have the problem.
The exception is instructive. Arm 1 MIPRO scores 35/42, because with the constraints handed to it directly it becomes over-eager and edits sentences it should leave alone. Telling a program exactly what must remain true apparently encourages it to demonstrate compliance by doing something.
9. What this shows, and what it does not
The headline is not “verifiable objectives make optimization work.” GEPA selected a beneficial child, MIPRO returned the unchanged baseline, and Bootstrap netted zero through an exact trade. Changing the objective did not uniformly unlock optimization.
It changed which search procedure received a useful signal.
That reframes the ladder as a question of signal type and search mechanism rather than a ranking by sophistication:
BootstrapFewShot filters demonstrations through the metric → demos teach both useful restraint and harmful under-editing
MIPROv2 uses a scalar to select program states → no bounded-search proposal beats the default
GEPA uses text to propose and a score to select → a named failure supplies an actionable target
That is an account of these runs, not a universal taxonomy of the optimizers. Now four honest limitations, because this result is easy to over-read.
GEPA selected against the development set. Its +0.0909 is measured on the valset it used for selection — the same status chapter 12 assigned MIPRO’s editorial gain. The selected result repeated across seven fresh sessions, all using the same frozen configuration. The magnitude is not out-of-search evidence, and the repeated sessions are not seven independent holdouts.
The gain is one case, in the least interesting class. The development baseline was already 10 of 11 perfect. GEPA fixes the eleventh, ed-032, by shortening it to fit max_words. Length is the shallowest of the six constraint types.
On the class that motivated the whole experiment, GEPA netted zero. Across the pooled train-plus-development diagnostics, it fixed seven qualifier failures and introduced seven. Bootstrap — the optimizer that achieved nothing on the governing development aggregate — netted +7 on that same diagnostic count. Because the training cases were optimizer-visible, neither number is a generalisation result. More importantly, the optimizer that “worked” did not improve the property ed-035 is about.
The winning instruction is partly memorised. It hard-codes “a maximum of 15 words” and quotes two phrases from a specific training case. It generalises to short development sentences and would misfire on a legitimately long one. GEPA learned a bound, not a principle.
None of that reverses the narrow finding. Across the development family, every program contributed 77 case-runs per arm, and none produced an entity, number, forbidden or negation hard failure. GEPA’s direction change repeated, its mechanism is visible in the selected instruction, and the objective made the previous chapter’s signature regression explicitly selectable against.
But the honest verdict is qualified: holding everything else fixed, changing the objective was load-bearing for GEPA’s proposal-and-selection behaviour in this experiment. It does not show that the optimization machinery was irrelevant, that GEPA generalised beyond its selection set, or that deterministic objectives are generally superior. The observed gain occupies a narrow band of headroom, is worth one development case, and comes from an optimizer measured on the set used to select it.
Which produces a prediction rather than a conclusion.
If GEPA improves in proportion to how much its feedback can say, then a task whose failures come with a test name, an expected value, a received value and a traceback should be its best case in this book. That is not an argument. It is something to go and measure, and the second half of this book measures it.
What Usually Goes Wrong
| Symptom | Likely cause | How to diagnose it | What to change |
|---|---|---|---|
| A new objective and a bigger budget improve results | Two variables moved | Check whether any configuration changed alongside the metric | Change one thing; a cheap metric is not a licence |
| A deterministic metric produces impossible scores | Empty-string matching, normalisation, or boundary bugs | Score known-good references and inspect every failure | Unit-test the metric before trusting a single run |
| Constraints quietly disappear during setup | A guard is editing the instrument to fit the answer | Look for code that mutates the metric based on the reference | Assert and require a recorded human decision instead |
| The aggregate says there is no headroom | It is diluted by checks the baseline always passes | Compute each candidate aggregate separately before running | Declare a governing instrument and justify it |
| An optimizer’s gain is flat and called noise | It may be an offsetting trade | Compare per-case, per-session, both directions | Count fixes and regressions separately |
| A reflective optimizer proposes nothing useful | Its feedback is a number | Read the feedback strings it actually received | Make failures nameable before changing the optimizer |
| The winning instruction contains literal values from training data | The optimizer memorised a bound | Read the selected instruction end to end | Test on cases outside the range it memorised |
| A selection-set gain is reported as an improvement | Selection evidence is being read as independent | Check which split the optimizer’s valset was | Report it as selection evidence, or spend a holdout |
Conclusion
We held the corpus, program, model, optimizer configurations and paired harness fixed, then changed the objective. MIPRO’s reproducible editorial gain became a state-level null. Bootstrap remained flat through an exact one-fix, one-regression trade. GEPA moved from rejecting every child to selecting a child that scored +0.0909 on the development set in all seven sessions, with no new development-family hard failures.
The mechanism is the central finding. Named feedback gave GEPA an actionable mutation target where the blended lexical scalar had not. But this is selection evidence, not generalisation: the gain is one max_words case, the winning instruction memorises a 15-word bound, and GEPA nets zero on qualifier failures across the wider train-plus-development diagnostics.
The objective itself also required repair before it deserved trust. Three implementation bugs falsely zeroed eight cases. A fourth protocol bug let the reference guard edit the measuring instrument. Twelve tests and six recorded human decisions made the metric inspectable; determinism alone did not make it correct.
The licensed conclusion is therefore narrower and stronger: DSPy optimization is objective-sensitive and optimizer-dependent. Changing the objective changed what search rewarded, what regressions became visible, and which optimizer received usable guidance. It did not make every optimizer effective, and it did not turn a development-set selection into an independently verified improvement.
The editorial task has now taken us as far as it can. The repair chapters move to a regime where parsing, scope checks, public tests, regression tests and sealed hidden tests can establish more of correctness externally. That does not eliminate judgement, but it gives the next prediction a much harder place to hide.
Does named, executable failure evidence produce better repair on optimizer-sealed cases?