Let the Program Reflect
Give the optimizer language instead of a scalar, watch GEPA propose three mutations and reject all three — then ask whether the rejected candidates were actually worse.
Chapter 12 ended on a specific limitation. MIPROv2 selected a candidate that deleted regularly from a regulated claim, and nothing in the pipeline could have told it otherwise, because the only signal it received was 0.967.
candidate
↓
metric
↓
0.967
A number can rank candidates. It cannot say which word was the problem.
The next mechanism replaces the scalar with language:
flowchart LR
P[prediction] --> M[metric]
M --> SF[score + diagnostic feedback]
SF --> RF[reflection]
RF --> IM[instruction mutation]
IM --> CA[candidate]
CA -->|re-run| P
GEPA is DSPy’s evolutionary, reflection-driven optimizer. It reads per-predictor feedback, uses a separate reflection model to propose instruction edits, maintains a population of candidate programs, and returns the best one it found. The engineering idea underneath is smaller than the API:
A failure should carry enough explanation to suggest a program change.
In our run, GEPA proposed three mutations and rejected all three. Whether that was good judgement is the most interesting question in this chapter, and the answer is not the one the numbers first suggest.
1. Score and feedback are different
The chapter 8 metric returns a number. For reflective optimization, the metric must also explain:
from dataclasses import dataclass
@dataclass(frozen=True)
class FeedbackScore:
score: float
feedback: str
def editorial_feedback(example, pred) -> FeedbackScore:
candidate = str(getattr(pred, "rewritten_text", "")).strip()
if not candidate:
return FeedbackScore(0.0, "The program returned no rewritten_text.")
for entity in example.required_entities:
if entity.lower() not in candidate.lower():
return FeedbackScore(
0.0,
f"The candidate removed required entity {entity!r}. Preserve named entities.",
)
for term in example.forbidden_terms:
if term.lower() in candidate.lower():
return FeedbackScore(
0.0,
f"The candidate used forbidden term {term!r}. Avoid entity substitutions.",
)
if candidate == example.sentence:
return FeedbackScore(
0.65,
"The candidate copied the original sentence and did not address the edit goal.",
)
overlap = reference_overlap(example, pred)
if overlap < 0.35:
return FeedbackScore(
round(0.3 + overlap, 4),
"The candidate is too far from the reference edit; check meaning preservation and scope.",
)
return FeedbackScore(1.0, "The candidate satisfies the deterministic checks.")
The contract that matters:
score: used for selection
feedback: used to explain what failed
Two rules about the feedback channel.
It must not leak the answer. “The candidate removed the required entity ‘Jalen’” is good feedback. “Use this exact sentence: …” is the reference rewrite, and putting it into the optimization loop turns the experiment into the leakage scenario Chapter 18 measures at about 0.386 of apparent score inflation.
It must not change the objective. The score returned alongside the feedback is still the frozen scalar from chapter 8. Feedback influences what gets tried. It must not influence what counts as winning, or the comparison against the baseline is no longer a comparison.
Notice also what this feedback cannot say. Look back at ed-035 — the water-filter sentence where the candidate deleted regularly. Every branch above passes: the entity is present, no forbidden term appears, the candidate differs from the original, the overlap is high. The feedback returned would be “the candidate satisfies the deterministic checks.”
Richer feedback does not mean truer feedback. It means more words wrapped around the same blind spot.
2. Reflection is not evaluation
Two model roles, and chapter 6’s provenance table applies to both:
task model produces program behavior
↓
metric evaluates that behavior
reflection model reads failure evidence
↓
proposes program changes
The reflection model does not decide what is true. It proposes an edit, and the frozen metric decides whether the edit survives.
failure → reflection proposes → candidate → evaluation → keep or reject
This is the same boundary chapter 10 established, one level up. The optimizer proposes. Evaluation selects. Promotion is chapter 19.
3. The compile shape
DSPy’s documentation illustrates GEPA with a strong hosted reflection model:
# Illustrative only — not the configuration this book ran.
reflection_lm = dspy.LM("openai/gpt-5", temperature=1.0, max_tokens=32000)
That is a sensible default and it is not what we did. Every measured result in this book runs on the canonical local dependency, and this chapter is no exception — both the task role and the reflection role are local Qwen3, differing only in generation settings:
import dspy
REFLECTION_CONFIG_OVERRIDES = {"temperature": 0.7, "max_tokens": 1024}
def gepa_metric(example, pred, trace=None, pred_name=None, pred_trace=None):
result = editorial_feedback(example, pred)
return dspy.Prediction(score=result.score, feedback=result.feedback)
def compile_reflective_candidate(trainset, devset):
reflection_lm = dspy.LM(**canonical_dspy_lm_kwargs(), **REFLECTION_CONFIG_OVERRIDES)
optimizer = dspy.GEPA(
metric=gepa_metric,
reflection_lm=reflection_lm,
max_metric_calls=66, # = max(24, 6 * len(valset)); GEPA needs an explicit budget
candidate_selection_strategy="current_best",
use_merge=False,
track_stats=True,
seed=13,
)
return optimizer.compile(
EditorialRewriteProgram(),
trainset=trainset,
valset=devset,
)
The higher reflection temperature is deliberate. Proposing mutations is a generative task that benefits from variety; executing the program is not.
Sharing weights between the task and reflection roles is the role collapse chapter 6 warned about, and it is a real limitation here. A reflection model with the same blind spots as the task model will propose mutations that share them.
The search boundary:
- train: 26 cases, 7 families
- validation: 11 development cases
- holdout: 7 cases, sealed and uninspected
- mutable predictors:
analyze,rewrite - fixed predictor:
assess - effective metric-call budget: 66, computed as
max(24, 6 × |valset|) - candidate selection strategy:
current_best - merge: disabled
- seed: 13
That budget is worth a note. An earlier version of this experiment set max_metric_calls=12 against a one-case validation set. With eleven development cases the budget scales to 66, and the run made 76 feedback calls. The budget is a function of your data, so changing the fixture changed the experiment’s size without anyone editing a configuration value.
4. What the search did
GEPA proposed three child candidates. The lineage, including the baseline as candidate 0:
| Candidate | Parent | Dev score, v1 | Outcome |
|---|---|---|---|
| 0 (baseline) | — | 0.7992 | selected |
| 1 | 0 | 0.7792 | rejected |
| 2 | 0 | 0.7932 | rejected |
| 3 | 0 | 0.7899 | rejected |
All three children scored below the baseline, so best_idx = 0. The returned program had candidate_state_changed = False and no instruction changes at all.
The component selector alternated between analyze and rewrite and never touched assess — which is correct, and slightly funny given chapter 5’s finding that assess cannot affect the outcome anyway. The one predictor the optimizer was forbidden to modify is the one that does nothing.
No feedback record contained an exact reference rewrite. The scalar score always matched the frozen metric. The holdout stayed sealed.
So the headline is: reflection proposed three concrete alternatives, evaluation measured all three, and the system kept the better parent. That is a working governance loop.
It is also, on closer inspection, not clearly true.
5. Were the children actually worse?
Apply chapter 6’s noise floor.
The run-to-run variation on this 11-case development mean is about ±0.015. Two identical runs of the unchanged baseline have measured 0.7893 and 0.7992.
GEPA’s baseline measurement in this session was 0.7992 — the top of that band.
Now look at the children again:
| Candidate | Score | Difference from this session’s baseline | Difference from the baseline’s typical value (~0.789) |
|---|---|---|---|
| 1 | 0.7792 | −0.0200 | −0.010 |
| 2 | 0.7932 | −0.0060 | +0.004 |
| 3 | 0.7899 | −0.0093 | +0.001 |
Candidate 3 scored 0.7899. The baseline scored 0.7893 in chapter 8’s first run. Those two numbers describe programs that are, on this instrument, indistinguishable.
GEPA rejected two candidates that were statistically identical to the baseline, because it compared them against a single baseline measurement that happened to land at the high end of its noise band.
This is not a criticism of GEPA specifically. It is a property of every optimizer in the last four chapters. None of them has any concept of measurement uncertainty:
| Chapter | Optimizer | Decision | Magnitude | Noise floor |
|---|---|---|---|---|
| 11 | BootstrapFewShot | accepted its candidate | +0.0112 | ±0.015 |
| 12 | MIPROv2 | accepted its candidate | +0.0298 | ±0.015 |
| 13 | GEPA | rejected three | −0.006 to −0.020 | ±0.015 |
Chapter 11 accepted on a difference inside the noise. Chapter 13 rejected on differences inside the noise. Only chapter 12’s decision rests on anything that clears it, and chapter 12’s gain decomposed to a single case.
Every one of these is a single-run comparison. The optimizers evaluate once, compare two numbers, and decide. They do not repeat, they do not compute a spread, and nothing in their APIs suggests they should.
That is the honest state of the art, and it is why chapter 19 exists. The promotion policy there requires a margin of 0.05 — more than three times the noise floor — precisely because the optimizer’s own accept/reject decision is not evidence of anything on a fixture this size.
6. What GEPA did not do, and why we should not credit it
GEPA is the only optimizer in this book that did not break ed-035.
Resist the obvious conclusion. GEPA did not avoid the failure through better judgement, richer feedback, or reflective insight into regulated claims. It avoided the failure by shipping nothing. It kept the baseline, and the baseline scores 1.000 on ed-035 because the baseline never touched it.
An optimizer that changes nothing cannot introduce a regression. That is not safety, it is inaction, and inaction is only the right answer some of the time.
The genuinely interesting question — whether GEPA’s reflective loop is more conservative than a Bayesian search, or whether it simply drew three unlucky mutations at a small budget — is not answerable from one run. Three children is not a sample. We ran the experiment once, with a seed, on 11 development cases, comparing against a baseline measured once.
Everything this chapter can honestly claim is that on this run, at this budget, reflection proposed three changes, all three measured below a baseline measurement that was itself near the top of its noise band, and the system kept the parent.
7. Reflection has a measurable cost
| Resource | Measured |
|---|---|
| Search elapsed time | 437.7 s |
| Task-model history entries | 198 |
| Task-model tokens | 90,679 |
| Reflection-model history entries | 10 |
| Reflection-model tokens | 6,276 |
| Feedback-metric calls | 76 |
Just under 97,000 tokens and roughly seven and a half minutes, to produce three rejected mutations and retain the program we started with.
The cost split is worth noticing. Reflection itself is cheap — 6,276 tokens across ten calls, about 6% of the total. The expense is evaluation: every proposed child has to be run against eleven development cases to find out whether it helped, and that is where 90,679 tokens went.
That has a practical consequence. If you are trying to make reflective optimization affordable, a cheaper reflection model saves you almost nothing. Reducing the number of candidates evaluated, or the size of the validation set each candidate is measured against, saves you almost everything.
And put the three optimizer budgets side by side:
| Chapter | Optimizer | Search tokens | Outcome |
|---|---|---|---|
| 11 | BootstrapFewShot | 6,195 | candidate accepted, regression introduced |
| 12 | MIPROv2 | 85,607 | candidate accepted, same regression |
| 13 | GEPA | 96,955 | all candidates rejected |
Roughly 189,000 tokens of search across three optimizers. Not one of them produced a program this book will promote.
What Usually Goes Wrong
| Symptom | Likely cause | How to diagnose it | What to change |
|---|---|---|---|
| Reflections are generic | Feedback says only “wrong” | Read the feedback strings the optimizer received | Name the concrete failure and the affected field |
| Feedback is detailed and the failure persists | The feedback path is blind to that failure class | Trace what feedback a known-bad case produces | Fix the metric; verbosity is not vision |
| Score and feedback disagree | The metric has diverging branches | Unit-test the metric against known cases | Return score and feedback from one decision path |
| A child is rejected on a small difference | Single-run comparison against a noise floor | Repeat the baseline evaluation and compute the spread | Require a margin larger than the noise before deciding |
| An optimizer that changes nothing looks safe | Inaction is being read as judgement | Check whether any candidate was actually selected | Distinguish “rejected a bad change” from “made no change” |
| Instructions grow indefinitely | Reflection keeps appending rules | Inspect saved candidate state across generations | Penalise verbosity or set an instruction budget |
| The gold answer appears in feedback | Reference text leaked into the optimization loop | Grep feedback records for reference strings | Rebuild the feedback path; treat the run as contaminated |
| Reflection cost is blamed for a slow search | The expense is candidate evaluation, not reflection | Track task and reflection tokens separately | Reduce candidates or validation size, not reflection quality |
| Budget changes when the fixture changes | The budget is a function of validation size | Record the effective budget, not the configured one | Log effective_max_metric_calls in the manifest |
Conclusion
GEPA read per-predictor feedback, proposed three concrete instruction mutations through a reflection model, evaluated each against the frozen objective, and kept the parent. No reference text leaked into the feedback channel, assess stayed fixed, and the holdout stayed sealed. As a governance loop, it worked exactly as designed.
As evidence, it is thinner than it looks. The three children scored 0.7792, 0.7932 and 0.7899 against a baseline measured at 0.7992 — and that baseline measurement sits at the top of a band whose bottom is 0.7893. Two of the three rejected candidates were indistinguishable from the program that beat them.
That points at something larger than this chapter. Across chapters 11, 12 and 13, three optimizers made accept-or-reject decisions on single-run differences of 0.011, 0.030 and −0.009, against a noise floor of 0.015 that none of them knows exists. One accepted on noise, one rejected on noise, and the third’s genuine gain turned out to be one case. Nearly 189,000 tokens of search across the three, and nothing this book will promote.
We removed the assumption that scalar scores are sufficient to guide improvement. We also removed a more flattering one: that an optimizer which declines to change anything has exercised judgement. GEPA is the only optimizer here that did not damage ed-035, and it earned that by doing nothing.
What none of these mechanisms could do is act. Every program in this book so far has worked from information handed to it — a sentence, a goal, a context string — and produced text. Real engineering programs have to look things up, read files, run commands, and decide what evidence they need.
What changes when the program can act on an environment instead of only reading its inputs?