← DSPy From First Principles

Let the Program Reflect

Give the optimizer language instead of a scalar, watch GEPA propose three mutations and reject all three — then ask whether the rejected candidates were actually worse.

Chapter 12 ended on a specific limitation. MIPROv2 selected a candidate that deleted regularly from a regulated claim, and nothing in the pipeline could have told it otherwise, because the only signal it received was 0.967.

candidate
   ↓
metric
   ↓
0.967

A number can rank candidates. It cannot say which word was the problem.

The next mechanism replaces the scalar with language:

    flowchart LR
    P[prediction] --> M[metric]
    M --> SF[score + diagnostic feedback]
    SF --> RF[reflection]
    RF --> IM[instruction mutation]
    IM --> CA[candidate]
    CA -->|re-run| P
  

GEPA is DSPy’s evolutionary, reflection-driven optimizer. It reads per-predictor feedback, uses a separate reflection model to propose instruction edits, maintains a population of candidate programs, and returns the best one it found. The engineering idea underneath is smaller than the API:

A failure should carry enough explanation to suggest a program change.

In our run, GEPA proposed three mutations and rejected all three. Whether that was good judgement is the most interesting question in this chapter, and the answer is not the one the numbers first suggest.


1. Score and feedback are different

The chapter 8 metric returns a number. For reflective optimization, the metric must also explain:

from dataclasses import dataclass

@dataclass(frozen=True)
class FeedbackScore:
    score: float
    feedback: str

def editorial_feedback(example, pred) -> FeedbackScore:
    candidate = str(getattr(pred, "rewritten_text", "")).strip()
    if not candidate:
        return FeedbackScore(0.0, "The program returned no rewritten_text.")

    for entity in example.required_entities:
        if entity.lower() not in candidate.lower():
            return FeedbackScore(
                0.0,
                f"The candidate removed required entity {entity!r}. Preserve named entities.",
            )

    for term in example.forbidden_terms:
        if term.lower() in candidate.lower():
            return FeedbackScore(
                0.0,
                f"The candidate used forbidden term {term!r}. Avoid entity substitutions.",
            )

    if candidate == example.sentence:
        return FeedbackScore(
            0.65,
            "The candidate copied the original sentence and did not address the edit goal.",
        )

    overlap = reference_overlap(example, pred)
    if overlap < 0.35:
        return FeedbackScore(
            round(0.3 + overlap, 4),
            "The candidate is too far from the reference edit; check meaning preservation and scope.",
        )

    return FeedbackScore(1.0, "The candidate satisfies the deterministic checks.")

The contract that matters:

score:     used for selection
feedback:  used to explain what failed

Two rules about the feedback channel.

It must not leak the answer. “The candidate removed the required entity ‘Jalen’” is good feedback. “Use this exact sentence: …” is the reference rewrite, and putting it into the optimization loop turns the experiment into the leakage scenario Chapter 18 measures at about 0.386 of apparent score inflation.

It must not change the objective. The score returned alongside the feedback is still the frozen scalar from chapter 8. Feedback influences what gets tried. It must not influence what counts as winning, or the comparison against the baseline is no longer a comparison.

Notice also what this feedback cannot say. Look back at ed-035 — the water-filter sentence where the candidate deleted regularly. Every branch above passes: the entity is present, no forbidden term appears, the candidate differs from the original, the overlap is high. The feedback returned would be “the candidate satisfies the deterministic checks.”

Richer feedback does not mean truer feedback. It means more words wrapped around the same blind spot.


2. Reflection is not evaluation

Two model roles, and chapter 6’s provenance table applies to both:

task model        produces program behavior
     ↓
metric            evaluates that behavior

reflection model  reads failure evidence
     ↓
                  proposes program changes

The reflection model does not decide what is true. It proposes an edit, and the frozen metric decides whether the edit survives.

failure → reflection proposes → candidate → evaluation → keep or reject

This is the same boundary chapter 10 established, one level up. The optimizer proposes. Evaluation selects. Promotion is chapter 19.


3. The compile shape

DSPy’s documentation illustrates GEPA with a strong hosted reflection model:

# Illustrative only — not the configuration this book ran.
reflection_lm = dspy.LM("openai/gpt-5", temperature=1.0, max_tokens=32000)

That is a sensible default and it is not what we did. Every measured result in this book runs on the canonical local dependency, and this chapter is no exception — both the task role and the reflection role are local Qwen3, differing only in generation settings:

import dspy

REFLECTION_CONFIG_OVERRIDES = {"temperature": 0.7, "max_tokens": 1024}

def gepa_metric(example, pred, trace=None, pred_name=None, pred_trace=None):
    result = editorial_feedback(example, pred)
    return dspy.Prediction(score=result.score, feedback=result.feedback)

def compile_reflective_candidate(trainset, devset):
    reflection_lm = dspy.LM(**canonical_dspy_lm_kwargs(), **REFLECTION_CONFIG_OVERRIDES)
    optimizer = dspy.GEPA(
        metric=gepa_metric,
        reflection_lm=reflection_lm,
        max_metric_calls=66,  # = max(24, 6 * len(valset)); GEPA needs an explicit budget
        candidate_selection_strategy="current_best",
        use_merge=False,
        track_stats=True,
        seed=13,
    )
    return optimizer.compile(
        EditorialRewriteProgram(),
        trainset=trainset,
        valset=devset,
    )

The higher reflection temperature is deliberate. Proposing mutations is a generative task that benefits from variety; executing the program is not.

Sharing weights between the task and reflection roles is the role collapse chapter 6 warned about, and it is a real limitation here. A reflection model with the same blind spots as the task model will propose mutations that share them.

The search boundary:

  • train: 26 cases, 7 families
  • validation: 11 development cases
  • holdout: 7 cases, sealed and uninspected
  • mutable predictors: analyze, rewrite
  • fixed predictor: assess
  • effective metric-call budget: 66, computed as max(24, 6 × |valset|)
  • candidate selection strategy: current_best
  • merge: disabled
  • seed: 13

That budget is worth a note. An earlier version of this experiment set max_metric_calls=12 against a one-case validation set. With eleven development cases the budget scales to 66, and the run made 76 feedback calls. The budget is a function of your data, so changing the fixture changed the experiment’s size without anyone editing a configuration value.


4. What the search did

GEPA proposed three child candidates. The lineage, including the baseline as candidate 0:

CandidateParentDev score, v1Outcome
0 (baseline)—0.7992selected
100.7792rejected
200.7932rejected
300.7899rejected

All three children scored below the baseline, so best_idx = 0. The returned program had candidate_state_changed = False and no instruction changes at all.

The component selector alternated between analyze and rewrite and never touched assess — which is correct, and slightly funny given chapter 5’s finding that assess cannot affect the outcome anyway. The one predictor the optimizer was forbidden to modify is the one that does nothing.

No feedback record contained an exact reference rewrite. The scalar score always matched the frozen metric. The holdout stayed sealed.

So the headline is: reflection proposed three concrete alternatives, evaluation measured all three, and the system kept the better parent. That is a working governance loop.

It is also, on closer inspection, not clearly true.


5. Were the children actually worse?

Apply chapter 6’s noise floor.

The run-to-run variation on this 11-case development mean is about ±0.015. Two identical runs of the unchanged baseline have measured 0.7893 and 0.7992.

GEPA’s baseline measurement in this session was 0.7992 — the top of that band.

Now look at the children again:

CandidateScoreDifference from this session’s baselineDifference from the baseline’s typical value (~0.789)
10.7792−0.0200−0.010
20.7932−0.0060+0.004
30.7899−0.0093+0.001

Candidate 3 scored 0.7899. The baseline scored 0.7893 in chapter 8’s first run. Those two numbers describe programs that are, on this instrument, indistinguishable.

GEPA rejected two candidates that were statistically identical to the baseline, because it compared them against a single baseline measurement that happened to land at the high end of its noise band.

This is not a criticism of GEPA specifically. It is a property of every optimizer in the last four chapters. None of them has any concept of measurement uncertainty:

ChapterOptimizerDecisionMagnitudeNoise floor
11BootstrapFewShotaccepted its candidate+0.0112±0.015
12MIPROv2accepted its candidate+0.0298±0.015
13GEPArejected three−0.006 to −0.020±0.015

Chapter 11 accepted on a difference inside the noise. Chapter 13 rejected on differences inside the noise. Only chapter 12’s decision rests on anything that clears it, and chapter 12’s gain decomposed to a single case.

Every one of these is a single-run comparison. The optimizers evaluate once, compare two numbers, and decide. They do not repeat, they do not compute a spread, and nothing in their APIs suggests they should.

That is the honest state of the art, and it is why chapter 19 exists. The promotion policy there requires a margin of 0.05 — more than three times the noise floor — precisely because the optimizer’s own accept/reject decision is not evidence of anything on a fixture this size.


6. What GEPA did not do, and why we should not credit it

GEPA is the only optimizer in this book that did not break ed-035.

Resist the obvious conclusion. GEPA did not avoid the failure through better judgement, richer feedback, or reflective insight into regulated claims. It avoided the failure by shipping nothing. It kept the baseline, and the baseline scores 1.000 on ed-035 because the baseline never touched it.

An optimizer that changes nothing cannot introduce a regression. That is not safety, it is inaction, and inaction is only the right answer some of the time.

The genuinely interesting question — whether GEPA’s reflective loop is more conservative than a Bayesian search, or whether it simply drew three unlucky mutations at a small budget — is not answerable from one run. Three children is not a sample. We ran the experiment once, with a seed, on 11 development cases, comparing against a baseline measured once.

Everything this chapter can honestly claim is that on this run, at this budget, reflection proposed three changes, all three measured below a baseline measurement that was itself near the top of its noise band, and the system kept the parent.


7. Reflection has a measurable cost

ResourceMeasured
Search elapsed time437.7 s
Task-model history entries198
Task-model tokens90,679
Reflection-model history entries10
Reflection-model tokens6,276
Feedback-metric calls76

Just under 97,000 tokens and roughly seven and a half minutes, to produce three rejected mutations and retain the program we started with.

The cost split is worth noticing. Reflection itself is cheap — 6,276 tokens across ten calls, about 6% of the total. The expense is evaluation: every proposed child has to be run against eleven development cases to find out whether it helped, and that is where 90,679 tokens went.

That has a practical consequence. If you are trying to make reflective optimization affordable, a cheaper reflection model saves you almost nothing. Reducing the number of candidates evaluated, or the size of the validation set each candidate is measured against, saves you almost everything.

And put the three optimizer budgets side by side:

ChapterOptimizerSearch tokensOutcome
11BootstrapFewShot6,195candidate accepted, regression introduced
12MIPROv285,607candidate accepted, same regression
13GEPA96,955all candidates rejected

Roughly 189,000 tokens of search across three optimizers. Not one of them produced a program this book will promote.


What Usually Goes Wrong

SymptomLikely causeHow to diagnose itWhat to change
Reflections are genericFeedback says only “wrong”Read the feedback strings the optimizer receivedName the concrete failure and the affected field
Feedback is detailed and the failure persistsThe feedback path is blind to that failure classTrace what feedback a known-bad case producesFix the metric; verbosity is not vision
Score and feedback disagreeThe metric has diverging branchesUnit-test the metric against known casesReturn score and feedback from one decision path
A child is rejected on a small differenceSingle-run comparison against a noise floorRepeat the baseline evaluation and compute the spreadRequire a margin larger than the noise before deciding
An optimizer that changes nothing looks safeInaction is being read as judgementCheck whether any candidate was actually selectedDistinguish “rejected a bad change” from “made no change”
Instructions grow indefinitelyReflection keeps appending rulesInspect saved candidate state across generationsPenalise verbosity or set an instruction budget
The gold answer appears in feedbackReference text leaked into the optimization loopGrep feedback records for reference stringsRebuild the feedback path; treat the run as contaminated
Reflection cost is blamed for a slow searchThe expense is candidate evaluation, not reflectionTrack task and reflection tokens separatelyReduce candidates or validation size, not reflection quality
Budget changes when the fixture changesThe budget is a function of validation sizeRecord the effective budget, not the configured oneLog effective_max_metric_calls in the manifest

Conclusion

GEPA read per-predictor feedback, proposed three concrete instruction mutations through a reflection model, evaluated each against the frozen objective, and kept the parent. No reference text leaked into the feedback channel, assess stayed fixed, and the holdout stayed sealed. As a governance loop, it worked exactly as designed.

As evidence, it is thinner than it looks. The three children scored 0.7792, 0.7932 and 0.7899 against a baseline measured at 0.7992 — and that baseline measurement sits at the top of a band whose bottom is 0.7893. Two of the three rejected candidates were indistinguishable from the program that beat them.

That points at something larger than this chapter. Across chapters 11, 12 and 13, three optimizers made accept-or-reject decisions on single-run differences of 0.011, 0.030 and −0.009, against a noise floor of 0.015 that none of them knows exists. One accepted on noise, one rejected on noise, and the third’s genuine gain turned out to be one case. Nearly 189,000 tokens of search across the three, and nothing this book will promote.

We removed the assumption that scalar scores are sufficient to guide improvement. We also removed a more flattering one: that an optimizer which declines to change anything has exercised judgement. GEPA is the only optimizer here that did not damage ed-035, and it earned that by doing nothing.

What none of these mechanisms could do is act. Every program in this book so far has worked from information handed to it — a sentence, a goal, a context string — and produced text. Real engineering programs have to look things up, read files, run commands, and decide what evidence they need.

What changes when the program can act on an environment instead of only reading its inputs?