Few-Shot Optimization
Compile a few-shot candidate, watch it improve one metric and damage the other, and find out why compiling against the better metric changes nothing at all.
Chapter 10 produced a candidate and refused to say whether it was any good. This chapter says.
The mechanism is few-shot optimization, and the question is practical:
Can we improve the program by selecting better demonstrations, rather than rewriting the program ourselves?
The answer, on our fixture, is the most interesting result in this book. In one run, the candidate appears to improve the metric we optimized against while damaging the metric we did not. Under seven paired fresh-process sessions, the apparent v1 gain collapses to essentially zero while the v2 loss remains large and systematic.
Then we try the obvious repair: compile against v2 instead. The resulting candidate is byte-for-byte identical to the v1-compiled candidate.
So this chapter gives us three different facts that must not be collapsed into one score: a noisy aggregate gain, a reproducible semantic regression, and an objective change that has no effect on the candidate because the relevant failure never enters this optimizer’s compile loop.
1. What BootstrapFewShot does, and what its metric sees
BootstrapFewShot builds demonstrations for each predictor in a program. Some come from labeled examples; the interesting ones are bootstrapped — produced by running the program on a training case and keeping the trace only if the metric approves.
flowchart TD
TE[training examples] --> AT[student attempts each case]
AT --> CT[candidate traces]
CT --> MF{metric approves?}
MF -->|yes| DEM[trace becomes a demonstration]
MF -->|no| DROP[discarded]
DEM --> CS[compiled student]
The filter is the whole mechanism:
Only traces that satisfy the metric become behavior guidance.
Chapter 9 already told us what that implies here. v1 can assign very high scores — about 0.85 to 0.96 on the three canonical attacks — to candidates that reverse important parts of the source meaning.
A bootstrapped demonstration is therefore evidence that a trace passed the compile metric. It is not independent evidence that the trace is semantically safe, representative of the target domain, or acceptable to a human reviewer.
But notice something more specific about where that metric is applied, because it becomes the point of section 7. The filter runs over training traces only. The metric never sees a development case during compilation. BootstrapFewShot has no validation set — chapter 10 noted that compile takes student, teacher, and trainset, and nothing else.
So the compile metric’s admissible field of view is the 26 training cases. The actual field of view of this run is narrower still: five training traces are scored, four clear the threshold, and those four become the source cases for the saved demonstrations.
2. Compiling the candidate
import dspy
def compile_few_shot_candidate(trainset):
student = EditorialRewriteProgram()
optimizer = dspy.BootstrapFewShot(
metric=dspy_editorial_metric_v1,
metric_threshold=0.7,
max_bootstrapped_demos=4,
max_labeled_demos=0,
max_rounds=1,
)
return optimizer.compile(student=student, trainset=trainset)
Then evaluate both programs under the frozen protocol from chapter 8 — same model, same cases, same metric, holdout untouched:
baseline = EditorialRewriteProgram()
candidate = compile_few_shot_candidate(trainset)
evaluate_dev = dspy.Evaluate(devset=devset, metric=dspy_editorial_metric_v1)
baseline_dev = evaluate_dev(baseline)
candidate_dev = evaluate_dev(candidate)
The original measured v1 compile took 30.9 seconds and 6,195 task-model tokens. The metric was called five times during bootstrapping; four traces cleared the 0.7 threshold.
Keep both counts. Five traces were inspected; four were admitted. Candidate provenance and optimizer-consumption provenance are not the same statistic.
3. What it selected
Four accepted traces, distributed across three predictors:
| Predictor | Demonstrations |
|---|---|
analyze | 4 |
rewrite | 4 |
assess | 4 |
Twelve demonstrations in total, all bootstrapped, all traceable to one of four training cases: ed-001, ed-002, ed-006, ed-007.
Four cases out of twenty-six supplied the saved demonstrations, but the optimizer inspected five training traces to get them.
In canonical train order those are ed-001, ed-002, ed-005, ed-006, and ed-007. Four cleared the threshold: ed-001, ed-002, ed-006, and ed-007. ed-005 was evaluated and rejected.
That gives the consumption funnel Chapter 10 asked us to record:
- 26 training cases were admissible;
- 5 were inspected by the compile metric;
- 4 supplied demonstrations;
- 21 were never reached in this compile;
- 22 supplied no demonstration to the candidate.
Raising max_bootstrapped_demos from 2 to 4 therefore expanded the inspected prefix, but it still left most of the available training evidence outside candidate construction.
The provenance checks passed cleanly. No development case became a demonstration. No holdout case appeared in the compile evidence. Every saved demonstration resolved to a known training case with a complete trace.
We can therefore distinguish three questions precisely:
- Was the evidence admissible? Yes.
- Which evidence was inspected and selected? Five inspected, four selected.
- Were the selected demonstrations good guidance for every development family? That remains an empirical question.
Split correctness is necessary. It is not demonstration quality.
4. The result
First, preserve the original single-run result because it exposes the concrete mechanism. On that one 11-case development evaluation, baseline and candidate scored:
| Program | Metric v1 | Metric v2 |
|---|---|---|
| Frozen baseline | 0.7893 | 0.7893 |
| Few-shot candidate | 0.8005 | 0.7399 |
| Delta | +0.0112 | −0.0494 |
On that measurement occasion, the candidate is higher under the metric used for compilation and lower under the semantic guardrail.
Do not promote those deltas into final effect estimates yet. Section 6 repeats the comparison across seven fresh-process paired sessions and changes the interpretation of the positive number completely.
What this single run is excellent for is diagnosis: it gives us exact cases and exact sentences to inspect.
The per-case breakdown shows where:
| Case | v1 | v2 | Note |
|---|---|---|---|
ed-032 | 0.782 → 0.818 | 0.782 → 0.818 | higher in this run |
ed-033 | 0.818 → 0.927 | 0.818 → 0.927 | higher in this run; known session-sensitive case |
ed-036 | 0.956 → 1.000 | 0.956 → 1.000 | higher in this run |
ed-003 | 0.877 → 0.877 | 0.877 → 0.877 | canonical case, unchanged |
ed-035 | 1.000 → 0.967 | 1.000 → 0.300 | qualifier dropped; semantic violation |
ed-038 | 1.000 → 0.967 | 1.000 → 0.967 | small structural loss |
ed-029, ed-030, ed-031, ed-034, ed-037 | 0.650 → 0.650 | 0.650 → 0.650 | unchanged at the structural floor |
Three cases move upward in this run. Five stay pinned at 0.650. ed-038 loses 0.033 under both metrics. And ed-035 loses only 0.033 under v1 but 0.700 under v2.
That is why the two metrics disagree so sharply. v1 treats the ed-035 edit as one small lexical loss among several ordinary case movements. v2 treats it as a disqualifying semantic event.
5. ed-035
ed-035 belongs to the marketing-copy family. The canonical source sentence is:
Using the filter regularly can help reduce limescale build-up in your kettle.
The editorial goal is Make the benefit statement more direct. The context says this is a regulated category and the claim must remain at helps reduce, not prevents.
The reference rewrite is:
Using the filter regularly helps reduce limescale build-up in your kettle.
Two semantic constraints are recorded with the case:
- keep the claim at
helps reduce, notpreventsoreliminates; - the benefit depends on regular use.
The frozen baseline produced the reference rewrite exactly and scored 1.000. The few-shot candidate returned:
Using the filter helps reduce limescale build-up in your kettle.
One word removed. regularly.
As line editing, the cut is plausible. regularly can look like expendable modifier language, and the remaining sentence is cleaner and fully grammatical. This is not a model producing nonsense.
It is also worth using the case’s actual goal. The model was asked to make the benefit statement more direct, not generically to maximize concision. Removing can from the source into the reference satisfies that goal while preserving the condition. Removing regularly goes one step further and changes what the claim depends on.
But regularly is not decoration in this fixture. It carries the condition named explicitly by the semantic constraint.
With it, the claim is about the benefit of regular use. Without it, the sentence attributes the benefit to use of the filter without preserving that condition.
We do not need a jurisdiction-level legal claim to establish the regression. The candidate violates the frozen task requirement on its own terms.
The metrics disagree about how much that matters.
v1 charged 0.033. The candidate still shares almost all of its vocabulary with the reference, still changed something, still stayed in scope. Deleting one adverb barely registers as a lexical event, so the score falls from 1.000 to 0.967 — which reads, in a results table, as noise.
v2 charged 0.700. The judge evaluated the candidate against the recorded regular-use constraint, returned violated with reason code meaning_weakened, and the violation cap took the score from a structural 0.967 to 0.300.
As Chapter 9 established, the verdict is evidence from a validated but same-model semantic guardrail rather than an independent authority. Here the disputed text and the constraint are simple enough to inspect directly.
That is closely related to ed-025 in Chapter 4, where Predict changed a deadline trigger from delivery to receiving the item.
The surface operations differ — ed-035 drops a scope-bearing word, while ed-025 substitutes one event description for another — but the metric failure is the same: a lexically small edit changes a condition or scope that the structural score cannot represent.
Both examples are dangerous precisely because the resulting prose remains fluent and plausible.
Now recall where the saved demonstrations came from. The four source cases — ed-001, ed-002, ed-006, and ed-007 — belong only to chapter-opening and close-third-narration.
No marketing, legal, technical, news, or instructional family contributes a demonstration to this candidate. That concentration gives us a plausible mechanism for cross-domain style transfer: examples that reward tightening in literary prose may bias later rewrites toward similar compression.
But we did not run a demonstration-family ablation, so do not claim that those four demos caused the deletion of regularly. The measured claim is narrower: a candidate built from demonstrations concentrated in two literary families produces the ed-035 regression on an unseen marketing family.
Chapter 10’s coverage finding and this chapter’s failure therefore connect without requiring a causal story we have not tested.
6. Is the gain even real?
The single-run +0.0112 result is no longer the right evidence for the aggregate effect. We repeated the comparison in seven fresh-process sessions, pairing candidate and baseline case by case and rotating execution order.
The result:
| Paired candidate − baseline delta | Mean | SD | Range | Interpretation |
|---|---|---|---|---|
| v1 | +0.0003 | 0.0016 | −0.0020 to +0.0013 | no measurable aggregate improvement |
| v2 | −0.0603 | 0.0016 | tightly negative | large systematic regression |
The v1 candidate’s aggregate delta is positive in 5 of the 7 sessions and negative in the other two, and the mean sits well inside the paired SD. The original +0.0112 was therefore not a small effect hiding near a heuristic noise threshold; it was a measurement-occasion result that collapses under the stronger paired design.
ed-033 explains why the earlier run was especially vulnerable to this mistake. Across earlier sessions, the unchanged baseline itself can land on alternative rewrites for that case. Within the seven-session paired batch the baseline happens to be stable, but pairing prevents that historical session effect from being misread as a candidate effect.
The semantic result behaves differently. The candidate drives ed-035 to v2 0.300 in every run, and the paired aggregate v2 loss is about thirty-eight times the observed paired SD.
So the defensible statement is:
The few-shot candidate has no measured aggregate advantage under v1 in the paired experiment, while it introduces a reproducible semantic regression that v1 prices as a minor lexical change.
That is the reason this book insists on separating candidate generation from promotion. If we had stopped after the first run, reported +0.0112, and promoted on sign alone, we would have had a numerical justification for accepting a candidate that reproducibly drops a condition from ed-035.
The failure is not that the first number was fabricated. It was a real measurement. The failure would have been treating one measurement occasion as sufficient evidence for replacement.
7. “Then why not compile against v2?”
This is the obvious objection, and we ran the experiment.
The setup: the same trainset, the same optimizer, the same threshold and demo budget, but with the compile metric changed from v1 to v2.
If the regression were caused by v1 admitting a training trace that v2 would reject, this intervention should change the selected demonstrations. That is the mechanism the experiment actually tests.
The result:
The two compile arms produced identical candidate state. Not merely the same demo IDs: the saved candidate files are byte-identical and share the same Git blob.
Both contain twelve demonstrations, sourced from the same four accepted cases and distributed identically across the three predictors.
The objective change did have a cost. In the controlled arm, v1 compilation took about 26.4 seconds, made 5 metric calls and no judge calls; v2 compilation took about 41.5 seconds, made the same 5 metric calls plus 12 semantic-judge calls, while consuming the same 6,195 task-model tokens.
The extra semantic evaluation changed the bill but not the candidate.
In the original one-shot dev evaluation the two byte-identical candidates naturally produced the same scores. In the later interleaved seven-session experiment they show tiny score differences despite identical saved state. That is a useful negative control: those differences are measurement-order/runtime variation, not program differences.
ed-035 still breaks under the candidate state.
The explanation is Section 1. BootstrapFewShot applies its compile metric to training traces. Five traces were scored in each arm; the same four cleared the 0.70 threshold under v1 and v2.
The semantic judge therefore had twelve per-constraint opportunities to alter the v2 arm’s accept/reject decisions and changed none of them. The accepted demonstration set remained identical.
And ed-035 is a development case. It never enters the compile loop at all. The compile metric was never asked about it, under either version, because BootstrapFewShot has no mechanism for asking.
compile may access: 26 training cases
traces actually scored: 5
traces accepted as demos: 4
regression observed on: development case ed-035
ed-035 compile exposure: none
The lesson is narrower than a universal claim about compile metrics:
Changing only the compile metric from v1 to v2 did not close this regression under this corpus and
BootstrapFewShotconfiguration.
Why? ed-035 is development-side evidence and never enters this compile loop, while the training-trace decisions that do enter the loop are unchanged under v2.
A different training corpus, ordering, threshold, demo budget, optimizer, or metric could change candidate generation. Chapter 12 matters precisely because MIPROv2 uses development evidence during search, so v2 would have a different opportunity to act there.
The practical consequence is that a better metric is not a substitute for evaluating outside the loop. Whatever you compile against, something independent has to look at what came out.
8. Labeled demonstrations and bootstrapped demonstrations
Worth being precise about the two kinds, because they carry different risks.
A labeled demonstration comes from the dataset. It shows the reference output:
Here is an accepted rewrite.
A bootstrapped demonstration is generated by running the program and keeping the trace if the metric approves:
Here is a rewrite this program produced that passed our metric.
The second statement inherits every weakness of the metric used to admit the trace. Four traces scored above 0.7 and became behavior guidance for all three predictors. The resulting candidate later drops a scope-bearing condition on ed-035.
We have not isolated which demonstration, if any, caused that behavior, so “the demos taught the program to delete qualifiers” would overstate the evidence.
The defensible principle is still important: metric design governs both which candidate wins and, for bootstrap-style optimization, which generated traces are allowed to become future context. Admission therefore deserves its own audit, because a demonstration does not announce which behaviors a downstream case will imitate.
One more asymmetry. In the measured run, ed-001 scored 0.92 and ed-002 about 0.964 during bootstrapping — both comfortably above threshold. A threshold tells you a trace was good enough. It does not tell you the trace was representative, and with demonstrations selected by list order, nothing in this pipeline is asking that question.
What Usually Goes Wrong
| Symptom | Likely cause | How to diagnose it | What to change |
|---|---|---|---|
| The compile metric improves and users complain | Optimized against a proxy the users do not share | Score the candidate under an independent metric | Evaluate outside the compile loop, always |
| A small gain is reported as an improvement | One execution schedule is being treated as an effect estimate | Run paired comparisons across fresh processes with rotated order | Report the paired delta distribution and sign counts |
| The gain comes from one unstable case | Aggregate movement was never localized | Break the result down case by case and compare repeated outputs | Separate candidate effects from session/order effects |
| Upgrading the compile metric changes nothing | The regression is outside the metric’s field of view | Check which split the compile metric evaluates | Add independent evaluation; do not expect the filter to catch it |
| Demonstrations come from one domain | Selection is by list order, not diversity | Group demo sources by family | Shuffle, stratify, or select deliberately |
| The candidate learns a habit nobody asked for | Demonstrations encode an accidental style | Read the selected demonstrations end to end | Diversify sources or reduce the demo count |
| Prompt cost grows with no quality gain | Demonstrations consume context without earning it | Compare baseline and candidate token usage per case | Use fewer demos, or reject the candidate |
| A holdout case appears in the demos | Split boundary violated at the call site | Compare demo source IDs against holdout IDs | Rebuild the split and rerun from scratch |
Conclusion
BootstrapFewShot inspected five training traces, admitted four above the 0.7 threshold, and turned those four traces into twelve demonstrations across three predictors.
The original one-shot evaluation reported +0.0112 under v1 and −0.0494 under v2. The seven-session paired experiment gives the result we should actually carry forward: +0.0003 mean paired delta under v1 and −0.0603 under v2, with paired SD about 0.0016.
So the apparent v1 improvement disappears under repetition. The semantic regression does not. ed-035 reaches v2 0.300 in every run after the candidate deletes regularly. On that case v1 prices the deletion at only 0.033; v2 prices the violated constraint at 0.700.
Then we tried changing the compile metric. The v2 arm made 12 additional judge calls and returned a byte-identical candidate, because the same four training traces cleared the threshold and ed-035 never enters BootstrapFewShot’s compile loop.
The important lesson is not that a better compile metric is powerless. It is that an objective can only change search through evidence the search procedure actually consults.
We removed the assumption that adding accepted demonstrations is automatically an improvement. And we removed something larger: the assumption that improving the thing you optimize against is sufficient protection against optimizing the wrong thing. It is not, because the optimizer and your evaluation do not necessarily look at the same evidence.
What we have not yet tested is whether this is a property of BootstrapFewShot or of optimization generally. This optimizer only selects demonstrations — it never touches an instruction, and it never uses the development set for anything.
The next one does both.
What happens when the optimizer can rewrite the program’s instructions and select against the development set directly?