← DSPy From First Principles

Examples Are Experimental Data

Build a real evaluation fixture: structured rows with provenance, families rather than isolated cases, splits that respect them, and explicit boundaries around what optimization may consume.

Chapter 6 gave the program a dependency boundary and, more importantly, evidence that identical recorded program state can still produce different outputs. We can now say what the program is, what ran it, and why a small score delta cannot be interpreted without repeated measurement.

What we still cannot do is say whether the program is any good, because we have nothing to measure it against.

Examples in DSPy are not decoration around the program. They are experimental data with several possible roles. Optimization consumes that evidence: training data may shape candidate state, development data may select among candidates, and holdout data must remain outside both processes if it is to support an independent claim.

program
   ↓
no disciplined evidence
   ↓
"seems better"

The next seven chapters build one experiment: a fixture, a baseline, a metric, an attack on that metric, four attempts to optimize the program against it, and a final run that changes only the objective. This chapter builds the fixture, which is the least glamorous and most load-bearing part.

A warning about ordering. Chapters 4 and 5 already used this fixture — that is where the 37 training and development cases came from. It is described formally here.


1. The tempting shortcut

A handwritten few-shot prompt looks like this:

Example 1:
Sentence: She opened the door and then she laughed.
Goal: Sharpen rhythm.
Rewrite: She opened the door and laughed.

Example 2:
...

That may improve a single prompt. As evidence it is close to worthless.

The examples have no identity, so you cannot refer to one later. Inputs and labels are mixed, so nothing distinguishes what the program should see from what the evaluator should see. Provenance is missing, so you cannot tell whether two examples came from the same source. And an example that carries a note like “this was accepted by the reviewer” has just handed the model the answer.

A dataset row needs structure:

case_id              stable identity
program inputs       what the model sees
reference fields     what the evaluator compares against
evaluation metadata  constraints, expected failure mode
provenance           where this came from
split                what role it plays

The distinction between rows two and three is the one that matters most, and section 5 makes it mechanical rather than a matter of discipline.


2. The corpus

The fixture is built for this book. It is a diagnostic corpus, not a representative sample of editorial traffic, so it demonstrates mechanisms and failure classes rather than production performance.

Forty-four cases are enough to expose behaviors the four-case fixture could not expose. They are still far too few to support fine-grained statistical claims from aggregate means. The rest of this chapter treats both facts as true at once.

44 cases across 12 source families.

FamilySplitCases
chapter-openingtrained-001, ed-002, ed-005, ed-006
close-third-narrationtrained-007 – ed-010
first-person-memoirtrained-011 – ed-014
academic-abstracttrained-015 – ed-018
news-reporttrained-019 – ed-022
legal-plain-languagetrained-023, ed-024, ed-025
instructional-recipetrained-026, ed-027, ed-028
dialoguedeved-003, ed-029, ed-030, ed-031
marketing-copydeved-032 – ed-035
childrens-fictiondeved-036, ed-037, ed-038
technical-explanationholdouted-004, ed-039 – ed-041
email-correspondenceholdouted-042, ed-043, ed-044

That is 26 training cases, 11 development cases, and 7 holdout cases.

Several of these you have already met. ed-025 is the legal plain-language case where Chapter 4’s direct predictor changed the deadline trigger from delivery to receipt. ed-035 is the marketing-copy case that Chapters 11 and 12 will both damage by dropping the qualifier regularly. ed-033 is the coffee sentence from Chapter 6 whose output can move between plausible alternatives under unchanged recorded program state.

Those cases are useful precisely because the family table gives them context: they are not anonymous rows created after a failure was discovered. Their identities and split roles were frozen before the optimizer chapters interpreted them.

Each row carries:

case_id               # stable row identity
source_id             # provenance handle
source_group          # family used to derive the split
sentence              # program input
goal                  # program input
context               # program input
reference_rewrite     # evaluation only
required_entities     # evaluation only — must survive the edit
forbidden_terms       # evaluation only — must not appear
semantic_constraints  # evaluation only — meaning that must be preserved
target_failure        # primary failure mode this case is designed to exercise

The split is not stored on EditorialCase itself. It is derived from source_group through the frozen family-to-split mapping. That is why the dataset and the split receive separate fingerprints: row content can stay fixed while split policy changes.

One canonical row in full makes the shape concrete. ed-003 is the dialogue development case that supplied the original one-case 0.8769 baseline. Later chapters retain it as a useful walkthrough even after the aggregate evaluation expands to the full development family:

case_id:             ed-003
source_id:           dialogue-a
source_group:        dialogue
sentence:            I do not think we should go there, Anna said, because it is unsafe.
goal:                Improve dialogue punctuation and rhythm.
context:             Anna is warning her brother.
reference_rewrite:   "I do not think we should go there," Anna said. "It is unsafe."
required_entities:   ["Anna"]
forbidden_terms:     ["Anya"]
semantic_constraints:
  - do not change the warning
  - do not change the speaker
target_failure:      semantic_drift
split:               dev   # derived from source_group for this display

Notice what the reference rewrite does and does not do. It adds quotation marks and splits one sentence into two. It does not contract “do not” to “don’t.” The editorial goal is punctuation and rhythm, and the reference stays within it. A model that also modernises the diction has done something the reference did not ask for, and the metric in chapter 8 will notice.


3. A fixture is an instrument, not a sample

The corpus was not collected from a representative traffic stream. It was designed as a diagnostic instrument, and the target_failure field records the primary failure mode each case exists to exercise.

A row can still contain several constraints and can fail in more than one way. target_failure is a coverage label, not a claim that every observed error on that row must belong to one category:

Failure modeCasesWhat it tests
semantic_drift12Meaning quietly shifts while the prose improves
constraint_violation8A stated requirement is broken
style_failure7Voice or register changes
unnecessary_edit6The sentence was already correct
entity_change6A name, number, or reference is altered
invalid_output5The result is structurally unusable

A representative traffic sample and a diagnostic fixture answer different questions. If the sample is genuinely representative, it can estimate how often failures occur in deployment. A diagnostic fixture deliberately over-samples the failures you most need to understand.

This book needs the second instrument. If the corpus contains no case where a deadline trigger can shift, no experiment on that corpus can reveal that failure class. If it contains only one such case, that case can still dominate an aggregate while telling you almost nothing about prevalence.

So design coverage around the failures you are worried about, then keep frequency claims separate from diagnostic findings.

Two of these categories were included knowing they would cause trouble, and both did.

The six unnecessary_edit cases make restraint part of the task rather than treating change as automatically good. Some should remain unchanged; others permit only a minimal edit. The label therefore means “do not invent work merely to satisfy the existence of a rewrite field,” not “reference_rewrite always equals sentence.”

That distinction matters. Chapter 2 showed that the current contract has no explicit semantic state for “no edit needed.” Chapter 8 will show that an exactly unchanged candidate receives only the 0.65 structural floor because the metric’s changed component is zero even when restraint is correct. Chapter 18 then finds the same design tension from the security side: identity-reference cases collide with a literal rule that treats emission of the reference text as leakage.

We keep restraint cases anyway. A corpus in which change is always rewarded teaches both the program and the metric the wrong editorial prior.

What this fixture can detect

Here is arithmetic that almost nobody does before running an experiment, and everybody should.

There are two kinds of resolution to think about, and they should not be collapsed into one number.

First is fixture granularity: with only 11 development cases, one case can move the mean materially.

Second is execution variability: Chapter 6 already showed that unchanged program state can move under fresh sessions and different interleavings, even at temperature zero.

Start with the arithmetic that is exact:

cases in the dev split:                    11
weight of one case in the mean:              1 / 11 ≈ 0.0909

one case moving from 0.65 to 1.00:
case-level change:                            0.35
effect on the dev mean:                       0.35 / 11 ≈ 0.0318

one case changing by 0.11:
effect on the dev mean:                       0.11 / 11 = 0.010

This tells you how coarse an 11-case aggregate is before any model nondeterminism enters the picture. A single 0.35 case-level swing moves the mean by about 0.032. A much smaller 0.11 movement on one case already moves the mean by 0.010.

Now compare that with Chapter 6: unchanged program state has produced aggregate movement on roughly the same scale, and byte-identical candidate artifacts have provided a zero-effect negative control that still measured non-zero differences under interleaving.

So when Chapter 11 initially reports +0.011, the correct statement is not “small improvement.” It is “one observed aggregate difference that is not, by itself, evidence of improvement.” One modest case movement or execution-state variation can explain a delta of that size.

Keep full precision in machine artifacts so reruns can be compared exactly. In prose, do not let four decimal places imply four decimal places of evidential resolution.

Run this granularity calculation before the experiment. Then design repeated, paired measurements around the runtime variability you actually observe. If the effect you care about cannot be separated from either source, collect more independent cases or stronger repeated evidence instead of collecting more interpretations of the same point estimate.


4. Convert rows into DSPy examples

DSPy uses dspy.Example as the row object. The fields named in .with_inputs(...) are the fields this book supplies as ordinary task inputs. The remaining fields stay attached to the example for evaluation, provenance, and optimizer-side code.

That projection is a boundary, not a universal secrecy guarantee. An optimizer or custom metric can still inspect fields attached to the example if its implementation chooses to. Chapter 18 therefore audits evidence access in addition to checking .with_inputs(...).

import dspy

INPUT_FIELDS = ("sentence", "goal", "context")

def to_dspy_example(row: dict) -> dspy.Example:
    return dspy.Example(
        case_id=row["case_id"],
        source_id=row["source_id"],
        source_group=row["source_group"],
        sentence=row["sentence"],
        goal=row["goal"],
        context=row["context"],
        reference_rewrite=row["reference_rewrite"],
        required_entities=row["required_entities"],
        forbidden_terms=row["forbidden_terms"],
        semantic_constraints=row["semantic_constraints"],
        target_failure=row["target_failure"],
    ).with_inputs(*INPUT_FIELDS)

One line carries the entire boundary:

.with_inputs("sentence", "goal", "context")

The task program receives three decision-time fields through this projection. Evaluation code can still access the attached reference, constraints, failure label, and provenance.

This is the mechanism behind Chapter 3’s third input-design question — could this value contain, imply, or be derived from the answer? The projection makes the ordinary generation boundary machine-checkable. It does not by itself prove that every optimizer, tool, or metric is unable to reach the other fields; that broader admissibility boundary is audited later.

LEGAL PROGRAM INPUT
  sentence = "I do not think we should go there, Anna said, because it is unsafe."
  goal     = "Improve dialogue punctuation and rhythm without changing the warning."
  context  = "Anna is warning her brother."

EVALUATION-SIDE ONLY
  reference_rewrite = "\"I do not think we should go there,\" Anna said. \"It is unsafe.\""
  semantic_constraints = [...]

WOULD BE FATAL
  human_accepted = true

A program that sees human_accepted or answer-derived reference material is no longer being evaluated under the same task boundary. Chapter 18 measures the scale of the problem: across 288 constructed attack executions, all 288 are blocked; 224 would otherwise inflate the apparent score, by an average of about 0.3856.

That does not mean 0.3856 can simply be added to every honest score; the metric is bounded and the attacks differ by case. It means leakage effects are large enough to dominate the ordinary program deltas later chapters are trying to interpret.


5. Split by family, not by row

The split roles are conventional:

SplitRole
trainMay be used to construct demonstrations or other candidate state, according to the experiment
devDevelopment evidence: may be reserved for comparison or may be optimizer-visible for candidate selection, but the protocol must say which
holdoutNever optimizer-visible in this experimental cycle; spent once after candidate selection is frozen

The design decision that is not conventional is splitting at the family boundary rather than the row boundary.

The chemistry augmented generation TPSA project is the clearest illustration of why. Its prediction target is numerical and its splits are scaffold-aware, because closely related molecules make a random row split far too optimistic — the model appears to generalise when it has really seen a near-sibling of the answer.

The editorial analogue is exact. If two sentences come from the same chapter, the same author, or the same accepted-edit history, a random split can put near-duplicates in training and holdout. The program then looks like it learned to edit when it actually learned that author’s cadence.

So all four dialogue cases are in dev. All four marketing-copy cases are in dev. All four technical-explanation cases are in holdout. No family straddles a boundary.

    flowchart LR
    F[12 source families] --> T[train, 7 families, 26 cases]
    F --> D[dev, 3 families, 11 cases]
    F --> H[holdout, 2 families, 7 cases]
    T --> O[optimizer-visible evidence]
    D --> O
    H -.->|sealed until candidate selection is frozen| S[spent once]
  

A whole family lands in one split. That is what makes a holdout result an independent claim: no near-sibling of a holdout sentence was ever in training or development.

It is worth noting how weak the previous version of this rule was. The fixture this replaced had only four cases, and each case effectively stood alone at the family boundary. No family had enough members for cross-split leakage to be possible.

So “families do not cross splits” was technically true but almost vacuous. The rule starts doing real work only when several related cases could otherwise be scattered across train, dev, and holdout.

The split is frozen and audited before any optimizer runs:

InvariantResult
Case IDs uniquePass
Canonical cases ed-001 – ed-004 presentPass
Train / dev / holdout IDs disjointPass
Source families disjoint across splitsPass
Holdout visible to the optimizerNo
Holdout inspected by the runnerNo
reference_rewrite exposed as a program inputNo
Any evaluation field exposed as a program inputNo
Dataset fingerprint matches canonical helperPass

Both the dataset and the split carry their own fingerprints:

dataset fingerprint:  40ab192bc45590803fcf84ea168a8fa26764f82147e4e38b8334677f93d4fe46
split fingerprint:    ebe3b001711ea92d6ddc12c983b2ef2dacf559a9a014e044038a7084a35c8567

Two fingerprints rather than one, because they can change independently. Adding a case changes the dataset. Moving a family between splits changes only the split. Either invalidates a comparison, and you want to know which happened.

A promise about the holdout

Seven cases, two families, sealed here in chapter 7.

Books that build this apparatus often never spend it — the discipline gets described and the final measurement quietly never happens. We spend it in chapter 19, exactly once, on whichever program the promotion policy selects, and we report what it says regardless of whether it flatters the preceding ten chapters.

If you want to know now: it does not go the way the optimizer chapters set you up to expect.


6. Four words that are not synonyms

TermMeaning in this book
exampleA structured row, usually a dspy.Example
demonstrationAn example placed into program state as behavior guidance
training caseAn example the optimizer may inspect
evaluation caseAn example used to measure a frozen program

One row can play different roles in different experiments, never in the same one. A holdout case used as a demonstration has stopped being holdout evidence, permanently, for that experiment — and no amount of care afterwards restores it.

Optimizer visibility is deliberately not identical across Chapters 10 through 13.

BootstrapFewShot in Chapters 10 and 11 sees the training cases for demo construction while the development family is reserved for evaluation. MIPROv2 and GEPA in Chapters 12 and 13 use development cases inside their search or selection loop. In every case, the seven holdout rows remain outside the optimizer-visible set.

That distinction is part of the experimental protocol, not a property of the word dev. The holdout overlap is always empty, and the runners verify that mechanically rather than trusting this paragraph.

Where this goes next

The mapping to chapter 20’s repository-repair capstone is direct, which is why the fixture is built this way rather than as a bag of sentences:

Editorial programRepository repair
sentencerepository issue
contextrepository evidence
reference rewriteevaluation oracle or expected behavior, when one exists
semantic constraintsengineering constraints
target failurefailure class the case exercises
accepted or rejected editpromoted or rejected intervention

What Usually Goes Wrong

SymptomLikely causeHow to diagnose itWhat to change
The optimizer’s results look too goodEvaluation labels reached the inputsInspect the .with_inputs projection; run a leakage suiteRemove label and metadata fields from program inputs
A result cannot be traced back laterNo stable identity or provenanceLook for case_id and source_id in stored rowsAdd persistent identity before the first experiment
Train and holdout overlapThe split was created or mutated after optimizationCompare IDs, families, and split fingerprintsFreeze and fingerprint splits before compiling
Performance drops sharply on new dataRelated rows crossed a split boundaryCompare source_group values across splitsSplit by family, not by row
Every experiment produces tiny, unstable deltasThe fixture is coarse relative to case-level and runtime variationCompute one-case contribution to the mean; repeat with paired controlsAdd independent cases or stronger repeated evidence before interpreting the delta
Only average performance is measuredThe fixture has no failure-mode structureGroup scores by an expected-failure fieldDesign cases against specific failures
The corpus never rewards restraintEvery reference rewrite changes the sentenceCount rows where reference equals inputInclude no-edit cases, and fix the metric that penalises them

Conclusion

We have a fixture: 44 cases, 12 families, split at family boundaries, fingerprinted twice, and audited against nine invariants before any optimizer was permitted near it. Every case is labelled with the failure mode it exists to exercise, and the six categories are populated deliberately rather than by whatever the sampling happened to produce.

Two ideas from this chapter are worth more than the corpus itself.

The first is that a fixture is an instrument with a granularity you can compute. With 11 development cases, one case owns about 9.1% of the mean: a 0.35 case-level swing moves the aggregate by about 0.032, and even a 0.11 swing moves it by 0.010.

Runtime variability is a separate empirical quantity and must be measured rather than multiplied into that arithmetic. Knowing both scales is what stops you from writing up a single-run +0.011 as evidence of improvement.

The second is that the split is a claim about generalisation, and splitting by row rather than by family quietly weakens the claim to nothing. Our previous fixture obeyed the family rule by having only one case per family, which is the kind of compliance that passes an audit and protects nobody.

We removed the assumption that any useful example is safe to show an optimizer, and the assumption that writing “holdout” beside a row is a control rather than a note.

What we still cannot do is score anything. We have references, constraints, and expected failure modes, and no function that turns a rewrite and a row into a number.

That function is going to be more consequential than the program.

What exactly are we measuring, and what happens when the measurement is wrong?