← DSPy From First Principles

You Cannot Optimize What You Cannot Measure

Define the metric, freeze the evaluation protocol, and measure the program — then see why a measurable objective can make the wrong behavior systematically optimizable.

Chapter 7 produced material. This chapter produces evidence.

That makes optimization possible. It does not yet make optimization safe.

no metric
    → no systematic optimization

metric
    → behavior can be optimized

wrong metric
    → the wrong behavior can be optimized systematically

This chapter establishes the surface on which search can operate. Chapter 9 asks what that surface rewards when the metric becomes the target.

    flowchart TD
    P[frozen program] --> E[evaluate]
    C[frozen cases] --> E
    M[metric] --> E
    E --> PC[per-case results]
    PC --> AG[aggregate result]
    PC --> FI[failure inspection]
  

No optimization happens here. Note that the per-case results feed two readings — the aggregate and the failure inspection — and this chapter’s finding comes from the second. The question is narrow:

How well did this frozen program perform under this frozen protocol?

The answer turns out to be less interesting than what we find while computing it. The aggregate number is fine. The per-case breakdown contains a defect that shapes the next nine chapters.


1. A baseline is an identity, not a number

A baseline is not “whatever the code did yesterday.” It is a program identity, a model identity, a dataset identity, a metric identity, and a protocol — and if any of those is missing, the number cannot be compared to anything later.

BASELINE_IDENTITY = {
    "program_id": "EditorialRewriteProgram",
    "module_strategy": "analyze_predict__rewrite_predict__assess_predict",
    "model_provider": "Ollama",
    "model_name": "qwen3:latest",
    "dataset_id": "book-editorial-corpus",
    "split_role": "dev",
        "metric_name": "editorial_metric",
    "metric_version": "v1",
    "optimization_allowed": False,
}

In a real system this record is generated, never typed. The generated version covers program fingerprint, dataset fingerprint, split fingerprint, provider and LM-configuration fingerprints, metric fingerprint, DSPy version, seed, budget, and source revision.

That collection has a name in this book: the evaluation protocol fingerprint. It exists to make one sentence meaningful.

A candidate that reports a higher number under a different protocol has not beaten this baseline. It has run a different experiment.

That sounds pedantic until chapter 12, where an optimizer reports a gain, and the only way to know whether it is real is to confirm that the baseline and the candidate were scored by the same function over the same cases under the same model configuration.


2. The metric: a hard gate and a structural score

The metric has two stages, and the first one is not a score at all.

Stage one is a deterministic hard gate. In the frozen v1 implementation it checks two non-negotiable lexical constraints: every required entity must still be present, and no forbidden term may appear. Its answer is binary.

That is narrower than “is the output usable?” Empty output, sentence count, excessive length, and other shape problems are handled elsewhere in the scoring or validation path. Keeping the exact boundary visible matters because Chapter 9 will attack what the gate does not know how to inspect.

import re

def contains_term(text: str, term: str) -> bool:
    if " " in term.strip():
        return term.lower() in text.lower()
    return re.search(rf"\b{re.escape(term.lower())}\b", text.lower()) is not None

def passes_hard_gate(example, candidate: str) -> bool:
    required_ok = all(
        contains_term(candidate, entity)
        for entity in example.required_entities
    )
    forbidden_ok = all(
        not contains_term(candidate, term)
        for term in example.forbidden_terms
    )
    return required_ok and forbidden_ok

A rewrite that drops a required entity or emits a forbidden term is not treated as slightly worse under this protocol. It scores zero.

That is a policy choice, not a universal property of metrics. The useful distinction is that a gate encodes disqualification, while a weighted score ranks candidates that remain eligible. If you decide that an entity failure is disqualifying, do not let unrelated fluency points buy it back.

Stage two scores what survives the gate, as a weighted sum of three components:

score = 0.35 · changed    (did the program edit at all?)
      + 0.25 · scope      (did it stay proportionate to the task?)
      + 0.40 · overlap    (how close is it to the reference rewrite?)

The weights sum to 1.0, but the component names are more generous than their implementations.

changed is binary: after token normalization, the candidate either differs from the source or it does not.

scope is also binary. It returns zero for an empty candidate or for one longer than max(original_tokens + 4, 1.35 × original_tokens). It does not penalise aggressive shortening unless the candidate becomes empty.

overlap is asymmetric reference-token recall:

|tokens(reference) ∩ tokens(candidate)| / |tokens(reference)|

Extra candidate tokens do not directly reduce it, and word order does not matter.

This is editorial_metric at version v1. It is deterministic, cheap, inspectable, and therefore useful. It is also a deliberately incomplete proxy. Chapter 9 will exploit exactly the asymmetries you can already see in these definitions.

Note one thing now, because it becomes important in section 6. Drop the changed term and you get:

0.25 + 0.40 = 0.65

0.65 is the score for an unchanged candidate when scope passes and reference overlap is perfect. In other words, if the reference itself says leave this sentence exactly as it is, the correct identity output forfeits the entire 0.35 changed component by construction.

That number is going to come up a lot.


3. Running it

DSPy’s evaluator takes a devset and a metric and returns an aggregate plus per-case results.

import dspy
from common.metrics import editorial_metric

def dspy_editorial_metric_v1(example, pred, trace=None) -> float:
    return editorial_metric(example, getattr(pred, "rewritten_text", ""), version="v1").score

def evaluate_baseline(program, examples):
    evaluator = dspy.Evaluate(
        devset=examples,
        metric=dspy_editorial_metric_v1,
        display_progress=True,
        display_table=False,
        failure_score=0.0,
    )
    return evaluator(program)

Two naming decisions are worth explaining, because both were mistakes we had to correct.

The metric is called editorial_metric and takes version as a required argument with no default. Earlier drafts had four names for what were really two functions — editorial_metric, editorial_metric_v1, score_editorial_output, score_editorial_output_v1 — and a default version, which meant a call site could silently change meaning when the default moved. Version is now explicit at every call.

The DSPy-facing wrapper is dspy_editorial_metric_v1, not dspy_editorial_metric. When chapters 10 to 13 hand a metric to an optimizer, the version being optimized against is the single most consequential fact about that run, and it should be legible in the call:

optimizer = dspy.BootstrapFewShot(metric=dspy_editorial_metric_v1)

Nobody reading that line has to guess what is being maximised.


4. What we measured

The frozen EditorialRewriteProgram, on the 11-case development split, under the canonical local Qwen3/Ollama dependency.

On the canonical development case ed-003, the program produced:

“I don’t think we should go there,” Anna said, “because it’s unsafe.”

Score: 0.877, stable to three decimal places across every run we made. Latency around 5.45 seconds.

The development aggregate:

RunDev mean, v1Dev mean, v2
First0.78930.7893
Second0.79920.7992

The baseline is approximately 0.79.

That phrasing is deliberate. Two reported executions of unchanged program state produced development means of 0.7893 and 0.7992, an observed span of about 0.010. Chapter 6 shows why we should not promote that span into a universal ± noise constant: fresh-session drift and within-session order effects are both part of the measurement process.

Reporting 0.7893 as the baseline would therefore give the point estimate more authority than it has.

Compare that with the fixture this replaced, which reported 0.8769230769 — ten decimal places describing one sentence about Anna. The digits were real in the sense that arithmetic produced them. They were not ten decimal places of evidence.

The book carries forward about 0.79, while the machine artifacts retain full precision and the repeated runs retain their exact values. A candidate earns the word improvement only after a paired comparison shows that its delta survives the variability of the evaluation process.

Two deterministic sanity baselines, on ed-003:

BaselineScore
No edit0.6500
Mechanical and then → comma rule0.6500
EditorialRewriteProgram0.8769

On ed-003, the mechanical rule is a no-op because the sentence contains no and then, so it collapses to the same output as the no-edit baseline.

These are harness checks, not serious competitors. They show that the deterministic metric path executes as expected and that the LM program can score above an unchanged candidate on this particular row. They do not establish that an arbitrary mechanical edit is worth 0.65 or that 0.65 universally means poor quality.

One more result, which we did not expect and should have: in the two reported development-baseline runs, v2 equals v1 case for case. The semantic judge recorded no violated constraints on those baseline outputs.

Recall what v2 changes. It keeps the same hard gate and the same structural score, then applies a semantic factor from the judge; if any stated semantic constraint is violated, the final score is capped at 0.30. So agreement with v1 means the semantic component had nothing to penalise in these particular baseline runs.

That makes v2 useful in a different way from a finer-grained ruler. In the measurements so far, it behaves like a tripwire: ordinary baseline output leaves the structural score untouched, while specific semantic regressions such as ed-025 and later ed-035 make the two metrics separate sharply.

That is an empirical pattern in this experiment, not a guarantee that v2 only diverges after optimization. Any future baseline output that violates a named semantic constraint would trigger it too.


5. The aggregate hides everything interesting

Two programs can both score 0.75:

Program A: mostly good, one invalid output
Program B: always valid, quietly renames entities

The mean cannot distinguish them, and the difference is the only thing that matters. So the runner stores per-case evidence — output, score components, provider and configuration identity, latency, tokens, and whether the result is eligible as model evidence at all.

def unpack_evaluation(result) -> list[dict]:
    return [
        {
            "case_id": example.case_id,
            "target_failure": example.target_failure,
            "score": float(score),
            "prediction": {
                "rewritten_text": getattr(prediction, "rewritten_text", None),
                "rationale": getattr(prediction, "rationale", None),
                "risk": getattr(prediction, "risk", None),
            },
            "reference": {
                "reference_rewrite": example.reference_rewrite,
                "required_entities": list(example.required_entities),
                "forbidden_terms": list(example.forbidden_terms),
            },
        }
        for example, prediction, score in result.results
    ]

The target_failure field from chapter 7 now earns its place. Grouping by it turns one number into a diagnosis:

Target failureDev casesMean score
semantic_drift20.938
entity_change20.825
style_failure30.808
constraint_violation10.782
invalid_output10.650
unnecessary_edit20.650

The aggregate of about 0.79 hides where the score comes from. This particular run’s grouped table shows two development categories at exactly 0.650.

But 0.650 does not have one universal interpretation. For one category it reflects an output that failed to make useful progress. For the restraint cases below, the same number can be the arithmetic penalty imposed on a correct identity decision.

That ambiguity is itself evidence that the aggregate needs the per-case record beside it.


6. The floor, and what the metric cannot represent

Two categories are pinned to the do-nothing floor, for completely different reasons. One is a program failure. The other is a metric failure, and it is ours.

invalid_output at 0.650 is a program-side failure in this development run. The case is designed to exercise output-shape risk, but the frozen v1 hard gate does not actually test general output shape; it only checks required entities and forbidden terms.

So a candidate can survive the gate and still land at the structural floor. The target_failure label tells us what the case was designed to probe; the component record tells us what the metric actually observed.

unnecessary_edit at 0.650 is different. The two development cases in this category, ed-034 and ed-037, have identity references: the correct reference output is exactly the source sentence.

For any such identity-reference case, a program that correctly leaves the sentence unchanged scores:

changed:  0     (nothing changed — correct)
scope:    1.0   (perfect — correct)
overlap:  1.0   (identical to the reference — correct)

score = 0.35·0 + 0.25·1 + 0.40·1 = 0.65

For identity-reference cases, the metric caps a correct unchanged output at 0.65 by construction. The changed component is 35% of the score and it pays for editing regardless of whether editing was required.

The full corpus contains six cases labelled unnecessary_edit, but that label does not mean all six references are byte-identical to the source. Some ask for restraint or minimal intervention rather than literal identity. The defect is therefore better stated as a property of the metric: whenever no change is the correct answer, v1 has no way to award the missing 0.35.

This is the Chapter 2 problem arriving on schedule. The contract has no explicit semantic state for “no edit needed,” and the metric independently encodes a prior that changing text is always worth points. Those two design choices can compound: an unnecessary edit may outscore a correct decision to leave the sentence alone.

Chapter 18 finds the same tension from another direction when literal reference-blocking collides with cases whose admissible answer can equal the reference. The lesson propagates because the representation, metric, and leakage policy all made assumptions about what a valid edit must look like.

We did not silently repair v1 after seeing this defect. Once Chapters 10 through 13 optimize against a frozen metric, changing its weights or adding a special case would create a different experiment and invalidate direct comparison with the runs already recorded.

That does not mean the defect is comparison-neutral. A candidate that edits an identity-reference case can gain the 0.35 changed credit and therefore outrank a candidate that correctly declines. v2 inherits the same structural blend, so its semantic judge does not fix this particular incentive either.

A proper repair would require a new metric version — ideally one that scores the decision whether to edit explicitly rather than assuming change is always desirable — followed by a new experimental cycle.

For this cycle we freeze the defect, publish it, and keep its direction visible when interpreting later scores.

A metric with a documented defect is more useful than a metric with an undocumented one. But notice what has happened. Before Chapter 9 attacks the metric deliberately, routine inspection of the per-case table has already found a place where the objective encodes the wrong preference.


7. What deterministic checks can catch

Not every failure needs a model to detect it. A surprising number of production text failures are cheap, boring string problems, and catching them before a judge is involved saves cost and removes noise:

helper text leaked into the output
prompt fragments in the response
TODO or placeholder markers
markdown artifacts in plain-text fields
unbalanced quotation marks
terminal punctuation changed
more than one sentence returned
length wildly disproportionate to the original

Every one of those has a deterministic component, so none needs an LM judge merely to detect its surface form.

That does not mean the frozen v1 hard gate already checks all of them; it does not. Its gate is limited to required entities and forbidden terms. Additional deterministic validators can live before scoring or become future hard-gate rules, but adding one changes the evaluation protocol and must be versioned.

The broader rule is: if a property can be checked deterministically and policy says failure is disqualifying, enforce it in software rather than asking a model to rediscover it.

The failure taxonomy for the editorial program is the target_failure vocabulary from chapter 7. Its repository-repair analogue in chapter 20 has the same shape:

patch does not apply
tests still fail
unrelated files changed
validation unavailable
candidate produced by fallback

Evaluation is where these become visible, which is only true if you store enough per-case detail to see them.


What Usually Goes Wrong

SymptomLikely causeHow to diagnose itWhat to change
The baseline cannot be reproducedProgram, model, data, or protocol identity was not recordedCompare run manifests and fingerprintsPersist the full evaluation protocol
A candidate looks better after the metric changedBaseline and candidate ran under different protocolsCompare protocol fingerprintsTreat the comparison as invalid until they match
Evaluation crashes on one malformed outputThe metric assumes well-formed predictionsRe-run with the tracebackMake the metric defensive and assign a failure score
The aggregate looks fine and users complainThe metric is too narrow to see the failureGroup per-case scores by expected failure modeAdd components that can see what users notice
A whole category sits at one exact valueA metric component or floor may dominate the categoryDecompose the value into component weights and inspect the casesFix it in a new metric version, or publish the defect and its direction
Scores are quoted to many decimal placesArithmetic precision is being confused with evidential precisionRepeat the same frozen program under the actual execution scheduleKeep full precision in artifacts; use appropriately rounded prose plus the repeated values
Optimization starts before a baseline existsEvaluation and optimization were blurred togetherLook for candidate state predating the baseline recordFreeze the baseline first

Conclusion

We have a frozen baseline: a named program, a named metric with an explicit version, and a fingerprinted protocol. Two reported executions put its 11-case development mean at 0.7893 and 0.7992, so the prose-level baseline is about 0.79, not one sacred four-decimal number. ed-003 remains a useful walkthrough at about 0.877, and v2 agrees with v1 case for case in those reported baseline runs.

Two things from this chapter travel further than the number.

The first is that the aggregate was the least informative thing we computed. Grouping by expected failure mode turned about 0.79 into a diagnosis, and the diagnosis exposed a defect: the two development unnecessary_edit cases sit at exactly 0.650.

That is not a coincidence. On identity-reference cases, an unchanged candidate gets full scope and overlap credit but loses the entire 0.35 changed component. The broader six-case unnecessary_edit category is not uniformly identity-reference, so the correct general claim is structural rather than numerical: v1 systematically under-rewards cases where restraint is the right decision.

The second is what that implies about the whole enterprise. We built v1 from deterministic components and put two explicit disqualifying checks in front of its weighted score. That still leaves properties the gate does not inspect and a changed-term incentive that can reward the wrong behavior.

We found one of those defects by reading a grouped table, not by any clever adversarial technique.

So the obvious question is what else is in there.

We removed the assumption that optimization can begin before measurement, and the assumption that a larger number beats a baseline regardless of the protocol that produced it.

What remains is the metric itself, which so far we have mostly used rather than challenged. The quality comparisons in Chapters 4 and 5 depend on its v1 structural score, and the optimizer chapters will deliberately search against that frozen objective.

We now know one place where the objective is wrong before the optimizer has even touched it. Its only remaining qualification cannot be that we wrote it.

That is not a good enough reason to trust it.

What happens when the measurement stops being an observer and becomes the target of search?