← DSPy From First Principles

Define the Contract

Treat a DSPy signature as a task contract: choose field names deliberately, decide what belongs inside the boundary, use types where they surface failure, and delete fields nothing consumes.

Chapter 2 built a boundary and then exposed a weakness in it on purpose. We designed the most natural output field available — rewritten_text, a replacement sentence — but failed to represent “no edit needed” as an explicit decision.

The model can technically return the original sentence unchanged. What the contract cannot tell us is whether that identity output means deliberate restraint or an attempted rewrite that happened to reproduce the input. That distinction matters on six of the forty-four later fixture cases — about 13.6% of the corpus. Having a signature did not prevent the mistake.

So the boundary is in the right place and the contract at that boundary is not yet any good. This chapter is about the difference.

The claim worth stating up front: a signature is not a prompt template with types bolted on. It is a task contract, and like any contract, the expensive mistakes are the ones you cannot see when you sign it. The later optimization experiments are largely a record of paying for decisions made in this chapter.

    flowchart LR
    C[which distinctions get a field] --> E[what the model can express]
    E --> O[what the metric can observe]
    O --> T[what an optimizer score can be trusted to mean]
  

Chapter 2’s missing “no edit needed” field is this chain in miniature: a distinction the contract never named becomes a distinction the metric cannot see, and six cases are scored as if restraint and a failed rewrite were the same thing.


1. The compact form, and what it costs

DSPy accepts a compact string form:

rewrite = dspy.Predict("sentence, goal, context -> rewritten_text, rationale")

That is a real signature. It works, it produces the right fields, and for a quick experiment it is the correct amount of ceremony.

Compare it with the class form:

class RewriteSentence(dspy.Signature):
    """Rewrite one sentence to satisfy an editorial goal
    while preserving meaning, entities, and voice."""

    sentence: str = dspy.InputField(desc="The exact sentence to rewrite")
    goal: str = dspy.InputField(desc="The local reason this sentence is being edited")
    context: str = dspy.InputField(desc="Nearby prose needed to preserve continuity and voice")

    rewritten_text: str = dspy.OutputField(desc="A replacement sentence, not a paragraph")
    rationale: str = dspy.OutputField(desc="Short explanation of the edit")

Both forms produce the same five field names in this example. They do not carry the same amount of declared meaning.

The compact form is a real DSPy signature and it is still ordinary source code: you can assign it to a variable, diff it in a pull request, and save the resulting program state. But this particular string supplies only the field names. DSPy has to generate default instructions, default field descriptions, and implicit string types from those names.

The class form makes those choices explicit. It gives the contract a reusable symbol, a deliberate instruction, field-level descriptions, and a place for richer types. That makes review easier because a reader can see which semantics were chosen rather than inferred.

Neither source form is sufficient as experimental identity by itself. A class named RewriteSentence can change between commits just as a string literal can. Later chapters therefore fingerprint the actual program and signature state that executed.

Use the compact form when the semantics really are obvious or when you are still probing the idea. Move to the class form when field meaning, type, or provenance begins to matter.


2. Field names are the durable part of the contract

Field names are not labels for humans. DSPy renders them into the prompt, so the model reads them, and they change its behavior. You can confirm this directly:

result = program(sentence=..., goal=..., context=...)
dspy.inspect_history(n=1)

That prints the prompt DSPy actually constructed, and it is worth doing early and often — it is the single most useful debugging habit in DSPy, and it dissolves most of the mystery about what the framework is doing on your behalf.

Once you accept that the model reads the names, naming becomes design work:

Weak fieldStronger fieldWhat the change does
textsentenceFixes the unit of work. The model stops returning paragraphs.
instructiongoalAsks for intent, not method. “Make this tighter” rather than “delete adverbs.”
contextcontext with a description bounding itSignals nearby prose, not the whole manuscript
answerrewritten_textGives downstream code a stable field and names what it contains
notesrationaleIdentifies an explanation rather than arbitrary metadata
outputrewritten_text + rationaleTwo things the caller wanted separately, returned separately

Now the part that is easy to miss, and that changes how you should allocate effort.

Durability is a property of the optimization surface.

In the optimizer configurations used in this book, field names and field descriptions are held fixed while other program state is allowed to move. BootstrapFewShot changes demonstrations. MIPROv2 searches demonstrations and signature instructions. GEPA proposes instruction mutations. In the measured Chapter 12 candidate, the analyze and rewrite instructions changed while every field name and description remained unchanged.

That is an empirical property of the optimization surface we chose, not a law of DSPy. Another optimizer — or ordinary application code — can choose a wider surface. DSPydantic, for example, deliberately optimizes Pydantic field descriptions. A description is therefore not intrinsically immutable merely because it sits next to a type annotation.

The allocation rule is stronger when stated this way:

  • decide explicitly which parts of the program an optimizer is allowed to change;
  • fingerprint that surface before and after compilation;
  • put stable task semantics in fields and descriptions only if that surface is frozen;
  • put correctness properties that must survive any optimization in deterministic validation or the evaluation criterion.

An instruction such as “preserve character names” is vulnerable if instructions are mutable. A field description is safer in the experiments we run here because descriptions are outside their search surface. A hard entity gate is stronger still because candidate generation cannot rewrite it.

The useful distinction is not docstring versus field description. It is mutable proposal surface versus independently enforced invariant.


3. What belongs inside the boundary

The rule has two halves. Inputs should carry the semantic information a competent human would need to do the task. Outputs should carry the information downstream software needs to act or to validate. Operational metadata stays outside unless it genuinely changes what the model must do.

For the measured rewriting program in this book, the generation inputs are deliberately narrow:

sentence
goal
context

The fixture also carries required_entities, forbidden_terms, semantic_constraints, and a reference_rewrite. Those fields are available to evaluation and provenance, but the canonical dspy.Example exposes only sentence, goal, and context to generation.

The distinction between contract and preference matters before optimization begins:

PropertyRole
Preserve the number 30Contract: dropping it changes the claim
Preserve negationContract: reversing polarity changes the claim
Preserve regularlyContract when it limits the scope of the claim
Sound slightly smootherPreference: legitimate, but not an invariant
Prefer a shorter sentencePreference unless the application defines a hard limit

Preferences can trade against one another. Contract violations should not disappear inside an average. If a property is both computable and non-negotiable, it belongs in deterministic validation or a hard evaluation gate rather than merely in prose the optimizer may rewrite.

That separation is deliberate. Data that an evaluator may inspect is not automatically data the model should receive. If a real application needs an explicit constraint at decision time, add it to the generation contract and record that change. Do not smuggle evaluator-side information into context merely because it exists on the case.

Operational metadata is another category. These are not task inputs:

database row id
HTTP request id
current retry count
UI tab name
optimizer run id

Those matter to the application. They are not task inputs, and putting them in the signature means the model reads them and may act on them. A retry count in the prompt is an invitation to behave differently on the second attempt for no principled reason.

Then there is a third category, which is the dangerous one.

Never let evaluation information become a task input.

The reference rewrite, the human reviewer’s decision, the validation outcome, the score from a previous run — these may all live legitimately in your evaluation record. If any of them crosses into the program’s inputs, the experiment is contaminated, and it is contaminated in the direction that looks like success.

This is not a theoretical caution. In Chapter 18 we build a deliberate leakage suite for this book’s own program: nine attack variants across 32 eligible cases, for 288 attack executions in total. The firewall blocks all 288 contaminated paths. Without that boundary, 224 of the attacks would inflate the apparent score, by an average of about 0.3856.

That number does not mean we can add 0.3856 mechanically to an honest aggregate score — the metric is bounded and the attacks differ by case. It means something more useful: access to answer-derived information creates an effect large enough to dominate the ordinary program differences we spend entire chapters measuring. Once that contamination enters the program, downstream comparison is no longer evidence of model quality.

Nothing about that leakage is exotic. The most common route is a helpful engineer adding “similar approved edits” to the context field, drawn from a table that happens to include the case being evaluated.

So when adding any field, three questions:

Would a competent human need this value to perform the task?
Would downstream code need this output to act safely?
Could this value contain, imply, or be derived from the answer?

If the third is anything other than a confident no, the value is not admissible as a generation input until you can explain its provenance and why it cannot leak outcome information. The question is not whether the field is convenient. It is whether the experiment would still mean the same thing after the model sees it.


4. Types surface failures that prose hides

DSPy signatures use ordinary Python type annotations, and its adapters use the declared structure when formatting requests and parsing responses. Literal, Optional, lists, and Pydantic models are all available as part of the task interface.

The best example in this book is not the rewriting program. It is the judge we build in chapter 9 to evaluate semantic preservation:

from typing import Literal

import dspy

class SemanticConstraintJudge(dspy.Signature):
    """Decide whether a one-sentence rewrite still satisfies a single stated
    requirement about meaning."""

    original_sentence: str = dspy.InputField(
        desc="The sentence before editing."
    )
    candidate_rewrite: str = dspy.InputField(
        desc="The proposed one-sentence rewrite."
    )
    editorial_goal: str = dspy.InputField(
        desc="What the edit was asked to achieve."
    )
    requirement: str = dspy.InputField(
        desc="The single meaning requirement to check."
    )

    verdict: Literal[
        "preserved",
        "violated",
        "unclear",
    ] = dspy.OutputField(
        desc="Whether the rewrite satisfies the requirement."
    )

    reason_code: Literal[
        "meaning_preserved",
        "meaning_reversed",
        "meaning_weakened",
        "meaning_strengthened",
        "entity_altered",
        "scope_changed",
        "register_changed",
        "output_shape_invalid",
        "cannot_determine",
    ] = dspy.OutputField(
        desc="The primary reason for the verdict."
    )

    explanation: str = dspy.OutputField(
        desc="One sentence explaining the verdict."
    )

Every typed field here earns its type. verdict has three values so uncertainty is represented explicitly instead of being coerced into success or failure. The metric can then decide what unclear should cost rather than letting a parser make that policy decision accidentally.

reason_code has a closed vocabulary because the whole point is to aggregate across runs and ask which failure mode is recurring. Free-text explanations are useful for diagnosis but poor grouping keys. explanation therefore stays a string for a human reading the record, while the enum provides the machine-readable category.

The extra editorial_goal input is also deliberate. A rewrite can be semantically faithful to the source and still violate the thing the edit was asked to accomplish; the judge needs enough task context to interpret the named requirement without seeing the reference answer.

Compare that with a field like risk: Literal["low", "medium", "high"] on the rewriting program. It looks disciplined. Ask what reads it. If nothing does, it is three tokens of cost per call and a field that will drift for years without anyone noticing, because nothing ever checks it.

Types are most useful when downstream code branches on the value, when an invalid value would otherwise be silently accepted, when the vocabulary is small and closed, or when the natural shape is a list, number, or structured object. They are least useful on open-ended prose.

Two honest caveats. Types do not make the model honest — a well-typed verdict can still be wrong, which is why chapter 9 spends most of its length validating the judge rather than trusting it. And the exact failure behavior when a model produces "pretty safe" for a Literal["low", "medium", "high"] field depends on the adapter and provider. It may surface at parse time; it may be coerced. Any field that controls persistence, ranking, or policy still needs validation in your own code.

A related case study. DSPydantic is useful here precisely because it is a counterexample to the idea that field descriptions are inherently fixed. It explicitly treats Pydantic prompts and field descriptions as an optimization surface for structured extraction.

That reinforces Section 2 rather than contradicting it. A description survives only when your chosen optimizer is not allowed to rewrite it. The Python type may remain unchanged while the natural-language semantics around that type move substantially.

Descriptions are semantics, not decoration — and whether they are stable or optimizable must be an explicit design choice.

None of which makes a typed output true. A Pydantic-valid invoice date can still be the wrong date, and a valid enum can still be the wrong classification. Structure gives you a place to validate. It does not do the validating.


5. Every field needs a job

Chapter 2 promised this question, so here it is: should the rewriting program emit a confidence score?

The instinct is yes. It is one float, it looks cheap, and downstream code might eventually want to threshold on it.

Chapter 1 gives us a warning, not a calibration study. Its illustrative failure outputs attach 0.88 to a renamed protagonist and 0.94 to an overwrought rewrite. Those examples do not prove that model confidence is statistically uncorrelated with correctness. They prove something more immediate: a number emitted by the same generative call is not independent evidence that the output is safe.

The model also does have enough information to notice that Jalen became Jason; the original sentence is in its input. The failure is not lack of access. The failure is that nothing has established what the self-reported number means, how it is calibrated, or whether 0.8 has the same interpretation across cases, prompts, models, or program versions.

That is enough reason to remove the field. Until confidence has an external calibration procedure, persisting it invites downstream code to manufacture authority from decoration:

if result.confidence > 0.8: auto_approve()

The rule that generalizes is:

Every output field needs a job. Any field that can control an action also needs independent validation or calibration.

Applying that rule to the running program:

FieldJob / consumerIndependent checkVerdict
rewritten_textapplication output; evaluation targethard gates, metric, semantic judgeKeep
rationalediagnostic context for humansnone; explicitly non-controllingKeep, but never use it as acceptance evidence
confidencewould invite thresholdingno calibrated meaning establishedRemove
riskno current consumerno validation policyDo not add until the application defines both

rationale therefore does not survive on a loophole. It survives because its role is observational rather than controlling. A human may use it to understand what the program attempted, but the system does not promote, reject, persist, or route a candidate because the rationale sounds persuasive.

That distinction is the useful one: informational fields may help explain behavior; control-bearing fields must earn authority.


6. Contracts can be too strong

The opposite failure is real and less discussed.

class OverSpecifiedRewrite(dspy.Signature):
    """Rewrite the sentence in twelve to sixteen words using active voice,
    one comma at most, no semicolons, no adverbs, and a stronger final verb."""

    sentence: str = dspy.InputField()
    rewritten_text: str = dspy.OutputField()

As a copyediting constraint test, fine. As the general sentence-rewriting contract, it has baked one editorial theory into the task boundary, and it will now be wrong for every author who does not write that way.

Three forms of over-specification, all of which look like rigor:

FailureExampleWhat it costs
Method leakage“Think step by step and compare three alternatives”You can no longer swap the module cleanly, so chapter 4’s comparison becomes uninterpretable
Policy leakage“Always accept if confidence is above 0.8”Application decision policy is now inside the model’s instructions, where it cannot be tested or audited
Inherited-judgment leakage“Use the previously approved house style”A prior system’s opinion is encoded as task truth, and nobody can say what it was

The test is the same one from chapter 2: would this clause still be true if you changed the execution strategy, the model, or the author? If not, it is not a contract term.

There is also a subtler version. In the optimizer surface used later in this book, signature instructions are mutable. A constraint placed only in those instructions can therefore be weakened or removed by a compiled candidate if doing so improves the search objective.

Field descriptions remain fixed in our Chapter 10–13 experiments, but Section 2 deliberately does not treat that as a universal guarantee. If a requirement is truly non-negotiable, the strongest home is an independently enforced gate or metric rather than any natural-language instruction the optimization process may be allowed to rewrite.


7. Two contracts, not one

A DSPy signature governs the LM boundary. Your application has a different boundary, and conflating them is a common source of trouble.

from pydantic import BaseModel, Field

class RewriteInput(BaseModel):
    sentence: str = Field(min_length=1)
    goal: str = Field(min_length=1)
    context: str = ""
    constraints: list[str] = Field(default_factory=list)

class RewriteOutput(BaseModel):
    rewritten_text: str = Field(min_length=1)
    rationale: str = ""

def validate_prediction(raw: dict) -> RewriteOutput:
    return RewriteOutput.model_validate(raw)

This validates the application payload around the program, not a dspy.Prediction object directly. That distinction is the point. DSPy’s signature and adapter define and parse the model-facing shape. Pydantic enforces the application-facing payload before persistence or downstream use.

Neither layer proves semantic correctness. A perfectly valid RewriteOutput can still contain the wrong entity or a distorted claim. Shape validation protects the interface; the hard gates, metric, and semantic judge protect different properties.

The two boundaries overlap, but they are not substitutes for one another.


What Usually Goes Wrong

SymptomLikely causeHow to diagnose itWhat to change
rewritten_text comes back as a paragraph with commentaryOutput scope is declared semantically but not enforcedCount sentences and newlines across a runNarrow the field description; enforce shape at the application boundary
A constraint you wrote in the docstring stops being honoured after compilingInstructions are rewritten by the optimizer; docstrings are not durableDiff the compiled program’s instructions against the originalMove the constraint into field descriptions or the metric
Optimizer score shows a large, sudden jumpEvaluation information reached the inputsCheck whether any input field contains, implies, or derives from the referenceRemove it, then run a leakage suite (chapter 18)
Downstream code is full of fallback branchesThe output contract is too looseCount missing or invalid fields per hundred callsAdd typed fields or split coarse ones
A field has drifted for months and nobody noticedNothing consumes it and nothing validates itFor each output field, name the consumer and the validator out loudDelete it
The program follows stale product policyPolicy was written into the signatureSearch docstrings for deployment rules and thresholdsMove policy into deterministic application code

Conclusion

The contract now names what the model consumes and produces, and the names reflect the application’s concepts rather than the convenience of whoever wrote the first prompt.

Two things in this chapter are worth carrying forward, because neither is standard advice.

The first is that durability depends on what optimization is allowed to mutate. In the experiments ahead, demonstrations and signature instructions are search variables while field names and descriptions are held fixed. That makes the fields more stable than the instructions in this program, not universally immutable. Anything that must remain true regardless of optimization belongs outside the mutable proposal surface, in deterministic validation or an independently versioned metric.

The second is that output fields do not acquire authority merely by being typed. confidence looked free, but no experiment established a calibrated meaning for the number. A field that can influence an action must earn that influence through an independent check; otherwise its apparent precision is a liability.

The assumption we removed is that the model will understand what we mean. It will do something with what we wrote, which is a different thing.

What remains untouched is execution. We have specified the task carefully and said nothing about how a model should attempt it — deliberately, because that separation is what makes the next question answerable at all.

We can now hold the contract completely fixed and change only the strategy. So:

Does asking the model to reason before it answers actually produce a better rewrite, and what does it cost?