AIBussin Book Published

DSPy From First Principles

Turn language-model behavior into an experimental variable, then learn when optimization is measurable, when improvement is trustworthy, and what DSPy contributes to the process.

A language-model application does not have to remain a collection of prompts. It can become a program whose behavior can be varied, measured, and compared.

That is DSPy’s deepest contribution. The important shift is not from one way of writing instructions to another. It is from prompt wording as craft to language-model behavior as an experimental variable.

Once behavior becomes an experimental variable, a harder problem appears. What may change? What evidence guides the change? What counts as success? Which evidence must remain independent of optimization? And when a candidate scores better, what justifies calling it an improvement?

This book builds that experimental system from first principles.

We begin with an ordinary handwritten prompt and keep asking what is missing. Each answer introduces another mechanism: a contract, a module, a dataset, a metric, an optimizer, a tool boundary, a context strategy, an experimental firewall, a promotion policy.

Then we test the machinery in two different regimes. Editorial rewriting gives us a task whose success is partly judgment: a metric can approximate what we care about while missing a load-bearing word. Repository repair gives us a task with stronger external evidence: code can parse, scope can be checked, and tests can establish observable behavior.

The contrast is the book’s empirical question:

How does optimization change when correctness moves from a proxy for judgment to evidence produced by the environment?

The aim is not to memorize the DSPy API.

It is to understand what DSPy makes possible well enough that optimizers, agents, reflective search, and compiled programs stop looking like framework magic and become inspectable engineering machinery.

The recurring question throughout the book is:

What exactly is the program, what evidence says it is better, what was the optimizer allowed to see, and what justifies replacing the version that already works?

The question we will test

The book does not assume that optimization and improvement are synonyms.

    flowchart LR
    A[Editorial rewriting] --> B[Success is partly judgment]
    B --> C[Metric is an approximation]
    C --> D[Optimization can exploit omissions]

    E[Repository repair] --> F[Success has executable consequences]
    F --> G[Parsing, scope, and tests]
    G --> H[Optimization can be checked externally]
  

Neither regime gives perfect truth. Human editorial judgment is real even when it is difficult to compress into a scalar, and a passing test suite can still be incomplete. The difference is the strength and independence of the evidence available.

To make that comparison possible, the progression begins simply:

    flowchart LR
    A[Task] --> B[Signature]
    B --> C[Module]
    C --> D[Language model]
    D --> E[Structured output]
  

Then we add the machinery required to improve that program empirically:

    flowchart LR
    A[Program] --> D[Baseline evaluation]
    B[Examples] --> D
    C[Metric] --> D
    D --> E[Optimizer]
    E --> F[Candidate program]
    F --> G[Untouched evaluation]
    G --> H[Comparison]
    H --> I{Decision}
    I -->|Better evidence| J[Promote]
    I -->|Insufficient or worse| K[Reject]
  

By the end of the book, the same experimental discipline has grown into something much more substantial:

    flowchart TD
    A[Task] --> B[Program contract]
    B --> C[Execution strategy]

    C --> D[LM modules]
    C --> E[Deterministic code]
    D --> F[Tools and context]
    E --> F

    F --> G[Behavior]
    G --> H[Observation]
    H --> I[Metric]
    I --> J[Score and feedback]
    J --> K[Optimizer]
    K --> L[Candidate program]

    L --> M[Independent holdout]
    M --> N[Comparison]
    N --> O{Promotion decision}
    O -->|Promote| P[Production evidence]
    O -->|Reject| Q[Retain active version]
    P --> R[Next experiment]
    Q --> R
  

A language model still performs some of the work.

But it is no longer the whole system.

The rest is software, data, evaluation, experimental discipline, and evidence.

That distinction is the foundation of this book.

What this book is designed to teach

By working through the chapters, you will learn how to turn the hidden parts of an LLM application into explicit engineering objects.

You will learn how to:

  • distinguish a handwritten prompt from a language-model program;

  • express a task as a semantic contract rather than burying the interface inside prose;

  • use DSPy Signatures to define what a component consumes and produces;

  • separate what a program should do from how it attempts to do it;

  • compare execution strategies such as direct prediction and explicit reasoning under the same contract;

  • compose several LM operations into one inspectable program;

  • decide which parts of the system belong in deterministic Python rather than in another model call;

  • treat the underlying language model as a configurable dependency instead of an invisible assumption;

  • record provider, model, configuration, fallback state, and program identity as part of the evidence for a run;

  • turn examples into structured data rather than copying demonstrations into prompts by hand;

  • distinguish program inputs, labels, metadata, training examples, demonstrations, development cases, and holdout evidence;

  • establish a baseline before allowing an optimizer to change anything;

  • distinguish optimization, generalization, and improvement rather than using the words interchangeably;

  • design metrics that reflect the task rather than merely rewarding convenient proxies;

  • attack metrics with adversarial examples before trusting an optimizer to pursue them;

  • use deterministic constraints, reference evidence, LM judges, and human preference without confusing one form of evidence for another;

  • understand DSPy compilation as candidate-program generation rather than mysterious prompt improvement;

  • use few-shot optimization to construct or select demonstrations;

  • use instruction optimization to search over LM-facing program state;

  • use reflective feedback to explain failures and propose better instructions;

  • understand why reflection proposes changes while evaluation still decides whether those changes helped;

  • treat tool-using agents as language-model programs whose trajectories, tools, and outcomes can also be measured and optimized;

  • design bounded tool interfaces instead of granting an LM uncontrolled access to its environment;

  • distinguish direct context, retrieval, memory, agent-directed search, and programmatic exploration of large contexts;

  • prevent training labels, future repository revisions, solution patches, evaluator outputs, or repeated holdout inspection from contaminating an experiment;

  • version programs, datasets, metrics, providers, tool surfaces, retrieval systems, and validation policies so an experimental result remains interpretable;

  • distinguish an optimizer score from independent evidence that a candidate is actually better;

  • separate evaluation, comparison, promotion, activation, observation, rejection, and rollback;

  • treat production activity as a source of future evidence rather than automatically converting every interaction into a training label; and

  • assemble the mechanisms into a repository-repair program that can improve through controlled experiments without silently promoting its own changes.

The objective is not to make prompt optimization more elaborate.

It is to make language-model behaviour more understandable, measurable, reproducible, and improvable.

One argument, built progressively

The chapters build one engineering argument across two task regimes. The editorial program makes subjective success measurable enough to optimize and then exposes the boundary of the proxy. Repository repair moves the same machinery into an environment where important parts of correctness can be established externally.

01  Why Are We Still Hand-Writing Prompts?

    Start with a reasonable prompt and discover why a string
    becomes an inadequate software abstraction.


02  A Prompt Is Not Yet a Program

    Give language-model behaviour a program boundary:
    inputs, declared behaviour, execution, and outputs.


03  Define the Contract

    Use Signatures to express what the program consumes
    and what downstream software can rely upon.


04  Separate What From How

    Keep the task contract stable while changing the
    execution strategy used to satisfy it.


05  Build Programs From Programs

    Compose smaller LM components with deterministic
    Python into an inspectable multi-stage program.


06  Your Model Is a Dependency

    Make provider, model, configuration, failure,
    and fallback behaviour explicit.


07  Examples Are Experimental Data

    Treat examples as evidence with identities, provenance,
    roles, permissions, and frozen family-level splits.


08  You Cannot Optimize What You Cannot Measure

    Establish a named baseline and record per-case
    evidence before attempting optimization.


09  When the Metric Becomes the Target

    Attack the objective, expose what it cannot represent,
    and separate optimization success from improvement.


10  Compile the Program

    Treat optimization as the production of a candidate
    program under a frozen experimental boundary.


11  Few-Shot Optimization

    Let the optimizer select and construct demonstrations
    instead of hand-pasting examples into prompts.


12  Optimize the Instructions

    Search instruction and demonstration space under
    an explicit metric, development set, and budget.


13  Let the Program Reflect

    Turn scalar failure into diagnostic feedback that
    can guide reflective program improvement.


14  Change One Thing

    Hold corpus, program, model, and optimizers fixed and
    swap only the objective, once failure can be named.


15  Agents Are Programs Too

    Give the program tools and let it act on an environment
    without abandoning contracts, measurement, or validation.


16  Search, Memory and Long Context

    Decide what evidence the program should inspect when
    the relevant environment is larger than the prompt.


17  Search the Reasoning Space

    Search over reasoning states, not just program states,
    and see why an oracle-guided search is not evidence.


18  Don't Let the Optimizer Cheat

    Build an experimental firewall around holdouts,
    tools, memory, retrieval, history, and evaluation.


19  From Experiment to Production

    Compare candidates independently, promote deliberately,
    observe active versions, and retain the ability to rollback.


20  Build a Self-Improving Engineering Program

    Assemble the mechanisms into a repository-repair
    system whose improvement claims are backed by evidence.


21  Beyond the Book: Production DSPy Systems (Appendix)

    See where the book's machinery goes next: budgeted
    reasoning search, non-regressing memory, and a small
    case-based reasoning module, drawn from production code.

Each chapter adds a mechanism because the previous system has exposed a specific limitation. The task changes only when the limitation itself becomes the subject of the experiment. Chapter 20 is the capstone; Chapter 21 is an appendix that points beyond the book’s evidence gates to production systems, with full source linked rather than claimed as book evidence.

That progression matters.

An optimizer makes little sense before we have a metric. A metric is difficult to trust before we have examples and a baseline. Holdout isolation becomes much more important once an optimizer can search instructions, retrieve memories, call tools, inspect repository history, and learn from diagnostic feedback.

The final chapter therefore does not introduce another optimizer.

It asks whether everything we have already built can actually compose.

The resulting system looks roughly like this:

    flowchart TD
    A[Engineering request] --> B[Freeze repository revision]
    B --> C[Inspect evidence]
    C --> D[Search, retrieve, use tools]
    D --> E[Diagnose]
    E --> F[Propose intervention]
    F --> G[Generate candidate patch]
    G --> H[Apply in isolation]
    H --> I[Validate]

    I -->|Fail| J[Inspect failure]
    J --> K[Revise when justified]
    K --> G

    I -->|Pass| L[Record outcome evidence]
    L --> M[Evaluate program behavior]
    M --> N[Build audited dataset]
    N --> O[Controlled optimization]
    O --> P[Candidate program version]
    P --> Q[Independent holdout]
    Q --> R[Compare with active baseline]
    R --> S{PROMOTE / REJECT / INSUFFICIENT}
    S --> T[Operator decision]
    T --> U[Observe production]
  

The point is not autonomous self-modification.

It is controlled empirical improvement.

Three claims that must remain separate

This book uses three words deliberately:

optimization
    a candidate scores better on the objective used by search

generalization
    the candidate performs better on unseen cases drawn from the task

improvement
    independent evidence says the system is actually better for its purpose

An optimizer can succeed without producing generalization. A candidate can generalize under a proxy without improving the real application. The editorial experiments make those separations visible; the repair experiments test how much stronger the claim becomes when the environment can check behavior directly.

The central engineering principle

Language models are probabilistic.

That is useful.

They can generate rewrites, diagnoses, plans, patches, search queries, tool selections, explanations, hypotheses, and alternatives that would be difficult to encode as ordinary deterministic logic.

But not every part of an LLM application should therefore become probabilistic.

When ordinary software can establish something reliably, let ordinary software establish it.

Does this output parse?                 → parser

Does it satisfy the schema?             → validator

Is this tool allowed?                   → capability policy

Which examples belong to holdout?       → dataset manifest

Did the patch apply?                     → patch engine

Did the test pass?                       → test runner

Did the named entity survive?           → deterministic check

Which program version generated this?   → artifact registry

Did the optimizer see this case?        → experiment audit

Did the candidate beat the baseline?    → evaluator + comparison policy

Which version is active?                → deployment registry

That gives us one of the book’s central rules:

Use the language model where uncertainty and generative reasoning are useful. Move correctness, identity, boundaries, measurement, and governance into inspectable software wherever possible.

DSPy then occupies a very specific place.

It gives us machinery for expressing and improving the probabilistic parts of the program.

It does not remove the need for the deterministic system around them.

Optimization is not promotion

One distinction will recur throughout the book:

    flowchart LR
    A[Optimizer] --> B[Candidate]
    B --> C[Evaluation]
    C --> D[Evidence]
    D --> E[Comparison]
    E --> F[Recommendation]
    F --> G[Promotion]
    G --> H[Active version]
  

Those are different operations.

An optimizer may discover a program that performs better on development examples.

That does not entitle the candidate to replace the active program.

A reflective optimizer may explain why an example failed.

That does not make its proposed fix correct.

An agent may produce a convincing trajectory.

That does not prove that the requested outcome occurred.

A compiled program may achieve a higher scalar score.

That does not erase a new hard regression.

The book therefore keeps returning to another rule:

The optimizer proposes. Independent evaluation establishes evidence. Promotion is a separate decision.

That separation becomes more important as the system becomes more capable.

Evidence has a boundary

Optimization introduces another problem: the system can learn the experiment instead of learning the task.

A program may accidentally receive a human decision as an input.

A demonstration may contain a holdout case.

A retrieval index may contain the eventual solution.

A repository agent may inspect a later commit containing the exact patch.

Reflective feedback may reveal the gold answer.

A developer may inspect holdout failures, modify the program, and then call the same cases an untouched test set.

All of these can create impressive-looking results.

None provide trustworthy evidence of general improvement.

The experimental boundary therefore becomes part of the software:

allowed:

training examples
development examples
decision-time repository evidence
approved tools
approved memory
approved retrieval sources


not allowed:

holdout labels
future outcomes
solution patches
post-decision metadata
hidden evaluator state
future repository revisions

By the end of the book, a program version is meaningful only together with the evidence describing how it was created and evaluated.

What you should be able to do after reading

The goal is not that you memorize every DSPy optimizer or every current API.

DSPy will change.

The more durable goal is that unfamiliar DSPy and LLM-programming systems become legible.

You should be able to open one and ask:

What is the actual program?

What is its task contract?

Which parts are LM behaviour?

Which parts are deterministic software?

What execution strategy is being used?

Which model and provider executed it?

What examples can the program see?

What examples can the optimizer see?

Which cases are genuinely untouched?

What does the metric reward?

Can the metric be gamed?

What can the optimizer change?

What artifact did optimization produce?

What feedback caused a reflective mutation?

What tools are available?

What evidence can retrieval or memory expose?

Could any of those sources leak the answer?

How is a candidate evaluated?

Under what conditions is it compared with the baseline?

Who decides whether it is promoted?

Which version is active?

Can it be rolled back?

What production evidence becomes eligible for the next experiment?

Those questions are more durable than the current name of an optimizer.

They also make debugging much more precise.

Instead of saying:

DSPy did not improve my program.

we can ask whether the failure occurred in the contract, examples, split, metric, model, demonstrations, instruction search, reflection feedback, retrieval, tool use, evaluation protocol, or promotion policy.

That is a much more useful way to engineer language-model systems.

You will also learn when not to optimize

DSPy makes optimization available.

That does not mean every program should be optimized.

If a direct Predict reliably solves the task, adding a multi-stage module may only add latency and failure modes.

If deterministic code can enforce a constraint, searching for an instruction that persuades the model to follow it may be the weaker design.

If the dataset is tiny or noisy, an optimizer may mostly learn accidents.

If the metric does not represent the real objective, better optimization can make the system worse faster.

If a retrieval failure is causing the model to miss the necessary evidence, instruction optimization may solve the wrong problem.

If an existing program already satisfies the real requirement, searching for a numerically better one may have no practical value.

Throughout the book we therefore start from a baseline and add machinery only when a specific failure justifies it.

The aim is not maximum optimization.

It is useful improvement that survives measurement.

Who this book is for

You should be comfortable with Python and with the basic idea of calling a language model.

You do not need prior DSPy experience.

You do not need to know its optimizer APIs in advance.

You also do not need to have built a full agent system.

The examples begin with ordinary Python, explicit data structures, small DSPy programs, and visible evaluation code. More advanced mechanisms arrive only after the pieces they depend upon have been established.

If you have already built production LLM applications, the later chapters are intended to be equally useful.

They focus increasingly on the problems that appear after the first demo works:

How do we know this version is better?

What exactly changed?

Can we reproduce it?

Did the optimizer cheat?

Can the program act safely?

Can it find the right context?

What evidence earns deployment?

What happens when production proves us wrong?

Those are engineering questions rather than prompting questions.

What this book does not try to cover

This is not an encyclopedia of every DSPy class.

It does not attempt to document every optimizer option, adapter, provider, experimental module, deployment integration, or API parameter.

Those details change too quickly to form the conceptual spine of the book.

Instead, the book focuses on the mechanisms underneath them:

contracts
execution
composition
data
measurement
optimization
reflection
tools
context
experimental isolation
versioning
comparison
promotion
evidence

We will use current DSPy APIs to make those mechanisms concrete.

But the aim is that the reasoning survives an API change.

This is also not a claim that every language-model application should become a self-improving system.

Sometimes the right solution is one prompt.

Sometimes it is one Predict.

Sometimes it is ordinary Python.

Sometimes it is a fixed workflow.

The machinery in this book should be earned by a problem it solves.

The promise

By the end of DSPy From First Principles, you should be able to look at an unfamiliar language-model application and work out:

what the program actually is, what behaviour it specifies, what evidence measures that behaviour, what the optimizer is allowed to change, what information it was allowed to see, and what evidence would justify replacing the current version.

You should also be able to ask one final question:

When this system says the new program is better, what exactly makes that claim trustworthy?

Once you can answer that question, DSPy stops being a collection of prompt-optimization techniques.

It becomes something much more useful:

a way to make probabilistic language-model behaviour part of an empirical engineering process.

Contents

Chapters

01

Why Are We Still Hand-Writing Prompts?

Start with a reasonable handwritten prompt, watch it fail three different ways, and frame DSPy as machinery for turning language-model behavior into an experimental variable.

02

A Prompt Is Not Yet a Program

Define the properties a language-model component needs before it deserves to be called a program — and see why those boundaries are prerequisites for controlled comparison.

03

Define the Contract

Treat a DSPy signature as a task contract: choose field names deliberately, decide what belongs inside the boundary, use types where they surface failure, and delete fields nothing consumes.

04

Separate What From How

Hold the task contract completely fixed, change only the execution strategy, and measure what the extra reasoning actually bought across 222 local calls.

05

Build Programs From Programs

Compose several DSPy modules with deterministic Python, then measure which intermediate stages actually influence the result and whether the decomposition earns its cost.

06

Your Model Is a Dependency

Treat the language model as a controlled experimental dependency with explicit identity, honest fallback state, and reproducibility evidence — including why temperature zero does not give you determinism.

07

Examples Are Experimental Data

Build a real evaluation fixture: structured rows with provenance, families rather than isolated cases, splits that respect them, and explicit boundaries around what optimization may consume.

08

You Cannot Optimize What You Cannot Measure

Define the metric, freeze the evaluation protocol, and measure the program — then see why a measurable objective can make the wrong behavior systematically optimizable.

09

When the Metric Becomes the Target

Attack the objective on purpose, separate scoring from diagnosis, and decide what to do once optimization can exploit an instrument you have proved is incomplete.

10

Compile the Program

Treat DSPy compilation as candidate generation under a frozen boundary — then look at what the optimizer actually read, and find that expanding the corpus bought this chapter nothing.

11

Few-Shot Optimization

Compile a few-shot candidate, watch it improve one metric and damage the other, and find out why compiling against the better metric changes nothing at all.

12

Optimize the Instructions

Let MIPROv2 search over instructions and demonstrations, decompose the improvement it finds down to the single case that produced it, and watch a second optimizer break the same sentence as the first.

13

Let the Program Reflect

Give the optimizer language instead of a scalar, watch GEPA propose three mutations and reject all three — then ask whether the rejected candidates were actually worse.

14

Change One Thing

Hold the corpus, program, model, optimizer configurations and evaluation harness fixed, replace only the objective, and test whether optimization changes direction when failure becomes deterministic and nameable.

15

Agents Are Programs Too

Give a program bounded environment access, attack the boundary you built, and measure what happens when safe tools deliver correct evidence to reasoning that still stops one step short.

16

Search, Memory and Long Context

Hold the task, rewrite program, model configuration and metric fixed, vary how supplemental evidence is acquired, and measure six context policies — including adaptive selection and a failed RLM path.

17

Search the Reasoning Space

Derive MCTS over reasoning states from first principles, build it in DSPy, audit whether the tree actually branches, then run the budgeted seven-strategy comparison — with an oracle-leakage control — on two models and separate what it establishes from what it cannot.

18

Don't Let the Optimizer Cheat

Build an experimental firewall for DSPy optimization so tools, memory, feedback, retrieval, and holdout evidence cannot quietly leak answers.

19

From Experiment to Production

Spend the sealed holdout exactly once, then build the ordered governance gate that decides what an offline win is allowed to authorize — independent evaluation, promotion, activation, and rollback as five separate records.

20

Build a Self-Improving Engineering Program

Assemble the book's mechanisms into a repository repair program with contracts, tools, context strategy, validation, optimization, promotion, and production evidence.

21

Beyond the Book: Production DSPy Systems (Appendix)

Appendix: the book's patterns — budgeted reasoning search, a non-regressing memory champion, a parser that outranks its type signature — running in a real production codebase, with the deterministic mechanics reproduced by a committed fixture and the full source linked.