Your Model Is a Dependency
Treat the language model as a controlled experimental dependency with explicit identity, honest fallback state, and reproducibility evidence — including why temperature zero does not give you determinism.
Chapter 5 built a composed program and measured which of its stages earned their cost. Every number in that chapter, and in chapter 4 before it, has an unstated qualifier:
Which model executed the program?
That is not operational trivia. A language-model program is not fully described by its signatures and modules. Its behavior depends on what sits behind the LM boundary and how that thing is configured.
The model is therefore both a dependency and an experimental variable. Holding it fixed is often as important as swapping it:
model fixed + objective changed
→ interpretable objective experiment
model changed + objective changed
→ mixed intervention
The controlled optimization experiments later in the book retain the same local Qwen3 dependency while varying program state or the objective. Model comparison is a different experiment and deserves a separate protocol.
flowchart TD
P[DSPy program] --> B[LM boundary]
B --> L[local model]
B --> H[hosted model]
The signatures and modules stop at the boundary; behind it sits a model whose provider, identity, and configuration are all experimental variables. This chapter makes the dependency explicit: configured, fingerprinted, tested for failure, and — in the section that surprised us most — not as deterministic as temperature=0 tempts us to assume.
1. Configure the LM explicitly
import dspy
lm = dspy.LM(
"ollama_chat/qwen3:latest",
api_base="http://127.0.0.1:11434",
temperature=0.0,
max_tokens=256,
cache=False,
think=False,
)
dspy.configure(lm=lm)
That is the canonical dependency for every measured experiment in this book: Qwen3 8.2B, Q4_K_M, served locally by Ollama through the DSPy/LiteLLM route. Native Qwen thinking is disabled so that model-level reasoning does not blur the DSPy execution strategies chapter 4 compared.
Seven settings, and each one is part of the experimental condition:
model identifier
provider route
base URL
temperature
max tokens
native thinking mode
cache behavior
Two runs both described as “Qwen3” can differ in every one of those and produce different results. The friendly model name is not the dependency. The configuration is.
2. Temperature zero is not determinism
Here is the finding that changed how the rest of this book reports its numbers.
An early baseline audit evaluated the same unchanged program on the same eleven development cases in five separate sessions, at temperature 0, with caching disabled and the recorded configuration held fixed.
Nine of the eleven cases produced the same output in every one of those sessions. Two did not.
Case ed-033 — a marketing sentence about a coffee roaster — exposed the clearest split. One run produced a rewrite scoring 0.927; another produced a plausible alternative scoring 0.818. The program state and declared LM configuration had not changed, yet the output had.
That is the important observation. The model string, temperature, prompt contract, and input fingerprint are necessary provenance, but they are not a complete guarantee of repeated output.
The consequence for the aggregate in those baseline runs:
same program, same cases, same recorded config, five sessions
↓
dev-family mean observed from 0.7893 to 0.7992
↓
observed end-to-end span: about 0.010
That span has teeth, but it is not a universal statistical “noise floor.” Five sessions are too few to estimate one, and later interleaved experiments expose an additional order-dependent source of variation.
The defensible rule is narrower: a small delta from one execution schedule is not self-interpreting on this setup. Chapter 4’s v1 chain-of-thought delta is +0.014; Chapter 11’s initial few-shot delta is +0.011. Both are on the same order as variation already observed from unchanged program state, so neither should be treated as a stable effect without paired repetition and an appropriate negative control.
We are not going to pretend this experiment isolates the cause. Plausible contributors include numerical nondeterminism in model-serving kernels, request scheduling and batching, process or cache state, and small logit perturbations introduced by quantization or other serving details.
Those are hypotheses, not findings from this experiment. What we did observe is consistent with a system sitting near alternative continuations: small execution-state differences can become visible when several next-token paths are competitive. But we did not run the controlled precision, batching, cache, or kernel experiments required to attribute the behavior to any one mechanism.
The first block-style repeats made the behavior look deterministic within a warm session: repeated calls in the same sequence returned the same strings. That turned out to be an artifact of the execution schedule, not a guarantee.
The later paired experiment supplied a stronger negative control. Two Chapter 11 candidate files were byte-identical, yet when several programs were interleaved in one session they produced different aggregate scores (0.8005 versus 0.7972). Because the program artifacts were identical, that difference cannot be a candidate effect. It is measurement variation induced somewhere in execution order or runtime state.
So the variability has at least two observable forms:
- cross-session drift — the same program can land on different outputs after a fresh process or model state;
- within-session order effects — interleaving otherwise identical program artifacts can also move an output.
A naive design can miss both. Back-to-back repetitions may look perfectly stable precisely because they repeat the same call history.
The stronger design is paired and repeated across fresh processes:
session 1: baseline, candidate A, candidate B
session 2: candidate A, candidate B, baseline
session 3: candidate B, baseline, candidate A
...
Rotate order by case and session so that runtime state is not systematically assigned to one condition. Compute the delta per case, per session and aggregate those paired differences rather than comparing two independently collected means.
When possible, include a negative control whose true delta is known to be zero. The byte-identical Chapter 11 candidates give us exactly that: any measured difference between them is variation in the measurement process, not program improvement.
Then report the distribution of paired deltas — mean, median, spread, and sign counts — rather than a single point estimate.
This is the instrument the optimizer chapters depend on. Without it, a small candidate delta can be indistinguishable from execution-state variation.
3. What should change when the model changes?
The task contract should stay stable:
sentence, goal, context → rewritten_text, rationale
That is the core task-facing contract. The composed implementation also emits diagnostic fields such as risk and assessment, but Chapter 5 showed that those are observability fields unless some downstream policy consumes them.
The implementation may need to adapt:
| Model change | Likely effect |
|---|---|
| Hosted to local | Different latency profile, different reliability, different output formatting |
| Larger to smaller | More schema failures, weaker constraint following |
| Temperature raised | More variety, less reproducibility |
| Provider change | Different timeouts, errors, rate limits, structured-output behavior |
| Context limit change | Different truncation, silent input loss |
| Quantization change | Different logits or output behavior; remeasure rather than assuming equivalence |
The wrong response is to quietly rewrite the signature for each model until the output looks acceptable. That is chapter 1’s failure mode wearing a different hat — several things changing at once, with no way to attribute the difference.
same program
same cases
different LM dependency
↓
behavior comparison
Hold the contract steady, observe which failure mode actually appears, and only then decide whether the contract, the module, the validation, or the model is the thing to change.
4. Silent fallback is dangerous
When the provider fails, something has to happen. The tempting shape:
try:
return program(**inputs)
except Exception:
return {"rewritten_text": sentence, "rationale": "unchanged"}
That keeps the interface alive and corrupts every downstream claim. The record now says a rewrite was produced. It does not say that no model was involved, so an evaluation run over these records will happily score the fallback as model output — and because the fallback returns the original sentence, it will score around 0.65, the structural floor. Your program appears to have had a mediocre day rather than an outage.
The better shape separates availability from evidence:
from dataclasses import dataclass
@dataclass
class ProgramRunResult:
output: dict
provider: str
model: str
used_real_lm: bool
eligible_as_model_evidence: bool
fallback_state: str | None
error: str | None
def run_with_explicit_fallback(program, inputs, model_name: str) -> ProgramRunResult:
try:
pred = program(**inputs)
return ProgramRunResult(
output={
"rewritten_text": pred.rewritten_text,
"rationale": pred.rationale,
"risk": pred.risk,
"assessment": pred.assessment,
},
provider="dspy",
model=model_name,
used_real_lm=True,
eligible_as_model_evidence=True,
fallback_state=None,
error=None,
)
except Exception as exc:
return ProgramRunResult(
output={
"rewritten_text": inputs["sentence"],
"rationale": "Fallback preserved the original sentence.",
"risk": None,
"assessment": None,
},
provider="deterministic_fallback",
model=model_name,
used_real_lm=False,
eligible_as_model_evidence=False,
fallback_state="lm_unavailable",
error=str(exc),
)
We tested this by injecting a dependency failure. The wrapper raised RuntimeError: injected LM dependency failure, returned degraded output preserving the original sentence, and recorded used_real_lm = false, eligible_as_model_evidence = false, fallback_state = "lm_unavailable".
The invariant that matters: output existed, and no claim was made that the model produced it.
real LM response
↓
eligible as model evidence
deterministic fallback
↓
useful degraded behavior
↓
NOT evidence that the LM performed the task
Availability and evidence are different concerns, and a system that conflates them will eventually report an outage as a quality regression.
5. Reproducibility needs model identity
If a run matters, store enough to explain it later:
program id and version
signature version
module strategy
LM provider, model name, and model version
temperature, max tokens, timeout and retry policy
input fingerprint
output fingerprint
fallback state
import hashlib
import json
from datetime import datetime, timezone
def fingerprint(payload: dict) -> str:
blob = json.dumps(payload, sort_keys=True, default=str)
return hashlib.sha256(blob.encode("utf-8")).hexdigest()
def make_run_record(inputs: dict, output: ProgramRunResult) -> dict:
return {
"created_at": datetime.now(timezone.utc).isoformat(),
"program": "editorial_rewrite_program",
"program_version": "0.1",
"lm_provider": output.provider,
"lm_model": output.model,
"used_real_lm": output.used_real_lm,
"eligible_as_model_evidence": output.eligible_as_model_evidence,
"fallback_state": output.fallback_state,
"input_fingerprint": fingerprint(inputs),
"output_fingerprint": fingerprint(output.output),
"error": output.error,
}
We ran one ed-001 execution under the canonical dependency. The composed program completed in 5.78 seconds and returned:
Jalen opened the door, peered into the room, and felt a wave of fear wash over him.
Then we mutated one configuration field at a time, without making another model call:
| Mutation | Configuration fingerprint changed? |
|---|---|
| model | Yes |
| endpoint | Yes |
| temperature | Yes |
| native thinking | Yes |
| max tokens | Yes |
Every single-field change produced a different fingerprint. As a hashing result that is unsurprising. As a discipline it is the point: the fingerprint is what lets you say two runs were the same experiment, and “we used Qwen3 both times” does not.
One thing worth noticing about that rewrite, because it separates three claims this experiment is often assumed to establish. The dependency worked: the call completed against the configured provider. The run record captured what executed and what came back.
Whether the rewrite is good is a different claim. peered into the room is not a new event — the source already says Jalen looked into the room — but felt a wave of fear wash over him is longer and more embellished than the restrained source. This experiment did not run the editorial acceptance protocol needed to decide whether that trade is acceptable.
So the dependency experiment establishes execution and provenance, not quality. That distinction is the whole chapter in miniature. A healthy dependency and a complete run record are entirely compatible with an output that later evaluation may reject.
6. One program, several model roles
A DSPy program does not imply one global model dependency. Different responsibilities can, and often should, use different models:
| Role | Responsibility | Should it differ from the task model? |
|---|---|---|
| task model | produce the result | — |
| teacher | generate traces for bootstrapping | Often |
| prompt model | propose candidate instructions | Often |
| reflection LM | reflect on feedback and propose mutations | Often |
| judge | evaluate a candidate | Ideally, yes |
| router | choose a cheaper or stronger path | Yes |
STORM is a clean reference for this: its modular article-generation workflow assigns different LMs and retrieval modules to different responsibilities. CodeSpy shows the runtime version — scoped search, PR summary, specialist review, and audit each carry their own dependency profile and cost.
The provenance rule follows directly. If the judge model changes, the score changed even though the candidate did not. If the teacher changes, the selected demonstrations may change even though the student did not. If the prompt model changes, the search path changed. Record role, provider, model, adapter, and configuration together, or you will not be able to say which of them moved.
An admission about this book
The judge that provides metric v2’s semantic component runs on qwen3:latest — the same model, at the same endpoint, as the task program it evaluates. So does the reflection LM in chapter 13’s GEPA run.
That is the exact role collapse the table above warns about, and it is a real limitation of this book’s evidence. A judge built from the same model family can have correlated blind spots with the task program. Different prompting and a different role can still catch errors — Chapter 9 demonstrates that — but shared weights are not independent evidence in the strong sense.
Two things partially defend the design. First, Chapter 9 validates the judge against an adversarial suite rather than assuming that a typed verdict is correct: it catches all three canonical semantic reversals and produces zero false positives across 44 clean reference rewrites. Second, using one local dependency makes the whole experiment rerunnable without API cost or provider access.
Neither turns the judge into an independent authority. A separately validated judge built from a meaningfully different model would provide stronger evidence. Where the current judge’s verdicts matter most — ed-025 in Chapter 4 and ed-035 in Chapters 11 and 12 — we therefore print the sentences and the named constraint so the reader can inspect the disputed semantic change directly.
7. Local and hosted are engineering choices
Neither is more serious than the other. They expose different constraints.
| Choice | Advantages | Costs |
|---|---|---|
| Local (Ollama) | Control, privacy, offline work, no per-call cost, cheap repetition | Machine setup and serving state become part of the experiment; quantization and runtime details must be recorded |
| Hosted API | Access to stronger models, managed infrastructure, broad provider support | Cost, rate limits, external dependency, and provider-side version or serving changes you may not control |
| OpenAI-compatible gateway | Routing flexibility, centralized policy and logging | Another layer to observe and another place to fail |
The choice this book made was local, and Section 2 is part of the bill. A stronger hosted model might produce better rewrites or fewer hard-gate failures, but that is an empirical question rather than a property we can assume from model class alone.
Hosted execution would also attach real API cost to the 222 calls of Chapter 4 and the much larger optimizer runs later in the book, and some providers can change serving behavior or model aliases outside your experiment. Local and hosted systems therefore have different reproducibility envelopes; neither removes the need to record the dependency actually used.
DSPy helps because the program can often survive a change of dependency. “Often” is doing work in that sentence. You still have to run the same program under the target model and look at what breaks.
What Usually Goes Wrong
| Symptom | Likely cause | How to diagnose it | What to change |
|---|---|---|---|
| Every output is empty | Provider call failing, or wrong model identifier | Health-check the provider directly, outside DSPy | Fix the model string or base URL |
| Outputs ignore the declared fields | The model or adapter struggles with structured output | Log the raw response and any parse warnings | Simplify the output contract or change model |
| Direct prediction works, the composed program does not | Context or token budget exceeded | Compare stage-by-stage raw outputs and token counts | Shorten context, raise max_tokens, or use fewer stages |
| Results vary between runs at temperature 0 | Cross-session nondeterminism, not a bug | Repeat in fresh processes and measure the spread | Establish a noise floor before quoting any delta |
| A quality regression appears overnight | An outage is being scored as model output | Count records where used_real_lm is false | Exclude fallback from quality claims |
| Two “identical” runs disagree | The configuration differed in a field nobody recorded | Compare configuration fingerprints, not model names | Fingerprint the whole dependency |
| The judge’s score moved and the program did not | The evaluation dependency changed | Check the judge’s recorded model and configuration | Version the judge as carefully as the program |
Conclusion
The model is now a dependency rather than an assumption. It has an identity you can hash, a failure mode you can distinguish from a bad result, and a role that has to be recorded alongside the program’s.
The measurements backed that up in three ways. Every single-field configuration change produced a different fingerprint, which is what makes “the same experiment” a checkable claim. An injected dependency failure produced usable degraded output that was explicitly marked ineligible as model evidence. And repeated execution of unchanged program state at temperature zero produced different outputs and an observed dev-mean span of about 0.010 in the first baseline audit.
The later paired work makes the lesson stronger rather than cleaner. Variation can appear across fresh sessions and under different interleavings inside one session. Temperature zero therefore removes one intentional source of sampling randomness; it does not make the whole inference stack deterministic.
We do not turn those observations into a universal scalar noise floor. From here on, small improvements are treated as hypotheses that need paired repetition, per-case inspection, and—where available—a zero-effect negative control. That standard matters more than any single threshold.
We removed two assumptions: that a DSPy program is fully described by its signatures and modules, and that a returned value proves a model produced it.
What remains is improvement. We can specify a program, execute it, and record exactly what executed it. We still have no principled way to teach it from examples, and no way at all to say whether one version is better than another — because we have no cases we trust and no criterion worth optimizing.
Both of those are data problems.
Examples are data, not decoration.