Why Are We Still Hand-Writing Prompts?
Start with a reasonable handwritten prompt, watch it fail three different ways, and frame DSPy as machinery for turning language-model behavior into an experimental variable.
Almost every language-model application starts as one string.
That is not a mistake. A prompt is the fastest way to find out whether a model can do the job at all, and it lets you work directly on the behavior instead of building a framework around a problem you do not yet understand. Most good LM systems begin this way and should.
The trouble is that the prompt usually survives longer than its usefulness. It stops being a probe and becomes the specification, and by the time anyone notices, the string is four hundred words long, nobody remembers why the third paragraph is there, and changing it feels dangerous.
This chapter builds that string. We will write a genuinely reasonable prompt for a real task, run it, and watch it fail in three different ways. The point is not that the prompt is bad. It is that all three failures arrive through the same channel, look the same in the logs, and have nothing in common underneath.
The book begins with editorial rewriting deliberately. “Is this sentence better?” is a real question, but part of the answer is judgment. That makes the task a useful stress test for the distinction the rest of the book will keep making:
Can we program the behavior? yes
Can we measure the behavior? partly
Can we optimize the measurement? yes
Does that prove real improvement? not by itself
Later we will repeat the engineering process on repository repair, where a candidate can be parsed, constrained to an allowed scope, executed, and tested. The two tasks ask one empirical question from opposite sides: what changes when correctness moves from judgment toward externally checkable behavior?
The task is small and concrete: given one sentence from a manuscript, an editorial goal, and the surrounding context, produce a better sentence without changing what it means or how it sounds.
flowchart TD
I[sentence + goal + context] --> P[handwritten prompt]
P --> M[model]
M --> R[candidate rewrite]
That is the shape we start from. Three inputs collapse into one handwritten string, the model runs, and a rewrite comes back. Nothing between the prompt and the output is a separate object you can inspect, version, or change on its own.
We will use one sentence throughout the chapter, and we will make it the first canonical case in the book’s experimental fixture:
Jalen opened the door and then he looked into the room and felt afraid.
The editorial goal is: Sharpen the sentence without changing the event. The decision-time context is: A tense scene in a realist novel. Later evaluation will also require that the rewrite not add a new event and not make the tone comic.
That sentence is doing more work than it looks like. It carries a named character, an event sequence, an emotion, and a tone, and a rewrite can damage any of them independently. It is a useful first case precisely because “better” has several dimensions that can move in opposite directions.
Using the canonical fixture from the beginning also gives us continuity. When later chapters add metrics, optimizers, provenance, and split boundaries, we are not quietly changing the task underneath the experiment. When the task eventually changes to repository repair, that change will be explicit and motivated: it is a move to a different evidence regime, not a quiet substitution of an easier example.
1. A reasonable first prompt
Here is a first version — the kind of thing a competent team ships during exploration:
from dataclasses import dataclass
@dataclass
class RewriteRequest:
sentence: str
goal: str
context: str
def build_prompt(request: RewriteRequest) -> str:
return f"""
You are an expert line editor.
Rewrite the sentence to satisfy the editorial goal.
Preserve meaning, entities, point of view, and the author's voice.
Avoid adding new facts.
Context:
{request.context}
Editorial goal:
{request.goal}
Sentence:
{request.sentence}
Return JSON with:
- rewritten_text: the improved sentence
- rationale: a short explanation of what changed
- confidence: a number from 0.0 to 1.0
""".strip()
This is not a strawman. It assigns a role, supplies context, names its constraints explicitly, and asks for structured output. It is a competent exploration prompt.
That matters. If the chapter began with an obviously bad prompt, every later failure could be dismissed as prompt-writing incompetence. We want to see what remains difficult after the prompt is already reasonable.
Calling it is ordinary Python:
import json
def parse_json_object(raw: str) -> dict:
start = raw.find("{")
end = raw.rfind("}")
if start == -1 or end == -1:
raise ValueError("no JSON object found in response")
return json.loads(raw[start:end + 1])
def improve_sentence(model, request: RewriteRequest) -> dict:
prompt = build_prompt(request)
raw = model(prompt)
return parse_json_object(raw)
Note that parse_json_object already exists. Nobody planned it. The model returns prose that contains JSON rather than JSON, so somebody wrote a function to find the braces, and it has been there ever since.
The shape is easy to hold in your head:
flowchart LR
PO[Python object] --> PS[prompt string]
PS --> LR[LM response]
LR --> JP[JSON parser]
For an experiment, this is fine. As an engineering unit, the prompt is already carrying more than it can hold.
2. Three failures, one string
A prompt like this can fail in several independent ways. The three examples below are representative failure shapes rather than measurements from the later experimental corpus. Their purpose is to separate problems that a single prompt string makes look deceptively similar.
The first is a renamed character.
{
"rewritten_text": "Jason opened the door, afraid of what waited inside.",
"rationale": "Tightened the sentence and moved the fear into a participial phrase.",
"confidence": 0.88
}
The rewrite is better prose. It is also about a different person. The prompt said “preserve entities” and the model agreed in principle and renamed the protagonist anyway. Nothing in the system noticed, because nothing in the system was looking. The confidence field reports 0.88, which tells us what the model felt about a sentence it got wrong.
The second is a broken interface.
{
"rewrite": "Jalen opened the door, afraid of what he would find.",
"why": "Removed the coordinating conjunction.",
"confidence": "high"
}
The rewrite itself is good. The keys are wrong and confidence is a string. parse_json_object succeeds — it found valid JSON — and then the caller reaches for result["rewritten_text"] and raises a KeyError two functions away from the cause.
The usual repair is to patch the parser:
def normalize_response(raw: dict) -> dict:
return {
"rewritten_text": raw.get("rewritten_text") or raw.get("rewrite"),
"rationale": raw.get("rationale") or raw.get("why"),
"confidence": float(raw.get("confidence", 0.0)),
}
That patch works, and it is worth pausing on what just happened. An interface problem was discovered at runtime and repaired outside the artifact that supposedly defines the task. The prompt still says “Return JSON with,” and it is still not a contract. It is a natural-language request embedded in a longer natural-language instruction. The model may comply. The parser may recover. Neither of those is an interface.
Also notice that float("high") raises. The patch fixes one failure and introduces another.
The third failure is the one that should worry you.
{
"rewritten_text": "Jalen hurled the door aside, his heart hammering against the cage of his ribs.",
"rationale": "Heightened the tension and added visceral physical detail.",
"confidence": 0.94
}
The JSON is valid. The keys are right. The name is preserved. The fear is preserved — amplified, even. And it is unusable, because the author does not write like that, and a book where one sentence in forty has been ghostwritten by a different novelist is worse than a book with a flabby sentence in it.
This one carries the highest confidence score of the three.
There is no exception to catch here, no schema violation, no missing key. The only thing that can detect this failure is a judgement about voice, and no such judgement exists anywhere in the system.
Three failures, one channel:
| What we asked for | What came back | Surface symptom | The actual problem |
|---|---|---|---|
Preserve the name Jalen | “Jason opened the door…” | Entity changed | A constraint stated in prose was never independently checked |
| Return these JSON fields | {"rewrite": ..., "why": ...} | Parser fallback needed | The output schema was requested, not enforced |
| Keep the restrained voice | “…hammering against the cage of his ribs” | Fluent and wrong | No definition of voice preservation exists to check against |
These are not three instances of one problem. The first is an entity-constraint violation. The second is an interface violation. The third is a style failure that requires an explicit evaluation criterion before it can even be named.
That distinction will matter later when the fixture grows into an explicit failure taxonomy. A longer prompt can mention every failure mode, but length does not classify failures, and once failures have different causes they need different handles.
The prompt has quietly absorbed at least seven separate concerns:
| Concern | Where it lives in the handwritten version |
|---|---|
| Task definition | prose inside the prompt |
| Input names | implied by the headings |
| Output schema | a prose list near the end |
| Behavioral framing | “You are an expert line editor” |
| Model-specific tuning | wording that happened to work for one provider |
| Evaluation criteria | outside the prompt, if anywhere |
| Version identity | the entire string |
So when you edit one sentence of the prompt, what changed? Maybe the task contract. Maybe only the wording. Maybe how much explanation the model produces. Maybe you improved it for a hosted model and broke it for a local one. The string does not tell you, and it cannot, because all seven concerns are encoded in the same medium.
3. Version history that cannot answer questions
In a small application the prompt version is a filename:
sentence_rewrite_prompt_v4.txt
After a few months of real use, the history usually looks something like this:
v1: basic rewrite
v2: stricter JSON
v3: add examples
v4: fix over-compression
v5: mention voice
v6: local model variant
v7: do not rename characters
That is better than nothing. It is still weak evidence, because each version changed several things at once. The examples moved with the instructions. The output fields moved with the role description. The provider-specific phrasing moved with the task definition. Every version is a bundle.
Now ask a question you will genuinely need to answer: is v7 better than v6?
prompt string changed
↓
behavior changed
↓
unknown cause
You cannot attribute the difference, because you cannot isolate a variable you never separated. And v7 is the interesting case. It was written to stop the renaming, and the natural way to write it is to add a line like “never change character names.” That line may have worked. It may also have made the model more conservative in general, producing rewrites that change less and therefore improve less. Both effects would show up as “v7 feels different.”
This is where prompt engineering has to become something else. The questions that matter are no longer about wording:
- What behavior did we specify?
- Which implementation executed it?
- Which examples influenced it?
- What criterion judged it?
- Which model and configuration produced this output?
- Can we compare this version against the last one?
Those are software questions, and a string cannot answer any of them.
4. Evaluation cannot be an afterthought
Suppose we inspect ten rewrites by hand and conclude that v6 is better. That conclusion may well be right. The evidence is still fragile, because of what we actually evaluated:
prompt v6
whichever model was configured that afternoon
ten sentences somebody picked
a human memory of what v5 used to do
Change any one of those and the conclusion may not survive. Nothing binds them together into a thing that could be re-run.
The obvious next move is to write down a metric — some function that scores a rewrite so the comparison stops depending on memory. That is the right instinct, and most of this book is downstream of it. But it is worth saying early, because it shapes everything that follows: a metric is itself a program, and it can be wrong.
It can be wrong in the ordinary way, by giving noisy scores. It can also be wrong in a more dangerous way, by being confidently blind to the thing you actually care about — scoring the ghostwritten sentence highly because it changed the right number of words, or scoring a rewrite that quietly dropped a qualifier as an improvement because everything it did keep matched the reference closely.
A metric can even encode the wrong preference about whether change is desirable. The expanded fixture later includes cases where restraint is correct and leaving the sentence alone is the right behavior. A scoring rule that rewards change by construction can punish exactly that decision.
None of this remains hypothetical. In Chapter 9, three canonical semantic reversals score between about 0.85 and 0.96 under the original structural metric even though they reverse the intended meaning. The semantic v2 guardrail drives those same attacks down to about 0.29.
Then Chapters 11 and 12 produce an even more useful failure. Two different optimizers independently emit the same development-case rewrite, changing Using the filter regularly can help reduce... to Using the filter helps reduce.... The original metric still scores that candidate at about 0.97; the semantic guardrail caps it at 0.30 because the qualifier was lost.
Whether a small aggregate score delta survives run-to-run noise is a separate question. The candidate-level semantic regression does not depend on that delta.
An optimizer does not know which unmeasured properties you intended to preserve. It searches a space of programs, and the selection rule rewards what the metric can see. If the metric omits something important, optimization can expose the gap.
That gives us three claims that must not be collapsed:
optimization
the candidate scores better on the search objective
generalization
the candidate performs better on unseen cases
improvement
independent evidence says the application is actually better
The first is a fact about search. The second is a fact about a sampling boundary. The third is a claim about the world in which the program is used.
So evaluation cannot be bolted on at the end, and it cannot be trusted merely because it produces numbers. It has to be part of the program’s shape, versioned alongside it, and periodically attacked. We are getting ahead of ourselves — but you should know from the first chapter that the metric is not the safe part of this system.
The book will eventually hold the sentences, program, model, splits, optimizers, and budgets fixed while changing only what success means. That controlled comparison is the hinge between blaming an optimizer and diagnosing its objective.
5. DSPy enters after the problem is visible
DSPy is usually introduced as a framework with signatures, modules, optimizers, and a set of examples. That description is accurate and it buries the point.
The move DSPy makes is to stop treating the prompt as the artifact. Instead you declare what the task is — its inputs, its outputs, and the semantic job connecting them — and separately choose how that task gets attempted. DSPy then constructs the prompt.
For our sentence task:
import dspy
class RewriteSentence(dspy.Signature):
"""Rewrite one sentence to satisfy an editorial goal
while preserving meaning, entities, and voice."""
sentence: str = dspy.InputField(desc="The exact sentence to rewrite")
goal: str = dspy.InputField(desc="The local reason this sentence is being edited")
context: str = dspy.InputField(desc="Nearby prose needed to preserve continuity and voice")
rewritten_text: str = dspy.OutputField(desc="A replacement sentence, not a paragraph")
rationale: str = dspy.OutputField(desc="Short explanation of the edit")
rewrite = dspy.Predict(RewriteSentence)
result = rewrite(
sentence="Jalen opened the door and then he looked into the room and felt afraid.",
goal="Sharpen the sentence without changing the event.",
context="A tense scene in a realist novel.",
)
print(result.rewritten_text)
This assumes an LM has been configured. Every measured result in this book was produced on a local model, so that is what we will configure:
lm = dspy.LM(
"ollama_chat/qwen3:latest",
api_base="http://127.0.0.1:11434",
temperature=0.0,
max_tokens=256,
cache=False,
think=False,
)
dspy.configure(lm=lm)
That is a deliberate choice and worth stating plainly. The canonical dependency is a local Qwen3 8.2B model, quantized as Q4_K_M, with native thinking disabled. It is not the strongest model available, and several results would probably look different on a frontier model.
The advantage is experimental access: LM-backed runs cost no API money, can be repeated enough to study failures, and keep the full dependency under our control.
But temperature=0.0 is not a promise of bit-for-bit determinism. Repeated sessions later in the book show that some inputs still admit competing outputs. That means tiny score deltas must be treated cautiously, run identity must be recorded, and important comparisons should be repeated rather than trusted from one execution.
The same discipline applies in the other direction: the final holdout is intentionally not regenerated over and over just because the model is local. Cheap execution does not remove the need for an evidence boundary.
RewriteSentence establishes the core contract the rest of the book keeps. The decision-time task fields are sentence, goal, and context, and the primary outputs are rewritten_text and rationale.
When we decompose the program in Chapter 5, the rewrite stage gains two intermediate inputs — issue_summary and preservation_notes — produced by an earlier analysis stage. That extends the implementation without renaming the core task fields.
The same contract is then measured in Chapter 8, attacked in Chapter 9, and optimized under a frozen protocol. Keeping the core field names stable is what makes those comparisons meaningful. The later repair regime uses a different task contract but preserves the experimental discipline.
The important shift is not that the code is shorter. It is that the task is now inspectable:
Signature
inputs: sentence, goal, context
outputs: rewritten_text, rationale
Module:
Predict
LM:
a configured dependency, not an assumption
Each of those is a separate thing you can change, log, version, and compare. The confidence field is gone, incidentally — it was never measuring anything. A model’s self-reported confidence in a rewrite it got wrong was 0.88, and in a rewrite it got badly wrong was 0.94.
Beginning in Chapter 8, confidence is replaced by something external to the candidate: explicit evaluation evidence. That evidence will turn out to have its own failure modes, but at least it can be inspected, versioned, attacked, and compared.
The prompt has not disappeared. DSPy still constructs one, and you can print it. But it is no longer the primary abstraction anyone edits by hand.
6. What this does not solve
It would be convenient to end the chapter here, with structure defeating strings. That is not what happens, and pretending otherwise would make the next seventeen chapters dishonest.
Declaring a signature does not tell you whether your field boundaries are right. It does not tell you whether a reasoning step is worth its cost. In Chapter 4, the expanded fixture shows that ChainOfThought costs more but also avoids a meaning-changing deadline slip made by plain Predict, so the answer is not simply “reasoning is wasted tokens.”
It does not tell you whether decomposing the program helps either. Chapter 5 separates composition from evidence: the larger ablation shows that real intermediate analysis helps relative to blank or shuffled analysis, while the surrounding program still has to earn its additional calls and latency.
And structure certainly does not tell you whether the program is good. That requires a fixture, an evaluation protocol, and a metric — and later chapters show that all three can fail in ways that the program itself cannot diagnose.
What structure buys is the ability to ask precise questions. That is all, and it is a great deal, because the handwritten prompt could not be asked anything at all.
What Usually Goes Wrong
| Symptom | Likely cause | How to diagnose it | What to change |
|---|---|---|---|
| The model returns plausible prose but the parser fails | Output structure is requested in prose, not enforced | Log raw responses and count invalid parses | Make outputs explicit fields and validate them |
| A prompt edit improves one model and breaks another | Provider-specific wording is fused with the task definition | Run the same inputs against both models | Separate the task contract from LM configuration |
| Nobody can say whether v7 beat v6 | Prompt versions are not tied to a frozen dataset or metric | Look for saved examples, a scoring function, and run records | Make evaluation part of the program’s lifecycle |
| The prompt keeps growing | Every failure is patched by appending another sentence | Group past failures by cause: interface, semantics, reasoning, model, data | Move structure into code and reserve prose for semantics |
| Output looks fine and is subtly wrong | The failure mode has no detector | Ask what would have caught this automatically; usually nothing would | Define the criterion explicitly, then test the criterion |
| The model edits a sentence that was already fine | “Improve this” has no way to express restraint | Check whether leaving the input unchanged is ever the correct output | Make the do-nothing case a first-class expected result |
Conclusion
The handwritten prompt was a good discovery tool and a poor engineering unit.
Its real defect was not that the writing was bad. It was that one artifact carried the task contract, the execution tactic, the output schema, the provider assumptions, and the version identity all at once, in a medium that cannot distinguish between them. That is why the three failures in section 2 arrived looking identical and needed three different repairs.
We have removed one assumption: that improving a language-model application means editing a prompt until it feels better.
We have also named the book’s larger problem. Turning behavior into an experimental variable is easier than deciding what an improvement claim is allowed to mean. Editorial rewriting will expose that gap; repository repair will let us test the same machinery against stronger external evidence.
What we have not done is design the replacement. We have a signature with five fields, and no account of why those five, why those names, or where the boundary between an input and a piece of context should fall. A weak contract makes every later stage worse — the metric measures the wrong thing, the optimizer optimizes the wrong thing, and the failures get harder to see rather than easier.
So the next question is the obvious one:
What properties does a language-model component need before it deserves to be called a program?