Intelligence in the Wrong Direction
Capability is a scalar; you need a vector. The deterministic measurement playbook does not transfer, evals are experiments, and the critic model that scores your work has to be validated, frozen, and never optimized against. Part 1 turns from argument to something to build.
Part 1 — Where You Stand
Capability without an objective is magnitude without direction
The thing the marketing does not cover
Every vendor will sell you more capability. None of them can tell you whether that capability moves your process toward its objective without a measurement you supply.
A stronger model can execute the wrong objective more effectively just as it can execute the right one more effectively. Capability alone does not tell you whether the gap between what you wanted and what you got narrowed or widened. That is an empirical question.
Capability gives you magnitude. Measurement tells you whether it points where you want.
Chapter 6 ended on a purchasing question — what are you buying and does it ever stop. This chapter is the prior question, and it is the one that determines whether any of that spending was worth it: how would you know?
How do you measure whether a stochastic process is getting better at something you cannot define precisely?
What Part 1 has actually been about
The last three chapters have been circling one problem from three directions.
Chapter 4 said that without a cheap verifier you have a demo rather than a process. Chapter 5 said you cannot tell when you have become a rubber stamp. Chapter 6 said you cannot tell whether the expensive model earned its price.
Those are not three problems. They are one problem wearing three coats: you cannot see the quality of what you are producing. Verification, review discipline, and cost control all bottom out in measurement, and if measurement is absent, all three degrade silently and simultaneously.
That is why this chapter belongs in Part 1 rather than in a technical appendix. Everything before it assumed you would be able to tell whether it worked.
Why operational metrics are not enough
You already know how to measure software operations: uptime, latency percentiles, error rate, throughput, test pass rate. Those metrics still matter for an AI system. What they do not tell you by themselves is whether a semantically plausible output was good.
| Operational question | Output-quality question |
|---|---|
| Did the request complete? | Was the result acceptable for the task? |
| One execution establishes what happened on that execution | One generation is one observation from the process |
| Exceptions and status codes expose many execution failures | A semantic failure can return 200 OK and fluent prose |
| Latency and throughput are directly observable | Quality needs a criterion, check, reference outcome, or judgment |
| A deterministic check can establish the property it implements | An open-ended quality claim has to be operationalized before it can be estimated |
| Re-running a deterministic function against the same relevant state should reproduce its result | Re-running a model call may legitimately produce a different candidate |
The hard row is the one about definition. A 99.9% uptime target is measurable because down is defined. “The review was good” is not a measurement until you specify what evidence would count for or against it. Defining that criterion is part of measurement, not paperwork before measurement.
And semantic degradation usually does not announce itself through the operational layer. A process can move from 94% to 88% acceptable output while HTTP success, latency, and throughput remain normal. Unless quality is measured separately, nothing has to page you — which is Chapter 5’s audit problem arriving through a different door.
Evals are experiments
The most useful reframe available here is Evan Miller’s: evaluations are experiments, and the literature on evaluation has largely ignored a century of work on how to analyze and plan experiments (Miller, 2024). Treating an eval as a number you print, rather than as an estimate with uncertainty, is the root of most bad decisions in this area.
His recommendations are concrete, and they are the methodology this book uses:
- Report standard errors on eval scores, computed from the Central Limit Theorem, treating the questions as drawn from an unseen super-population.
- Use clustered standard errors when questions come in related groups — several questions about one document, several tasks from one repository — because those questions are not independent draws.
- Reduce variance by resampling. Running each question K times reduces the conditional variance proportionally. Where you have access to next-token probabilities, using them removes sampling variance entirely.
- When comparing two models, do inference on question-level paired differences, not on the two population-level summary statistics. The models are correlated across questions, and exploiting that correlation shrinks the standard error substantially.
- Run a power analysis to find out whether your eval is even capable of detecting the effect you care about.
Number five is the one that should change behavior this week. The standard industry practice is to compare two models on twenty examples, observe that one won fourteen to six, and adopt it. Two equally good models split at least that lopsidedly about one time in nine, and a power analysis would have shown in advance that twenty examples cannot reliably detect the kind of difference usually being claimed. An enormous amount of confident model-selection folklore is built on evals that had no power to support it.
The arithmetic fits in a few lines, and it is worth running once by hand:
from math import comb, sqrt
# Two models, 20 head-to-head items, one wins 14. If they were equally good,
# how often would a split at least this lopsided happen?
n, wins = 20, 14
p_two_sided = 2 * sum(comb(n, k) for k in range(wins, n + 1)) / 2**n # ≈ 0.115
# A pass rate is an estimate, not a fact: 40 passes out of 50 runs.
rate, runs = 40 / 50, 50
standard_error = sqrt(rate * (1 - rate) / runs) # ≈ 0.057
A pass rate of 80% with a standard error near six percentage points has substantial uncertainty; under a simple normal approximation, about two standard errors gives a rough 69% to 91% interval. But overlap between two separately computed intervals is not the test of whether two models differ. When the same tasks are run through both models, analyze the paired task-level differences — and cluster them where the design requires it. (The fourteen-six calculation above assumes independent head-to-head items.)
And number three is the direct consequence of Chapter 3. If temperature=0 does not give you caller-visible determinism, one generation is one observation from the process, not an estimate of its expected performance. Repeated runs are what let you measure the run-to-run component of variation.
The ladder of what you can measure
Work down this list from the cheapest evidence that can actually answer your question. The ordering is a cost heuristic, not a claim that every higher row is stronger than every lower one.
| Level | What it is | Cost | Weakness |
|---|---|---|---|
| 1. Mechanical checks | Does it parse? Do cited sources resolve? Do quoted spans exist in the document? Does it compile? | Low, repeatable | Exact only about the property implemented |
| 2. Frozen task set | N fixed tasks with reference outcomes or criteria; score with uncertainty | Moderate, repeatable | Corpus drifts from production |
| 3. Downstream signal | Was the edit accepted? Did the ticket reopen? Was the commit reverted? | Often low incremental cost | Delayed, confounded, sometimes gameable |
| 4. Human pairwise judgment | Is A better than B, on a blinded sample? | Expensive, slow | Judgment noise and disagreement depend on task and protocol |
| 5. Critic model | A model scores the output | Cheap at scale | A model-produced proxy with its own biases and errors |
Two things about this ordering.
Level 1 should be exhausted for properties that really are mechanical, because it gives repeatable evidence without spending model or human judgment. Level 3 is valuable because it observes what happened downstream, but “real-world” does not mean “ground truth”: acceptance, reopening, and reversion can all be confounded by factors other than output quality.
Level 5 is therefore a scalable supplement, not the automatic destination of the ladder. A critic is useful only after you know what it is standing in for and how well it tracks that reference.
The critic model, engineered properly
A critic model can be a useful scalable instrument when mechanical checks and downstream outcomes do not cover the property you need to measure. It is not independent evidence merely because it is a separate call; it is another model-produced judgment whose relationship to the target has to be established.
Five constraints, each with a reason.
Separate it from the generator. LLM judges favor their own outputs, and the strength of that preference scales with the model’s ability to recognize its own text (Panickssery et al., 2024). Using a different model family removes the most direct same-model self-evaluation path, but it does not guarantee independent errors. Panels can help: Verga and colleagues found that a jury of smaller models from disjoint families outperformed a single large judge with less intra-model bias at over seven times lower cost in their evaluated settings (Verga et al., 2024). Whether diversity actually buys complementary error coverage is something to measure, and Part 5 does.
Validate it against a reference you trust, and carry that validation with its scores. Where known outcomes exist, use them. Where the target is judgment, build a blinded human-labeled or adjudicated reference set whose size and sampling match the decision you intend to make. Measure the critic’s agreement on that set, with its denominator and uncertainty. A critic whose relationship to its reference has not been measured is still a generator with a numeric output format.
PaperBench shows the industrial form of this discipline: replicating 20 research papers is graded against rubrics broken into 8,316 individually gradable tasks, co-developed with the papers’ own authors, and the automated judge is assessed on a separate benchmark of its own (Starace et al., 2025). Decompose the judgment until the authors of the ground truth would recognize it, then validate the judge separately.
Zheng and colleagues found GPT-4 reached roughly 85% agreement with human preferences on open-ended questions, while humans agreed with each other about 81% of the time (Zheng et al., 2023). Those two agreement rates do not create a hard ceiling. Pairwise human agreement measures disagreement in the labeling process; a judge can agree with an aggregate or majority reference more often than two individual humans agree with each other.
The useful lesson is to report the reference process itself: who labeled, how disagreement was resolved, how much disagreement remained, and how the critic performed against that reference. Near-perfect agreement may mean the task is narrow and objective; it may mean the set is too easy; it may mean leakage or a flawed protocol. The number alone cannot tell you which.
It must be frozen and versioned. This is the practical failure that ruins real measurement programs, and it gets almost no attention. You want a time series — is the process getting better over months? A critic that silently changes underneath you destroys that series. Providers update model snapshots. Endpoints get deprecated. A prompt tweak someone made on a Thursday shifts every score after it.
So pin the critic: model, version, prompt, parameters, all of it, recorded with each score. When you must change it, re-score the historical archive with the new critic to establish the offset before comparing across the boundary.
Which means you must have kept the archive. This is the point where Part 3 stops being administrative hygiene: re-scoring requires the preserved outputs (Chapters 11 and 17), and rebuilding the time series requires the log of what was scored, when, and by which critic (Chapters 16 and 18).
Never optimize against your only score. Gao, Schulman, and Hilton studied this in a controlled proxy setting, using a fixed “gold” reward model to stand in for human judgment and a proxy reward model trained from its labels. As optimization pressure on the proxy increased — through RL or best-of-n sampling — performance under the fixed gold model first rose and then turned and degraded (Gao et al., 2023). The “gold” model was itself a proxy for human preference, so the experiment establishes overoptimization against one learned proxy relative to another fixed reference, not access to literal true quality.
The practical consequence is to keep a held-out reference evaluation that is never used to select or tune the system. A held-out critic can be one such reference, but it remains another proxy; human-labeled items or downstream outcomes can provide a different kind of check. Divergence between the optimized score and the held-out reference is evidence that the proxy relationship has changed. It does not, by itself, locate a universal Goodhart threshold.
The shape is worth seeing once, because the trap is in the divergence:

Schematic of the overoptimization turn measured in Gao, Schulman, and Hilton’s proxy-versus-fixed-reference setup. The curves shown here are illustrative rather than the paper’s measured curves.
The critic must not see provenance that can bias the comparison. Strip model and arm identity. In pairwise evaluation, randomize candidate order. If verbosity or length may affect judgments, measure that effect or control it in the evaluation design rather than silently rewriting the artifacts being scored. This is the same isolation discipline Part 5 applies to competing generations, applied to the thing doing the judging.
flowchart LR
G["generator<br/><i>the work</i>"] --> O["output<br/><i>preserved, blinded</i>"]
O --> M1["1 · mechanical checks"]
M1 --> C["critic panel<br/><i>pinned version</i>"]
C --> S["score + critic id<br/>+ agreement figure"]
GOLD["gold set<br/><i>human-labeled</i>"] -.->|"validates"| C
HELD["held-out critic<br/><i>never optimized against</i>"] -.->|"detects the turn"| S
style HELD stroke-dasharray: 4 4
Build this before you build anything else
Part 1 has been argument. Here is the part that is not, and it is deliberately the first thing you make.
Build the measurement before you build the feature. Not because it is virtuous, but because of a simple asymmetry: if the harness exists first, every subsequent change in this book is testable the day you make it. If it exists last, you have a year of undocumented decisions that the harness can no longer test.
The smallest useful version can fit in an afternoon, as long as you do not mistake the seed harness for a powered experiment:
Do this now.
- Freeze a first set of real tasks from your actual domain. Thirty to fifty is enough to make the machinery concrete; it is not a universally adequate sample size. Preserve the inputs, and where a reference outcome exists, preserve that too. For open-ended tasks, record the acceptance criteria and how they will be judged.
- Write the mechanical checks that genuinely apply — output parses, cited sources resolve, quoted spans appear verbatim, numbers sum. They are usually cheap and repeatable, and they establish only the properties they test.
- Record a baseline. Today’s model, today’s prompt, and item-level outcomes. For stochastic operations, repeat enough runs to expose meaningful run-to-run variation; three draws can be a useful smoke test, not a general statistical rule. Report the denominator and uncertainty appropriate to the design. If you compare models, run them on the same tasks so the comparison can be paired.
- Write down the date, model identity, model version where available, parameters, and prompt hash. Attach them to every score, or later comparisons drift.
- Store the raw item-level results where a later process can read them — a directory of JSON is entirely sufficient today.
That artifact is the reference point for the rest of this book. Chapter 11 will add the instrumentation that feeds it. Part 5 turns the same discipline into preregistered comparisons with frozen tasks, preserved draws, explicit denominators, and promotion rules — including a result that did not replicate, which is exactly the outcome a harness exists to preserve.
If you build nothing else from Part 1, build this.
Where this gets hard
- Frozen sets rot. Your corpus stops representing production as production drifts, and the score stays flat while quality falls. Refresh on a schedule, and keep the old set alongside the new one so the comparison survives.
- Small corpora cannot detect small effects. Power analysis will tell you your eval is underpowered. The honest response is to stop claiming small wins — not to reach for a bigger model.
- The things that matter most are often unmeasurable. “Did this book get better” lies outside any frozen task set. For those, measure the process rather than pretending to measure the output: was it reviewed, against what criteria, by whom, with what evidence attached. A recorded process is a weaker claim than a measured outcome and an enormously stronger one than a feeling.
- Measurement costs money from the same budget as the work. Chapter 6’s arithmetic applies to critics, human labels, reruns, and deterministic checks too. Start with the lowest-cost evidence that actually covers the property; “cheap” is not the same as free.
- Optimization can erode a proxy. Gao, Schulman, and Hilton show one measured form of this under increasing optimization pressure. Do not promote that result into a law that every optimized metric must immediately degrade. Held-out references and downstream signals help reveal divergence; they do not make the optimized metric ground truth.
- A critic is a model, and everything in this book about models applies to it. It is stochastic, it has a jagged frontier, and it will be confidently wrong in its own characteristic ways.
Failure modes
- Buying capability without direction. More capability does not tell you whether the process moved toward its objective.
- Treating one generation as expected performance. One run is one observation. Report the denominator and uncertainty appropriate to the question you are estimating.
- Choosing a sample size by habit. Twenty examples may be enough for a very large effect and useless for a small one. Power depends on the effect size, variance, design, and decision threshold.
- Unpaired comparison. When two models see the same items, comparing only their aggregate scores throws away information in the paired differences.
- An unvalidated judge. A critic whose relationship to its reference has not been measured is a model-produced opinion with a numeric output format.
- An unpinned judge. Silent critic drift breaks comparability across time unless the archive can be re-scored across the change.
- Optimizing against your only score. A learned proxy can improve while a held-out reference stops improving or degrades.
- Calling a different model family independent evidence. It avoids one obvious self-evaluation path; correlated biases and shared blind spots still have to be measured.
- Starting at level 5 of the ladder. A critic was added before cheaper mechanical or downstream evidence that already covered part of the question.
- Waiting until after the build to measure. Then earlier decisions cannot be evaluated against a baseline that was never recorded.
What this chapter established
- Capability without an objective is magnitude without direction; measurement is what tells you whether a change moved toward the objective you actually care about.
- Verification, review discipline, and cost control share one dependency: evidence about output quality. Without it, all three can degrade while the operational system still looks healthy.
- Ordinary software metrics still transfer to AI systems, but they do not measure semantic output quality. One generation is one observation; quality claims need explicit criteria, appropriate sampling, and uncertainty.
- Evals are experiments. Report uncertainty; cluster when items are related; use repeated draws when run-to-run variance matters; compare models on paired item-level differences when the design is paired; use power analysis before interpreting small differences.
- The measurement ladder runs mechanical checks → frozen task set → downstream signal → human judgment → critic model. Start with the lowest-cost evidence that actually covers the property; none of the rows is automatically ground truth.
- A critic is a model-produced proxy. Separate it from the generator, validate it against a trusted reference with denominator and uncertainty, pin and version it, blind irrelevant provenance, and do not make it the sole optimization target.
- Human-human agreement is not a hard ceiling on critic agreement. It is evidence about ambiguity and noise in the reference process, which must be reported alongside critic performance.
- Gao, Schulman, and Hilton measured overoptimization against a learned proxy relative to a fixed reference model. Keep a held-out reference evaluation that optimization never sees, and treat divergence as a warning rather than as a direct measurement of true quality.
- Build the harness before the feature. A small frozen seed set, applicable mechanical checks, item-level baseline results, appropriate replication and uncertainty, and recorded model/prompt identity are enough to make later changes testable.
Next
There is one more question, and it is the one this whole part has been building toward without saying so.
If the capability is real — and it is — and if you measure properly, then a straightforward prediction follows: things should be finishing. Backlogs should burn down. Projects should reach the terminal state where there is genuinely nothing left to do, sooner and more often. A tireless colleague working at a hundred times your rate should close a ten-year project in about five weeks.
That is not happening, or at least nobody has instrumented the thing that would show it. The last chapter of Part 1 asks why, and the answer turns out to be arithmetic rather than opinion — and it explains why the remaining five parts of this book are shaped the way they are.
Continue with Where Are the Finished Projects?.
References
- Evan Miller. Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations. arXiv:2411.00640, Anthropic, 2024. https://arxiv.org/abs/2411.00640
- Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems 36 (NeurIPS) Datasets and Benchmarks Track, 2023. https://arxiv.org/abs/2306.05685
- Leo Gao, John Schulman, and Jacob Hilton. Scaling Laws for Reward Model Overoptimization. Proceedings of the 40th International Conference on Machine Learning (ICML), PMLR 202, 2023, pp. 10835–10866. https://proceedings.mlr.press/v202/gao23h.html
- Pat Verga, Sebastian Hofstatter, Sophia Althammer, Yixuan Su, Aleksandra Piktus, Arkady Arkhangorodsky, Minjie Xu, Naomi White, and Patrick Lewis. Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models. arXiv:2404.18796, 2024. https://arxiv.org/abs/2404.18796
- Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, Johannes Heidecke, Amelia Glaese, and Tejal Patwardhan. PaperBench: Evaluating AI’s Ability to Replicate AI Research. arXiv:2504.01848, 2025. https://arxiv.org/abs/2504.01848
- Arjun Panickssery, Samuel R. Bowman, and Shi Feng. LLM Evaluators Recognize and Favor Their Own Generations. Advances in Neural Information Processing Systems 37 (NeurIPS), 2024. https://arxiv.org/abs/2404.13076