Regressions: Did the Model Change, or the Measurement?
Establish whether a regression exists before explaining it by measuring baseline variation, using paired comparisons when valid, fingerprinting the evaluation path, controlling randomness and determinism, bisecting interacting changes, and recording runs so model changes can be separated from measurement changes.
Here is a refactor. The old training script set the seed at the top; the new one moves that line below model construction, where it reads more naturally, and draws the batch-shuffle permutation from a local torch.Generator() instead of the global RNG, so the shuffling is “self-contained.” The training procedure is mathematically the same and the model architecture is unchanged. But the concrete initialized parameters are no longer coupled to the old run, because model construction now consumes randomness before manual_seed is applied. Run it once:
baseline validation loss 0.9726
after the refactor 0.9980
The refactor cost 0.026 validation loss. That is the sentence an engineer writes in the pull request, and it is not yet supported by anything.
Run the unchanged baseline at twelve different training seeds:
seed 0 val_loss 0.9726 seed 6 val_loss 0.9757
seed 1 val_loss 0.9847 seed 7 val_loss 0.9598
seed 2 val_loss 0.9762 seed 8 val_loss 0.9871
seed 3 val_loss 0.9778 seed 9 val_loss 0.9771
seed 4 val_loss 0.9690 seed 10 val_loss 0.9706
seed 5 val_loss 0.9731 seed 11 val_loss 0.9754
median 0.9756 stdev 0.0071 min 0.9598 max 0.9871 spread 0.0273
The spread of runs that differ only in their seed is 0.027. The refactor’s 0.026 is inside it. Run the refactored code at twelve seeds and it produces median 0.982, range [0.959, 0.999] — a distribution that overlaps the baseline’s almost entirely. Moving the manual_seed line changed which random numbers reached the model’s initialization, so every run is different; whether the distribution shifted at all is not resolved by twelve runs. The “regression” was a single sample from a noisy process, reported as a fact. There was a number, and there was a story, and the story arrived first.
Now the other half. Take one trained model — fixed, saved, not retrained — and evaluate it four ways, changing only the evaluation code:
evaluation variant val loss val acc
canonical 0.9726 0.6760
model left in train() mode (dropout on) 1.0711 0.6290
a different 500-example validation subset 1.0348 0.6580
inputs scaled x1.15 on the eval path only 1.0881 0.6330
The model is byte-for-byte identical in every row. Every number moved — some by four times the seed noise — and nothing regressed. The meaning of the number changed.
A metric that moved is not yet a finding. It is a symptom with at least three explanations, and only one of them is about the model.
Where we are
Chapters 11, 12 and 13 all investigate inside a single run. Is the model learning — which link in the chain is broken? Where is the time and memory going? What did the compiler capture, and what invalidated it? Every one of those questions has an answer you can extract from one execution.
This chapter is the first whose unit of analysis is a pair of runs. That is its entire reason to exist. A regression is not a property of one run; it is a claimed difference between two, and the claim is only as good as the comparison behind it.
Chapter 13 handed this over precisely: compilation adds dimensions along which two runs can differ invisibly — a different sequence of shapes seen, a warm cache in one run and cold in the other, the recompile limit reached this week and not last week. None of that raises. None of it shows in a unit test. The question this chapter answers:
When this run is worse than that one — did the model change, did the measurement change, or did nothing change and the difference is noise?
Three explanations, one symptom:
The technique
Every chapter earns an investigation method. Chapter 11: prove the chain one boundary at a time. Chapter 12: measure, localize, intervene, remeasure. Chapter 13: find the first assumption that stopped holding.
Chapter 14 adds:
Establish the noise floor before attributing a cause. Before explaining a difference, prove there is one.
The hidden structure this makes visible is the causal chain behind any reported number:
code + configuration + data + randomness + environment + measurement
↓
result
A regression means something in that chain changed enough to move the result. The investigation is to find which link. The link that people reach for first is code, because that is the one they were working on. The links that are actually broken are usually measurement (the benchmark changed) or randomness (nothing changed; it is a different seed). And measurement is a link like any other — a changed evaluation path and a changed model produce exactly the same symptom.
The environment these numbers came from
Every number, table and trace in this chapter was produced by running the code shown, on one machine:
Python 3.11.4
PyTorch 2.6.0+cu118
OS Windows 11
CPU 24 cores; torch.set_num_threads(1) for every training run in the chapter
GPU present but used only for the cross-device comparison in E3
This chapter is almost entirely CPU work, and cheaply so. The controlled workload trains in about 0.7 seconds; a twelve-seed sweep is under fifteen seconds. Every experiment below, including the development runs that tuned the workload, is roughly fifteen minutes of CPU in total. There is no honest excuse for an unmeasured number in a chapter about not trusting unmeasured numbers.
One thread is the canonical setting, chosen deliberately: on one thread, with the data fixed and the seed fixed, two runs of this workload are bit-identical, which makes every later difference attributable to something. It is also, for a model this small, the fastest option.
One controlled workload
A single synthetic classification task, used for every experiment.
DEFAULT_CONFIG = {
"n_train": 3000, "n_val": 1000, "in_dim": 32, "hidden": 32, "classes": 4,
"noise": 0.30, # 30% of labels randomised: val loss has a real floor
"data_seed": 0, # the DATA is fixed; only the training seed varies
"epochs": 20, "batch_size": 128, "lr": 0.02,
"weight_decay": 3e-3, "dropout": 0.3, "optimizer": "adamw",
}
The task is a random linear rule with 30% label noise, so no model can drive validation loss to zero — there is a floor to sit at, around 0.97. The model is a three-layer MLP. The data is generated once from data_seed and never varies; the seed passed to a run controls only initialization, batch order and the dropout mask. That separation is what makes “run it again at a different seed” a clean experiment.
The shared instruments are in explab.py: train_run(seed, config) returning a run record, seed_sweep, a median/spread summariser, paired and unpaired comparison functions, and an eval-path fingerprint. Every table in this chapter is reproducible from them.
E1 — The noise floor
This is the first experiment, before any change is evaluated, and skipping it is why most reported regressions cannot be argued about.
QUESTION how much do two runs of the IDENTICAL configuration differ, from the
training seed alone?
METHOD run DEFAULT_CONFIG at 12 seeds; report the distribution.
identical config, 12 seeds
val loss median 0.9756 stdev 0.0071 min 0.9598 max 0.9871 spread 0.0273
train loss median 0.8911 stdev 0.0098 spread 0.0370
val acc median 0.6795 stdev 0.0065 spread 0.0200
examples/s median 196143 spread 75794
Across these twelve observed seeds, validation loss spans 0.0273 and has standard deviation 0.0071. That is an empirical scale for ordinary seed-to-seed variation in this experiment, not a universal detection threshold. A single-run difference of 0.026 is therefore not compelling evidence of a regression: it is the same size as variation already produced by the unchanged configuration. Establishing a smaller systematic shift requires repeated runs — preferably paired when the pairing remains meaningful.
This baseline distribution is the yardstick for the rest of the chapter. It is a property of this workload under this protocol — a bigger model, a real dataset, a different environment or a different metric needs its own baseline distribution. The throughput spread is enormous here — nearly 40% — because this desktop’s CPU load swung during the sweep; that is the observed variation of whole-run throughput under this measurement protocol, not a threshold for every performance benchmark on the machine. Chapter 12 deliberately used tighter timing windows and repeated measurements for that reason. The validation-loss path, by contrast, is bit-identical when seed and execution conditions are held fixed in E2; across E1 the variation comes from deliberately changing the training seed.
When a candidate’s single-run delta is comparable to ordinary baseline variation, the regression has not been established. Compare distributions or paired deltas before explaining it.
E2 — Same seed, same code, twice
What does setting the seed actually buy?
run A (1 thread) vs run B (1 thread), same seed:
val curves bit-identical: True
final val loss A 0.972560 B 0.972560
On one thread, with the data, code and seed fixed, these two executions produce bit-identical validation curves. That is a strong reproducibility result for this controlled workload under these recorded conditions, and it is what makes later interventions interpretable.
Now change one thing — the thread count — and nothing else:
E3 — Nondeterminism you can execute
The thread-count difference in E2 has one mechanism, and it is worth stating plainly because it explains a large fraction of what people call flakiness.
Floating-point addition is not associative.
(a + b) + c = 0.0 with a = 1.0, b = 1e16, c = -1e16
a + (b + c) = 1.0
That is not a pathology invented for this chapter; it is a consequence of finite-precision arithmetic. Floating-point sums can change when their accumulation order changes. Sum the same 200,000 float32 values three ways:
sequential, forward order : -730.639587
sequential, backward order : -730.646729
pairwise (torch.sum) : -730.646667
max gap: 7.14e-03
A gap of 7e-3 from accumulation order alone. Put that effect inside training, where reductions feed later floating-point operations, and a small numerical difference can propagate. Change the thread count:
torch.use_deterministic_algorithms(True)
This forces deterministic implementations where they exist and raises where they do not:
torch.use_deterministic_algorithms(True)
x = torch.randn(2, 3, 8, 8, device="cuda", requires_grad=True)
grid = torch.rand(2, 8, 8, 2, device="cuda") * 2 - 1
torch.nn.functional.grid_sample(x, grid, align_corners=False).sum().backward()
RuntimeError: grid_sampler_2d_backward_cuda does not have a deterministic
implementation, but you set 'torch.use_deterministic_algorithms(True)'. You can
turn off determinism just for this operation, or you can use the 'warn_only=True'
option, if that's acceptable for your application.
That is the feature working. Deterministic algorithms can also change performance, and some CUDA matrix operations require CUBLAS_WORKSPACE_CONFIG (:4096:8 or :16:8) for deterministic execution or PyTorch raises.
What the setting does not provide is universal bit identity across environments:
E4 — Worker seeding
DataLoader workers are Chapter 6’s subject. The reproducibility question is narrower: with num_workers > 0, which randomness in the pipeline is tied to the run’s seed, and which floats free? Put a draw from four different RNGs inside __getitem__, run the loader twice from the same torch.manual_seed(0), and compare:
num_workers=0 num_workers=2 with the repair
data order (shuffle) reproduces reproduces reproduces
torch.rand in __getitem__ reproduces reproduces reproduces
np.random.rand (legacy global) DOES NOT reproduces reproduces
random.random DOES NOT reproduces reproduces
np.random.default_rng() object DOES NOT DOES NOT DOES NOT
The surprising result is version- and implementation-specific, so keep its boundary visible. In this PyTorch 2.6 experiment, the multi-process worker path caused the legacy NumPy global RNG and Python’s random stream to reproduce, while the single-process num_workers=0 path did not. An independently created np.random.default_rng() object remained outside that mechanism in every case.
For portable code, do not rely on this observed worker-startup side effect. Pass an explicit generator= to the DataLoader; seed external libraries explicitly in worker_init_fn when workers exist; and, when num_workers=0, seed those external RNGs in the main process because worker_init_fn is never called. A default_rng() object must itself be constructed or reseeded from a controlled seed.
Reproducibility of pipeline randomness depends on the exact RNG object that produces it. Control that object explicitly rather than assuming the DataLoader will discover it.
E5 — Paired comparison
E1 gave the noise floor. When a change’s real effect is comparable to it, how the two configurations are compared decides whether the effect is visible at all.
Configuration B lowers the validation set’s label noise from 0.30 to 0.29 — a “the data got slightly cleaner” change, the kind a dataset revision makes. Run A and B, twelve seeds each.
Unpaired — summarize the two sweeps without using the seed correspondence:
Paired — same seed on both sides, compare per-seed:
seed 0: 0.9726 -> 0.9456 -0.0270
seed 1: 0.9847 -> 0.9532 -0.0315
seed 2: 0.9762 -> 0.9692 -0.0070
seed 3: 0.9778 -> 0.9353 -0.0425
seed 4: 0.9690 -> 0.9587 -0.0102
seed 5: 0.9731 -> 0.9864 +0.0133
seed 6: 0.9757 -> 0.9606 -0.0151
seed 7: 0.9598 -> 0.9541 -0.0057
seed 8: 0.9871 -> 0.9605 -0.0265
seed 9: 0.9771 -> 0.9605 -0.0166
seed 10: 0.9706 -> 0.9268 -0.0437
seed 11: 0.9754 -> 0.9450 -0.0304
mean per-seed delta: -0.0203 stdev of deltas: 0.0166
11/12 seeds improved |mean| / stdev = 1.2
Eleven of twelve paired deltas are negative. Using the same seed on both sides creates a common-random-number comparison: when the two runs consume randomness in corresponding ways, seed-specific difficulty is partly shared and the variance of the difference can be much smaller than the variance of the raw outcomes.
This chapter deliberately stops short of turning that pattern into a formal significance claim. |mean| / stdev = 1.2 is a descriptive effect-to-spread ratio, not a hypothesis test or confidence interval. The evidence is directional and consistent enough to justify the next experiment; if the decision matters, predeclare an inferential criterion or collect more paired seeds rather than converting one descriptive ratio into “proof.”
Pairing is strongest when a shared seed induces comparable randomness on both sides. A refactor that changes RNG consumption — an extra draw, a reordered stochastic pipeline, a different sampler — can weaken that coupling even though the seed numbers still match. In that case inspect the paired-delta variance itself; if the coupling disappeared, the paired design may buy little over an unpaired comparison.
E6 — The broken benchmark
The MEASUREMENT-CHANGED branch. Take one trained model, fixed, and evaluate it while changing only the evaluation path:
evaluation variant val loss val acc fingerprint == canonical?
canonical 0.9726 0.6760 7176cf3fce66 -
model in train() mode (dropout on) 1.0711 0.6290 f8dce02c6c6b False
different validation subset (500 of 1000) 1.0348 0.6580 0e6c39db2124 False
inputs scaled x1.15 on eval path only 1.0881 0.6330 9fb90744a142 False
cross_entropy reduction='sum' 1074.8778 - 22aacc4965b5 False
Every reported metric moved, three by much more than the baseline seed-to-seed variation and the last by a factor of a thousand. The last value is simply the mean loss summed over one thousand examples, but without the aggregation contract it looks like the model exploded. The model parameters are identical in every row; the evaluation procedure is not.
The defence is an eval-path fingerprint over the components this experiment has chosen to make part of the measurement contract:
Leaving evaluation in train() mode is the specific case the previous chapter’s semantics — eval() versus no_grad() — exist to prevent; here it is a cause whose symptom is a 0.10 metric shift, and the fingerprint is what surfaces it.
E7 — A real regression, diagnosed
Now the MODEL-CHANGED branch, with a genuine cause: the learning rate is ten times too high.
Detect — seed sweeps:
baseline median 0.9744 range [0.9598, 0.9847]
lr x10 median 1.3900 range [1.3880, 1.4082]
shift = +0.4156 noise floor spread = 0.027 -> 15x the noise floor
The ranges do not come close to overlapping; the shift is fifteen times the noise floor. Detection is trivial. The work is the explanation.
Explain 1 — learning curves and gradient norms, one seed:
epoch 0 3 6 9 12 15 18
baseline 1.016 0.996 0.969 0.965 1.003 0.978 0.984
lr x10 1.393 1.400 1.388 1.392 1.399 1.393 1.401
baseline gradient norm: median 0.510 max 0.983
lr x10 gradient norm: median 0.084 max 6.581
The broken run’s validation loss is flat from the first epoch and its gradient-norm distribution is dramatically different from baseline: much smaller in the median with occasional large spikes. Those observations tell us when the trajectories separated and give us a numerical symptom to pursue. They do not, by themselves, prove that “most units saturated” or that the step size “destroyed the parameters.” To make that mechanistic claim, inspect the corresponding activations/parameter updates or run a targeted intervention such as restoring the learning rate from the same initialization. The chapter’s own rule applies here too: do not let a plausible story outrun the measurement.
Explain 2 — a different cause, per-example evidence. The training labels are shifted by one in the loader, so every input is paired with the next input’s label:
baseline val loss 0.9726 shifted val loss 1.4034 (+0.4308)
examples correct before, wrong after: 528 / 1000
shifted-model accuracy: 0.2320 (baseline 0.6760)
Accuracy 0.232 is close to the 0.25 chance rate on four classes; being slightly below chance on one finite validation set is not evidence by itself that the model learned an anti-rule. What localizes this failure is the known intervention — labels were shifted — together with the large aggregate degradation and the per-example flips. The model has been trained against a broken input-target relation, so the data pipeline is the first boundary to inspect. This is Chapter 11’s shuffled-target failure arriving as a regression.
The aggregate metric detects a change. Curves, per-example failures and numerical signals narrow the mechanism — but none should be asked to prove more than it measured.
E8 — Curves, not endpoints
Two configurations, baseline versus a lower learning rate run for twice as long. Their final validation losses differ by less than the baseline single-run spread:
A: 1.02 0.99 1.00 1.00 1.00 0.97 0.97 0.98 0.97 0.96 0.96 0.98 1.00 0.98 ...
min 0.9648 at epoch 10, final 0.9726
B: 1.13 0.98 0.97 0.97 0.95 0.95 0.96 0.96 0.96 0.95 0.95 0.95 0.96 0.95 0.94 0.94 ...
min 0.9433 at epoch 22, final 0.9824
B descends to 0.943 by epoch 22 and then rises again. Across eight shared seeds, B’s lowest recorded validation loss is lower than A’s in all eight, by about 0.015 on average. That is interesting trajectory evidence, but comparing post-hoc minima deserves care: B runs for twice as many epochs and therefore gets more opportunities to produce a low validation point. If the real training procedure uses early stopping or best-checkpoint selection, define that checkpoint policy before comparing the configurations and evaluate the policy across seeds. The chapter’s durable finding is that equal-looking endpoints can hide materially different trajectories.
“When did the runs diverge?” is a question the final number cannot answer. If deployment selects a best checkpoint, make the checkpoint-selection rule part of the experiment rather than choosing the minimum after seeing the curves.
E9 — Bisection
“Change one thing at a time” is sound advice and useless after someone has already changed twelve. The candidate branch here makes six plausible improvements at once, and it regressed badly:
baseline (0 changes) val loss median 0.9747
candidate (6 changes) val loss median 2.6609 (+1.69, 62x noise floor)
Order the six changes and binary-search on how many of them, applied as a prefix, are enough to cross the regression criterion:
When both the code and the data changed
If a commit changed the model and a dataset revision landed the same week, the four cells of a 2 × 2 experiment separate main effects from interaction:
E10 — A policy guardrail from baseline variation
A regression gate needs a decision boundary, but twelve baseline seeds do not magically turn one formula into a confidence bound. This chapter uses a deliberately conservative policy threshold derived from the observed baseline variability:
new healthy run (seed 999): val_loss 0.9924 threshold 1.0574 pass=True margin +0.065
new run with lr 10x (seed 999): val_loss 1.3902 threshold 1.0574 pass=False margin -0.333
The illustrative test:
The run record
Every recent chapter has a reusable instrument: Chapter 7’s stage report, Chapter 10’s attention ledger, Chapter 11’s learning ledger, Chapter 12’s performance ledger, Chapter 13’s compilation ledger. Chapter 14’s is the run record — and it is the one that makes all the others comparable across time.
The comparison is run at a shared set of seeds so it can report the distribution and the per-seed delta, and it reports across several dimensions because a change routinely helps one and hurts another — the multi-dimensional habit from Chapter 12’s performance ledger:
RUN candidate-lr10x BASELINE main-baseline (10 shared seeds)
IDENTITY
config hash 587cf3ca1f03 -> 91b44992b14c CHANGED
eval fingerprint 7176cf3fce66 -> 7176cf3fce66 SAME
env hash 0838b7da7ddc -> 0838b7da7ddc SAME
CORRECTNESS (median [min..max])
val loss 0.9760 [0.9598..0.9871] -> 1.3900 [1.3880..1.4082]
val acc 0.6785 -> 0.2460
paired val-loss delta: mean +0.4173 stdev 0.0121 (10/10 worse)
-> 15x the noise floor
COST (median)
examples/sec 94178 -> 95149 (+1.0%)
DECLARED BEFORE THE RUN
hypothesis: raising the learning rate speeds convergence without
materially worsening validation loss
success: val-loss delta within the noise floor
result: NOT MET
Read top to bottom, the report narrows the comparison before debugging starts. The eval fingerprint is SAME, so none of the recorded evaluation-contract fields changed. The env hash is SAME, so none of the fields included in that environment record changed. The config hash changed, and the paired delta is enormous with every seed moving the same way. That combination is strong evidence that the candidate configuration changed model behavior rather than merely changing the recorded benchmark or environment.
A failed experiment recorded this way is not a wasted run. “Raising the learning rate destroyed convergence” is a real, durable fact about this system, worth more than a run quietly deleted because it did not confirm an intuition.
A worked pass
Put the sequence together on one regression. A commit reorganised the data loading, and validation loss on a single run went from 0.97 to 1.40.
1 — Baseline variation. The baseline configuration at eight seeds: median 0.974, range [0.960, 0.985], spread 0.025. This is the empirical reference for this version of the workload; if the data, evaluation contract or environment changes materially, re-establish it.
2 — Is the difference plausibly ordinary seed variation? 1.40 - 0.97 = 0.43, about 17× the observed baseline range. That is far too large to dismiss as the ordinary variation seen in step 1, so the investigation continues.
3 — Eval-path fingerprint. a4c1... == a4c1... on both runs. The recorded evaluation indices, preprocessing, aggregation and eval() state are identical. That rules out changes in those recorded fields; it does not prove no unrecorded measurement code changed.
4 — Environment. Same env_hash: the recorded Python/PyTorch/thread settings match. No recorded environment difference explains the shift.
5 — Pair it. Run the new code at the same eight seeds as the baseline. Every seed lands near 1.41; the paired delta is +0.44 ± 0.01, 8/8 worse. The effect is large and consistent.
6 — Now suspect model/data behavior. The recorded measurement and environment contracts match and the repeated comparison is unambiguous, so something in the changed training system genuinely moved the behavior.
7 — Curves. The new run’s validation loss is flat from epoch 1. It shows no useful improvement under this training recipe.
8 — Individual failures. Validation accuracy is 0.23, close to the four-class chance rate of 0.25. That number alone does not prove the model learned a coherent wrong rule. It tells us the degraded model is near chance on the real validation task, so per-example and data-contract evidence matter next.
9 — Numerical signals. Gradient norms remain finite, so there is no simple “backward is broken” explanation. That narrows the search; it does not identify the cause.
10 — Isolate. The commit diff has one change that touches label handling: an enumerate that now starts at 1. Reverting only that line restores 0.97.
11 — Test. Add an assertion that the training set’s known-rule agreement is 1.0 before the loop starts — Chapter 11’s contract promoted to a guardrail — plus the separately calibrated metric gate from E10 as a backstop.
Steps 3 and 4 are cheap places where the investigation could have ended as measurement or environment drift, and steps 1 and 2 determine whether repeated model-level diagnosis is worth paying for. The model/data investigation comes later, after the comparison itself has survived those checks.
Forensic mode versus everyday mode
Chapter 12 kept forensic instrumentation (hooks, the profiler, anomaly detection) separate from the production workload because the instruments cost time. The run record is the exception: it costs almost nothing — a few hashes and a JSON write — and its entire value is being there before you knew you would need it. Write it on every run. The forensic escalation in this chapter is not more instrumentation; it is more seeds — and the ten-second sweep is the reason that is affordable.
Using AI on a suspected regression
Given two numbers and a diff, an assistant will produce a fluent causal story — a learning-rate effect, a precision effect, a data effect — and it will do this whether or not the difference is real, because narrative explanation is what the request pattern-matches to. That fluency is precisely the hazard, and it is the same hazard this chapter teaches you to resist in yourself.
The prompt below refuses the explanation and demands the comparison.
Run B looks worse than run A. Do not propose a cause yet.
First reconstruct the COMPARISON:
1. What exactly are the two runs -- code version, config, data version, seed?
2. What is the baseline distribution of run A's configuration across several seeds?
If I have not measured it, say so and stop -- that is the first missing
piece of evidence.
3. Is B - A large relative to that baseline variation? If A and B are single
runs, the honest answer may be "not resolved"; a finite observed range is
context, not a formal significance threshold.
4. Were the seeds paired (same seeds on both sides), and did the code preserve
enough RNG coupling for pairing to reduce variance? Report the paired deltas.
5. Is the recorded evaluation contract identical -- same eval indices,
preprocessing, aggregation, and model.eval() state? Compare fingerprints,
while remembering they cover only the fields included in the hash.
6. What else changed: code, config, data, environment (PyTorch / device /
thread count), and compilation (shapes seen, recompiles, cache state)?
7. Was the success criterion written before or after the result was seen?
Identify the FIRST link in {code, config, data, randomness, environment,
measurement} that lacks evidence, and stop there.
Do not propose a mechanism -- learning rate, precision, normalization, data --
until (a) repeated evidence establishes a difference worth explaining and
(b) the recorded evaluation contract has been checked.
Three clauses carry the weight. Demanding the seed sweep before anything else puts the noise floor first, where the chapter’s argument says it belongs. Asking whether the seeds were paired is what separates “cannot resolve” from “resolved with a sensitive comparison.” And forbidding a mechanism until the difference is established and the eval path is verified enforces the order: rule out NOTHING-CHANGED and MEASUREMENT-CHANGED before spending a day on MODEL-CHANGED.
Ask AI whether there is a difference before asking it to explain one.
The regression-debugging sequence
The early steps are intentionally cheap: many apparent regressions can be resolved as ordinary variation or a changed measurement contract before model-level forensics begin.
What you should now be able to answer
“Validation loss went from 0.97 to 1.01. The change hurt us.” The unchanged workload produced a 0.027 range across twelve seeds, so a 0.04 single-run gap is large enough to investigate but not enough to justify a causal claim by itself. Run both configurations at shared seeds and inspect the paired deltas.
“I set the seed, so the run is reproducible.” A seed controls one source of randomness. E2 showed this workload stopped being bit-identical when the thread configuration changed, and E4 showed that independent RNG objects require their own control.
“It’s nondeterministic, that’s just how GPUs are.” Floating-point non-associativity explains why changing reduction order can change numerical results, including many parallel reductions and atomic update patterns. cuDNN benchmarking is a different mechanism: benchmark noise can select a different algorithm on a later run. Name the actual source instead of calling every difference “GPU nondeterminism.”
“I changed twelve things and it’s better.” Which one? Bisect only when the good/bad predicate behaves suitably, then confirm the candidate alone; interactions can make a clean prefix boundary misleading.
“Accuracy dropped from 68% to 66%.” The unchanged workload showed a two-point min-to-max accuracy range across twelve seeds. That is context, not a formal threshold. Repeat the comparison — preferably paired — before naming a regression; if it is real, inspect which examples flipped.
“Both runs ended at 0.97, so they’re equivalent.” Similar endpoints do not establish equivalence. E8 showed trajectories can differ substantially before converging to similar final values. Compare curves and apply a predeclared checkpoint-selection policy if best checkpoints matter.
“A 5% drop should fail CI.” A round number is not automatically better than an empirical rule, but neither is median + 3 × observed range a statistical law. Calibrate the gate from baseline behavior, then validate its false-alarm rate on held-out healthy runs.
“The compiled version regressed.” Same eval fingerprint? Same shapes seen, same recompile count, same cache state (Chapter 13)? Compilation is a link in the causal chain, and it changes between runs more easily than the code does.
Exercises
These build on the current chapter’s “five deliberate regressions” idea — each one is a different meaning of “the new version is worse,” and the exercise is to tell which.
Baseline variation. For a training workload of your own, run the unchanged configuration across multiple seeds. Report validation-loss, accuracy and throughput distributions. State what the observed variability tells you — and explicitly state why a finite min/max range does not define the smallest defensible single-run regression.
Bit-identical, and where it stops. Run the same seed twice and confirm whether the full curve matches. Then change, one at a time: thread count, device, and PyTorch minor version if available. For each, report whether the run is bit-identical and how large any difference is relative to the baseline variation.
Which RNG. Put a draw from
torch, NumPy legacy RNG, Pythonrandom, and an independentnp.random.default_rng()insideDataset.__getitem__. Testnum_workers=0and2. Then build an explicit seeding strategy for the main process, worker processes, and independent generator object rather than relying on worker-startup side effects.Unpaired versus paired. Make a subtle controlled change. Compare it once as two unpaired distributions and once using shared seeds. Report the actual paired-delta variance and whether pairing improved precision; do not design the exercise so a particular conclusion is guaranteed.
The broken benchmark, five ways. Take one trained model and evaluate it under five deliberately changed measurement contracts. Record the fields that your fingerprint covers, and name at least one plausible evaluation change that your fingerprint would currently miss.
Detect, then explain. Introduce a learning-rate intervention large enough to produce a reproducible effect. First establish the shift across seeds. Then use curves, updates, activations or gradients to test a specific mechanism rather than inferring one from gradient norm alone.
Curves and checkpoint policy. Find two configurations with similar final metrics but different trajectories. Compare them under a checkpoint-selection or early-stopping rule declared before examining the test run; explain why post-hoc minima from unequal numbers of epochs are biased evidence.
Bisect a diff. Construct a config change with six simultaneous edits, one of which is sufficient to regress the metric. Bisect an ordered predicate, confirm the suspected change alone, then add one interaction that violates monotonicity and show how the simple bisection story changes.
A policy gate. Use baseline runs to propose a regression threshold, then test it on additional healthy seeds that were not used to choose the threshold. Report the observed false alarms before calling the gate calibrated.
The run record, later. Save a run record and comparison report for one experiment. Reconstruct later what changed, what was declared beforehand, what evidence supported the verdict, and which causal fields your record failed to capture.
Next: the capstone
The book’s investigative toolkit is now complete. Tensor geometry and the first wrong value. The gradient path and where it breaks. Registered structure and the four things that can disagree about what the model owns. The input pipeline and where it waits. The representation contract and where a legal tensor stops being a meaningful one. Convolutional and attention geometry, derived rather than guessed. The learning chain, one boundary at a time. Performance, localized to a phase and an operator. Compilation, and the assumption that stopped holding. And now comparability — whether a difference between two runs is real, and which link in the chain produced it.
The capstone builds a small GPT-style language model from scratch. Its premise is that success is no longer the script finished without an exception. Every instrument in this book becomes a question that can be asked of that model as it is built: what does this tensor represent, and what should its shape be? Does the loss reach every parameter? Is the input contract intact? Where does the step spend its time? Did compilation help? And — the question this chapter added — when the next change lands, is the model actually better, or does it just look that way?
The capstone owns the model. This chapter’s job was to make sure that when it is built, you can tell whether it works.
A result is not a number. It is a comparison between two runs that share enough structure to support an explanation. Establish the difference before explaining it, and check the measurement before blaming the model.