← Applied AI

Replicate Before You Believe

Does the promising prompt intervention survive a matched replication? Counterfactual wording tied normal on coverage and candidate passes at 46% more tokens, so the frozen rule keeps the default. The tie still moved success between tasks, and a post-hoc check shows why that movement cannot be promoted either.

Part 5 — More Intelligence Is Not Automatically Better

A signal is not a result, and a result is not a default

Somewhere in your own records there is an encouraging pattern: a prompt that seemed to work better, a model that looked stronger on the cases you happened to inspect, a setting that went with success more often than not. The temptation is to adopt it.

The problem is where the pattern came from. It was found by looking at data you already had, which means the data chose the hypothesis. Anything strong enough to notice in a small sample is also the kind of thing chance produces regularly, and hindsight makes it feel predicted rather than discovered.

a signal  ≠  a result  ≠  a default

A signal suggests a question. It becomes a result only when a frozen question, with a decision rule fixed in advance, meets new data. It changes a default only if that rule says so — and a rule worth having can tell you to keep what you already had.

By the end of this chapter you will be able to put a promising signal from your own records through a matched replication: the same tasks, the same exposure on both sides, one variable changed, the decision rule written down before the run, and the resource cost in the headline rather than a footnote. You will also see what to do when the replication ties, which is the ordinary outcome and the one that most needs the discipline.

The challenger gets its matched fight

This book has such a signal. Chapter 26 found a subgroup that looked like a better prompt, and refused to promote it. Here is the test that refusal called for, with its answer stated before any interpretation: at twelve draws per task on both sides, counterfactual wording covered 9 of 12 tasks, matching normal wording, passed 44 of 144 candidates, also matching, while using 46% more tokens. Normal stays the default.

Does the promising prompt intervention survive a matched replication?

What the signal actually was

It is worth being exact about what P3 was built to test, because the signal had two readings and both were weaker than they looked.

The P3 preregistration states the signal in its own words: within one P2 run, counterfactual framing reached 16 of 36 candidates and 9 of 12 tasks, while normal framing in the same portfolio reached 11 of 36 and 7 of 12. That comparison is matched — three draws per task each — but it comes from one twelve-task sample, and its per-stance attribution is not a recorded field in the frozen P2 rows; Chapter 26 recovered it from input-token signatures.

The other reading, the one Chapter 26 warned against, sets counterfactual’s 9 tasks from three draws beside the normal arm’s 10 tasks from twelve: attributable, but unmatched. 1 2

P3 removes the exposure mismatch and turns the discovered subgroup into a named experimental arm. Each arm gets twelve draws per task, on the same twelve tasks, with the same model, starting state, verifier, sealing, and budget accounting. The only intended difference is the prompt suffix. That variable is attributable in the preregistered design and runtime record, but the frozen export later dropped the suffix field, so the export alone does not prove which wording each arm received. The preregistration asks whether the prompt distribution is better, not whether a portfolio is. 1

Signal, result, default

Each step of that chain has to be earned separately.

The discipline as a one-way flow — postdiction never flows backward into prediction:

    flowchart TD
    SG["exploratory signal<br/><i>found in frozen rows: postdiction</i>"] --> FR["freeze hypothesis + rule<br/><i>preregister before new draws</i>"]
    FR --> ND["new draws<br/><i>same tasks, one variable changed</i>"]
    ND --> CMP{"comparison<br/>against the frozen rule?"}
    CMP -->|"rule met"| PRO["promote<br/><i>result moves the default</i>"]
    CMP -->|"rule not met"| REJ["reject<br/><i>signal stays a signal</i>"]
  

Preregistration literature supplies the vocabulary for that cycle. Analyses chosen before seeing outcomes test hypotheses; analyses shaped by the data generate them. Treating the second as the first, helped along by ordinary hindsight bias, is how attractive subgroups become false findings (Nosek et al., 2018). The counterfactual subgroup was discovered in P2’s rows, so it was postdiction; P3 is the prediction built from it. That mapping is the book’s own.

The second discipline is reporting the budget behind a comparison. When conclusions shift with the computation spent, a test score alone cannot say which method is better, and the answer can depend on the budget chosen (Dodge et al., 2019). That paper is about model comparisons, not prompt wordings, but the principle transfers: coverage parity at unequal token spend is not parity of methods, so the premium belongs in the headline.

The experiment as code

In CodeAI, a matched prompt replication is two arms that differ in one field. Expressed in the current ArmDef API, the P3 design is:

cf = stance_suffix(COUNTERFACTUAL)
arms = (
    ArmDef(name="P3C",  models=("qwen-local",), samples=12),
    ArmDef(name="P3CF", models=("qwen-local",), samples=12,
           prompt_suffixes=(cf,) * 12,
           stance_labels=(COUNTERFACTUAL,) * 12),
)

The suffix is a fixed paragraph appended to the unchanged base prompt. It asks the model to set aside the existing implementation strategy, describe the simplest implementation it would write from the observable contract and failing behavior, and then make the smallest change. The normal arm’s prompt is byte-identical to the control prompts of earlier P-series runs. 3 4

That block is the design restated for readability, not the frozen configuration, and the difference matters. The frozen P3 export records the two arms as {"name": "P3C", "models": ["qwen-local"], "samples": 12, ...} and {"name": "P3CF", ...} — identical except for their names. The suffix text appears nowhere in the export. CodeAI’s ledger stores prompt_suffixes in the immutable experiment.created event, but the exporter never copied that field, so the one variable the experiment manipulated was dropped on the way out. 5 6

The rows still carry an indirect trace of it. On every one of the twelve tasks, every counterfactual call consumed exactly 44 more input tokens than every normal call on the same task. A fixed suffix produces precisely that pattern. It is consistent with the preregistered design; it is not the suffix text itself. 5

Current CodeAI contains the exporter repair: exported arms include prompt_suffixes and stance_labels, exported call rows carry each call’s prompt_variant, and a regression test builds a normal-versus-counterfactual pair and asserts the export can tell them apart. The frozen P2 and P3 exports are unchanged and still lack the field. 6 7

The decision, then the interpretation

The rule’s inputs, recomputed from the frozen rows: 8

Normal × 12Counterfactual × 12
Task coverage9/12 (0.75)9/12 (0.75)
Candidate passes44/144 (0.306)44/144 (0.306)
Input tokens20,24426,580
Output tokens8,08114,723
Total tokens28,32541,303 (+45.8%)
Median call latency3.1 s3.7 s

Solve sets were identical with zero rescues either way across 288 calls and 286 checks. 8 5

The primary hypothesis was higher verified coverage for counterfactual. Coverage held at 9/12, so the primary hypothesis fails. The preregistered interpretation rules then apply as written. Rule 2 — counterfactual approximately equal to normal — reads the P2 difference as sampling noise and keeps normal as the default. Rule 6 fires alongside it: retry-once-accepted and size-format remain unsolved in every arm, so framing variants stop being aimed at them. 1 9

One sentence in the P3 report goes further than the rules: that the token premium alone would reject counterfactual as a default. The preregistered rules contain no token threshold, so that is the report’s judgment rather than a frozen clause. It leaves the decision unchanged here, because the coverage tie already decides, but a reader applying the same method should put the resource multiple into the rule before the draws, as Dodge’s argument implies. 9

The premium itself has a mechanism the rows can show. Of the 12,978 extra tokens, 6,336 are the suffix (44 tokens × 144 calls) and 6,642 are longer outputs: counterfactual responses were longer on all twelve tasks. The prompt cost more than its own length; it changed how much the model wrote. 5

Keeping the default is not proving normal superior. The challenger failed the condition for replacing the incumbent in this test — one local model, twelve tasks, one wording — and nothing more.

The tie that moved

Equal aggregates do not mean identical rows. Per task, recomputed from the frozen export and matching the report cell for cell: 5

TaskNormal passes / 12Counterfactual passes / 12
unsound-cache39
dt-roundtrip106
env-timing108
lsp-square68
half-up-rounding65
single-append12
splitlines-cr32
stable-priority43
registry-pollution11
retry-once-accepted00
size-format00
strip-query00

Both columns sum to 44 over the same nine covered tasks, yet the passes sit in different places: up six on one task, down four on another, smaller shifts both ways. On this run, success appeared in different places between the two arms without improving the metric required for promotion. The rows therefore rule out the simple description “nothing differed”, but they do not by themselves establish that the wording caused the redistribution.

It is fair to ask the next question, though: is that movement more than draw-to-draw noise? Nothing was preregistered to answer it, so what follows is a post-hoc diagnostic over the frozen rows, labeled as exactly that. Under an exchangeability null, condition on each task’s total number of passes and treat its twenty-four draw outcomes as exchangeable between two groups of twelve. Shuffle those outcomes between the arms and ask how often total movement is at least the observed 18. That assumption is part of the diagnostic; the rows do not independently establish exchangeability.

import random

rows = [(3, 9), (10, 6), (10, 8), (6, 8), (6, 5), (1, 2),
        (3, 2), (4, 3), (1, 1), (0, 0), (0, 0), (0, 0)]   # (normal, counterfactual)
observed = sum(abs(a - b) for a, b in rows)                 # 18

def shuffled_movement():
    total = 0
    for a, b in rows:
        draws = [1] * (a + b) + [0] * (24 - a - b)
        random.shuffle(draws)
        x = sum(draws[:12])
        total += abs(x - (a + b - x))
    return total

trials = 200_000
p = sum(shuffled_movement() >= observed for _ in range(trials)) / trials
# p ≈ 0.31

About 31% of shuffles move success at least as much as the real arms did. The single largest shift — unsound-cache at 3 against 9 — gives an uncorrected Fisher exact p of about 0.04. All twelve task rows were available for inspection, and a simple Bonferroni correction across those twelve comparisons would not retain that result. The P3 report reached the same qualitative reading in words: task-by-prompt interaction noise, not a stable framing effect. 5 9

So the redistribution is observed, but this post-hoc exchangeability check does not distinguish it from chance-like allocation of the recorded passes between arms. It stays unpromotable. The chapter claims observed redistribution only — not specialization, not a task-wording affinity, not a routing rule. The reading rule for ties is symmetric: unfold the aggregate before concluding that nothing happened, and test the unfolded rows before concluding that something did.

The two missing checks

Two hundred eighty-eight calls produced 286 checks. The frozen rows identify both gaps: one EXECUTION_ERROR candidate per arm with no check ID — a normal draw on stable-priority and a counterfactual draw on strip-query. Candidates that fail in execution never reach the hidden tests, the same accounting the P1.1 report states for its compile-gate errors. Coverage arithmetic is unaffected, because an execution error is a failed draw either way. 5

A second quirk stays logged rather than normalized: every P3 call row, like P2’s, labels its prompt version seeded-code-v1 on a semantic-repair-v1 run. It affects no arm identity, task mapping, or metric, and it is one more reason the suffix trace above matters — the version field could not have told the arms apart either. 5

What this is not

  • Not proof the P2 signal was noise. The rule’s output is that the challenger did not earn promotion, and Rule 2 reads that outcome as noise; the mechanism behind P2’s difference remains unidentified.
  • Not a vindication for normal. One incumbent survived one test, nothing broader.
  • Not evidence the wording did nothing. Success moved between tasks without improving the promotion metric, though whether the wording caused that movement is not established.
  • Not out-of-sample. Nothing here travels to other tasks, models, or wordings.

Where it is still weak

  1. Twelve tasks, one model, one wording. This is a small within-corpus replication with no out-of-sample generality. The permutation and Fisher calculations are post-hoc diagnostics, not preregistered confirmatory tests. 9
  2. Movement without a test that could see it. Twelve draws per task cannot separate redistribution from chance; the post-hoc check was not preregistered. 5
  3. The manipulated variable is not in the frozen export. The suffix is attested by the preregistration, the report, and a constant input-token delta, not by a recorded field. 5
  4. Checks trail calls by two. Both gaps are identified execution errors. 5
  5. The token rejection is not a frozen clause. The rules had no resource threshold. 1

Do this now

Thirty minutes. Run one replication whose answer could embarrass you.

  1. Take an attractive subgroup finding of yours. Write the matched question it would need to survive: same exposures, same metric, one variable changed.
  2. Write the decision rule before collecting anything: what margin promotes, what retains, what drops — and the resource multiple at which a tie still loses.
  3. Check that your export records the variable you are changing. If two arms export identically except for their names, fix that first.
  4. Run it. Score the rule in one sentence before any per-task inspection.
  5. Then unfold the aggregate task by task, and run a shuffle test like the one above before you believe any single row.

If you are building with an assistant:

Give every promising signal a matched replication before promotion: same
draw counts, same tasks, one variable changed, rule frozen first, including
any resource threshold. Verify the exported configuration records the
changed variable. Score the rule before interpreting; a tie at higher cost
does not promote. Then unfold aggregates task by task, report movement in
both directions, and test it against a shuffle null before describing it as
an effect. Account for every missing check by row and rename no historical
label.

Failure modes

  • Calling the tie a near-win. Same coverage at +46% tokens fails the promotion condition.
  • Claiming nothing happened. Success moved; totals hide it.
  • Claiming something promotable happened. Movement that a shuffle reproduces a third of the time is not a routing rule.
  • Adding the threshold afterward. A resource clause invented after the draws is judgment, not preregistration.
  • Trusting the export’s arm names. If the changed variable is not in the record, the record cannot show the experiment was the one you meant.

What this chapter established

  • A signal is not a result, and a result is not a default. A pattern found in records you already have is a hypothesis. Only a frozen question meeting new draws can answer it, and only the rule written first can change what you do.
  • Match the exposure. Nine tasks from three draws beside ten from twelve measures draws as much as wording. Same tasks, same number of draws per side, one variable changed.
  • Put the cost in the headline. Coverage parity at a 46% token premium is not parity of methods.
  • A tie that moves results is still a tie. Success can shift between tasks without improving the metric the rule names, and that movement deserves a check rather than a promotion.
  • Keeping the incumbent is a real outcome. Most honest replications end here, and a rule that can never say “keep what you have” is not a rule.

What CodeAI measured. The matched replication tied on the rule’s metrics: 9/12 coverage and 44/144 passes each, identical solve sets, zero rescues, at +45.8% tokens — half suffix input, half longer outputs. 8 Under the frozen rules, normal remains the default, and framing attacks on the two universally unsolved tasks stop. 1

Per-task success frequencies differed between the arms without improving the promotion metric; under the post-hoc exchangeability diagnostic, a shuffle tail of p ≈ 0.31 did not distinguish that observed redistribution from chance-like allocation of the recorded passes. It does not establish a task-by-wording effect. 5 The frozen export omits the manipulated prompt suffix; a constant 44-token input delta is the rows’ only trace of it, and the exporter is repaired in current source with a regression test. 5 6

Evidence notes

Independent verification. The analysis verifier asserts the P3 identities — 9/12 against 9/12, 44/144 against 44/144, identical solve sets, zero rescues — and catches its seeded corruption; the export verifier checks all four frozen bundle hashes. Those cover the aggregates the rule consumes. 8

Everything else in this chapter is recomputation over the same frozen rows, done for this chapter and labeled as such: the per-task table, the input/output token split, the constant 44-token input delta, the execution-error identification, and the permutation check. None of it is a new bundle, and none of it required a model call. 5

The report’s limits are kept whole: one local model, twelve tasks, one counterfactual wording, within-corpus replication, no out-of-sample validity. With no effect to carry outward, the planned out-of-sample follow-up is moot by the rules’ own logic. Declining to fund another replication is a resource decision the evidence supports. 9

The same instincts show up outside laboratories. In one self-reported public thread, an author running AI-rewritten trading rulebooks against a static control found most AI accounts trailing it and called their best account luck; a reply named the multiple-comparisons problem (thread). It evidences nothing here — only that a control, and a suspicion of the winner, are cheap and usually available. 10

Next

Generation policy now has its discipline: signals earn matched tests, tests earn decisions only through frozen rules, and rules may keep the incumbent. But every experiment in this part spent model calls freely in order to ask its question. The remaining question is whether spending another call should itself be governed — not which model answers, but whether any model needs to be asked at all.

Continue with What Should Happen Next?.

References

  • Brian A. Nosek, Charles R. Ebersole, Alexander C. DeHaven, and David Thomas Mellor. The Preregistration Revolution. Proceedings of the National Academy of Sciences 115(11):2600–2606, 2018. PNAS.
  • Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A. Smith. Show Your Work: Improved Reporting of Experimental Results. EMNLP-IJCNLP, 2019. ACL Anthology.

Implementation sources: P3 ran on the P-series CodeAI lineage (src/codeai/stances.py: stance_suffix, stance_prompt; src/codeai/experiments.py: ArmDef, run_arm; src/codeai/analysis.py: arm_metrics, unique_rescues, build_report, export_experiment). Current CodeAI contains the exporter repair: arm prompt_suffixes and stance_labels, per-call prompt_variant, and the regression test tests/test_stances.py::test_export_carries_the_prompt_variable_of_each_arm. That current repair does not alter the frozen P2 or P3 exports. The ArmDef block restates the preregistered design in the current API; it is not the frozen configuration. The permutation check was executed over frozen rows during editing and is not preregistered. Evidence: experiments/applied-ai/evidence/p-series/ (frozen exports, hashes verified), experiments/applied-ai/evidence/p-series-analysis/ (analysis.json, TABLES.md, analyze_pseries.py, verify_analysis.py), experiments/P3-prereg.md, and experiments/P3-results.md (frozen preregistration and report). Nothing was rerun and nothing frozen was modified.


  1. Preregistration: experiments/P3-prereg.md. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  2. Frozen report: experiments/P2-results.md. ↩︎

  3. Source inspection: src/codeai/stances.py (stance_suffix). ↩︎

  4. Source inspection: src/codeai/experiments.py (ArmDef). ↩︎

  5. Measured run: experiments/applied-ai/evidence/p-series/p3. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  6. Source inspection: src/codeai/analysis.py (export_experiment). ↩︎ ↩︎ ↩︎

  7. Source inspection: tests/test_stances.py (test_export_carries_the_prompt_variable_of_each_arm). ↩︎

  8. Measured run: experiments/applied-ai/evidence/p-series-analysis. ↩︎ ↩︎ ↩︎ ↩︎

  9. Frozen report: experiments/P3-results.md. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  10. Report: docs/applied-ai/plans/capstone-selective-intelligence.md. ↩︎