Blind Before You Compare
Several calls are not automatically independent, and different models are not automatically diverse. Before anyone can measure whether extra proposals add anything, they must be collected blind, none able to see another. Blind is not independent and independent is not diverse. A seal can keep declared sibling information out of a proposal's inspected request path, and that boundary is exactly as strong as the provenance it is given.
Part 5 — More Intelligence Is Not Automatically Better
Once a system can produce several answers, adding more intelligence looks easy: call the model again, try another model, change the prompt, build a portfolio. None of that guarantees more useful coverage.
This part builds the comparison in stages. Proposals first have to be collected without influencing one another, because agreement after exposure means something different from agreement without it. Then variety has to be measured by what actually gets solved rather than by how many model names were involved. When the first benchmark turns out to be too easy to tell the methods apart, the test itself has to be made more discriminating. A promising signal found inside a failed experiment then has to survive a matched replication before it is allowed to change anything.
The question running through the part is not “can we generate more?” It is “did the extra intelligence buy something we can measure and justify?” — asked four times, each under a matched test with its decision rule written down before the run.
Blind is not independent, and independent is not diverse
Ask three reviewers for an opinion on the same paragraph and you appear to have three pieces of evidence. Now suppose the second reviewer read the first reviewer’s answer before writing their own, and the two agree. That agreement no longer means what it seemed to. It could be the same insight reached twice, or the first answer speaking twice. The three opinions are no longer three separate looks at the paragraph.
The same thing happens with AI. Send one task to several calls or several models and it is tempting to assume that separate calls are independent, that different models are diverse, and that several proposals add up to several pieces of evidence. None of that follows from making several calls. If one proposal could see another, through a shared transcript, a shared context file, or a summary someone pasted into the prompt, every comparison built on them is contaminated before it starts.
So collection comes first, and it has to be blind:
blind ≠ independent ≠ diverse
Blindness is an information condition: no proposal’s input contained another proposal’s output. A runtime can construct that condition, and a reader can inspect it. Independence and diversity are different claims about what the proposals then do, and they can only be measured afterwards, on outcomes. This chapter earns only the first. It builds the experimental condition that Chapters 24–27 need before they can ask whether extra candidates, extra models or extra wordings add anything.
By the end of this chapter you will be able to collect several proposals so that none can have seen another through the inputs you control, prove that from the recorded request path rather than from how similar the answers look, and say exactly where that proof stops. The last part is the most practical lesson in the chapter: isolation is only as strong as the information-flow model beneath it. A seal can exclude only information whose origin it knows. Undeclared provenance, text copied by hand, and channels nobody modeled, such as shared files, caches and tools, pass straight through.
Three reviewers, one paragraph
Here is the procedure concretely. Three reviewers receive the same versioned paragraph and the same source pack. Each writes a proposal. Every proposal is collected before any is revealed.
Contrast that with a second procedure: show the second reviewer the first answer and ask for agreement. It costs less, runs faster, and sounds more confident, while answering a different experimental question.
Chapter 22 made repetition explicit so that independent proposals could be compared fairly. This chapter builds the condition those comparisons require.
How can several proposals start without seeing one another?
Blind, independent, diverse
Each of the three words makes a different claim:
| Term | Meaning here | What it does not mean |
|---|---|---|
| Blind | Sibling proposal output is absent from the inspected boundary the claim names — selection at the weaker level, rendered/prepared input at the stronger one | Statistical independence, or isolation through every channel |
| Independent | A statistical claim about joint behavior | Anything established by sealing or byte inspection alone |
| Diverse | Outputs or errors that differ in ways that matter for the task | Anything established by listing different model names, prompts, branch counts, or request digests |
This chapter earns only the first row. Any sentence in it that drifts toward “independent samples” or “diverse reasoning” without a measurement behind it has already failed its own standard.
Three pieces of research frame the distinction, each within limits. Self-consistency samples many reasoning paths from one model and keeps the most consistent answer, with large gains on arithmetic and commonsense benchmarks (Wang et al., 2023). That supports homogeneous repeated sampling as the baseline any fancier fan-out must beat. It does not establish statistical independence, and it does not establish the audited sibling-exclusion property defined in this chapter; its samples begin from the same prompt. The application of that baseline to CodeAI’s fan-out experiments is the book’s own.
Cobbe and colleagues generate many candidate solutions and keep the one a trained verifier ranks highest (Cobbe et al., 2021). The relevant lesson is architectural, not statistical: candidate production stays distinct from selection. More candidates do not always add coverage, and nothing in their result says they must. That boundary — production distinct from selection — is what this chapter constructs the conditions for.
Lorenz and colleagues show the other direction’s danger: in laboratory estimation with real stakes, even mild social influence narrowed opinion diversity without improving accuracy, reduced the centrality of the truth, and boosted confidence anyway (Lorenz et al., 2011). Humans, not models — the mapping is the book’s analogy, not their claim. But the analogy motivates the procedure: collect first, expose later, because exposure collapses exactly the spread that aggregation needs.
One implementation of the blindness boundary
CodeAI implements blindness with a seal: a declaration of what one proposal’s compiled input may not contain. A Seal holds four forbidden sets: event IDs, call IDs, artifact IDs, and lineage IDs. That is the whole enforcement vocabulary. 1
The compiler turns offered candidates into a package plus a trace. Events carry lineage derived from their IDs and payload references — call IDs, source calls, claim and artifact references, explicit lineage lists, and for completed calls the stream ID itself. Artifacts and claims carry whatever lineage the caller declares in the lineage maps; with no declaration they carry none. Each candidate is checked against the seal with a named reason, required candidates that collide fail loudly instead of degrading silently, and the package ID hashes the prompt, actor, selected IDs, budget, and the seal itself — so two branches over identical inputs still hash differently when their actors differ. 2
sealed_fanout runs the procedure: one shared base prompt and base inputs, a per-branch seal, and collection only after every branch has been attempted. Compilation and invocation now sit inside the same per-branch failure boundary. A branch that compiles records its package and trace and then uses the standard recorded-call path; a branch refused during compilation records a failed completion without manufacturing a package or manifest. There is no debate, voting, synthesis, or cross-branch feedback, by design. The heart of the current path is:
One shared base in, sealed branches out, comparison only afterwards:
flowchart TD
B["shared base<br/><i>one prompt, one input set</i>"] --> S["per-branch seal<br/><i>forbids exactly the siblings</i>"]
S --> B1["branch A package<br/><i>compiled from base only</i>"]
S --> B2["branch B package<br/><i>compiled from base only</i>"]
B1 --> C1(["proposal A"])
B2 --> C2(["proposal B"])
C1 --> V["verifier compares afterwards<br/><i>collection only, no synthesis</i>"]
C2 --> V
Two facts about that path matter more than the rest. First, because branches run in sequence from fixed base inputs, the seal usually has nothing to exclude: sibling outputs do not exist when a branch compiles, and later branches never read them. In the ordinary case isolation comes from construction; the seal guards against sibling-derived content arriving through the base inputs.
Second, the fake adapters used to establish the boundary never read their input at all: the deterministic fake returns canned responses keyed by call ID and records the spec it was given. Its blindness is trivially true and proves nothing about bytes reaching a model. What the fake-based evidence actually establishes lives one layer down, at compilation: which declared content entered each package. 3 4
At the Stage 15B point in the construction, bytes-level proof existed one layer further down and only where a call opted in. With context_render set and a prepare/send adapter, the runtime resolves every selected ID to bytes, lays them out canonically with per-item content hashes and offsets, composes the model input, refuses a prepared body that does not carry exactly that text, stores the rendered bytes as an artifact, and binds the rendered hash, item list, and input layout into the manifest before any provider effect.
The pinned fan-out run below still used the older selection-record path: its branches sent the prompt string with no renderer, so it could establish what the compiler selected but not the final composed model-input bytes. That scope statement is measured by the Stage 15B evidence and the frozen fan-out run; the later section “The fan-out path itself” describes the current repair. 5
So there are two blindness claims, and they must not be merged:
A seal plus recorded compilation can establish that declared sibling material was not selected. A seal plus rendered prepare/send binding can establish the stronger claim that declared sibling material was absent from the inspected prepared request input.
Both claims stop at the provenance and channels the runtime models. Selection traces, package contents, prompt hashes, and — where rendering is enabled — rendered bytes and request-body hashes are the inspected path. Shared files, process state, caches, tools, environment, and any channel the runtime never modeled are outside it.
What a seal excludes, and what it cannot
The preserved isolation demo runs three offline parts with fake adapters, and its independent verifier checks the summary plus a seeded corruption. 6
| Part | What ran | Recorded result |
|---|---|---|
| Fan-out | Three branches, one shared base prompt; branch C with invalid max_attempts=0 | 3 distinct calls and packages, A/B succeeded, C failed without touching siblings, nothing replayed, fanout.requested + fanout.completed present |
| Seal exclusion | Branch A’s seal over sibling B/C lineage against a base event, a sibling event, a sibling artifact and a sibling claim | All three sibling candidates excluded with excluded because seal forbids call lineage ['call-B']; base selected; package cda02013… |
| Boundary | The same artifact ID with its lineage map omitted | Included — the seal had nothing to match against |
The third row is the chapter’s load-bearing result: seals block declared channels, not every information-flow channel. The mechanism cannot exclude information whose provenance the caller never declared. Provenance is part of the enforcement mechanism; when it disappears, the seal loses something it needs. That lesson connects directly to Chapters 17–18: an observation without provenance is not evidence with an attributable origin, and a compiler cannot recover what was never supplied. The Stage 15 bundle measured the same boundary independently — supplied lineage excluded, omitted lineage admitted — with its own seeded corruptions rejected. 6 7
The negative controls below were executed against the working tree as teaching fragments, not as pinned runs. Each uses a sentinel string standing in for sibling content: 8
Declared sibling content is excluded, and the sentinel never reaches the package prompt:
package, trace = compiler.compile_with_trace(
task_id="t", actor=reviewer_A, prompt="review",
events=(base, sibling_event_with_sentinel),
artifact_ids=("artifact-sib-B",),
seal=sibling_forbidding_seal,
artifact_lineage={"artifact-sib-B": ("call-B",)},
)
assert trace_entry("artifact-sib-B").decision == "excluded"
assert SENTINEL not in package.prompt
The same artifact with the lineage map omitted is included:
package, trace = compiler.compile_with_trace(
task_id="t", actor=reviewer_A, prompt="review",
events=(base,), artifact_ids=("artifact-sib-B",),
seal=sibling_forbidding_seal,
)
assert trace_entry("artifact-sib-B").decision == "included"
And sibling text copied into the shared prompt is not filtered — seals match declared identifiers, never arbitrary semantics:
package, _ = compiler.compile_with_trace(
task_id="t", actor=reviewer_A,
prompt=f"review. A sibling once wrote: {SENTINEL}",
events=(base,), seal=sibling_forbidding_seal,
)
assert SENTINEL in package.prompt
The three outcomes together say what the mechanism is: an identifier filter with durable reasons, not a semantic firewall and not access control. The Stage 15 report adds the sharpest corollary from its own fixture: sealed content excluded from selection could still be read directly from storage. Selection does not dereference sources and does not enforce downstream access. 7
Testing the boundary on the fan-out path
The fan-out path itself was tested under a protocol frozen before execution, with six sealed_fanout invocations over one ledger and disposable fake-adapter branches. An independent stdlib-only verifier reconstructs every branch row — seal, context package, exclusions, prompt representation, call, output lineage — from the ledger alone, and rejects four seeded corruptions.
| Case | Outcome |
|---|---|
| Clean fan-out (A1, A2) | sentinel absent from both recorded packages; each seal forbids exactly its sibling |
| Declared artifact, directly forbidden | RequiredContextMissing raised loudly; zero branch executions; no package, call, or manifest recorded |
| Forbidden sibling event/call lineage | RequiredContextMissing raised loudly via call-lineage forbidding; zero branch executions |
| Omitted provenance | sentinel artifact INCLUDED by reference (boundary, kept) |
| Copied sibling text | caller-copied sentinel present in C2’s package (boundary, kept) |
One branch failure (OSError) | F1 succeeded with output intact; F2 failed with the error preserved; fanout.completed holds both |
| Ledger reopen | a second runtime recounts the identical per-task event graph |
The second and third rows changed the protocol before any result existed, and that change is the finding, because the frozen plan had expected silent exclusion of the forbidden base input. The first execution attempt instead raised RequiredContextMissing: on the sealed_fanout path every explicitly passed base event and artifact is a required candidate, so a seal conflict fails loudly before any branch executes. The amendment is recorded in the bundle’s preregistration with its reason. Silent exclusion lives one layer down, at the compiler with explicit required subsets — the behavior the unit tests pin and the fragments above demonstrate — not on the fan-out path itself.
The verifier establishes two precisions about the inspected boundary. First, each recorded call.manifest prompt_hash recomputes as SHA-256 over instruction plus package prompt, which means manifests preserve hashes rather than bytes, and the final composed model-input bytes are not preserved on this legacy path. Second, the omitted-provenance inclusion is by reference: the artifact ID sits in the package’s included set with its bytes retrievable from the store, not inlined into the prompt.
Down to the bytes
The fan-out run ends where the next question starts: hashes and package text are not the bytes the provider would have received. A follow-up extension closes that gap on the opt-in rendered path — genuine prepare/send calls with CONTEXT_RENDER_V1, an offline gateway, compiler-level seals with explicit required subsets, and an independent verifier that reads the content-addressed rendered artifacts and recomputes the composed input.
| Case | Rendered bytes | Sent body |
|---|---|---|
| Sealed exclusion | sentinel absent; manifest hash recomputes from the stored bytes; layout present | absent; body equals the recomputed composed input byte for byte, bound by request_body_sha256 |
| Omitted provenance | sentinel present | sentinel present — the boundary, proved at the byte level |
| Copied prompt | sentinel absent (render covers selected items, not the branch prompt) | sentinel present via the prompt — the limit, proved at the byte level |
| Tampered body | — | refused: RenderBindingError before any effect, call.preparation_failed with provider_effect: false, no manifest, no attempt, zero sends |
What the extension does not claim: outbound wire bytes (body identity is canonical-JSON hash, not wire capture), semantic prevention, sandboxing, or independence. 9
The fan-out path itself
That extension proved the byte-level boundary on the opt-in rendered path, and left the then-current
sealed_fanout path at selection evidence only. Current sealed_fanout now accepts context_render
and threads the renderer into each branch. A rendered branch resolves selected material to bytes, the
prepared body must carry exactly the composed rendered input, and the manifest records both the
rendered-context digest and request-body digest. A branch whose adapter cannot prepare a request
fails as a branch rather than falling back quietly. 10
There is one current-source boundary worth keeping explicit. fanout.requested records the fan-out
level context_render default and labels that request rendered_bytes or selection_record, while an
individual branch may override context_render. In uniform fan-outs the event says exactly what every
branch is doing. In a mixed-mode fan-out, the per-branch manifest is the evidence for that branch’s
actual render binding; the fan-out-level field alone is not sufficient. 10
So the chapter’s levels now stay separate:
selection blind no sibling was selected seal + compilation trace
final-input blind no sibling was in prepared input rendering + request binding
statistically independent not established here
usefully complementary later outcome measurements
And the middle level is exactly as strong as the provenance it was given, which is the next point.
Blind is not diverse, and the bytes say so
Run two branches over the same material, with the same question and the same model, and their prepared requests come out byte-identical — same rendered digest, same request body digest. The record proves that property of the inputs. 10
Those branches are blind through the inspected path, and their prepared inputs are identical. What the record does not establish is that their outputs are non-diverse, statistically independent, or statistically dependent: stochastic sampling can still produce different answers and different errors. Blindness creates an information condition. It does not manufacture outcome diversity, and a fan-out that reports “three independent reviews” from sealing or byte identity alone is upgrading an input property into a statistical claim.
Change the question and the request-body digests diverge while the rendered material stays the same. That proves which input dimension changed; it still says nothing by itself about whether the resulting errors will be complementary.
What failure does to a fan-out
The demo’s branch C failed on an invalid configuration without touching A or B. Current CodeAI holds that property more widely than the demo exercised it: the per-branch guard catches any exception, not just value and runtime errors, so a transport-level failure — an OSError from a real adapter, the TimeoutError family included — records that branch as failed and lets its siblings finish. The docstring always promised that no branch kills its siblings; in the version the demo ran on, the implementation kept that promise for only two exception types. 3
A regression test pins it: one branch raising OSError, two succeeding, the failure’s message preserved on its completion, the fan-out completion carrying all three statuses. The test also pins what the failure leaves behind — an attempt.started with no observation and no interpretation. That orphan is the honest residue of an interrupted attempt: inspectable through stream reads, resolved by nothing. It is the call-level crash gap of Chapter 16 appearing inside fan-out, and the chapter does not pretend otherwise. 11
Attribution otherwise survives per branch, but not every branch has the same depth of record. A branch that reaches normal completion has its manifest, attempts, observations, interpretations, decisions, and completion through the standard recorded path, with variant, seal, package, and trace recoverable after reopen. The interrupted OSError case above is deliberately thinner: it has a manifest and attempt.started, then a synthetic failed call.completed appended by the fan-out guard, with no observation, interpretation, or attempt completion. Partial success is therefore mixed per-branch evidence collected under one fan-out completion, not a claim that every branch traversed the same lifecycle. 11
What the runtime does not do is equally deliberate. Unknown-provenance inputs are allowed, not rejected or quarantined: a blanket policy would need to distinguish “no lineage exists” from “lineage was omitted,” and the compiler cannot tell those apart. Sealed experiments that need the stricter reading must supply complete lineage — the mechanism’s strength is exactly the provenance it is given, and the chapter leaves that as a stated limitation rather than engineering around it here. 8
Testing the current seam exposed another defect in the primitive’s stated contract. sealed_fanout
promised that one branch failing would not erase successful siblings, but the guard originally covered
the call and not compilation. A branch refused by its own seal could therefore propagate out of the
whole fan-out. Compilation now sits inside the per-branch guard: that branch is recorded as failed,
with the seal’s reason, and its siblings continue. This is current implementation/test evidence, not a
new pinned model run. 10
Sealed is not sandboxed
Even where every declared byte is sealed, branches may still share the filesystem, the repository, environment variables, external services, caches, and durable runtime state. A branch adapter that reads the ledger can observe completed siblings regardless of any seal; the fake adapters in the demo simply never exercise that ability. The fan-out branches here generate text and perform no tool effects — had they written files, no seal would have confined them. 3
The scope sentence from the proof boundary therefore extends one step: sealed prompt and context is not sandboxed execution, and blind collection is not information-flow security. The P-series ran its heterogeneity measurements under this machinery, and its first result is a useful warning label for the next four chapters: on a twelve-task corpus, three sealed model families solved the same eleven tasks and all failed the twelfth. That is a ceiling-bound result about one shared failure, and Chapter 24 reads it carefully.
It is still enough to show why this chapter refuses the stronger words: sealed branches from different models can make the same mistake, so neither separate calls nor separate model names establish diversity. 12
What this is not
Five denials hold the boundary in place:
- Not statistical independence. Sealed inputs do not make joint behavior independent; nothing here measures a joint distribution.
- Not diversity. Separate calls, separate models, and separate prompts are sources to be tested for diversity, not diversity already established.
- Not a quality claim. Whether extra candidates help is Chapters 24–27’s question; this chapter builds the condition their experiments need.
- Not synthesis. Collection happens after generation, with no debate, voting, critics, or cross-branch feedback.
- Not information-flow security. Undeclared channels, shared mutable state, and omitted provenance all pass through seals untouched.
Where it is still weak
- A seal excludes by identity, so unattributed material is admitted. Hand
sealed_fanouta sibling’s output without saying whose it is and that output can reach the other branch’s prepared bytes. The runtime records how much base material was offered and how much carried lineage, but it cannot tell material that never had lineage from material whose lineage was stripped. Rendering proves what was sent; only supplied lineage decides what the seal can exclude. 10 - Semantic duplication is invisible. Copied sibling text in a shared prompt, file, or tool result matches no identifier the seal knows. 13
- Shared mutable state is outside the boundary. Filesystem, repository, environment, services, caches, and the ledger itself remain readable to adapters that choose to read. 3
- Selection is not access enforcement. Excluded content can still be read directly from storage; rendering excludes because compilation did. 7
- Interrupted branch attempts leave orphans. A branch killed mid-attempt can record
attempt.startedwith no observation or interpretation; the fan-out can preserve that residue but does not reconcile it. 11 - The fan-out-level provenance label is a default, not necessarily every branch’s actual mode. Branches may override
context_render, so mixed-mode runs require the per-branch manifests to establish which branches were actually rendered and bound. 10 - Unknown provenance has no policy. Allow, exclude-under-seal, and explicit UNKNOWN classification remain unimplemented. A blanket rule could break ordinary compilation, so the current seam records the gap rather than silently changing the global default. 10
Do this now
Thirty minutes. Collect two proposals that cannot have met.
- Compile one shared prompt with two different actors and sibling-forbidding seals. Confirm two packages, two traces, and neither package naming the other — then find every place the two records still share fate (task, ledger, filesystem, your own eyes).
- Offer a sibling artifact with declared lineage and confirm its exclusion with the reason quoted. Remove the lineage map, recompile, and confirm inclusion. Write both decisions next to each other.
- Copy one sentence of sibling text into the shared prompt and confirm the seal leaves it alone. Decide, in writing, whether your experiment’s threat model includes a careless curator — and what would catch one.
- Fail one branch of a three-branch fan-out with a transport-style error. Confirm the siblings, the statuses, and the orphan attempt row.
If you are building with an assistant:
Collect proposals before revealing them. Seal each branch from its siblings
through declared identifiers and lineage, and record the seal, the candidate
inventory, the exclusion reasons, and the selected package per branch. Prove
blindness from the inspected request path — selection trace, package contents,
and, where available, rendered and prepared request bytes — never from answer
similarity. Demonstrate the boundary deliberately: omitted provenance included,
copied semantics unfiltered, shared mutable state outside the seal. Contain
branch failures without touching siblings, and attribute every proposal to its
task, variant, call, package, and seal after reopen. Claim blind, never
independent or diverse, until later chapters measure what the extra
candidates actually add.
Failure modes
- Calling separate calls independent. Two invocations are not two independent samples.
- Calling different models diverse. Labels are not error distributions; correlated sensors wear different names.
- Reading an exclusion trace as a proof of bytes. Selection says what was permitted, not what the adapter received through every path.
- Trusting answer similarity. Different outputs do not prove blindness; identical outputs do not prove leakage.
- Forgetting the curator. The seal enforces declared provenance; whoever assembles the base inputs decides what is declared.
- Letting one branch kill the fan-out. A transport failure is a branch outcome, not a collection failure.
- Confusing collection with synthesis. Gathering proposals answers what was proposed; it settles nothing about what is right.
What this chapter established
- Several calls are not automatically independent, and different models are not automatically diverse. Blind, independent and diverse are different claims. This chapter establishes only bounded blindness over the inspected path.
- Collect blind before you compare. If one proposal could see another, agreement between them cannot separate shared insight from shared influence, and every later comparison inherits the contamination.
- Prove blindness from the request path, not from the answers. A compilation trace can prove what was selected; rendered and prepared-request evidence can prove a stronger final-input claim. Similar answers do not prove leakage, and different answers do not prove blindness.
- Isolation is only as strong as the information-flow model beneath it. A seal can exclude only what the provenance it is given declares. Omitted lineage, hand-copied text, and unmodeled channels such as shared files, caches, tools and the ledger itself stay outside the guarantee. Blind collection is not sandboxing, and it is not information-flow security.
- Blindness is a precondition for later measurement, not the measurement itself. Chapters 24–27 test whether extra candidates produce useful complementary coverage under matched experiments. They do not turn sealing into proof of statistical independence.
What CodeAI showed. The executed demo collected three branches with contained failure, excluded declared sibling candidates with seal-naming reasons, and admitted the same artifact once its lineage was omitted. 6 The Stage 15 context evidence independently showed supplied-lineage exclusion with admitted omissions and, on opt-in rendered calls, binding down to stored rendered bytes. 7 5
The pinned sealed-proposals run established the fan-out’s selection-level boundary, including loud refusal of required forbidden input, omitted-provenance and copied-text limits, branch failure, and reconstruction after reopen. The separate rendered-byte extension carried the same boundary down to prepared request bytes on its opt-in path. 9 Current CodeAI then threaded rendering through sealed_fanout, recorded the fan-out-level provenance mode, demonstrated the unattributed-material hole on that path, and moved compilation inside the per-branch failure guard. Those are current source and regression-test results, not additions to the earlier pinned runs. 10 Unknown provenance still has no policy.
Evidence notes
The pinned runs. The sealed-proposals bundle is experiments/applied-ai/evidence/sealed-proposals/2026-09-14-a1b562a/. Seeded corruptions — a flipped manifest hash, a sentinel smuggled into a sealed package, a sibling dropped from a recorded seal, rewritten completion statuses — are each rejected with the failure named.
The rendered-byte extension is experiments/applied-ai/evidence/sealed-rendered/2026-09-14-1b3c7a2/. Four seeded corruptions — a flipped rendered hash, a sentinel injected into the captured body, a deleted exclusion entry, a deleted refusal record — are each rejected. Two corrections are recorded in that bundle’s preregistration (a missing sentinel constant; the copied-prompt rendered expectation), neither changing any result.
Independent verification. The demo’s independent verifier imports no CodeAI code. It requires three distinct calls with contained sibling failure and no replays, three distinct packages, exclusion of all declared sibling candidates with seal-naming reasons, and — as the boundary assertion — the undeclared artifact included. It also flips that boundary bit in a mutated copy and requires rejection. 14
Its limit is the familiar one: it checks the producer’s summary, not the ledger. A summary that misreported an exclusion would pass if it misreported consistently. The Stage 15 and 15B verifiers check harder — recomputed identities, byte inventories, rerendered bytes against stored artifacts — with seeded corruptions that insert sealed content, remove exclusion reasons, and reorder selections, each rejected for the claims it breaks. Those verifiers cover compilation and rendering. The fan-out collection step is covered by the pinned fan-out run reported earlier, whose verifier recomputes every branch row from the ledger rather than from a summary, searching the recorded package text and manifest hashes for sentinels itself.
Next
Proposals can now be collected without meeting through the declared path — blind, attributable, separately recorded, failures contained. Whether having more of them is worth anything is a measurement question, and the first measurement is the least flattering one: same question, different models, same mistakes.
Continue with The Models Were Different. Their Mistakes Weren’t.
References
- Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. Self-Consistency Improves Chain of Thought Reasoning in Language Models. ICLR, 2023. arXiv:2203.11171.
- Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training Verifiers to Solve Math Word Problems. arXiv:2110.14168, 2021. Paper.
- Jan Lorenz, Heiko Rauhut, Frank Schweitzer, and Dirk Helbing. How Social Influence Can Undermine the Wisdom of Crowd Effect. Proceedings of the National Academy of Sciences 108(22):9020–9025, 2011. PMC3107299.
Implementation sources: historical demos and pinned runs are evidence about the versions that produced them; current-source statements above were checked against the later CodeAI successor and are not projected backward. The sealed_fanout excerpt is reduced from current source. src/codeai/runtime.py: sealed_fanout, compile_and_record_context, invoke_recorded_call, _build_manifest, _build_manifest_from_prepared, _append_context_compiled, _call_spec_payload; src/codeai/context.py: ContextCompiler, event_lineage, candidate_blocked_by_seal, ContextSealViolation, RequiredContextMissing; src/codeai/domain.py: Seal, Variant, ContextPackage, RenderedContext; src/codeai/rendering.py: render_context, compose_model_input, RenderBindingError; src/codeai/adapters.py: FakeCognitionAdapter. Current tests include tests/test_fanout_isolation.py and tests/test_fanout_provenance.py, with earlier fan-out, epistemic, context, and rendering tests kept as surrounding regression coverage. Evidence: experiments/applied-ai/evidence/independent-calls/; experiments/applied-ai/evidence/context-selection/2026-09-13-54e32384/; experiments/applied-ai/evidence/context-rendering/2026-09-13-0f9a83b/; experiments/P1-results.md; experiments/applied-ai/evidence/sealed-proposals/2026-09-14-a1b562a/; experiments/applied-ai/evidence/sealed-rendered/2026-09-14-1b3c7a2/; and the later seam record docs/seams/fanout-provenance.md with preregistration experiments/W2-prereg.md. Footnote prefixes retain their evidence meaning: s inspected source or regression evidence, m pinned measured run, d preserved demo, r frozen report. The independent-calls bundle and earlier pinned runs remain unchanged by the later fan-out implementation repairs.
Source inspection:
src/codeai/domain.py(Seal). ↩︎Compiler source:
src/codeai/context.py(ContextCompiler, event_lineage, candidate_blocked_by_seal). ↩︎Runtime source:
src/codeai/runtime.py(Runtime.sealed_fanout). ↩︎ ↩︎ ↩︎ ↩︎Adapter source:
src/codeai/adapters.py(FakeCognitionAdapter). ↩︎Measured run:
experiments/applied-ai/evidence/context-rendering/2026-09-13-0f9a83b. ↩︎ ↩︎Unpinned demonstration:
experiments/applied-ai/evidence/independent-calls. ↩︎ ↩︎ ↩︎Measured run:
experiments/applied-ai/evidence/context-selection/2026-09-13-54e32384. ↩︎ ↩︎ ↩︎ ↩︎Compiler internals:
src/codeai/context.py(ContextCompiler). ↩︎ ↩︎Measured run:
experiments/applied-ai/evidence/sealed-rendered/2026-09-14-1b3c7a2. ↩︎ ↩︎Current seam evidence:
docs/seams/fanout-provenance.md; preregistrationexperiments/W2-prereg.md;src/codeai/runtime.py(Runtime.sealed_fanout); regression matrixtests/test_fanout_provenance.py. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎Frozen report:
experiments/P1-results.md. ↩︎Seal predicate:
src/codeai/context.py(candidate_blocked_by_seal). ↩︎Unpinned demonstration:
experiments/applied-ai/evidence/independent-calls/verify_isolation.py. ↩︎