One Process, End to End
Correct components can still fail where they hand off to one another. One disposable task runs end to end, from grant and model call through observed effect, bound check, acceptance and replay, plus six branches that stop. Then a frozen audit classifies every joint as enforced, derived, recorded, conventional or absent, and finds six that only looked strong when the components were tested one at a time.
Part 6 — Put Intelligence Into the Process
Correct parts can still fail where they meet
Every mechanism in this book was tested on its own, against a fixture built for its chapter. Context was compiled and recorded. Effects were observed by the runtime rather than taken from the worker’s report — or so each component’s own test reported. Checks were bound to the state they examined. Acceptance cited its evidence. Replays refused to act twice. Each passed its own test.
None of that proves the whole process works, because the failures that matter most in a composed system live in the handoffs. Some are easy to picture:
- the action records the hash of one state, and the verifier checks a fresh reading of a different one;
- the acceptance cites a check that ran on different bytes from the ones being accepted;
- an effect happens, the process dies before recording it, and the state in between is unknown;
- a branch is refused, but leaves a record that later projections read as silence rather than refusal.
Every component in those examples can be individually correct. The chain still breaks, because the bug is in the joint.
component correct ≠ composition holds ≠ deployment ready
This chapter is the book’s composition test. It runs one small task through a representative end-to-end path in the architecture the preceding chapters built — intent, context, cognition, claim, authority, effect, observation, verification, acceptance and replay — and asks whether each exercised boundary hands the next one the identity, state and evidence it actually established. It does not exercise every mechanism in the book; later sections keep those omissions explicit. It also runs the task into six walls on purpose, because a reliable process is defined partly by where it refuses to continue.
By the end of this chapter you will be able to trace one task through a complete AI process, name the joints between its mechanisms, classify each one by how strongly it actually holds — enforced, derived, recorded, conventional, absent — and say exactly why passing this test still does not make a system ready for production.
Why composition is its own test
Isolation proves that a component honors its own contract. Whether one component’s output matches the input the next one assumes stays untested, and that untested space is where the examples above live.
Chapters 14 through 28 earned the first term of component correct ≠ composition holds ≠ deployment ready. This chapter tests the second, for one exercised path, and leaves the third out of scope with the gap named.
The composition run had a protocol frozen before execution, a pass condition stated in advance — the independent verifier must reconstruct every named link from records plus bytes and catch every seeded corruption — and one explicit tolerance: a gap the run retains visibly does not fail composition; a gap it hides does. 1
That run is the first of two experiments in this chapter, and it answers the weaker question. It shows that the process can be reconstructed. A second experiment, run later against a successor CodeAI revision built from the same chapter mechanisms, froze its protocol before the harness existed and asked how strongly each joint actually held — and found that reconstructable and enforced are not the same property. 2
One paragraph, one process
The capstone fixture is deliberately small. A disposable file holds one line:
draft: models are stochastic, so review is hard [S1].
The frozen criteria for the task are narrow on purpose: the [S1] marker must be gone, and the sentence Models supply cognition. must be present. The “reviewer” is a fake model adapter that returns one canned revision. That is deliberate: Chapters 23 through 28 spent their evidence on what models and model policies do, and this chapter tests everything around the model.
By the end of the run, one authorized edit has happened exactly once, a check has examined the exact bytes the edit produced, a person-labeled acceptor has accepted them, and a repeated request has been answered from the record without touching the file. Six other branches stop: an edit under the wrong grant, a check aimed at a stale state, a key reused for a different instruction, a verifier that crashes, an effect whose outcome is unknown, and a proposal that fails its check. An independent reader, given only the ledger file and the preserved bytes, reconstructs all of it.
That is the book’s thesis in operational form. The model is not the process. It is one selectively invoked component inside an inspectable process — one that records why it acted, what it was allowed to do, what actually happened, what was checked, and what remains uncertain.
Can one task move from directive through authorized effect, observation, bound verification, and acceptance while leaving enough durable evidence for an independent reader to reconstruct what happened — and which joints between the mechanisms does the runtime actually enforce?
The architecture, joined
Read the path as the architecture the book built, one boundary at a time, each now required to hand its result to the next:
| Boundary | Where the book built it | What it must hand on in this run |
|---|---|---|
| Directive and grant | Chapter 20 | A child grant that narrows its recorded parent |
| Task and criteria | Chapter 14 | Frozen criteria that the acceptance cites by hash |
| Context | Chapter 15 | A compiled, recorded selection for the one call |
| Model call | Chapters 11–13 | A recorded call with its interpretation and completion state |
| Claim | Chapter 18 | An attributed statement that is not yet evidence (here through the older claim path) |
The boundaries after the model call — operation choice, authority, effect and observation, bound verification, acceptance, replay — are the ones this chapter measures rather than lists, and the matrix later in the chapter names each with the chapter that built it.
Here is the happy path as the preserved ledger records it. One line is not a ledger record, and it is marked: the scheduler is a pure function, and nothing persisted its answer. 1
directive.opened d-comp READ + WRITE + ACCEPT
directive.opened d-comp-edit READ + WRITE, max_tokens 1000
causation_id = d-comp's event (narrowing validated)
task.created t-comp criteria: remove [S1]; contain "Models supply cognition."
context.compiled selected from task.created
call.* call-comp one fake call; interpretation v1, generation complete
claim.recorded claim-comp-1 legacy claim path, attributed to call-comp
(unledgered) scheduler decide_next_step(has_required_verification) → CHECK
check.* check-comp-1 PASS on the proposal artifact
action.* act-denied DENIED under READ; observed 72bfd9f4… (pre-edit bytes)
action.* act-comp-1 SUCCEEDED under WRITE, requested_by human-approver
observed 77e25726… (post-edit bytes)
check.* check-comp-2 target_state_hash 77e25726… → PASS, observed 77e25726…
task.accepted acc-comp-1 human-reviewer under ACCEPT; names both check events
task.completed causation_id = task.accepted; artifact 77e25726…
action.* act-comp-2 same key, same fingerprint, WRITE → replay of act-comp-1
adapter invocations still 1
The success criterion was never whether the final paragraph reads well, but whether a reader holding only records can answer why the process acted, whether it was allowed to, what changed, which exact state was checked, why it was accepted, and whether repetition caused a second effect. The rest of the chapter shows how each answer is carried — and where an answer rests on convention rather than enforcement.
The whole run as one chain, in the visual vocabulary the book has built — stadium proposes, diamonds gate, cylinders endure:
flowchart TD
O["objective + grant<br/><i>directive: READ + WRITE + ACCEPT</i>"] --> T["task<br/><i>frozen criteria</i>"]
T --> CX["context<br/><i>compiled + recorded</i>"]
CX --> M(["model call<br/><i>the only cognition</i>"])
M --> CL["claim<br/><i>attributed, unresolved</i>"]
CL --> C1["check 1 · PASS<br/><i>independent command</i>"]
C1 --> G{"authority gate?"}
G -->|"READ: deny"| DN["DENIED recorded<br/><i>file unchanged</i>"]
G -->|"WRITE: permit"| A["action<br/><i>effect + observed hash</i>"]
A --> C2["check 2 · PASS<br/><i>bound to the post-edit hash</i>"]
C2 --> H(("human reviewer<br/><i>ACCEPT grant</i>"))
H --> AT["acceptance<br/><i>names call + checks</i>"]
AT --> CO[("completion<br/><i>caused by acceptance</i>")]
CO --> R["replay<br/><i>same key, no second effect</i>"]
G -.->|"cannot proceed"| ESC["stop branches<br/><i>recorded, not retried blind</i>"]
C2 -.->|"FAIL or ERROR"| ESC
The joints, in the real API
The producer script is ordinary Python against CodeAI’s public runtime. Reduced from the executed producer, with the joint comments added for this chapter: 3
rt = Runtime(ledger, artifact_store=store,
state_resolver=lambda: sha_file(paragraph)) # the runtime's only eyes
done = rt.execute_action(
ActionRequest(action_id="act-comp-1", task_id="t-comp",
directive_id="d-comp-edit", capability="write",
instruction="apply proposal", precondition_hash=None,
idempotency_key="comp:edit:1",
requested_by="human-approver", adapter="fixture-editor"),
authority=WRITE, adapter=editor) # grant passed in by the caller
after_hash = sha_file(paragraph)
assert done.observed_state_hash == after_hash # joint 1
check2 = rt.run_check(
py_check(file_code, cwd=str(ws), check_id="check-comp-2", task_id="t-comp",
target=artifact_target(sha), target_state_hash=after_hash),
verifier=verifier)
assert check2.observed_target_state_hash == after_hash # joint 2
completion = rt.accept_task(
AcceptanceRequest(acceptance_id="acc-comp-1", task_id="t-comp",
actor_id="human-reviewer",
criteria_sha256=criteria_sha256(CRITERIA),
artifact_sha256=sha, source_call_id="call-comp",
source_attempt_id=attempt.attempt_id,
source_interpretation_id=interp.interpretation_id,
check_ids=("check-comp-1", "check-comp-2")),
authority=ACCEPT) # joint 3
Joint 1: effect → observation. execute_action appends action.requested first, looks for a prior completion under the key, checks the capability, compares any precondition, calls the adapter, and then reads the state resolver itself before appending action.completed. The adapter reported succeeded; the runtime separately recorded 77e25726…, the hash of the bytes now on disk. On the denied path, nothing reaches the adapter and the runtime still records its reading — the pre-edit hash — so “denied” and “file unchanged” stand as two recorded facts rather than one inferred from the other.
That is Chapter 19’s reported ≠ observed ≠ expected, held in three different places: the adapter’s status, the runtime’s observation, and the task’s frozen criteria. 4 The question at this joint is whether the runtime preserved what actually changed, in a form the next mechanism can use.
It preserved both readings. What it did not do, at this revision, was compare them — which is the single most important thing the later audit found, and the section on that audit returns to it.
Joint 2: observation → verification. When a check names target_state_hash, run_check reads the resolver once, compares, and only on a match invokes the verifier. On a mismatch it records ERROR without running anything. Whatever happens, the runtime overwrites observed_target_state_hash with its own reading, so a verifier cannot supply the observation that authorized it. Requested target, runtime observation and preserved after-bytes all agree at 77e25726…. 5 1 A check that read its own fresh state would have verified something the action never recorded.
Joint 3: verification → acceptance. accept_task refuses unless the authority grants ACCEPT, the criteria hash matches the task, the source call succeeded with a complete generation on its final attempt, the named interpretation was that call’s decision basis, the artifact equals the preserved call output, every named check exists for this task, targets these bytes, and passed, and the acceptor is not the actor that produced the artifact. Only then does it append task.accepted and a task.completed whose causation_id points back at it. A completion without that causal pair completes nothing. 6 Chapter 14 separated a successful call from finished work; here acceptance closes that loop on bytes an authorized action wrote.
Read joint 3 closely and one link is missing: the acceptance names the call and the checks, but not the action. Nothing in AcceptanceRequest refers to act-comp-1. The effect joins the acceptance only because check-comp-2 was bound to the post-edit hash, and that hash equals the artifact hash — which holds here because the editor wrote the proposal verbatim. An edit that transformed the proposal on the way to disk would break that equality, and nothing in the acceptance path would say which action produced the accepted state. The capstone reconstructs this link; the runtime, at this revision, does not enforce it. 7
Later work measured how far that reconstruction actually goes, before changing anything. Given only the record — from an acceptance, to its checks, to the runtime observation each check carried, to the actions that observed that state — the action was derivable in one of six shapes. A second action that changed nothing, an action from another task, and a replay of the same operation each produced two candidates; and the ordinary shape produced none at all, because an acceptance-eligible check binds an artifact, and binding a state as well is optional. 8
What closed it is a distinction rather than a field:
accepting a verified result ≠ accepting an action as the producer of that result
An acceptance now says which basis it rests on. By default it records artifact_check with no effect basis at all — a task may legitimately accept something imported, hand-written or already correct, and forcing a fictional action into that history to satisfy a schema would manufacture a claim rather than record one. An acceptance that does name an action has it enforced: same task, the action exists, its effect state is OBSERVED, and a cited check examined that same observed state and passed. Absence is recorded as absence, so a later reader is not left inferring it from a missing field.
And the record carries its own limit in words, because this is a binding and not a cause: the action was followed by this runtime observation, and an accepted check examined it. Chapter 19’s proximity limit is unchanged. 8
The adequacy of the checks is equally specific. Check-comp-1 tested only that the proposal contained the required sentence; check-comp-2 tested the file for both the sentence and the marker’s absence. Together they cover the frozen criteria exactly, and nothing beyond them — not whether the paragraph is good, which no criterion asked. Independence is similarly modest: the checks are commands that read bytes and share no code with the fake proposer, not a second judgment about meaning. That is Chapter 21’s adequacy in its narrow, honest form. 1
Joint 4: completion → replay. After acceptance, the same request was sent again under the same key: act-comp-2, with the same fingerprint and a WRITE grant. The runtime returned the recorded result of act-comp-1, and the editor’s invocation count stayed at 1, so repetition returned history rather than causing a second effect. 1 A matching key alone would not have been enough: the replay path checks current authority first and the operation fingerprint second (Chapter 22), and the key-reuse branch below shows a changed instruction refused. What this run did not exercise is a replay under a grant that no longer allows the write; the later audit did, and found current recorded authority governing the replay. 9 10
Where cognition was, and was not
Exactly one model call ran, and it was a fake. Everything else in the ledger is deterministic bookkeeping or a labeled human gate: one context compilation, five check completions, six action completions and one replay refusal, one acceptance. 1
The scheduler’s contribution needs stating precisely, because it is easy to overstate. The producer asked decide_next_step a question — verification required — and got CHECK rather than another CALL. Separately, it asked with the model budget exhausted and verification pending, and still got CHECK: spent cognition does not stop a process that has a check to run. Both are real answers from the committed policy. But the flags were set by the producer, not projected from ledger state, and neither answer was persisted. The capstone shows the policy saying the right thing; it does not show the runtime deriving the question or recording the decision. 1 11
Both halves of that gap were later closed, in the two different ways they deserved. The facts the policy consumes are now projected from the ledger and the decision is recorded with the state it rested on — that is the durable state → decision joint, which the audit classifies as DERIVED, because it is reconstructable rather than a gate. Whether a decision then governs what actually happens is a separate joint, and a weaker one; the audit found it conventional. 12
The branches that stop
In each branch below, stopping is the correct outcome: continuing would have been unjustified, and the record says why. A capstone with only the happy path would hide every boundary the book built. Six branches stop, each in its own durable vocabulary: 1
| Branch | What the ledger holds | What it teaches |
|---|---|---|
| Edit under READ | denied, error capability denied: write, observed pre-edit hash; adapter not called | Capability ≠ authority; the grant is checked before any effect |
| Stale target | Check ERROR target state mismatch; verifier invocations +0 | The check never ran, so there is no verdict to misread |
| Same key, new instruction | action.replay_refused naming instruction; editor calls still 1 | A key is not an operation identity |
| Raising verifier | Check ERROR verifier raised RuntimeError: …; nothing promoted, nothing accepted | The checker failed ≠ the thing checked failed |
| Effect then crash | Two failed completions, the second replaying the first; marker present in preserved bytes | Failed ≠ ineffectual; the outcome is unknown |
| Bad proposal | Check FAIL; task projects INCOMPLETE; no acceptance | A failed check ends the line |
The crash branch is the composition’s honest unknown, so it gets a closer reading. The adapter appended a marker to crash.txt before raising, which the runtime recorded as FAILED. A retry under the same key returned that recorded FAILED without calling the adapter again. That is Chapter 22’s decision to cache failure rather than silently re-execute, and here it prevented a second marker.
But look at what the runtime observed for that action: 77e25726…, the hash of paragraph.txt. The resolver was wired to one file, so the recorded observation never saw the effect on the other. The only evidence that the crash changed anything is the preserved crash_final.txt, which the independent verifier reads directly. This is the warning Chapters 19 and 21 both recorded — the resolver is trusted configuration, and the runtime does not validate its coverage — turning up inside a composition that looked fully instrumented. 1
The run records the branch as UNRESOLVED and stops. Nothing in the ledger can say whether retrying is safe, and nothing in the runtime offers an action-level path to find out. Reliability, at this boundary, means refusing to infer “no effect” from a failure status. It also means admitting that the refusal is the whole recovery plan.
A second experiment: auditing the joints
The capstone has a blind spot: a joint that held because the producer cooperated looks exactly like a joint the runtime enforces, for as long as the producer cooperates. Reconstructing a link proves it is visible, not that anything would refuse a violation.
So the joints were audited as their own experiment. The protocol was registered and committed before the harness existed — including five predictions, so a surprise would be visible as a surprise — fixing one question per relationship and one vocabulary for the answers. 2
Is this arrow enforced by the runtime, derived from durable facts, merely recorded, conventional, or absent?
| Class | Meaning |
|---|---|
| ENFORCED | The runtime derives or checks the relation from durable facts and refuses a violation |
| DERIVED | Deterministically reconstructable from durable facts, but not an execution gate |
| RECORDED | Persisted, but trusted from the caller or producer rather than independently checked |
| CONVENTIONAL | Holds only because callers currently cooperate |
| ABSENT | The runtime does not represent the relationship at all |
That taxonomy is the transferable part of this chapter. Most architecture diagrams draw one kind of arrow. Real systems have at least five, and the difference between them is the difference between a guarantee and a habit.
Thirteen joints were probed, each with a deliberately hostile case: a caller widening its own grant, a decision that says CHECK while the caller writes, a worker that reports success and does nothing, a verifier that misreports what it examined, a check labelled with bytes it never opened, an acceptance citing a stale pass, a replay after the grant narrowed. No repairs were made while the evidence was gathered, and the baseline was frozen — classifications, hostile-case outputs, the subject commit — before a line of it was fixed. 10
Of those thirteen, six came back ENFORCED. Three were RECORDED, two CONVENTIONAL, two DERIVED. The answer to the governing question was not what the capstone’s successful reconstruction had suggested:
A second process could explain every transition from the durable record. It could not enforce five of them, and a sixth left a verification attempt undiscoverable from the claim it targeted.
The arrows that only looked real
Six joints came back weaker than the components they connected. All five registered predictions held, which is worth saying plainly: the audit was not vindicated by surprise. The most important finding was the one nobody had predicted at all.
A reported success was read as an observed effect. The runtime already held both facts. It read the state before an action and again after, recorded both, and never compared them. A worker that changed nothing and reported SUCCEEDED produced this:
worker report succeeded the worker's word
effect state OBSERVED derived from that word
runtime's readings pre == post held, and never compared
a check of the intended post-condition FAIL
The projection’s own explanation said it had completed “with the runtime’s own observation of the resulting state”. It had an observation. It had never compared it to anything. 10
That is Chapter 19’s principle failing inside the composition test that was built to look for exactly this — and failing in the one place a lie would show, since FAILED and DENIED were handled correctly. Isolated tests passed because each asserted what the component reported, not what two components together implied.
Authority was resolved from the record for effects, and taken from the caller for acceptance. Chapter 20’s grant resolution was real: a caller could not widen a recorded directive by passing a broader Authority, and a forged grant changed nothing. Acceptance used the other standard. A directive granting only WRITE, and a caller passing Authority({ACCEPT}), completed the task. The audit also found the prerequisite that made the gap hard to see: task.accepted recorded no directive at all, so an acceptance could not be re-resolved against any chain after the fact.
Three more followed the same shape — a label standing in for a reading, a record standing in for a gate. Acceptance compared the caller-written target string on a check with the artifact being accepted, so a check whose command was sys.exit(0) carried the artifact’s name, passed, and completed an acceptance.
The scheduler decided CHECK, the caller executed a write, and the two merely coexisted: an action’s record named no decision in either direction, and an operation with no decision at all ran identically.
And an ERROR, which correctly produced no evidence against a claim, produced no claim-side record either — so “no evidence was produced” and “no attempt is discoverable” were the same fact.
Which arrows moved
Each weak joint was then repaired on its own branch, and the frozen baseline was re-run after each one. The point of the frozen baseline is not the repairs; it is that an unrelated joint moving would have been visible immediately.
| Joint | Built in | Baseline | Now |
|---|---|---|---|
| directive → action authority | Ch 20 | ENFORCED | ENFORCED |
| directive → acceptance authority | Ch 14, 20 | RECORDED | ENFORCED |
| decision → execution | Ch 28 | CONVENTIONAL | ENFORCED |
| action request → effect | Ch 19 | RECORDED | DERIVED |
| effect → recovery / reconciliation | Ch 16, 22 | ENFORCED | ENFORCED |
| observed state → verification binding | Ch 21 | ENFORCED | ENFORCED |
| verification command → verdict semantics | Ch 21 | ENFORCED | ENFORCED |
| verification → claim linkage and admissibility | Ch 18, 21 | DERIVED | ENFORCED |
| verification → acceptance | Ch 14, 21 | RECORDED | ENFORCED |
| acceptance → completion | Ch 14 | ENFORCED | ENFORCED |
| completion / replay → current authority | Ch 22 | ENFORCED | ENFORCED |
| durable state → next-operation decision | Ch 28 | DERIVED | DERIVED |
| check → artifact identity | Ch 14, 21 | CONVENTIONAL | ENFORCED |
| human gate → acceptance authority | Ch 20, 28 | (raised later) | ENFORCED |
Seven arrows moved. Seven did not, and that is the other half of the result: no repair disturbed a joint it was not aimed at. 12
What the repairs did, in the vocabulary the book has been building: 12
- Report, observation and verification became three answers, not one. The record holds what the actor said (
SUCCEEDED/FAILED/DENIED), what the runtime read itself (CHANGED/UNCHANGED/UNAVAILABLE), and what it therefore supports (OBSERVED/UNKNOWN/REPORTED/NONE). A reported success whose observed scope did not change isUNKNOWNand asks for reconciliation; one nothing could observe isREPORTED, the actor’s claim and nothing more. CHANGED does not mean the intended change was made — diligent and careless workers both produce it, and only the check tells them apart. - Acceptance resolves its grant from the task’s own recorded directive, never the caller’s argument, which is recorded as an inert claim.
- A check is given the artifact it claims to examine, resolved from the store and digest-verified; a command check that never receives it is an ERROR, not a verdict. That establishes what the verifier was given, never what it read: artifact binding is not artifact adequacy.
- A scheduler-governed operation carries the decision that selected it, checked for task, operation class and freshness. External operations are still allowed and recorded as such; what they cannot do is pass for governed.
- Every verification attempt is navigable from the claim, as a reference to the check rather than a copy of its verdict. An ERROR is discoverable and still moves nothing.
One finding was not repaired, because refusing was the right behaviour. A task whose directive never granted ACCEPT reaches a state from which completion is unauthorized, and the scheduler says ASK_HUMAN. The gap was that no lawful move existed from there. The answer was not an accept-anyway call but a durable authority transition: a superseding directive, recorded with its actor and reason, after which the ordinary acceptance rule runs unchanged. Human intervention may change authority; it does not bypass authority. Delegation may only narrow a parent; a transition records a new external decision and may add or remove — and the runtime still cannot tell you that the named human was entitled to make it. 13
One protocol step in the capstone did not happen as written, and it is worth keeping visible. The frozen protocol ends the happy path with “reopen projects COMPLETED”, and the producer’s result field is named completion_after_reopen, but it projects completion on the same live handle: no fresh CodeAI handle was opened on the files. What that bundle shows across a process boundary is weaker — its stdlib verifier reads the checkpointed ledger and finds a completion caused by a valid acceptance.
A true restart was demonstrated by the earlier 2026-09-13 bundle, and again by the audit, whose lifecycle reopens the ledger and reprojects the same status and the same next operation. The label stays as found, and the claim shrinks to fit. 3 14 12
Checking it without trusting it
The independent verifier is stdlib only. It never imports CodeAI and never reads the producer’s summary for its predicates: it opens the preserved SQLite ledger and the fixture bytes and recomputes. Its predicates fall into four groups:
| Group | Recomputed from records and bytes |
|---|---|
| Grant | Both directives present; child’s causation_id is the parent’s event; child capabilities ⊆ parent’s |
| Effect | Replay fingerprint rebuilt from both request payloads (capability, instruction, canonical-JSON payload hash, precondition, adapter); exactly one original and one replay under the edit key; original’s observation = SHA-256 of the after-bytes; denial carries the pre-edit hash; both crash completions FAILED with the marker in the preserved bytes; one refusal naming instruction |
| Verification | Bound check requested and carried the after-hash; stale and raising checks are ERROR, the latter with its error preserved |
| Acceptance | One acceptance for t-comp referencing both check events; completion caused by it, carrying the after-bytes’ hash; no acceptance on the crash or bad-proposal tasks; call receipt present |
Nine seeded corruptions are applied to temporary copies: a forged completion cause, an altered bound observation, a deleted child directive, a changed instruction on the replay request, a deleted refusal, a replay rewritten as a second physical effect, the raising verifier’s ERROR rewritten as PASS, an acceptance forged for the crash task, and a deleted call receipt. All nine are caught, a re-run on a copy during editing printed the same result, and each corrupted copy’s problem list trips exactly the predicate it targets. 15
This verifier reads records, where Chapter 21’s demo checked a producer summary, and its limits still matter. It is written for this fixture’s identifiers and would need rewriting for another task. It confirms that the twelve-question report has twelve answers, but those answers are the producer’s prose rather than the verifier’s derivations. It checks that required event kinds are present, not their counts — the counts cited in this chapter come from reading the ledger — and it checks that a call receipt exists, not that no second call does. 1
The rest of the book, in its place
The capstone does not reprint Stage 29B or the P-series; it spends them where they belong.
Stage 29B is why the capstone holds its single call without apology — and why that restraint is not a verdict for cheap-first execution. On forty synthetic extraction items, the ladder accepted 33 outcomes (32 correct, 1 wrong) against top-first’s 31 (31 correct, 0 wrong). It was cheaper per accepted outcome under the preregistered person-cost scenario, and it still failed its frozen adoption rule because it added an accepted-but-wrong outcome, A05. Top-first avoided A05 through a token-exhausted, unusable output rather than demonstrated better judgment. The capstone inherits the discipline, not a winner. 16
A05 also explains why the capstone’s checks are so narrow. The grounding checker there passed an answer that was grounded and wrong; checks that test only exact byte criteria can be adequate to those criteria because they claim nothing else. The P-series travels as method — matched arms, frozen rules, no promotion without earning it. The capstone promotes nothing, selects nothing by oracle, and prices nothing, and it invokes no independent proposals, so nothing Chapters 23–27 measured about variety is retested here. The model-router challenger stays where Chapter 28 left it: specified with its falsifier, but still unfrozen and unrun. Nothing in this chapter depends on its outcome, and nothing here shows the deterministic policy would win it. 17 18
Two outside papers earn brief places, and only as interaction design. Chaining model steps so people can inspect, alter, and test intermediate work is the deployment form of this chapter’s trace: transparency through preserved intermediates rather than explanation (Wu et al., 2022). Guidelines for human–AI interaction ask systems to make clear what they can do, support efficient correction, and scope services when uncertain; the stop branches are one concrete way to meet that, since each ERROR, DENIED, and UNRESOLVED hands the person a named state instead of an automated guess (Amershi et al., 2019). Both mappings are the book’s; neither paper validates CodeAI.
What it does not establish
The third term of the chapter’s distinction stays out of reach: composition holding on this exercised path ≠ a production-ready system. Specifically:
- Neither autonomy nor model behavior. One disposable task, one artifact, canned cognition, and legacy claims leave real model variance and Stage 18 claim standing unexercised. 1
- Not correctness. PASS means the frozen byte criteria held; the criteria are the ceiling of what was checked.
- Not enforcement of every link, and ENFORCED is not truth. A joint classified ENFORCED means the runtime refuses a violation of that relationship; it says nothing about whether the world is as recorded. The durable state → decision joint stays DERIVED by design: reconstructable, not a gate. Effect → acceptance is no longer carried by a hash equality — an acceptance may now name the action it accepts, and that claim is enforced — but naming one remains optional, so an acceptance that stays silent about its effect basis is making no claim to enforce. 7 12 8
- Not recovery. There is no action reconcile path, and observation coverage is whatever the resolver was pointed at. The crash branch preserves disagreement; recovery is a person. 1
- Not exactly-once. Sequential replay suppressed a known duplicate; concurrency and crash windows remain outside what was shown. Reopen was not exercised in the capstone bundle, though the audit’s lifecycle does reproduce status and next operation from a reopened ledger. 1 12
- Not a routing result. In the capstone the deterministic policy answered two caller-posed, unledgered questions; the router challenger is still unrun. 11
- Not proof that a check was read, a person was who they said, or a decision still holds. Artifact binding establishes supply, not consumption or adequacy. Actor labels remain attribution, and an authority transition is one recorded decision, not an approval workflow with signatures, quorum, scope or expiry. Decision freshness is a whole-state digest — conservative enough that any recorded operation expires an outstanding decision, and coarse enough that it cannot tell a relevant change from an irrelevant one. 12 13
- Not identity or tamper-evidence.
human-approveris a string, the ledger is unsigned and single-writer, and the ledger is the trust boundary. Unknown-provenance policy and sandboxing (Chapter 23), and the pricing gaps of Chapter 28, ride along untouched. 19
Do this now
One hour. Run one task end to end, then find the joints you did not enforce.
- Pick a disposable file task with byte-checkable criteria. Freeze the criteria, the authority labels, and the acceptance rule before running.
- Walk it through grant, one model call, an authorized effect, the runtime’s own observation, a check bound to that observation, acceptance, and one exact-duplicate replay. Count model calls, verifier invocations, and physical effects separately.
- Break it four ways: remove the grant, stale the check’s target, reuse the key with a changed instruction, crash between effect and completion. Confirm each stops in its own durable record.
- For every pair of adjacent records — decision and action, grant and execution, effect and acceptance — ask whether the runtime refuses a mismatch or merely lets you reconstruct one. Write the second kind down. That list is your architecture’s real perimeter.
- Hand the ledger and the bytes to someone who has not seen your summary. What they cannot reconstruct is your next gap.
If you are building with an assistant:
Compose one task from directive to acceptance with every transition durable.
Invoke the model only where proposing earns it; decide, check, and gate with
deterministic code and labeled human authority. Record the runtime's own
observation of every effect, and state what the resolver covers. Bind each
check to one observed state. Replay only on matching fingerprint plus current
authority. Stop in durable vocabulary: denied, stale, conflict, error,
unresolved. Then classify each joint between records as enforced, derived,
recorded, conventional or absent, and verify the chain with a reader that
imports none of your code and must catch seeded corruptions.
Failure modes
- Trusting a joint because the records line up. Matching identifiers let you reconstruct a link; they do not prove anything enforced it.
- Reading enforced as true. An enforced joint refuses a violation of one relationship. It does not make the record an account of the world.
- Instrumenting one file and calling it observation. The resolver sees what it was pointed at, and a real effect elsewhere leaves no trace in the record.
- Resolving the crash gap in prose. UNRESOLVED is a state of the process, not a sentence awaiting polish.
- Letting a field name carry a claim.
completion_after_reopendid not reopen anything. Read the code that produced a result, not only its label. - Letting the verifier’s scope drift. A reader that checks twelve answers exist has not checked what they say.
What this chapter established
- The model provides cognition; it does not own the process. One selectively invoked call sat inside a process that recorded intent, context, claim, authority, effect, observation, verification and acceptance around it.
- Component correct ≠ composition holds ≠ deployment ready. Every mechanism passing its own test did not establish that the handoffs line up. Composition needed its own test, and passing it still says nothing about production.
- The arrows are contracts too. An effect must hand its observed state to the check, the check must bind to that state, acceptance must cite the checks that ran on the accepted bytes, and a replay must return history without a second effect. Every one of those arrows has a strength — enforced, derived, recorded, conventional, absent — and a diagram that draws them all the same way is hiding the difference.
- A reliable process stops honestly. Denied, stale, conflict, error, failed check and unresolved are correct outcomes when continuing would be unjustified.
What CodeAI showed. The capstone passed its frozen reconstruction criterion on the exercised path: six stop branches remained visible in durable records, and a stdlib reader reconstructed the named links from ledger and bytes and caught 9/9 targeted corruptions. 1 That is evidence of reconstructability under this fixture, not enforcement of every joint. The frozen audit then showed what neither isolation nor reconstruction could: of the thirteen joints it classified, five were recorded or conventional rather than enforced, a sixth hid errored verification attempts from the claims they targeted, and the runtime was reading a worker’s word as its own observation. 10
After five repairs and one authority-transition extension, seven arrows had moved and seven had not; each recheck showed only the targeted row or rows moving. 12 13 Deployment readiness — authenticated identity, containment, concurrency, scoped and expiring grants — remains a separate engineering problem.
Coda: what Applied AI means now
The book opened with a model in a chat box and an operator doing the engineering in their head. It closes with that engineering moved into software, and with a precise account of how far the move went.
Useful applied AI is ordinary software with intelligent boundaries placed selectively. Deterministic code where the rule is known. A model where proposing, interpreting, or exploring is actually useful. Chapter 28 keeps operation choice separate from model choice; Stage 29B then measured one escalation policy after the operation was already “get an answer”, under an adoption rule frozen in advance. Evidence around both. Authority around effects. Verification before acceptance, with independence, adequacy, and binding judged separately. People where intent, authority, or frontier judgment genuinely belong.
None of that makes the model less important. It makes the model replaceable, which is what lets the rest of the system keep its promises when the model changes underneath it.
The capstone shows those pieces joining on one path, and it shows the places where joining is still a convention rather than a guarantee. That is not an anticlimax. A process that can say denied, stale, conflict, error, and unresolved — and can tell you which of its own links it does not yet enforce — is one you can extend without guessing. The slider stays in the human hand: per operation, per risk, per evidence. The runtime’s job is to hold the process steady underneath it and to be honest about where its grip ends.
There is one more consequence, and it is not an engineering result. The architecture that makes the model replaceable is ordinary software, and ordinary software is becoming cheaper to build. The frame, the records, the checks and the policies do not have to be built for everyone. They can be built around you.
Next
The capstone answered the book’s question for one process. The last chapter asks the reader’s question: when software can be built around one person, what should that person build, and how does it keep improving without losing the discipline this book built?
Continue with Your Applied AI.
References
- Tongshuang Wu, Michael Terry, and Carrie J. Cai. AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts. CHI, 2022. arXiv:2110.01691.
- Saleema Amershi et al. Guidelines for Human-AI Interaction. CHI, 2019. Microsoft Research.
Implementation sources: the original composition run used the CodeAI baseline plus uncommitted changes whose file list is recorded in the bundle manifest; content identity was not hashed at run time, so its source identity remains the recorded base plus manifest rather than a clean commit hash. The code excerpt is reduced from that executed producer. Current-source statements about the later seams were checked against src/codeai/actions.py, src/codeai/authority.py, src/codeai/process_state.py, src/codeai/governance.py, src/codeai/verification.py, src/codeai/acceptance.py, and the corresponding entry points in src/codeai/runtime.py; scheduler behavior comes from src/codeai/scheduler.py. Frozen follow-on evidence is recorded in experiments/W1-composition-prereg.md, W1-composition-results.md and .json, W1-R1-effect-observation.md, W1-R2-authority-symmetry.md, W1-R3-artifact-binding.md, W1-R4-decision-execution-binding.md, W1-R5-claim-attempt-trace.md, W1-E1-authority-transition.md, and W2-2-acceptance-basis.md. Original capstone evidence remains under experiments/applied-ai/evidence/capstone-composition/ (frozen protocol, producer, results, report, preserved ledger, artifacts and fixture bytes, manifest, stdlib verifier with nine seeded corruptions) and experiments/applied-ai/evidence/capstone/ (2026-09-13 demo with fresh-handle restart). No historical evidence was modified. Footnote prefixes: s inspected source, m pinned run, d preserved demo, r frozen report. Open future work is stated in prose, not footnotes.
Measured run:
experiments/applied-ai/evidence/capstone-composition. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎Frozen protocol:
experiments/W1-composition-prereg.md, committed before the harness existed, with its five predictions. ↩︎ ↩︎Measured run:
experiments/applied-ai/evidence/capstone-composition/run_composition.py. ↩︎ ↩︎Source inspection:
src/codeai/runtime.py(Runtime.execute_action). ↩︎Source inspection:
src/codeai/runtime.py(Runtime.run_check). ↩︎Source inspection:
src/codeai/acceptance.py(accept_task). ↩︎Source inspection:
src/codeai/acceptance.py(AcceptanceRequest). ↩︎ ↩︎Measured run:
experiments/W2-2-acceptance-basis.md, case-A probeexperiments/W2-2-case-a-results.json;src/codeai/acceptance.py(effect_basis), matrix intests/test_acceptance_basis.py. ↩︎ ↩︎ ↩︎Source inspection:
src/codeai/runtime.py(Runtime._replay_action_result, Runtime._check_action_replay_fingerprint). ↩︎Measured run:
experiments/W1-composition-results.mdand.json(baseline; subject runtimef7d2910), harnessexperiments/wave1_composition_audit.py, branchaudit/wave1-composition. ↩︎ ↩︎ ↩︎ ↩︎Source inspection:
src/codeai/scheduler.py(decide_next_step). ↩︎ ↩︎Measured runs:
experiments/W1-R1-effect-observation.mdthroughW1-R5-claim-attempt-trace.md, each with its own recheck of the same harness against the frozen baseline. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎Measured run:
experiments/W1-E1-authority-transition.mdandW1-E1-composition-check.json(subject runtimec1dd92a); seam documentation underdocs/seams/. ↩︎ ↩︎ ↩︎Unpinned demonstration:
experiments/applied-ai/evidence/capstone. ↩︎Measured run:
experiments/applied-ai/evidence/capstone-composition/verify_composition.py. ↩︎ ↩︎Measured run:
experiments/applied-ai/evidence/execution-ladder/2026-09-14-7a0d43b. ↩︎Measured run:
experiments/applied-ai/evidence/p-series-analysis. ↩︎Report:
docs/applied-ai/ch28-router-experiment-design.md. ↩︎Source inspection:
src/codeai/acceptance.py. ↩︎