Restart Is Not Resume
Restarting reopens the files. Resuming means knowing, from recorded facts alone, what was attempted, what is unresolved, whether an external effect may already have happened, and what is safe to do next. Built in CodeAI, killed at durable checkpoints and reopened by other processes, and contrasted with a naive restart that sent the same provider request twice.
Part 3 — Give Intelligence a Runtime
Externalize working memory
When an AI process dies partway through, the usual fix is to restart it and re-run whatever did not finish. For a model call, that can mean paying for the same request twice. For an agent that sends email, opens tickets, charges a card or edits files, it can mean doing something twice that cannot be undone.
The capability you actually need is a different one:
restart ≠ resume
Restarting means the files survived and a process can open them again. Resuming means another process can decide, from recorded facts alone, what was being attempted, what inputs it had, what failed, what succeeded, what remains unresolved, whether an external effect may already have happened, and which next step is safe. The deciding question is not “did this step finish?” It is could the effect have happened?
By the end of this chapter you will be able to make a model-calling workflow resumable: mark the point just before each external effect, project a safe next operation from the record, resume only where no effect was possible, and refuse, rather than guess, where one might have been. This is memory for the process, not memory for the model, and the two are easy to confuse.
The step that didn’t finish
A process was reviewing a paragraph. It had compiled its context, recorded the call’s manifest, and sent the request. While the provider was still working on it, the process died.
The standard recovery is familiar: restart, find the step that didn’t finish, and run it again. Someone did exactly that in this chapter’s experiment, on a copy of the dead process’s files, and the step re-ran to completion.
The synthetic provider’s receipt log then showed two requests, byte-for-byte the same request under the same idempotency key, from two different processes. The experiment measured a duplicated provider-side effect, not a duplicated bill. On a provider that charges or consumes quota per served request, this is the kind of duplicate that can also duplicate cost.
The ledger had not been silent. Before the crash, it had recorded that an attempt started with nothing observed afterwards, a meaning the restart never asked about.
Can another process continue the work without reconstructing what happened from human memory?
Restart is not resume
Two different capabilities hide behind the word “recovery”.
Restart means the durable state survives the process. The ledger is still there, the artifacts are intact, and another process can open them. CodeAI has had that since Chapter 11.
Resume means another process can decide what to do next from what is recorded, without asking the person who was watching. It has to answer seven questions:
| A resuming process must know | Where CodeAI now records it |
|---|---|
| What was being attempted? | call.requested (the full call specification, including its context) and context.compilation_requested |
| What inputs were available? | The offered candidate inventory recorded before compiling |
| What was required? | Each candidate’s required flag and the explicit required IDs |
| What failed? | context.compilation_failed, call.preparation_failed, or a call’s decided status |
| What reached a recorded terminal state? | context.compiled; call.completed together with its adopted call status; and Chapter 14’s acceptance and completion |
| What remains unresolved? | Every compilation or call whose next operation is not “none” |
| Which next operation is safe? | A named next operation, whether a provider effect may have happened, and whether repeating risks a duplicate |
A database that survives a crash answers none of these by itself. They are questions about the work, not the storage.
There is a related idea that is easy to confuse with this one. MemGPT treats a model’s limited context like an operating system treats memory, moving information between fast and slow tiers so the model can work with more than fits in its window (Packer et al., 2023). That is memory for the model: what it can draw on while it thinks. This chapter is about memory for the process: whether the system knows what it was doing when it stopped. A model with excellent long-term memory can still be run twice by a process that doesn’t.
Effects that cannot be taken back
Elnozahy, Alvisi, Wang and Johnson’s survey of rollback recovery separates a system from what it calls the outside world (Elnozahy et al., 2002). Processes inside the system can be rolled back and replayed. The outside world cannot: in their examples, a printer cannot unprint a character, and a cash machine cannot take back money it has dispensed. In their words, the outside world cannot be relied on to roll back.
That produces what they call the output commit problem. Before a system sends something to the outside world, it must make sure the state it sent from will survive any later failure. Otherwise, after recovery, the system may send it again, or behave as if it never did.
A model provider is outside the world of a CodeAI process. The mapping is this chapter’s, not theirs. If a request was served, the provider may already have performed work, consumed quota or billable capacity, and produced a generation even if the response never arrives. Recovery on this side cannot undo that provider-side effect.
Lampson’s hints for system design state the same requirement from the logging side (Lampson, 1983). Log updates so the log records the truth about an object’s state, in entries that can be re-executed. Make actions atomic or restartable, where a restartable action can be partially executed any number of times without changing the result. His example: storing a set of values is restartable; adding one to a variable is not.
For recovery purposes, a provider call belongs on the non-restartable side unless some external contract actually guarantees duplicate suppression for the request. Reissuing the same bytes can otherwise create another request, another generation, and another consumption of whatever resource the provider meters.
So the question that governs resume is not “did this step finish?” It is: could the effect have happened? CodeAI’s call path answers it with an ordering chosen in Chapter 11. attempt.started is committed to the ledger before the request is sent, and the observation is committed after the response arrives. Between those two events, the honest answer is “unknown”.
What CodeAI already recorded
Earlier chapters had built most of the trail, one boundary at a time. Chapter 11 writes intent (call.manifest) before any attempt and records each attempt, and the interpretation work after it records interpretations and attempt decisions. Chapter 12 preserves the transport observation. Chapter 14 records acceptance and completion. And Stage 15B, built after Chapter 15, lets the selected context reach the request for calls that opt in, binding its rendered bytes to the manifest before any effect.
Two gaps remained. Chapter 15 had found that the offered inventory and failed compilations lived only in the experiment harness. And nothing turned the trail into an answer: a later process could read every event and still have to work out for itself what they meant.
The work-state projection below closes both, built around the question of whether an effect could have happened.
The projection
CodeAI now has Runtime.work_state(task_id). It is a pure projection: it reads the ledger and appends nothing. Each call is classified by the last evidence recorded for it:
| Last evidence for the call | Stage | Provider effect | Next operation | Acted on automatically? |
|---|---|---|---|---|
call.requested only | requested | none | start the call | yes |
call.manifest, no attempt started | manifest recorded | none | start the call | yes, if the recorded intent reproduces |
attempt.started, nothing after | effect unknown | unknown | reconcile the effect | never |
| Observation, no interpretation | observed | observed | reinterpret the preserved bytes | not yet |
| Interpretation, no decision | interpreted | observed | re-derive the decision | not yet |
| Decision to retry, next attempt not started | retry decided | observed | start the next attempt | not yet |
| Decision, no call completion | decided | observed | finalize the call | not yet |
call.preparation_failed | preparation failed | none | fix the inputs | no |
call.completed | completed | observed | none | — |
The table is a ladder in code: each rung asks whether a particular event exists. Reduced from the current source, with the reason strings shortened:
if "call.completed" in kinds:
stage, next_op = "completed", NONE
elif "call.preparation_failed" in kinds:
stage, next_op = "preparation_failed", FIX_INPUTS
elif "call.manifest" not in kinds:
stage, next_op = "requested", START_CALL # no manifest, so no attempt can have started
elif not attempts:
stage, next_op = "manifest_recorded", START_CALL # manifest recorded, no attempt: no provider effect
else:
seen = kinds_for(last_attempt)
if not seen & {"attempt.observed", "attempt.interpreted",
"attempt.retry_decided", "attempt.completed"}:
stage, next_op = "effect_unknown", RECONCILE_EFFECT
provider_effect, duplicate_effect_risk = "unknown", True
elif "attempt.interpreted" not in seen:
stage, next_op = "observed", REINTERPRET
...
It derives the stage solely from which events are present in the ledger, using no clock, provider query, or memory held outside it. The dangerous rung is the fifth: an attempt that started and left no trace after it. The projection does not guess whether the provider served it; it records “unknown” and a duplicate risk, and that is what resume_call refuses to override.
Compilations get the same treatment. A compilation that was requested and never finished projects “recompile”, which is always safe because compiling has no external effect. A compilation that failed projects “fix the inputs”, together with the inventory it was offered and any required IDs that were missing.
The task gets a next operation too. A completed task projects none. An acceptance whose matching completion has not yet been appended projects “repeat acceptance”. Otherwise the first unresolved call wins; if no call is unresolved and at least one completed call has adopted status succeeded, the task projects Chapter 14’s “check and accept”. With no succeeded call, it projects “start the call”.
Runtime.resume_call(call_id) acts on the two pre-effect call states whose projected next operation is “start the call”: call.requested with no manifest, and a recorded manifest with no attempt started. It rebuilds a new linked call under the same idempotency key. If a manifest exists, it reproduces the recorded request and checks the context package identity, rendered-context hash and request-body hash before acting; any drift is refused. If the process died before the manifest existed, there is no manifest against which to perform that comparison — one of the limits named later in this chapter.
For every other state it raises ResumeRefused, names the projected next operation, and appends nothing.
The handoff the chapter is about — one process dies, another continues from the record alone:
Kill it, reopen it, resume it
The demonstration was preregistered before it ran. Every phase ran as a separate operating-system process over one case directory: producing, inspecting, resuming, and restarting naively.
Two details make the evidence stronger than a test. Producers were really killed, with Popen.kill, which runs no cleanup, at durable checkpoints: immediately after a named event was committed, or while the provider held the request. And provider effects were counted outside the ledger: the synthetic provider appended an fsync’d line to its own receipt log for every request it received, so the ledger’s claims can be checked against something it didn’t write.
Outbound network connections were refused throughout, and no real provider was called.
| Case | Where the producer stopped | What a new process found | Then | Result |
|---|---|---|---|---|
| Killed during compilation | after context.compilation_requested | compilation interrupted; inventory recoverable | — | 4 events, 0 receipts |
| Compilation failed | clean exit (missing required input) | failed on never-offered; fix the inputs | — | 5 events, 0 receipts |
| Killed after manifest | after call.manifest | no effect possible; start the call | resume | 1 receipt in total; resumed call completed |
| Killed after manifest, drift | after call.manifest | same | resume with a different model | refused; nothing appended; 0 receipts |
| Killed during effect | provider holding the request | effect unknown; reconcile | resume | refused; nothing appended; receipts stay 1 |
| The same state, copied | (copy of the case above) | same | naive restart | 2 receipts, same request |
| Killed after observation | after attempt.observed | observed; reinterpret | resume | refused; receipts stay 1 |
| Clean run | clean exit | call completed; check and accept | — | 14 events, 1 receipt |
Resume once, refuse twice, and a restart that duplicated the request
Resumed. The producer was killed right after call.manifest, with 7 events in the ledger and nothing in the provider’s log. A new process projected “manifest recorded, no provider effect, start the call”. A third process called resume_call, rebuilt the request, found the same request body hash the manifest had recorded, and appended call.resumed from call-16 to a new call under key-16 before running it.
Afterwards the provider’s log held exactly one receipt, whose body hash equals the hash in both manifests. The old call now projects as superseded by the new one, the new call is completed, and the task’s next operation has moved from “start the call” to “check and accept”. The ledger grew to 17 events. Nobody had to remember which step was pending.
Refused for drift. The same interruption, resumed with a different model configured, was refused before anything was sent, because the recorded intent no longer reproduces its request body hash. The ledger stayed at 7 events and the provider saw nothing. Resuming does not mean running whatever the code would send today. It means running what was recorded as intended, or refusing.
Refused for an unknown effect. The producer was killed while the provider held the request, with one receipt already in the provider’s log and only attempt.started in the ledger. A new process projected “effect unknown, reconcile the effect, repeating risks a duplicate”. resume_call refused, the ledger stayed at 8 events, and the receipt count stayed at 1.
The naive restart. A copy of that same directory was handed to a process that did what restart loops usually do: take the recorded request and issue it again without consulting the work state. It succeeded, and the synthetic provider’s log now shows two receipts with identical request bodies, from two processes, under the same idempotency key.
Two idempotency mechanisms have to be separated. CodeAI’s local replay suppresses a request only when it can find a completed recorded call with the same key, and this interrupted call had no such completion. The synthetic provider used in the experiment did not deduplicate on the key either. So issuing the request again produced a second provider receipt. The key was recorded; no layer in this path had evidence that made the unknown first effect safe to collapse.
Observed, not automated. The producer was killed after the observation was committed. The projection says “reinterpret”: the response bytes are preserved, so the provider is not needed again. resume_call refused this too, for a different reason. It isn’t unsafe; CodeAI simply doesn’t yet automate re-running the interpreter over stored bytes. The distinction is recorded, not blurred.
Compilations that stopped or failed
Killed after context.compilation_requested, the ledger held 4 events: the task, the objective, the claim, and the compilation request. A new process recovered what had been offered (A, B and D, with their required flags, sizes and lineage) and projected “recompile”.
The compilation that failed on a missing required input was recorded, not just raised. The failure event names RequiredContextMissing. The inventory shows A (required, estimated size 12), B (required, size 0) and D (optional, size 0), and the projection derives the missing required ID, never-offered, from the difference between what was required and what was offered. Chapter 15 found exactly this information missing from the runtime.
What this is not
- Not crash atomicity. Every kill happened at a durable checkpoint — after a commit, or inside the transport — which leaves dying between an artifact write and its event, or during a database commit, still uncovered.
- Not exactly-once. A provider effect whose outcome is unknown is refused rather than resolved, since CodeAI has no way to ask the provider whether a request was served and no way to record a reconciliation decision.
- Not concurrency. Two processes resuming the same call at the same moment could both start it.
- Not automated recovery of post-effect states. Reinterpret, redecide, start the next attempt and finalize are named precisely, and left to the next chapters.
Where it is still weak
- One state resumes automatically. Only calls with no attempt started are resumed.
- No reconciliation. “Reconcile the effect” names the problem without a mechanism.
- Kills only at durable checkpoints.
- Single writer.
- No drift check when resuming from
requested. With no manifest recorded, there is nothing to compare against. - Supersession is by idempotency key only.
- Inventory sizes include estimates. A’s recorded size 12 is estimated from its serialized payload, and artifacts and claims count zero.
- Fanout compilations are not recorded through this path.
- A synthetic provider, and one fixture.
Do this now
Forty minutes. Find out whether your system resumes or merely restarts.
- Pick a pipeline or agent loop that calls a model. Kill it while a model request is in flight, restart it, and count the requests on the provider’s side: its usage page, logs or bill. Was the work done twice?
- For each step, ask whether a later process can distinguish “the effect may have happened” from “the effect did not happen”. If the record doesn’t mark the point just before the effect, the answer is no.
- List every effect outside your process: model calls, emails, payments, tickets, file writes. Mark each as restartable (setting a value) or not (an increment). The second kind must never be retried on a guess.
- Find the code that decides what to do after a restart. Does it read a record written before the crash, or does it re-run whatever “didn’t finish”?
If you are building with an assistant:
Make a model-calling workflow resumable, not just restartable.
- Record intent before any effect: the full request specification, and for
context, the offered inventory with requirements, sizes and lineage.
Record failures of compilation and preparation as events.
- Append "attempt started" durably before sending a request and "observed"
after the response is stored. Between them the effect is unknown.
- Add a pure projection that classifies each operation by its last recorded
evidence and names the next operation: start (no effect possible),
reconcile (effect unknown; never repeat automatically), reinterpret,
redecide, finalize, fix inputs, recompile.
- Resume only operations with no recorded attempt. Re-issue under the same
idempotency key as a new, linked operation, and refuse if the rebuilt
request no longer matches the recorded intent.
- Test with separate processes and real kills at each boundary. Count
provider effects with a log the provider writes, not the ledger. Include a
naive restart on a copy and show that it duplicates the effect.
Failure modes
- Treating restart as recovery. The files survived; the process still didn’t know what it was doing.
- Re-running the step that didn’t finish. An external call is not safely restartable merely because you still have the request bytes. Treat it as non-restartable unless the system that owns the effect supplies an idempotency guarantee you can rely on.
- Trusting an idempotency key without naming who enforces it. CodeAI’s local replay matches completed recorded calls; the synthetic provider in this experiment did not deduplicate the key. A key is a handle, not duplicate suppression by itself.
- No marker before the effect. Without it, “didn’t happen” and “unknown” look the same.
- Resuming what the code would send today. Where a manifest exists, reproduce and compare the recorded intent before acting. A requested-only interruption has no manifest to compare against, which is a current limit.
- Recording successes only. Failed compilations and preparations are part of the working state.
- Confusing model memory with process memory. A model that remembers is not a system that knows what it did.
- Claiming exactly-once. An unknown effect has been refused, not resolved.
What this chapter established
- Restart is not resume. A restarted program exists again. A resumed process can derive, from recorded facts, what was attempted, which inputs were offered, what was required, what failed, which calls reached terminal states, what remains unresolved and which operation the runtime permits next.
- Ask whether the effect could have happened. A served model request, like a printed page or dispensed cash, is outside the caller’s rollback boundary (Elnozahy et al.). The mapping onto model calls is the book’s: once a provider may have served a request, do not assume that issuing it again is restartable unless the external system supplies the required idempotency semantics. Between
attempt.startedand an observation, the honest answer is “unknown”. - Resume only where no provider effect was possible. CodeAI will resume a requested-only or manifest-recorded call only before any attempt started. On the manifest-recorded path it also rebuilds the request and refuses drift; on the requested-only path no manifest exists to support that check. Where an effect may have happened, it refuses and names reconciliation rather than retrying on a guess.
- Process memory is not model memory. A model that remembers more can still be run twice by a process that does not know what it was doing.
What CodeAI showed. CodeAI records compilation with its offered inventory and failures, projects each call’s next operation and whether a provider effect may have happened, and resumes only calls with no attempt started. In separate processes with real kills, and a provider receipt log kept outside the ledger, a resume after the manifest produced exactly one provider request with the recorded body hash, and the same resume with a different model was refused.
With the effect unknown, resume was refused, while a naive restart of an identical copy sent a second, identical request. After observation, resume was refused with “reinterpret”, and interrupted and failed compilations were recoverable with their inventories. Not established: crash atomicity, exactly-once effects, concurrent resume, reconciliation, or automated recovery after an effect.
Evidence notes
Independent verification. The bundle’s verifier imports neither CodeAI nor the producer. It holds its own implementation of the classification rules, re-derives every call and compilation state from the exported events, and compares them with the states each inspecting process recorded. It also requires that every inspection appended nothing, and that every provider receipt is covered by an attempt.started — an effect with no recorded intent would fail.
All 73 claims pass, including byte hashes for 424 files, and a fresh CodeAI process reopened three ledgers and projected exactly the recorded states.
Five seeded corruptions were run with the byte inventory bypassed, so only meaning could catch them:
| Corruption | Claims that failed |
|---|---|
| The projection claims the effect-unknown call is safe to start | call states re-derivation; the effect-unknown hypothesis |
attempt.started deleted from the exported events | inspection purity; call states re-derivation; receipts covered |
| A duplicate receipt added after the resume | resumed exactly once |
The call.resumed link deleted | inspection purity; call states re-derivation; resume link |
| A refusal that appended an event | refusal changed nothing |
Next
Four of the states the projection names — reinterpret, redecide, start the next attempt, finalize — share a premise. The provider’s response still exists, exactly as it arrived. Reinterpreting is only safe because the bytes were preserved before anything was concluded from them. Automating it means running a versioned interpreter over stored observations, without replacing them.
Continue with Preserve Before You Interpret.
References
- E. N. (Mootaz) Elnozahy, Lorenzo Alvisi, Yi-Min Wang, and David B. Johnson. A Survey of Rollback-Recovery Protocols in Message-Passing Systems. ACM Computing Surveys, 2002. https://doi.org/10.1145/568522.568525
- Butler W. Lampson. Hints for Computer System Design. Proceedings of the Ninth ACM Symposium on Operating Systems Principles, Operating Systems Review 17(5), 1983, pp. 33–48. https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/acrobat-17.pdf
- Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez. MemGPT: Towards LLMs as Operating Systems. arXiv:2310.08560, 2023. https://arxiv.org/abs/2310.08560
Implementation and evidence sources (CodeAI): current source inspection uses src/codeai/workstate.py (WORK_STATE_V1, project_work_state, NextOperation, CallState, CompilationState, resume_call, spec_from_payload, ResumeRefused), src/codeai/context.py (ContextCompiler.offered_candidates), src/codeai/runtime.py (compile_and_record_context, Runtime.work_state, Runtime.resume_call), and Stage 15B rendering in src/codeai/rendering.py. The pinned working-state evidence bundle is experiments/applied-ai/evidence/working-state/2026-09-13-68f4ba0/, containing the preregistration and execution record, seven cases plus the naive-restart copy, termination records, checkpoints, SQLite ledgers, synthetic-provider receipt logs, inspections, exported events, the separate verify.py, five seeded corruptions, test outputs, chapter-evidence-report.md and hashes.json. That bundle records tests/test_work_state.py with 14 tests and the then-full suite with 317 passing tests; those are evidence-run counts, not a claim about the current suite size. Producer: experiments/applied-ai/working_state_demo.py (the executed copy is pinned as the bundle’s run.py).