← Applied AI

One Runtime, Many Windows

Where does the work live when the interface closes? If state lives in the interface, you have as many processes as you have interfaces. One durable runtime should own the work's identity, record, context decisions and policy, with every chat, editor, terminal or phone a window onto it — shown by reopening the same work from separate processes, with an honest account of what is and is not built.

Part 2 — Get the Model Out of the Chat Box

A chat window is an interface with a person inside the loop doing the integration. This part takes the process out of it and gives the work somewhere else to live.

That happens in stages. The work needs a durable identity and a record that survives the interface closing. The model underneath has to be replaceable, because it will be replaced. A model call has to become a recorded process event rather than a string. That call then has to survive a change of provider dialect, its usage numbers have to mean something before anyone does arithmetic on them, and finally a successful call has to stop being mistaken for finished work.

By the end of this part the model is no longer a conversation partner. It is a component with a contract.


Tuesday, three times

Monday evening, on your laptop, you work through a design problem with a model. Forty minutes of back-and-forth. You settle three decisions and rule out two approaches for reasons that took a while to articulate.

Tuesday morning, on the train, you open the same app on your phone. It is a new conversation. You paste in what you can remember.

Tuesday afternoon, in your editor, a different assistant offers to implement something you explicitly ruled out last night, confidently, because it has never heard of last night.

Three surfaces with separate contexts and no shared memory. And the only thing carrying continuity between them is you, retyping.

        chat        editor
           \        /
  phone ── YOU ── terminal
             |
          browser

  every arrow into the middle is a person
  carrying objectives, decisions and evidence by hand

This is Chapter 1’s problem again, one level up. There, the human was the integration layer between a chat window and an application. Here the human is the integration layer between their own tools. We removed the person from inside a single exchange and left them holding everything between exchanges.

Where does the work live when the interface closes?

The diagnosis

The common reason continuity dies at a surface boundary is where canonical state lives.

Many chat apps, IDE assistants, and browser tools keep their own history or session state. When each surface owns the only durable copy of what it knows, moving between surfaces behaves like a process restart without a shared serialization of the work.

If your state lives in the interface, you have as many processes as you have interfaces.

That sentence is the architectural argument of this chapter, and everything else follows from inverting it. The surface should be a view onto a process, not the process itself. One runtime; many windows.

Notice what the problem is not. It is not that each tool has too small a context window. Give every tool in the Tuesday story a window a hundred times larger and you have three processes with larger private memories, still unaware of one another. The process itself is fragmented, and no amount of capacity inside one fragment joins it to the others.

The inversion looks like this:

BEFORE                                  AFTER

interface                               runtime
├── conversation                        ├── work identity
├── context                             ├── current state
├── state                               ├── decisions
├── decisions                           ├── history
└── history                             ├── authority
                                        └── evidence
switch interface
→ implicit process restart              interfaces
                                        → requests and views over the same runtime

The value of that move is not, primarily, “AI on more devices.” It is continuity of work independent of whichever surface you happen to be using. Devices are one consequence of that; so are restarts, crashes, a second assistant, a teammate reading the same work, and an automated job picking it up at night.

This is the same move the book has been making since Chapter 1, applied one level out: separate the thing that thinks from the thing that owns the work, then separate the thing that owns the work from the thing you happen to be looking at.

Continuation is not reconstruction

There is an obvious workaround, and it deserves to be taken seriously before it is rejected: summarize the old conversation and paste the summary into the new one.

RECONSTRUCTION                          CONTINUATION

new surface                             new surface
    ↓                                       ↓
summarize old conversation              open the same work by identity
    ↓                                       ↓
approximate state, rebuilt              authoritative record, read
                                        current state, projected

These are not two implementations of the same thing. A summary is a lossy reconstruction of a transcript, and a transcript was never a complete record of the work to begin with. A transcript records what people and models said. A process needs to know:

  • what was requested, and what counts as done;
  • what was decided, and on what basis;
  • what evidence exists, and where its bytes are;
  • what state changed as a result;
  • what remains unresolved;
  • what is permitted next, and by whom.

Some of that may be recoverable from a conversation by a careful reader. What a transcript does not give you by itself is an authoritative process record from which current state can be projected deterministically. Asking a model to summarize the transcript adds another inference step. Summarization is reconstruction. Reading the authoritative record and projecting state is continuation.

The current generation of coding tools shows the reconstruction pattern in its own engineering. A 2026 source-code study of eleven coding agents found that Codex had shipped an importer for Claude Code’s sessions and settings (Barbaste et al., 2026). Read that at the level it supports: the study documents the importer, not how much survives the import. But an importer exists because the session belongs to the tool it was created in, and moving work between tools means translating one tool’s state into another’s.

None of this makes transcripts worthless. A conversation is a genuine source of intent and observations, and it can be preserved as an artifact alongside everything else. It simply should not be the canonical state machine of the work.

What the runtime owns

Ask a precise question. When you move from laptop to phone to editor, what exactly is the same thing?

It is not the same conversation. It is the same piece of work, and that work needs a durable identity that does not belong to any interface:

project
  ↓
task / objective          — what is wanted, and what done means
  ↓
current state             — what is true about the work now
  ↓
recorded decisions        — what was settled, and why
  ↓
evidence and artifacts    — what supports it, preserved
  ↓
next permitted operation  — what may happen next, and under whose authority

A surface opens that work by its identity. It does not rebuild it from transcript fragments. Chapter 11 makes identity precise at the level of tasks, calls and attempts; this chapter needs only the principle: continuity requires something durable that identifies the work independently of the interface.

Once the work has an identity, five things that are easy to blur have to be kept apart:

TermWhat it isThe confusion it prevents
StateWhat is currently true about the workstate ≠ transcript
HistoryWhat happened previously, as recorded eventshistory ≠ context
MemoryDurable material that can inform future operationsmemory ≠ prompt
ContextThe subset deliberately supplied to one operationcontext ≠ everything remembered
UI stateWhat one interface happens to be displaying right nowUI state ≠ process state

Compactly: state is current, history is recorded, memory is durable, context is selected — and whatever a window is showing is only a window.

Be careful with the phrase “source of truth,” because it hides a distinction the runtime depends on. In CodeAI the ledger is an append-only record of events; its SQLite implementation deliberately has no update or delete path. Current state is not stored as a separate mutable truth. It is a projection: a deterministic reading of the recorded events. The work-state projection Chapter 16 builds reads the ledger and appends nothing. So the precise statement is: the record is authoritative; the state is derived from it. That is why two processes that open the same record see the same state.

And “one runtime” is a logical claim, not a physical one. It means one logical owner of canonical process state and policy. It does not mean one operating-system process, one server, one database node or one machine. A runtime can be distributed, replicated or split across services later. What must stay singular is the truth about the work, not the hardware that holds it.

What a surface becomes

If the runtime owns the work, a surface is left with two jobs, and neither of them is keeping the truth:

SURFACE SENDS       requests  — create this, run that, approve this
SURFACE RECEIVES    views     — projections of the work, rendered its own way
RUNTIME OWNS        the canonical record, the state derived from it, and the policy

That is what “many windows” means concretely. A terminal can print a run’s calls as lines of text; an editor can underline the sentence a claim is about; a phone can show only what needs your approval. Each renders the same durable process differently, and none of them owns a separate version of it. CodeAI’s command line already works this way in miniature: codeai run create and codeai call are requests, while codeai run show, codeai context show and the work-state projection are views computed from the ledger.

The less obvious consequence is about authority. An interface-centered system tends to derive permissions from where you are standing:

opened in the editor  → therefore editing is permitted
opened in the chat    → therefore only discussion

That is the wrong source of authority, and with several windows onto one process it becomes incoherent: the same work, opened in two places, would have two different sets of permissions. The runtime has to decide from what travels with the request — the operation requested, the actor, the work’s identity, the authority granted, and the policy in force — and only then may the effect be performed.

Permissions cannot be a property of a window when several windows operate on the same process. CodeAI’s policy layer refuses an operation whose granted authority lacks the capability, whatever asked for it; Chapter 20 builds the full mechanism, including the separation of capability from authority.

What lives where

One picture, answering one question:

    flowchart TB
    subgraph S["surfaces — requests in, views out"]
        S1["chat"]; S2["editor"]; S3["terminal"]; S4["phone"]; S5["browser"]
    end
    subgraph R["runtime — one logical owner of the work"]
        ID["work identity"]
        ST["current state<br/><i>projected from the record</i>"]
        LE["durable record<br/><i>events, artifacts, claims</i>"]
        PO["policy and authority<br/><i>review, model choice, permissions</i>"]
    end
    CC["context compiler<br/><i>selects, records exclusions</i>"]
    M["model call<br/><i>a replaceable participant</i>"]
    S --> R
    R --> S
    R --> CC --> M
    M -->|"result recorded"| LE
  

Read it once and notice what is not in the middle. There is no conversation. The transcript is not a component. The person supplies intent and holds authority through whichever window is open; the model is one participant, reached through a compiled context; everything that must survive the window closing lives in the runtime.

Context is selected, not forgotten

The earlier version of this design called one of its properties “never loses context.” That phrase is attractive and, read literally, wrong. A system with a finite, priced context window omits things from individual calls all the time. Retention and context selection are separate decisions: material can remain in the durable record without being supplied to a particular operation.

Memory is durable. Context is selected.

The difference is easy to see:

Three reasons make selection mandatory, and they get stronger in order.

The first is dated but instructive. Liu and colleagues measured how 2023-era models — GPT-3.5-Turbo, Claude-1.3, MPT-30B-Instruct and LongChat-13B, with Llama-2 in a follow-up analysis — used long inputs on multi-document question answering and key-value retrieval. Performance was often highest when the relevant information sat at the beginning or end of the input and degraded when it sat in the middle: a U-shaped curve. The shape depended on the model — the 7B Llama-2 models showed only recency bias, the larger ones the U — and on the task: Claude-1.3 retrieved key-value pairs nearly perfectly at every length tested (Liu et al., 2024). Newer models may do much better. What the result established, and what does not expire, is narrower: a larger context window is capacity, not a guarantee that everything placed in it will be used correctly.

The second is Chapter 6’s. Context is paid for per token, on every call, and a steady-state habit of sending everything is a recurring cost with no natural ceiling.

The third survives even perfect long-context models. Suppose a future model uses every token of an enormous window flawlessly. You would still need to know, later, what a particular call was given, what it was not given, and why — to reproduce a result, to explain a decision, to keep a sealed evaluation blind (Chapter 23), or to keep material a caller is not authorized to see out of a request entirely. Those are properties of the process, not of the model’s attention. That is why the context compiler records its decisions rather than merely making them, and why this chapter’s argument does not depend on context windows staying small.

The model will fill what you fail to record

There is an experience that seems to contradict everything above, and anyone who works with these tools has had it. You come back to a project weeks later, through a different tool, give the model a few sentences, and within a couple of turns it seems to be back on the project. Nothing was compiled for that call. Nobody opened the work by identity. It picked up anyway.

It is worth being exact about what happened, because “the AI remembered” is usually the wrong explanation, and the right one changes what a runtime should keep. A durable record does not influence a model merely by existing; this chapter has spent several sections establishing that. The causal path looks like this:

durable project record
  ↓
retrieval / context selection     what something chose to bring forward
  ↓
actual call context               the selection, plus your words and whatever the interface adds
  + model's learned defaults      what work of this kind usually looks like
  + post-training and policy      tuning, system instructions, provider policy
  ↓
a plausible continuation

The record is the source; the selected context is the input. The model also brings learned statistical regularities from training and post-training. When selected context omits a project-specific fact, the model may lean more heavily on those regularities rather than emit an obvious blank. Some apparent continuity can therefore come from recovered project material and some from inference. From the fluent answer alone, the two are not labeled.

This book’s own construction offers an anecdotal example. Assistant sessions sometimes resumed weeks-old work quickly after reading notes, outlines, concept files, and commit history, while also inferring unstated structure. That is the author’s experience, not a measurement, so the engineering conclusion stays narrow: do not treat a plausible continuation as evidence that the relevant project state was recovered.

Convention and exception

Imagine joining a conventional Python web service with much of its context missing. From familiar framework structure, ORM models, and tests, a model may infer likely conventions. That can be useful onboarding, but inferred convention is not recovered state.

Now invert the situation. Suppose the project deliberately rejected a conventional architecture. If that decision is absent from the call context, a conventional proposal becomes more likely precisely because the exception is missing.

That is the failure illustrated by the Tuesday editor at the start of this chapter. Missing state need not produce nonsense. It can produce a competent continuation that is wrong for this particular project.

context carries the project's facts   → the continuation follows the project
context misses them                   → the continuation partly follows the prior

conventional work                     → that is compression and onboarding
exceptional work                      → that is drift toward convention

The prior carries the convention. The record must carry the exception.

Two funnels, not one

Call the pull a funnel, but define it before using it: a shared model behaves like a restoring force toward high-probability solutions. There are two such forces, and merging them weakens the argument.

The first is statistical. A model tends toward continuations that are probable under what it learned. The second is post-training and policy. Tuning on human feedback, system instructions and provider policies shape which continuations are likely, and which are permitted at all. The two can push in the same direction. They are still different mechanisms. The studies below provide evidence that AI assistance can make outputs converge, and separate evidence that post-training can change how diverse a model’s outputs are. That system instructions and provider policies constrain outputs is not an empirical finding here but an architectural fact: they are written to do exactly that.

Doshi and Hauser gave 293 writers an eight-sentence story task under three randomized conditions: no AI, one GPT-4-generated idea, or up to five. Six hundred evaluators rated the stories without knowing the condition. With access to five ideas, stories were rated 8.1% more novel and 9.0% more useful than the human-only baseline, and the gains were concentrated among the less creative writers, largely closing their gap with the most creative. The same AI-assisted stories were 10.7% more similar to one another. The authors’ summary is the tension in one sentence: writers were individually better off, but collectively a narrower scope of novel content was produced (Doshi & Hauser, 2024).

Padmakumar and He had people co-write argumentative essays with GPT-3, with InstructGPT, or with no model. Writing with InstructGPT — but not with GPT-3 — produced a statistically significant reduction in diversity: essays by different authors became more similar, with lower lexical and content diversity. The reduction came from the model’s contributions; the text users wrote themselves was unaffected (Padmakumar & He, 2024).

Kirk and colleagues, comparing fine-tuning methods across two base models on summarization and instruction following, found that RLHF generalized better than supervised fine-tuning to new inputs, especially under larger distribution shift, while significantly reducing output diversity — a trade-off they name explicitly (Kirk et al., 2024).

Bound all three before leaning on them. None studied software projects or engineering teams. Two are writing tasks, one is a set of fine-tuning experiments, and all used models from 2023 or earlier. What they support is narrower and more durable than any claim about a particular model or viewpoint: better individual assistance and stronger convergence can arrive together, and feedback tuning is one plausible contributor to the convergence. They do not show that models steer people toward any particular opinions, and this chapter makes no such claim. The engineering statement is enough: models are not neutral reconstruction engines. Their priors and their post-training shape the solution space they tend to return.

Teams make it matter more

With one person, a wrong reconstruction costs one correction. With a team, the same mechanism can compound:

missing local decision
  ↓
the same broadly trained assistants
  ↓
similar conventional reconstructions
  ↓
twenty people receiving mutually reinforcing advice
  ↓
convergence

Both sides of that trade-off are plausible, and neither has been measured for software teams. A newcomer can reach a conventional project’s baseline faster, and a large team can align on shared conventions quickly — the direction Chapter 2’s support study points, where the least experienced gained most. The same shared prior can also reduce independence. People who would have disagreed usefully may receive versions of the same suggestion, and agreement that came from one source is easy to mistake for confirmation from many. The model can reduce coordination cost and reduce independence at the same time.

The closest formal result is adjacent, not direct. Kleinberg and Raghavan show in a theoretical model that decision-makers converging on one shared algorithm — even one more accurate for each of them alone — can lower the overall quality of decisions (Kleinberg & Raghavan, 2021). It concerns shared ranking, not AI-assisted teams; what carries over is only that shared decision machinery correlates behavior, and correlated errors behave differently at system scale than independent ones. Chapters 23 to 25 measure a small version of that question for portfolios of models. For teams working with assistants, it is argued here, not measured.

What is worth remembering, therefore

This changes the answer to a question the memory architecture would otherwise leave open: what is disproportionately valuable to remember?

The prior can produce a plausible version of a project’s ordinary parts. Plausible is not canonical, and the ordinary record is still needed to reproduce, audit, verify, replay and know what actually happened. But when deciding what deserves especially explicit recording — named, structured, and likely to be selected into context when it matters — prioritize the information the prior is most likely to reconstruct incorrectly:

  • approaches that were rejected, and why;
  • decisions that deliberately differ from conventional practice;
  • assumptions peculiar to this project;
  • local vocabulary and local architecture;
  • disagreements that are still unresolved;
  • experimental results that came back negative;
  • constraints that cannot be inferred from the artifact itself.

That list is exactly what a transcript summary tends to drop and a capable model tends to overwrite. This is a priority rule, not a retention rule: canonical state still belongs in the record even when the model could probably guess it. It is also a practical rule for the rest of this book:

Use the prior for the convention. Record the exception.

And it recasts why the runtime exists at all. The runtime is not there because the model becomes helpless without it. It is there because the model remains extremely capable without it — capable enough to paper over missing context, and so capable enough to continue, confidently, in the wrong direction.

Same work, different surface

The thesis deserves a demonstration, and CodeAI can support a small, truthful one. It is not an editor, a phone or a browser; none of those exist in this book. It is three separate operating-system processes, each of which knows nothing about the others except the identifiers it is handed:

A  terminal   codeai run create "Decide the cache TTL policy for the pricing API"
                --success "TTL chosen, with the condition that would change it"
run_id:  e447cfc0…
task_id: f0b9119a…

B  a separate client program — opens the work by task_id, compiles context from the ledger,
   records two offline calls, exits
call 1: succeeded
call 2: succeeded
package: 1ea4f2c8…
  included call.completed     included because required
  included task.created       included because required
  included directive.opened   included within budget
  excluded call.manifest      excluded because budget
  excluded attempt.completed  excluded because budget
budget 800, required 641, total 747

C  terminal, new process
codeai run show e447cfc0…
  objective: Decide the cache TTL policy for the pricing API
  calls:
    - 1741e6b4…: succeeded model=fake-model cost=None
    - 3805d0de…: succeeded model=fake-model cost=None
codeai context show 1ea4f2c8…
  trace: three included, two "excluded because budget"
work-state projection
  task_completion: incomplete
  next_operation: check_and_accept -- a call succeeded; the task needs a check and an acceptance

Read what happened. Process A created a piece of work and exited. Process B was not given any state — only a task identifier — and opened the work from the durable record: the objective, the success criterion, the directive. The first call drafted options; for the second, the compiler was offered the task, the directive, the first call’s outcome and two bulky bookkeeping events, required the task and the drafted options, and excluded the bookkeeping for budget, writing each reason down. Process C, again starting from nothing but identifiers, read back the same work: the objective, both calls, the exact selection record by its hash, and — from the projection — what the work needs next. Nothing was summarized. Nothing was retyped.

Now the boundaries, stated as carefully as the result. This was executed offline against CodeAI’s current source, with its fake model adapter, in a temporary directory; identifiers and hashes differ on every run, and the output above is abbreviated. It is a demonstration that ran, not a pinned evidence bundle; the script is experiments/applied-ai/surface_continuity_demo.py. What it shows is interface/process separation at the persistence boundary: different programs, in different processes, continuing the same work through its identity and record. It does not show multi-device synchronization, offline operation, concurrent writers, or any real editor, phone or browser integration. And the context trace records what the compiler selected for the call; binding that selection to the exact bytes a real provider request contains is Chapter 15’s work, where it is made opt-in and checked.

Five consequences of one move

Five design consequences organize what this move buys: surfaces can become views, context can become explicit input, history can become durable evidence, model choice can become runtime policy, and review can become runtime policy. Earlier slogans such as “works wherever you are,” “never loses context,” and “reviews every contribution” are aspirations that the sections below deliberately narrow; they are not literal guarantees of the current implementation. The five consequences are not independent wishes. Each follows from the work no longer living in the interface:

ONE RUNTIME OWNS THE WORK
        ↓
1. surfaces become replaceable views
        ↓
2. context becomes compiled input
        ↓
3. history becomes durable evidence
        ↓
4. model choice becomes runtime policy
        ↓
5. review becomes runtime policy

The rest of this section follows that chain, and at every step the question is the same: why does moving the work out of the interface force this next decision?

1. Surfaces become replaceable views

If surfaces do not own canonical state, adding one is primarily a client and integration problem — an adapter that sends requests and renders views — rather than the creation of another independent process with its own memory. That is the claim, and it is strong enough without inflating it.

The target has been articulated before, and well, from a different starting point. Kleppmann, Wiggins, van Hardenberg and McGranaghan set out seven ideals for local-first software: no waiting on the network to do work, work not trapped on one device, the network optional, seamless collaboration, longevity beyond any vendor, security and privacy by default, and ultimate ownership and control (Kleppmann et al., 2019). Read that list against the AI tools you use; most fail on longevity, privacy and user control immediately.

But be exact about the relationship, because local-first is not what this book builds. Local-first inverts the cloud arrangement: the copy on your device is primary and servers hold secondary copies, with conflict-free replicated data types (CRDTs) as the enabling technology — whose implementations the authors themselves described in 2019 as still experimental. This book’s design keeps one logical canonical record that every surface reads and writes through the runtime. It borrows local-first’s questions — who owns the data, does it outlive the vendor, can it be private, can it follow you across devices — without claiming its architecture.

Local-first concernStatus in this book
Ownership and longevityPartially implemented: the record is a local SQLite ledger and content-addressed files the user holds, with no vendor account; no general export or migration tooling for the work record is built
Work not trapped on one surfaceImplemented at the persistence boundary: separate programs continue the same work by identity (demonstrated above)
Multi-device continuityNot built: no synchronization, replication or remote access
Offline operationNot attempted as a design goal beyond “the ledger is local”
Collaboration and conflict resolutionNot attempted: no CRDTs, no designed concurrent writers
Privacy and user controlPartial: credentials are scrubbed from preserved provider payloads; no retention, redaction or access-control policy exists

2. Context becomes compiled input

Once the record outlives any surface, it quickly holds far more than any single call should see, so context stops being “what this window happens to remember” and becomes a compiled input: offered candidates, required items, a budget, a selection, and a recorded reason for every exclusion. That case, and a real trace, appear earlier in this chapter. The compiler itself is built later, in Chapter 15: its seals, its identity rules, and the gap between a selection record and the bytes actually sent.

3. History becomes durable evidence — and retention becomes policy

A runtime that owns the work needs enough recorded history to resume after a surface or process dies, explain a decision, reconstruct the evidence behind a claim, audit effects on the world, measure outcomes over time, and replay where replay is authorized.

Chapter 5 adds the reason in failure terms: when a defect is finally discovered, a system with a record has a bounded blast radius and one without has an unbounded one. Chapter 8 adds it in measurement terms: arrival — started versus finished, rework, backlog age — cannot be measured without a history to measure it over. And there is a reason that only appears once the system exists: a critic you cannot re-score is a critic you cannot change. Chapter 7 requires critics to be pinned and versioned, with the archive re-scored when one changes; that is impossible without preserved artifacts.

The earlier section on the prior adds a priority to that list, without removing anything from it. Explicit recording pays most exactly where a capable model would otherwise guess wrong: rejected approaches with their reasons, deliberate departures from convention, negative results, and constraints the artifact does not reveal. A history that faithfully records what happened but not what was ruled out leaves the prior to fill in the part that mattered most.

It does not follow that every byte from every interface must be kept forever. Durability is not infinite retention. Retention is itself policy:

record
  ↓
classify        — what kind of material is this, whose is it, how sensitive
  ↓
retain          — according to a stated policy and purpose
  ↓
expire / redact — where permitted, leaving a record that it happened

This book proposes that shape and does not build it. CodeAI’s ledger is append-only by construction, and the only removal it performs today is scrubbing credential-like fields from preserved provider payloads. Reconciling an append-only record with deletion obligations is a real design problem — one known approach encrypts material per subject and destroys the key rather than the bytes — and it is named here rather than solved. The principle is the part to keep: a runtime that centralizes memory must also centralize responsibility for what gets remembered.

A lens from cognitive science can help organize that memory, as long as it stays a lens. Sumers, Yao, Narasimhan and Griffiths’ CoALA framework, which draws on production systems and cognitive architectures such as Soar, describes language agents with a short-term working memory holding active information for the current decision cycle, and long-term episodic (experience from earlier cycles), semantic (knowledge about the world) and procedural (how the agent acts, in code and in model weights) memories (Sumers et al., 2024). The correspondence with this book’s runtime is the author’s reading, not the framework’s claim:

CoALA memoryNearest part of this runtimeBuilt in
WorkingThe compiled context package for one callChapter 15
EpisodicThe append-only ledger of events, plus content-addressed artifacts of what was returnedChapters 16–17
SemanticClaims with evidence levels — what was concluded, and how well supportedChapter 18
ProceduralGrants, budgets, required checks, and the policy that chooses the next operationChapters 20, 21 and 28

The runtime has these parts because resuming, auditing, verifying and deciding require them, not because a taxonomy predicted them. The mapping is useful for recognizing what kind of thing each part is; it is not the architecture’s source of truth, and the fit is loose at the edges — procedural memory in CoALA includes the model’s own weights, which this runtime deliberately treats as a replaceable occupant.

4. Model choice becomes runtime policy

When every surface owned its own conversation, each one also chose its own model, and those choices drifted independently. Once the runtime owns the work, which model runs an operation can become a policy the runtime applies the same way from every window, instead of a setting buried in whichever app you opened.

“Best, cheapest, most available” names three considerations, but not yet a complete policy:

  • Best is measured capability for this operation: does this candidate clear the criterion that matters at an acceptable rate?
  • Cheapest is process cost, not price per call. Chapter 6 keeps cost per accepted outcome separate from accepted-but-wrong outcomes, and computes cost per correct acceptance only where correctness is independently known.
  • Most available is reachability, and it is not one signal. Configured availability, observed reachability, recent failures, rate limits, and region or policy constraints can disagree with one another; the runtime has to preserve those distinctions rather than collapse them into one health bit.

Keep this decision separate from the one Chapter 1 put in front of it. The runtime must first decide which operation the process needs. Only when that operation is CALL does escalation or model selection become relevant. Chapter 28 source-inspects a small deterministic next-operation scheduler and measures an escalation ladder; its separate model-router challenge remains unrun. Those are three different evidence states, not one routing result.

That is as far as this chapter goes. Chapter 10 makes the model a replaceable occupant of a named slot. Chapter 28 later tests pieces of escalation and operation policy. The point here is only the architectural one: once the runtime owns the work, model choice can become shared policy rather than an interface-specific setting.

5. Review becomes runtime policy

The same logic applies to review, and it resolves a tension in the fifth property. Read strictly, “reviews every contribution” calls for an independent critic panel on everything, and the book’s own evidence does not support that as a standing requirement. A critic panel on every contribution is expensive — it can cost more than the contribution — and it is another stochastic subsystem, not ground truth. Chapter 24 found a three-model portfolio covering exactly the tasks that repeated draws from one model covered, and failing the hard task the same way. That was three small local models on twelve tasks, not a general law, but it is enough to say that independence between a critic and a generator has to be measured, not assumed.

The defensible policy is layered by what each layer can actually establish:

LayerWhen it runsWhat it is good for
Mechanical checksWhenever the property is mechanicalRepeatable evidence about exactly the property implemented: parsing, compilation, source existence, quoted spans, arithmetic
Model-based reviewWhere validation against a trusted reference shows useful signal, or policy explicitly requires itScalable judgment on properties mechanical checks cannot establish; still model-produced evidence, pinned and versioned
Human reviewRouted by consequence, novelty, unresolved judgment and authorityDecisions whose remaining criteria or consequences require human judgment

The practitioner literature points the same way: Böckeler argues that a good harness should not aim to eliminate human input but to direct it to where it matters most (Böckeler, 2026).

The durable principle is architectural, not a particular stack: review policy belongs to the runtime, not to the goodwill of whichever interface generated the artifact. If the editor checks citations and the phone does not, the work’s quality depends on where you happened to be standing — the Tuesday problem again, one level up.

Review itself should also be tested. Seed known failures that a particular review mechanism claims to detect and record which seeded failures survive. That measures sensitivity to those seeded defect classes; it is not a complete estimate of review quality.

Chapter 5 recommends the pattern for human review. From Chapter 15 onward, this book applies it to several automated evidence verifiers by running them against deliberately corrupted copies of their evidence and requiring the targeted corruption to be rejected. For production review of every workflow, it remains a recommended mechanism rather than something this book has built — and not every workflow justifies it.

Where the design fights itself

A design chapter that lists only benefits is marketing. Moving the work into one runtime buys continuity, and it concentrates costs that have to be decided rather than discovered.

Review versus budget. Review is paid from the same budget as the work. What fraction of spend is evaluation must be an explicit decision, made up front — not an emergent property of whatever check someone added last.

Many surfaces versus one record. One logical record reachable from several devices is a synchronization problem, and synchronization is where distributed systems go to be difficult. Every honest multi-device or offline ideal costs engineering to honor.

Durable memory versus the context budget. An unbounded history cannot be carried into a bounded, priced context. Selection is mandatory; the commitment is to record every exclusion, not to keep everything in view.

Remembering versus privacy and retention. A system that records what you do across every surface you use is a more revealing artifact than any single tool it replaces. Where does it live, and who can read it? What happens on breach, subpoena, departure or acquisition? What is kept, for how long, and what is expired or redacted? Privacy and user control have to be design inputs from the first record, not a compliance conversation in year two.

One runtime versus concentrated risk. Centralizing canonical state is what makes continuity possible, and it concentrates:

  • availability risk — when the runtime is unreachable, every window is;
  • privacy risk — one place holds everything;
  • corruption risk — a bad write damages the work for every surface;
  • authorization risk — one policy mistake applies everywhere;
  • schema migration risk — changing the record’s shape touches all history.

Those require backup, a documented recovery path, auditable migration, access control, and retention — none of which this book builds in full, and all of which belong to the design rather than to operations trivia. Centralizing truth centralizes responsibility.

What CodeAI actually implements

The reference implementation is CodeAI, and the distance between this chapter’s design and its current source should never be blurred. Read against the source used for the demonstration above:

Part of the designStatus
Durable work identity (runs and tasks) in an append-only ledgerImplemented
Current state as a projection of the record, including the next operationImplemented (Chapter 16)
Context compilation with recorded offers, selections and exclusion reasonsImplemented (Chapter 15)
Selection bound to the exact request bytesPartial: opt-in rendering path (Chapter 15)
Content-addressed preservation of what models returnedImplemented (Chapter 17)
Claims with evidence levels, decisions resting on themImplemented (Chapter 18)
Authority checked from the request, not the interfaceImplemented for actions and acceptance (Chapters 19–20)
Model choice as runtime policyPartial: replaceable occupants are built in Chapter 10; Chapter 28 measures an escalation ladder, but a general model-selection policy is not established
Review as runtime policyPartial: independent verification bound to state (Chapter 21); no standing model-review or human-routing policy
SurfacesCommand line and Python API only; no editor, phone or browser surfaces are built in this book
Synchronization, offline, collaborationNot built
Retention, redaction, backup and recoveryNot built, beyond credential scrubbing

The surfaces are left to later labs — and, as Chapter 30 argues, to you: the surfaces most worth building are the ones shaped around your own work.

The field is working on the same layer

In 2026 “harness engineering” became the industry’s name for the layer around the model, and it is worth placing this chapter against it precisely.

Böckeler defines a coding agent as a model plus a harness — everything except the model — and sorts a harness’s controls two ways: guides, which steer the agent before it acts, versus sensors, which observe afterward so it can correct itself; and computational controls (deterministic and fast) versus inferential ones (model-based review, slower, costlier and non-deterministic) (Böckeler, 2026).

Barbaste and colleagues read the source code of eleven production coding harnesses — Claude Code, Codex CLI, Gemini CLI, Mistral Vibe, OpenHands, Aider, Mini-SWE-Agent, Hermes, Pi, OpenCode and OpenClaw, about four million lines — and mapped seven canonical subsystems, one of which rations the context window and persists knowledge across turns and sessions. Across the corpus no agent runtime imports a general-purpose agentic framework and none retrieves code with vector embeddings, and over one quarter they observed behavioral policy migrating from prompt prose into configuration (Barbaste et al., 2026). The study describes and compares how the systems are built; it does not benchmark them.

That is evidence of relevance, not validation. Independent convergence on “the layer around the model matters” says nothing about whether this chapter’s particular decisions are right. What this chapter adds is an insistence about what must live in that layer: not only loops, tools and context management, but the durable identity and state of the work, authority, evidence, and the decision that work is complete — owned once, outside any one harness or window, so that switching tools is opening the same work rather than importing a session.

Do this now

Twenty minutes, one diagram. Find where your state lives.

  1. List every surface you currently use to work with AI — chat app, editor plugin, terminal, phone, browser.
  2. For each, write what it remembers and for how long. Be exact: does it survive a restart? A week? A device change?
  3. Draw the arrows that represent you carrying context between them. Count them. It will look something like this:
CURRENT

chat ──you── editor
 │             │
you           you
 │             │
phone ──you── terminal
  1. Pick one active piece of work. If you closed every AI interface right now, where would another program find its objective, the decisions already made, the unresolved questions, the evidence, and the next operation?

  2. For that same piece of work, write down three things a capable newcomer — or a model — would get wrong by assuming the conventional approach. Then check whether any of the three is written down anywhere a program could read it.

If the answer to step 4 is “nowhere”, “in the scrollback of one application” or “in my head”, you have located the missing runtime. If none of step 5’s exceptions is recorded, you have located the first thing it should hold. Now draw the target:

TARGET

chat ────┐
phone ───┼── durable work identity and state
editor ──┤
terminal ┘

Keep both diagrams. By the end of the book you should be able to replace every arrow that was you with a record another program can read.

Failure modes

  • The interface owns canonical state. Guarantees one process per surface and a person carrying context between them.
  • A new surface reconstructs state from a transcript. Summarization is lossy reconstruction, not continuation.
  • Memory and context are conflated. Either amnesia or enormous bills.
  • Context exclusion happens silently. Exclusion is fine; unrecorded exclusion is a bug you cannot reproduce.
  • Mistaking a smooth continuation for recovered context. A capable model fills gaps from its prior, and the result reads just as fluently as recovered state.
  • Recording what happened but not what was ruled out. The prior will reintroduce the conventional answer your project deliberately rejected.
  • Treating agreement among assistants as independent confirmation. Several people consulting the same kind of model may be hearing one prior several times.
  • Everything is retained forever with no retention policy. Durability is not infinite retention, and the record is a privacy artifact from its first write.
  • Authority is inferred from the interface. The same work opened in two windows would carry two sets of permissions.
  • Routing or review policy differs by surface. Quality then depends on where you happened to be standing.
  • One logical runtime is mistaken for one physical server. The truth must be singular; the machinery need not be.
  • Centralized state has no backup or recovery path. Every window loses the work at once.
  • Aspirational design is described as implemented behavior. A design chapter needs status labels, not confidence.

What this chapter established

  • If canonical state lives only in an interface, you have as many isolated processes as you have interfaces. The architectural move is one logical runtime that owns the work, with interfaces acting as views and request surfaces.
  • Continuity needs a durable work identity and authoritative record — objective, decisions, evidence, unresolved state and permitted next work — that a surface opens rather than reconstructs. Current state is projected from that record. Summarization is reconstruction; projection from the record is continuation.
  • State is current, history is recorded, memory is durable, context is selected, and UI state is only what a window shows. In CodeAI the append-only record is authoritative and current state is a deterministic projection of it.
  • A surface sends requests and receives views. Authority is evaluated from the request, actor, work and policy rather than inferred from which window is open.
  • “One runtime” means one logical owner of canonical process state and policy, not one operating-system process or server.
  • Context is selected from retained material. CodeAI records the candidate inventory offered to a compilation and the inclusion or exclusion decision for each offered candidate; it does not claim every ledger item was considered for every call. A larger context window is capacity, not evidence that every relevant fact was selected or used correctly.
  • A plausible continuation can mix recovered project material with model inference. When project-specific exceptions are absent, learned conventions can produce a competent but wrong continuation. The prior can supply convention; the record must carry the exception.
  • Evidence from writing and fine-tuning studies shows that improved individual assistance and increased convergence can occur together; the claim about reduced independence in software teams remains an argument, not a measured result.
  • Rejected approaches, deliberate departures, local assumptions, unresolved disagreements, negative results and hidden constraints deserve especially explicit recording because they are difficult to reconstruct safely from convention. That is a priority rule, not a retention rule.
  • An executed offline demonstration reopened the same work from separate processes by identity, including a recorded context selection with budget exclusions and a projected next operation. It demonstrates interface/process separation at the persistence boundary, not multi-device synchronization.
  • Moving work into the runtime makes shared model-selection and review policy possible, but the evidence stays separated: Chapter 28’s deterministic next-operation policy, measured escalation ladder and unrun model-router challenge are different things. Review is likewise layered: mechanical checks establish only their declared properties, model review needs validation against a trusted reference, and human judgment is routed where consequences or unresolved criteria require it.
  • Centralizing truth centralizes responsibility: retention, privacy, backup, recovery, migration and access control belong to the design. Local-first supplies useful questions; this book does not claim to implement its architecture.

Next

This chapter gave the design a dimension in space: the interface can change without losing the process. Many windows, one runtime.

It also has a dimension in time, and that one is easier to miss. The component in the one stochastic slot will not stay put. Later models will be substantially different machines, the “same” named model drifts between snapshots, and an upgrade that raises the average can break the specific thing you depended on. The next chapter makes the complementary move: the model can change without rebuilding the process. Many model generations, one durable frame.

Continue with A Revolver, Not a Foundation.

References

Implementation sources (CodeAI): src/codeai/ledger.py: SQLiteLedger (append-only). src/codeai/workstate.py: project_work_state, NextOperation. src/codeai/context.py: ContextCompiler.compile_with_trace, CompilationTrace. src/codeai/runtime.py: compile_and_record_context, invoke_recorded_call, show_run. src/codeai/policy.py: authority check raising AuthorityDenied. src/codeai/cli.py: run create, run show, context show, build_runtime, ensure_task. src/codeai/providers.py: credential scrubbing of preserved payloads. Demonstration: experiments/applied-ai/surface_continuity_demo.py, executed offline with FakeCognitionAdapter; not a pinned evidence bundle.