← Applied AI

One Operation, Several Model APIs

On a real gateway, choosing a model can also choose its wire dialect, so an occupant swap may be a protocol swap. The adapter's job is not to hide the differences but to contain them, and the conformance test asks whether matched fixtures reach the same canonical runtime outcomes across several dialects.

Part 2 — Get the Model Out of the Chat Box

Choosing a model can change the whole wire contract

Right model, wrong dialect

Chapter 11’s second live run failed with an HTTP 500. The gateway, the model, and the credential were all right. The request was shaped for OpenCode’s Responses endpoint, and mimo-v2.5 is served on Chat Completions.

That looks like a configuration slip. It is a property of the ground you are building on.

OpenCode Go’s catalog was inspected for the live run on 13 September 2026 and rechecked for publication after its 18 September update. It still lists each model on one of three endpoints (OpenCode). Three of them carry this chapter:

ModelEndpointAPI family
gpt-5.6-luna/zen/go/v1/responsesOpenAI Responses
mimo-v2.5/zen/go/v1/chat/completionsOpenAI Chat Completions
minimax-m2.7/zen/go/v1/messagesAnthropic Messages

On this gateway, choosing a model also selects that model’s protocol route. Chapter 10 made the occupant of each chamber replaceable. A swap between occupants on different routes is therefore a protocol swap too; a swap between two occupants on the same route is not. The three occupants in this chapter were chosen specifically to exercise all three dialects.

How does one logical operation survive a change of wire dialect, without the runtime knowing, and without hiding anything that matters?

The textbook answer is an adapter that hides the differences. That is half right. The part it gets wrong is the word hides. An adapter that hides everything also hides the differences that change your results. The job is to contain them: translate what is equivalent, record what is not, and refuse what it does not understand.

By the end of this chapter you will be able to keep one logical model operation stable across several wire dialects: know what happens to every control you pass, keep the record of a request from drifting away from the request itself, read each dialect’s answer and completion signal, and say which layer of “success” a call actually reached. Underneath all of it is one distinction, model ≠ protocol. A gateway may couple them, but they change for different reasons and break in different ways.

One swap, many changes

Here is what actually differs between the three routes, all on one gateway under one subscription.

gpt-5.6-lunamimo-v2.5minimax-m2.7
Endpoint/v1/responses/v1/chat/completions/v1/messages
Prompt goes ininputmessagesmessages
Output limit fieldmax_output_tokensmax_tokensmax_tokens, required
Extra header——anthropic-version
A requested reasoning_effortsent, as reasoning.effortomitted, and recordedomitted, and recorded
When no limit is requestednone sentnone sentCodeAI sends 1,024, recorded as defaulted
Where “finished” is reportedstatus, incomplete_details.reasonchoices[0].finish_reasonstop_reason
Where hidden reasoning comes backreasoning output itemsmessage.reasoninga thinking content block
Usage vocabularyinput_tokens / output_tokensprompt_tokens / completion_tokensinput_tokens / output_tokens
Monthly usage cap in the Go plan$15$60$60

Chapter 10’s release protocol said: change one thing, measure, then decide. Across the three occupants chosen here, changing the model also changes the endpoint, request shape, headers, the fate of at least one control, the completion vocabulary, the place reasoning comes back and the usage vocabulary. Sometimes the quota headroom changes too. Other occupant swaps on this gateway can remain within one dialect, so the protocol change is a property of the route pair, not of model replacement in general.

Two consequences follow, and the rest of the chapter is about both. First, the runtime must not care which row it is on, or every release becomes a rewrite. Second, a live comparison between two occupants can never tell you what the protocol did, because the protocol never changes alone. That is why the experiment at the end of this chapter has two layers.

One operation in, three dialects out, one contract back:

    flowchart TD
    OP["one logical operation<br/><i>review P, max 256 tokens</i>"] --> AR["adapter: Responses"]
    OP --> AC["adapter: Chat"]
    OP --> AM["adapter: Messages"]
    AR --> R1["renamed · nested"]
    AC --> R2["omitted + recorded"]
    AM --> R3["defaulted + recorded"]
    R1 --> CT["stable application contract<br/><i>same call, same record shape</i>"]
    R2 --> CT
    R3 --> CT
  

Six fates of a control

Ask all three routes for the same thing: max_tokens 256, temperature 0.2, reasoning_effort “low”. This is what CodeAI recorded as actually sent, from its offline semantic-request cases:

requested    {max_tokens: 256, temperature: 0.2, reasoning_effort: low}

responses    sent {max_output_tokens: 256, temperature: 0.2, reasoning: {effort: low}}
chat         sent {max_tokens: 256, temperature: 0.2}   omitted_unsupported: [reasoning_effort]
messages     sent {max_tokens: 256, temperature: 0.2}   omitted_unsupported: [reasoning_effort]

Every control a caller asks for meets one of six fates:

  1. Sent as is. temperature, on all three routes.
  2. Renamed. max_tokens becomes max_output_tokens on Responses.
  3. Nested. reasoning_effort becomes reasoning.effort on Responses.
  4. Omitted, with a record. reasoning_effort on Chat and Messages, and seed everywhere.
  5. Defaulted, with a record. Messages requires an output limit. If the caller supplies none, CodeAI sends 1,024 and writes that into defaulted_controls.
  6. Refused before any effect. An unknown control name, a malformed value, an attempt to override the model through parameters, or a credential passed as a parameter all raise before the adapter is invoked.

The first four are ordinary translation. Fates five and six are where systems quietly lie.

A default nobody records is a control the caller never chose and cannot see. If the Messages route silently supplies an output limit while the Chat route sends none, then any comparison between those two chambers carries a hidden variable. The difference in your results may be the limit, not the model.

Refusal is the unfashionable one. Protocol engineering spent decades under the robustness principle: be liberal in what you accept. RFC 9413 argues that this tolerance does long-term damage. As Thomson and Schinazi put it, “Tolerating unexpected inputs from another implementation might seem logical, even necessary” (Thomson & Schinazi, 2023). Their point is that tolerated deviations accumulate into de facto requirements nobody chose, and they recommend active maintenance in place of silent acceptance. A codec that accepts a control it does not understand, and simply does not send it, is exactly that tolerance.

This is not hypothetical, even inside CodeAI. The control audit recorded with its request-plan stage found two silent failures in the earlier code. A malformed max_tokens value such as "abc" was dropped from the request body while still being recorded as sent. And the Chat route accepted reasoning_effort, never sent it, and left no trace of the omission.

A synthetic probe pins the corrected behavior — no network, same three routes, two bad inputs:

InputOld behaviorCurrent behavior
max_tokens: "abc"dropped from the body, still recorded as sentInvalidControlError from prepare(), before any request exists
unknown control top_k2: 4accepted and silently never sentUnknownControlError from prepare(), with no invocation following
reasoning_effort: low on Chataccepted, never sent, no traceomitted and recorded in omitted_unsupported

The first two rows are refusals, not translations: prepare() in providers.py raises RequestPlanError subclasses before any provider effect, and regression tests pin that no invocation follows. The third row is the contrast worth keeping — a declared-but-unsupported control is legitimately omitted, provided the omission is in the record. Refusal and recorded omission are both honest; silent tolerance is the one that corrupts every comparison built on top.

Sculley and colleagues identified configuration as a characteristic source of hidden technical debt in machine learning systems (Sculley et al., 2015). A setting that silently fails to apply, while the record says it applied, is the worst form of that debt: every comparison built on the record inherits the error. The regression tests that pin the corrected behavior are test_invalid_values_rejected_before_effect and test_unknown_control_rejected_pre_effect_no_invocation.

Send what you record

The first of those defects had a structural cause: two code paths produced two views of one request. One path built the body that went over the wire. Another built the “effective parameters” that went into the record. They drifted, and nothing noticed.

The fix is to prepare once and use the result twice:

prepared = adapter.prepare(spec)   # validate, map, default, refuse, all before any effect
manifest = record(prepared)        # intent, written before the first attempt (Chapter 11)
for attempt in attempts:
    reply = adapter.send(prepared) # exactly the object that was recorded

That is simplified, but the shape is real. prepare() returns a PreparedCognitionRequest holding the endpoint, the body, the public headers, and the requested, effective, omitted and defaulted controls. The manifest records that object. send() transmits that object. A retry resends that object. Tests assert that the body sent equals the body prepared, and that a retry resends the identical prepared request.

Two precise limits. The request hash in the manifest is SHA-256 over canonical sorted JSON. It identifies the semantic request, not the outbound bytes, which the transport serializes separately. And credentials never enter the prepared object: they are applied inside send(). For Messages the key goes out as both a bearer Authorization header and x-api-key.

Parnas’s classic criterion for decomposing a system is to hide, inside each module, a design decision that is likely to change (Parnas, 1972). Which dialect a route speaks is such a decision, and on this gateway it changes whenever an occupant does. So it lives inside the adapter. CodeAI’s runtime.py contains no reference to any protocol name. The runtime decides retries, call status and task state from the canonical interpretation, and never asks which endpoint the bytes came from.

Read the answer where it lands

The three live captures show three different places to find the answer, and in two of them the answer is the smallest thing in the response.

  • Responses (gpt-5.6-luna): the text is an output_text part inside a message output item, and completion is status: "completed".
  • Chat Completions (mimo-v2.5): 112 characters of answer in choices[0].message.content, plus 780 characters of reasoning in message.reasoning. Completion is finish_reason: "stop".
  • Messages (minimax-m2.7): a thinking block of 1,094 characters, then a text block of 74 characters. Completion is stop_reason: "end_turn".

The canonical text CodeAI produces contains only the answer. The reasoning is neither folded into it nor thrown away. It remains in the preserved response bytes, stored under their content hash, where a later interpreter or a human can read it. Merging reasoning into the answer would make the answer wrong; deleting it would destroy an observation. The adapter does neither.

Completion signals arrive in three vocabularies and map onto one set of states. The mapping carries its own version, so it can be corrected later without rewriting history (Chapter 13):

Dialect signalCanonical generation state
stop, end_turn, completedcomplete
length, max_tokens, max_output_tokenstruncated
content_filterfiltered
anything elseunknown

Success has layers

Once the answer has been extracted from wherever it landed, it is tempting to treat the call as done. It is worth asking exactly what has been established, because the honest answer is: less than it looks.

Chapter 11 already met the smallest case. The transport worked, the provider returned a response, the response contained text, and the text read like a review. It was recorded as a success. The preserved bytes said finish_reason: "length", and the review stopped mid-sentence. “We received text” was one fact; “we received the answer” was another.

There are more boundaries above that one, and each answers a different question:

request prepared                controls were valid and the intended wire request was recorded
    ↓
response received               transport returned a response
    ↓
generation finished normally    the provider's completion signal says complete
    ↓
answer present                  there is answer text, not only reasoning or nothing
    ↓
syntax valid                    the answer parses
    ↓
shape valid                     required fields present, types and allowed values respected
    ↓
values meaningful               the content holds up against the input
    ↓
task criteria satisfied         it meets what the task asked for
    ↓
verification performed          a declared check evaluated it; independence and adequacy are separate questions

Success at one boundary does not imply success at the next:

HTTP success ≠ generation success ≠ parse success ≠ schema success
             ≠ semantic success   ≠ task success  ≠ verified success

The “return JSON” problem has moved. For years the common pattern was to ask a model, in the prompt, to “reply with JSON only”, and then to handle a steady trickle of malformed objects, missing fields, surrounding prose and wrong types. That made syntax a stochastic requirement. JSON modes improved it by guaranteeing parseable JSON but not its shape. Schema-constrained structured output goes further, and the major APIs now offer it. OpenAI’s Structured Outputs takes a JSON Schema with strict adherence, and the documentation distinguishes it from JSON mode, which only ensures valid JSON (OpenAI). Anthropic’s API accepts a JSON Schema as the output format and a strict flag on tool definitions, enforced by constrained decoding (Anthropic). Gemini’s structured output guarantees syntactically correct JSON against a supplied schema (Google).

That is real progress, and it is Chapter 3’s argument arriving from the provider side: much of syntax and structural conformance moves out of the stochastic box and into an enforced frame.

The same documentation marks where the guarantee ends, as it was on 17 September 2026. OpenAI notes that a response stopped by the token limit is incomplete and may not match the schema, and that refusals come back as refusals rather than as schema-shaped answers. Anthropic notes that output may not match the schema when generation stops for max_tokens or refusal, and that some schema constraints, such as numeric ranges and string lengths, are not enforced. Google is the most direct: “always validate values in your application.”

So structured output does not establish that generation completed, that the model did not refuse, that the values are sensible, that a quoted span exists, that an extraction is correct, that an allowed null is the right answer, that the task’s criteria are met, or that anything in the world ended up in the right state. Structured output can constrain the shape of an answer. It cannot establish the truth of the answer.

This book’s own evidence shows the failures moving upward. Chapter 28’s escalation ladder made 84 real calls across three models, 30 through the Responses dialect and 54 through Chat Completions, asking for a small JSON object through the prompt alone, with no structured-output mode. Recomputed from the preserved outputs with the experiment’s own parser and check:

BoundaryWhat happened in 84 calls
TransportAll 84 returned HTTP 200
Generation and answer present1 produced no answer text: the strong model spent its output budget on reasoning (Chapter 28), recorded as empty generation and a failed call
Syntax83 of 83 answers parsed as a bare JSON object; 0 were malformed
Shape83 of 83 had both fields with permitted types; 0 missing keys, 0 wrong types
Values: abstention20 answered "seconds": null, which the instruction allowed; every one was on an item with no single stated TTL
Values: grounding9 failed the check because the quoted words held zero or two durations, including answers to T07 and T10 that Chapter 21 shows the deliberately strict checker rejected
Passed the check54
Passed the check and wrong1, A05

That is one extraction task with a two-field object, three models and one week, not a malformed-JSON rate for anyone else’s workload. But in this run, broken JSON was not the failure that mattered. Completion mattered once, permitted abstentions mattered twenty times, grounding mattered nine times, and a well-formed, grounded, wrong answer got through once.

Record which layer failed. A syntax failure and a wrong factual answer are different failures and must not become the same telemetry. CodeAI records the lower layers as its own states, as the offline table below shows: transport error kinds, a generation state of complete, truncated, filtered, empty or unknown, and a call status. The upper layers belong to the task’s check, and Chapter 28’s check names each one: output_not_json for syntax; no_value_to_verify, seconds_not_a_whole_number and quote_missing for shape and permitted values; quote_not_in_input, quote_has_2_durations and seconds_do_not_follow_from_quote for grounding. Acceptance and completion are Chapter 14’s; independent verification is Chapter 21’s. One layer has no state at all yet: CodeAI does not recognize a refusal as its own outcome.

Two rules follow. Where a provider offers schema-constrained output, prefer it. Where a deterministic parser is needed instead, make it explicit, versioned and testable, as the ladder’s ttl-output-parser-v1 is: it accepts one optional code fence around one JSON object, and nothing else. Do not repair malformed output with a second model call; that adds a stochastic step exactly where Chapter 3 says a deterministic one belongs.

And stop at the right place. An answer that arrived, finished normally, parsed, matched its shape and passed a content check has established that the call produced a candidate that passed the declared check. It has not established that the check was adequate to everything the task requires; that distinction becomes explicit later in the book. Reading its completion signal, parsing it, extracting its fields and deciding that null means abstention are all interpretations of a preserved observation, and Chapter 13 makes such interpretations versioned. Whether the work itself is complete is Chapter 14’s question.

The same prompt is not the same input

Every live call sent the same one-sentence request. Here is what each route reported:

RouteInput tokensOutput tokensVisible answerOther usage fields
Responses / gpt-5.6-luna3826125 charscached 0, reasoning 0, total 64
Chat / mimo-v2.5279166112 charscached 192, reasoning 0, total 445
Messages / minimax-m2.77322074 charsnone reported

Three facts sit in that table, and each breaks a comparison people routinely make.

Identical requests produced input counts from 38 to 279. Tokenizers differ, and routes can add material you did not send. The Chat route reported 192 of its 279 input tokens as served from cache. Chapter 11’s truncated call, a different request to the same route on an earlier day, also reported exactly 192 cached tokens. Two different requests sharing an identically sized cached prefix suggests the prefix was supplied by the route rather than by us. The record makes that likely; it does not prove it.

The reported output counts are not explained by the visible answer alone. The visible answers were 112 and 74 characters while the routes reported 166 and 220 output tokens. Character count and token count are different units, so that comparison is only a clue. The reasoning returned beside the answers is another obvious candidate for what the output count includes, but neither route’s usage attributed tokens to it. MiMo’s usage reports reasoning_tokens: 0 while returning 780 characters of reasoning. A provider’s own accounting breakdown is an observation to be interpreted, not a fact to be believed.

One route gave no breakdown at all. Messages reported input and output, and nothing about cache or reasoning.

So tokens are not a unit that transfers across routes. That is the concrete reason Chapter 10 defined a chamber’s utility as cost per passing item rather than cost per token. On a per-token basis these three routes cannot even be placed on one axis.

CodeAI’s canonical usage today carries input and output only, labeled measured. Everything else stays in the preserved bytes. The rule to follow is short: normalize equivalent semantics, preserve non-equivalent semantics, never force equivalence.

There is a latent trap too. In Chat usage, cached_tokens is a detail inside prompt_tokens. In the Anthropic Messages convention, cache reads and cache creation are separate components reported beside input_tokens. CodeAI’s Messages codec reads only input_tokens. This capture reported no cache fields, so nothing was undercounted here. But the first time a Messages route serves from cache, one canonical field will quietly mean two different things. Deciding what these numbers mean is Chapter 13’s job.

The experiment, in two layers

A live comparison cannot isolate the protocol, because no model on this gateway is offered in two dialects. So the evidence comes in two layers that answer different questions, and they are reported separately and never pooled. The full bundle is preserved under experiments/applied-ai/evidence/protocol-conformance/.

Offline: does the boundary hold?

The harness runs each dialect through the real CodeAI runtime with a patched transport and outbound sockets refused. Eight cases per dialect are synthetic, and labeled as such: a complete answer, a truncated one, an unrecognized completion reason, an empty body, a tool call with no text, an HTTP error, a malformed body, and no response at all. Three more replay the real response bytes captured by the live layer. The captured bodies replay byte for byte, with the same hashes as the originals.

27 of 27 cases passed, with zero network calls. The result that matters is not the pass count. It is this table:

CaseTransportGenerationError kindCall status
completeresponse receivedcomplete—succeeded
truncatedresponse receivedtruncated—unresolved
unknown reasonresponse receivedunknown—succeeded
emptyresponse receivedemptyempty_outputfailed
tool call onlyresponse receivedemptyempty_outputfailed
HTTP errorHTTP errorunknowninvalid_requestfailed
malformed bodyresponse receivedunknownmalformed_responsefailed
no responseno responseunknowntimeoutfailed

For each of the eight synthetic case types, all three dialects produced the same transport, generation, error and call-status outcome. The three captured live bodies were replayed separately and matched independently checked text and completion signals. That is the bounded claim: within this conformance matrix, dialect-specific response shapes did not change the runtime decision once they had been mapped to the same canonical facts.

Two checks guard against fooling ourselves. The expectations for the captured fixtures were first copied from the codec’s own output, which would have made replaying them circular. So the bundle includes independent parses of the raw bytes, written with no CodeAI imports. They agree with the codec on the text and the completion signal for all three captures, and a separate scan found no credential material in the bundle.

Live: does each route actually execute?

Three calls, one per route: one attempt each, a 60-second timeout, and an output limit of 1,024 tokens. A budget guard allowed at most three calls and required a recorded reason for proceeding with unknown cost. All three returned HTTP 200 with complete generation, and all three answers named the missing benchmark or measurement.

That supports exactly one claim: each configured route executed once, on 13 September 2026. It says nothing about protocols, because the models differ. No winner was computed and none should be.

The budget detail is Chapter 6 arriving in practice. The guard was configured with a $0.50 cost ceiling, but on a subscription the cost of a single call is unknown, so that ceiling could never trip. What actually bounded the experiment was the call count and the token limit. Budgets you cannot measure are not budgets.

One more detail: the harness asked a one-sentence question about a code comment rather than reviewing paragraph P, to keep live output short. The operation under test is the boundary, not the review.

Where it is still weak

The boundary holds for what was tested. These are the places an honest reader of the code and the bundle will find gaps:

  1. An unknown completion reason becomes success. Truncation now correctly leaves a call unresolved. An unrecognized reason, the most likely shape of a future API change, still produces a succeeded call. The attempt policy CodeAI uses by default today keeps that behavior: output with no error and an unknown generation state is accepted.
  2. A tool-call-only reply is classified as empty_output, which the retry policy treats as retryable. For a text-only chamber that is defensible. For a chamber whose occupant legitimately answers with a tool call, the runtime will retry a valid answer. The current default policy still retries empty generation.
  3. The live captures contain no response-header evidence. Every 13 September live capture recorded an empty preserved header map. Current transport code can preserve an allowlisted set of request-id, retry and rate-limit headers when they are present, and regression tests exercise that path, but these three live bundles provide no evidence that OpenCode supplied any of them. Rate-limit and retry-after data therefore cannot be reconstructed from this run.
  4. Canonical usage is two numbers. Cache, reasoning and totals stay in the raw bytes, and the Messages cache convention is a latent mismatch.
  5. The request hash is semantic, not byte-level. It proves which request object was intended, not which bytes left the machine.
  6. The Messages route sends the credential twice. The route documentation names neither header. The successful call shows the combination is accepted, not which header is required.
  7. A refusal has no state of its own. Anthropic signals it as a stop reason and OpenAI as a separate output part; CodeAI maps neither to a distinct outcome, so a refusal is not told apart from other generation results.
  8. One question, one call per route, one day. “Executed today” is not “available tomorrow”, and the catalog these calls relied on had been updated two days earlier.

None of these weakens the central result. Each is a place where the contract is narrower than it may look.

Do this now

Forty-five minutes. Find out what your controls actually do.

  1. For each chamber occupant you use, write the six-fate table: for every control you pass, is it sent, renamed, nested, omitted, defaulted, or refused? If you cannot answer for one control, that control is currently unobserved.
  2. Pass a deliberately malformed value, such as a string where an output limit belongs. Does anything raise before a request is sent, and does your record still claim the value was used?
  3. Take one real response you have from each dialect you use, parse it by hand without your codec, and compare the text and completion reason with what your code recorded.
  4. Answer in writing: if you swapped your most important chamber’s occupant today, how many of the rows in “One swap, many changes” would change?
  5. Take one production model call and mark which of these are checked separately: transport, completion, answer present, parse, shape, values, task criterion, verification. If several collapse into one success boolean, that is the defect.

If you are building with an assistant, this is the increment:

Keep one logical model operation stable across several wire dialects.
Inspect first: list every control the code accepts and what each dialect does
with it. Then implement, reusing the existing call/attempt records:
- one prepare step per call that validates, maps, defaults and refuses,
  producing a single request object that is both recorded and sent;
- record requested, effective, omitted and defaulted controls separately;
- reject unknown or malformed controls before any request exists;
- extract canonical text per dialect; keep reasoning and tool blocks in the
  preserved response, never in the answer;
- map each dialect's finish/stop signal to complete/truncated/filtered/unknown
  under a versioned map.
Prove it offline: the same eight cases through every dialect must produce
identical transport, generation and call outcomes, with outbound sockets
refused. Do not compare model quality across dialects.

Failure modes

  • Treating an occupant swap as one change. A swap across protocol routes can change endpoint, headers, controls, vocabularies and quotas along with the model. A swap that stays on one route does not necessarily change all of them.
  • Silent omission. A control that is accepted, not sent, and not recorded.
  • Silent defaults. A limit the caller never chose becomes a hidden variable in every comparison.
  • Tolerating what you do not understand. Unknown controls accepted “to be safe” become requirements nobody chose.
  • Two code paths for one request. The record and the wire drift apart unnoticed.
  • Folding reasoning into the answer, or deleting it. One corrupts the result; the other destroys an observation.
  • Calling parseable success. Valid JSON establishes syntax, and with a schema its shape; not correctness, usefulness or task completion.
  • Collapsing every failure into failed. A truncation, a parse error and a wrong answer need different responses, so they need different records.
  • Repairing malformed output with another model call. Constrain the output, or parse it deterministically under a versioned rule.
  • Comparing tokens across routes. The same request was 38, 73 or 279 input tokens depending on where it went.
  • Believing a provider’s usage breakdown. reasoning_tokens: 0 arrived beside 780 characters of reasoning.
  • Reading a live multi-route comparison as a protocol effect. The model changed too.
  • Branching on protocol in the runtime. Every new dialect then becomes a runtime change.

What this chapter established

  • On this gateway, each listed model is served on one dialect, so choosing an occupant also selects its protocol route. Swapping between occupants on different routes changes several variables at once; swapping within one route need not.
  • The adapter’s job is to contain differences, not hide them. Every control meets one of six fates: sent, renamed, nested, omitted and recorded, defaulted and recorded, or refused before any effect.
  • Silent tolerance is debt. RFC 9413’s case against liberal acceptance and Sculley et al.’s configuration debt both appear in CodeAI’s own history: a malformed limit recorded as sent, and an unsent control with no record. Both now fail before a request exists.
  • Prepare once, record it, send it: the same prepared object drives the manifest, the transport and every retry. Its hash identifies the semantic request, not the wire bytes.
  • The runtime contains no protocol branch. Dialect knowledge lives in the adapter, which is Parnas’s criterion applied to the part of this system most likely to change.
  • Answers arrive in three places, and hidden reasoning in three more. The canonical text holds only the answer; reasoning stays in the preserved bytes.
  • Success has layers. Record which layer failed. Transport, completion, answer presence, syntax, shape, values, task criteria and verification each answer a different question. Schema-constrained structured output from the major providers moves syntax and shape into an enforced frame; it does not establish completion, the absence of refusal, or truth. In Chapter 28’s 84 prompt-only JSON calls, 0 answers were malformed, while one produced no answer, twenty correctly abstained, nine failed grounding, and one well-formed grounded answer was wrong.
  • The same request was reported as 38, 279 and 73 input tokens, output tokens mostly paid for invisible reasoning, and one route reported zero reasoning tokens beside a reasoning field. Tokens do not transfer across routes, and non-equivalent usage must be preserved, not forced.
  • Offline, all 24 synthetic conformance cases passed with the network refused: each of eight case types produced the same canonical outcome across three dialects. Three captured live bodies also replayed to their independently checked text and completion signals, for 27 of 27 offline cases in total. Live, three routes executed once each. The layers support different claims and are never pooled.
  • Still weak: unknown completion counts as success, tool-only replies are retried, no headers or request ids were captured, usage is two numbers, the hash is semantic, and refusal has no state.

Next

The Chat route reported 279 input tokens for a request another route counted as 38, said 192 of them were cached, and reported zero reasoning tokens while returning 780 characters of reasoning. The observation is preserved exactly. What it means is not settled: whether cached tokens are part of the input or beside it, whether “zero reasoning” can be believed, whether two routes’ “input” can ever be added up.

Whoever settles that will get it wrong at least once, as Chapter 11’s classifier did. So the interpretation has to be versioned, replaceable, and unable to rewrite the observation it came from.

Continue with Token Counts Don’t Add Up.

References

Implementation and evidence sources: the preserved Stage 12 protocol-conformance run is historical evidence from 13 September 2026; current-source claims above were checked against the later CodeAI successor rather than projected backward into that bundle. src/codeai/adapters.py: PreparedCognitionRequest, transport-observation types; src/codeai/providers.py: OpenCodeCognitionAdapter.prepare, OpenCodeCognitionAdapter.send, OPENCODE_ENDPOINTS, _messages_text; src/codeai/interpretation.py: _COMPLETION_REASONS, ATTEMPT_POLICY_V2, decide_attempt; src/codeai/domain.py: GenerationState. Current regression coverage includes tests/test_request_plan.py, tests/test_messages_codec.py and tests/test_transport_observation.py. Stage 12 evidence remains under experiments/applied-ai/evidence/protocol-conformance/ (synthetic/, offline/, live/, independent expectation checks); no live Stage 12 call was rerun for this publication review. Layer counts in “Success has layers” were recomputed from experiments/applied-ai/evidence/execution-ladder/2026-09-14-7a0d43b/ (ladder/ and top-first/ events and provider logs, analysis.json) with that bundle’s own parse_answer and ttl_check in run.py.