← Applied AI

Token Counts Don't Add Up

Before you add, compare, price or route on the usage numbers a provider reports, decide what each one means. The same field name can count cached input inside the number on one API and beside it on another. Normalize the equivalent, keep the related apart with the relation declared, keep the unknown unknown — and know that this makes one route's numbers coherent without making two models' tokens one unit.

Part 2 — Get the Model Out of the Chat Box

A field name is not a unit: normalize at the boundary

Every AI application ends up doing arithmetic on numbers a provider reported. It adds tokens into a budget, divides an invoice by them, compares two models on them, and routes work to whichever looks cheaper per token.

Those numbers arrive in fields with shared names: input_tokens, cached_tokens, reasoning_tokens. The shared names hide different rules. On one major API, cached input is counted inside input_tokens. On another, it is reported beside it. Add the two and the result is wrong even though every number in it is correct. Nothing crashes. The budget, the cost comparison and the routing decision are simply built on a sum that means nothing.

This chapter is about the step that has to come before that arithmetic:

field name  ≠  meaning  ≠  a unit you can add across routes

By the end of it you will be able to read a usage record quantity by quantity. You will know which numbers are equivalent and can be normalized, which are related but must be kept apart with their relation declared, which have unknown meaning and must stay unknown, and when a missing value turns a total into a lower bound. You will also see where that work stops: normalization makes one route’s numbers coherent, and it still does not make two models’ token counts comparable.

Three numbers that will not add

Three calls in Chapter 12 performed the same small task, through the same gateway, and all three succeeded. One route reported 38 input tokens. Another reported 279. The third reported 73.

Put those in a spreadsheet and the next move writes itself: sum them, average them, divide the invoice by them, route to whichever chamber is “cheapest per token”. Every one of those moves assumes the three numbers measure the same thing.

They do not, and the reason is not that one of the providers is wrong.

In 1999 NASA lost the Mars Climate Orbiter. Its investigation board found the root cause was the “Failure to use metric units in the coding of a ground software file” (NASA, 1999). A program called SM_FORCES produced thruster impulse data in English units. The interface documentation required metric, so the trajectory modelers built to that requirement. Nothing crashed; the numbers flowed, were accepted, were used, and over nine months put the spacecraft about 170 kilometers lower than planned.

The field existed on both sides of the interface. The unit did not travel with it.

Which usage quantities are actually equivalent enough to normalize, and what should a system do with the ones that are not?

A name is not a unit

Stevens defined measurement as the assignment of numerals to objects or events according to rules, and showed that the rule decides which operations on the numerals mean anything (Stevens, 1946). Two numbers produced under different rules cannot be added just because they are both integers, and they certainly cannot be added because they share a field name.

Token usage is a textbook case. The field names are shared across API dialects. The rules behind them are not. Here is what the two documented dialect families say, taken from their own documentation:

QuestionOpenAI ResponsesAnthropic Messages
Is cached input part of input_tokens?Yes. Cache reads and cache writes are components of input_tokensNo. input_tokens excludes both; total input is cache reads + cache writes + input_tokens
Where cache activity is reportedinput_tokens_details.cached_tokens, .cache_write_tokenscache_read_input_tokens, cache_creation_input_tokens
Is reasoning part of output_tokens?Yes: output_tokens_details.reasoning_tokens, billed as output and not visibleYes: output_tokens_details.thinking_tokens, billed as output
Is a total reported?total_tokensNo
How cached input is pricedCache reads and writes can have rates distinct from ordinary input, with model-specific rulesCache reads and writes have separate rates from ordinary input; write rates depend on cache lifetime and read rates can vary by model

Sources: OpenAI’s prompt-caching and reasoning guides (OpenAI); Anthropic’s prompt-caching and extended-thinking documentation (Anthropic). Both were read again on 18 September 2026, and both change.

Read the first row again. The same field name, input_tokens, has opposite relationships to cached input in the two dialects. In one, the cache is inside the number. In the other, it is beside it. Code that reads input_tokens from both and adds them is not normalizing. It is doing exactly what the orbiter’s ground software did.

The last row makes the point sharper still. Even inside one dialect, ordinary input, cache reads and cache writes can carry different rates. A “total tokens” figure therefore cannot determine cost unless the accounting layer also knows which components the total contains and which rate applies to each one.

Four rules for a field

So normalization is a decision, made per quantity, and four rules cover it. Every system that consumes usage makes these choices somewhere, usually by accident. CodeAI’s usage interpreter applies exactly these:

  1. Equivalent: normalize. A route’s reported prompt count and another route’s input count are the same kind of quantity (tokens the model processed as input) and can share a canonical component, provided the dialect’s rule for what is inside that number is known.
  2. Related but not equivalent: preserve separately, with the relation declared. Cache reads are inside input in one dialect and added to input in another. They get their own component, tagged subset_of_input or additive_to_input. Reasoning gets its own component, tagged subset_of_output. Nothing is summed until the relation says it may be.
  3. Unknown semantics: preserve raw, report UNKNOWN. A field with no rule, such as audio_tokens inside a Chat usage detail, is kept by path and not interpreted, not dropped, and never added.
  4. A provider’s total is an observation, not an answer. A reported total_tokens is kept as reported and compared with the derived total. If they disagree, both survive and the disagreement is recorded as a conflict. Neither replaces the other.

And one rule that runs through all four, from Chapter 11: absence is not zero. When a component required by a formula is not reported, the derived quantity stays UNKNOWN rather than borrowing a zero. Where the known additive components support a meaningful lower bound, the interpreter records that bound explicitly; a lower bound is useful evidence, but it is not a total.

From those rules the interpreter derives three quantities, each under a named formula:

Derived quantityMeaningWhere cache is inside inputWhere cache is beside input
total_inputEverything processed as inputinputinput + cache_read + cache_write
fresh_inputInput neither read from nor written to cacheinput − cache_read − cache_writeinput
processed_totalEverything processedinput + outputinput + cache_read + cache_write + output

In code, the decision lives in a small table: for each dialect, where each component is found and what relation it has to its parent. Reduced from CodeAI’s usage interpreter:

_RULES = {
    "responses": {
        "input_includes_cache": True,
        "input":          (None, "input_tokens", PRIMARY),
        "cache_read":     ("input_tokens_details", "cached_tokens", SUBSET_OF_INPUT),
        "reasoning":      ("output_tokens_details", "reasoning_tokens", SUBSET_OF_OUTPUT),
        "total_reported": (None, "total_tokens", REPORTED_TOTAL),
        ...
    },
    "chat_completions": {   # rule_source: Responses semantics "applied by analogy ... (not separately verified)"
        "input_includes_cache": True,
        "input":          (None, "prompt_tokens", PRIMARY),
        "cache_read":     ("prompt_tokens_details", "cached_tokens", SUBSET_OF_INPUT),
        ...
    },
    "messages": {
        "input_includes_cache": False,
        "input":          (None, "input_tokens", PRIMARY),
        "cache_read":     (None, "cache_read_input_tokens", ADDITIVE_TO_INPUT),
        "total_reported": None,
        ...
    },
}

Every rule also records its rule_source, so an interpretation can say which documentation it relied on, and the Chat rule’s source admits in its own text that it is an analogy. A protocol with no entry gets no derived quantities at all; the interpreter reports no_semantic_rule_for_protocol rather than guessing.

Testing the rules on real responses

Testing the rules needs no new model calls. The raw material is the response bytes Chapter 12 already preserved: three live captures, one per dialect, plus the decoded payload of Chapter 11’s truncated call.

Two interpreters read those bytes. usage-semantics-v1 reproduces the historical adapter view exactly: two numbers, input and output, extracted as the adapters always extracted them. usage-semantics-v2 applies the four rules. Both are pure functions of the preserved observation. Neither writes anything, and the runtime’s recorded usage and past decisions are untouched.

Before any code ran, the expected v2 result for every case was written by hand from the dialect documentation, and that file’s hash is recorded in the report. Then the harness ran with outbound sockets refused:

Casev1: input / outputtotal inputfresh inputprocessed totalWhat v2 also reported
Responses / gpt-5.6-luna38 / 26383864—
Chat / mimo-v2.5279 / 16627987445reasoning tokens reported as zero beside reasoning content
Messages / minimax-m2.773 / 220UNKNOWN (≥ 73)73UNKNOWN (≥ 293)thinking content present, thinking tokens not reported
Chat / mimo-v2.5 (Chapter 11, decoded)264 / 25626472520reasoning tokens reported as zero beside reasoning content

Alongside those four real cases, a synthetic specification corpus of twenty cases covered edge conditions outside the four captures: absent usage, partial usage, measured zeros, a conflicting reported total, cache reads inside and beside input, and cache writes without reads. The remainder stressed reasoning inside output, an estimate, a gateway cost field, an unrecognized detail, malformed counts such as negative, string and null values, subsets larger than their parents, and a protocol with no rule at all.

All 24 cases matched the hand-written expectations. Every source byte hash was identical before and after. No network calls were made. The same projection was then run through Runtime.interpret_usage_as over copies of the actual recorded Stage 12 attempts. It read the preserved transport bodies, appended no ledger events, changed no artifacts, and left each attempt’s recorded usage in its historical two-number form.

What the numbers say once they are interpreted

Look at the fresh-input column.

Apply each declared rule within its own route first. In the Responses capture, the cache components used by the rule were reported as zero, so v2 projects 38 fresh-input tokens. Chat reported 192 cached tokens inside its 279-token input count, so under v2’s explicitly analogical Chat rule the projection is 87 fresh-input tokens. Under the documented Messages rule, input_tokens excludes cache activity, so the reported 73 is the fresh-input component even though total input remains UNKNOWN because the cache-read and cache-write fields were absent.

Chapter 11’s different MiMo request also reported exactly 192 cached tokens. The repeated count suggests a route-level cached prefix of the same reported token length, but the record holds the counts, not enough evidence to establish that the cached bytes were identical. More importantly, 38, 87 and 73 are clearer interpretations under their respective rules. They are not a cross-model ranking, and usage-semantics-v2 correctly keeps route conformance unverified.

Now look at what did not become a number. Under the Messages rule, cache reads and writes are added to input. The MiniMax route left both fields out of the payload. The interpreter keeps 73 as a floor with lower bounds of 73 and 293 for total input and processed total. v1 would have given you 73 and let you believe it was the whole input. v2 gives you 73 as a floor and says so.

That is the difference between a normalizer and a renamer. A renamer maps prompt_tokens to input_tokens and moves on. An interpreter asks what is inside the number first.

What v2 caught that v1 could not

Four things surfaced that the historical view had no way to express.

A reported zero beside visible reasoning content. Both MiMo calls reported reasoning_tokens: 0 while returning reasoning content: 780 characters in Chapter 12, 1,128 characters in Chapter 11. v2’s Chat rule treats reasoning tokens as a subset of output by analogy to the documented OpenAI usage semantics; Chapter 13 does not claim that this OpenCode route is a conforming OpenAI implementation. The interpreter therefore does not “correct” the zero, because it has no evidence for the right count. It records reasoning_tokens_zero_with_reasoning_content, preserves the reported zero, and marks route conformance unverified. The payload does not fit the assumed usage mapping cleanly on either call; that is evidence to retain, not a number to repair.

Content with no count. MiniMax returned a 1,094-character thinking block and no thinking_tokens field. v2 marks reasoning as not reported, not zero.

Silent coercion in the old path. Fed a count of -5, v1 recorded −5. Fed the string "100", v1 recorded 100. That is what the adapters’ historical extraction does: it coerces whatever it finds into an integer. v2 marks both values invalid, derives only a lower bound from what remains, and records which path was wrong. A bad count that becomes a plausible number is Chapter 12’s silent-configuration debt arriving in the measurement layer.

Unrecognized detail and gateway cost. The Chat routes reported audio_tokens in both detail objects. v2 keeps them by path and interprets nothing. All four responses carried a gateway field "cost": "0", a string. v2 preserves it with its path and marks accounting as unknown. The same zero appears on free-model and subscription routes alike (Chapter 11), so on its own it is not a price.

Even interpreted, tokens are not a shared unit

It would be tempting to read 38, 87 and 73 as three comparable fresh-input counts. They are not, and this is the limit of normalization.

Each count came from a different route serving a different model. The preserved responses establish what each route reported; they do not establish that all three routes used the same tokenizer, the same preprocessing, or even an identical definition of the token unit beyond the dialect semantics already examined. Petrov and colleagues showed how sharply tokenized length can vary across tokenizers, including disparities of up to fifteen times across languages (Petrov et al., 2023). That result is a warning about assuming commensurability, not proof of which tokenizer OpenCode used on these three calls.

So normalization makes a route’s usage internally coherent under a declared rule. It can state how reported quantities relate, which parts are cached or reasoning, and which quantities remain unknown. It does not by itself establish that the route conforms to that rule, and it cannot establish that two routes’ token counts share one unit without additional evidence about the tokenization and preprocessing behind them. That is why the book compares chamber economics at the outcome level rather than using cross-model token counts as a shared denominator: cost per accepted outcome, with accepted-but-wrong outcomes kept visible, and cost per correct acceptance only where correctness is independently known.

Observation, interpretation, accounting

The chapter’s architecture is three layers, but the current implementation connects them asymmetrically:

observation     the preserved response bytes, content-addressed         never rewritten
     ↓
interpretation  usage-semantics-v1 / v2: replaceable readings           pure projections
     ↓
accounting      bill, quota draw-down, budget charge                    separate policy

usage-semantics-v2 deliberately stops before accounting. CodeAI already has a narrower, versioned per-token price table for models whose rates it knows, but that mechanism does not make this projection understand subscription allowances, free routes with conditions, or component-specific pricing from provider usage fields. v2 is not consumed by recorded attempt totals or experiment budgets yet, so richer usage-derived accounting remains unbuilt here rather than being silently inferred.

The same architecture as preserved bytes plus replaceable readings:

    flowchart TD
    RAW["raw response bytes<br/><i>content-addressed, never rewritten</i>"] --> V1["usage-semantics-v1<br/><i>reproduces the historical two-number view</i>"]
    RAW --> V2["usage-semantics-v2<br/><i>components · relations · UNKNOWNs</i>"]
    RAW -.->|"new docs, new evidence"| V3["usage-semantics-v3<br/><i>another projection over the same bytes</i>"]
    V1 -.-> HIST["historical recorded usage<br/><i>left unchanged</i>"]
    V2 -.-> AC["richer accounting<br/><i>not wired into the runtime; needs a pricing shape</i>"]
    V3 -.-> AC
    style AC stroke-dasharray: 4 4
  

The versioning protects the reading without rewriting the evidence or the history. v1 remains available because it reproduces the adapter extraction that produced the historical two-number view; v2 is a non-mutating alternative projection over the same observation. Retry decisions and call status are not consequences of usage-semantics-v1 or v2: they belong to the separately versioned attempt interpretation and decision policy. If usage documentation changes, or a route is shown not to conform to its declared dialect, a later usage interpretation can reread the same bytes without moving prior records.

The same discipline covers every layer Chapter 12 separated above the transport: mapping a finish signal, parsing an answer, extracting its fields and deciding whether a null means abstention or failure are all interpretations of one preserved observation, and each can be wrong and need replacing.

Where it is still weak

  1. The Chat rule is an analogy. OpenAI’s documentation establishes the relations for Responses. v2 applies them to Chat detail fields by analogy, and says so in its rule source.
  2. Route conformance is unverified everywhere. These are OpenAI- and Anthropic-shaped dialects serving MiMo, GPT and MiniMax models through OpenCode. A dialect rule describes what an API family documents; it does not prove that a gateway route implements those semantics exactly. The two MiMo responses — visible reasoning content beside a reported reasoning count of zero — show that the route does not fit the assumed mapping cleanly. That is enough to keep conformance unverified; it is not enough to diagnose exactly where the mismatch originates.
  3. The real corpus is thin. Four real responses, two from the same model. None exercised a Messages cache read, a non-zero reasoning count, a cache write, or a conflicting total. Those rules are tested only synthetically.
  4. Nothing consumes v2 yet. Recorded attempts, call totals and experiment budgets still use the two-number v1 view, so a budget still counts input plus output, with cached prefixes included. The recorded path has a failure of its own that no interpretation can repair: Chapter 28 found an adapter that discarded usage the provider had reported whenever a response arrived with no text, before anything downstream could read it.
  5. A missing Messages cache report blocks a total. That is correct, and it means budgets on Messages routes lose precision until either the route reports cache fields or a documented omission convention justifies treating absence as zero.
  6. The consistency check is narrow. It catches a zero reasoning count only when reasoning content is visible in the response. An under-reported non-zero count, or reasoning a route does not return, passes unnoticed.
  7. Unrecognized-path detection looks one level deep, and estimates are carried but never produced.

Do this now

Forty minutes. Find out what your token numbers contain.

  1. For each dialect you consume, fill in the relation table from its documentation: is cached input inside or beside input_tokens, is reasoning inside output, is a total reported? If you cannot fill a cell, that relation is currently a guess in your code.
  2. Take your last twenty usage records. Count how many carry cache, reasoning or detail fields that your code ignores or adds.
  3. Search your code for the point where “not reported” becomes 0. There is almost always one.
  4. Take any two routes you compare and compute fresh input for each under its own dialect’s rule. Does the ranking you were using survive?

If you are building with an assistant:

Add a versioned usage interpretation over preserved provider responses.
Do not change recorded usage, past decisions or stored bytes.
- Keep the historical two-number view as v1, reproduced exactly.
- In v2, read each component by documented path: input, output, cache read,
  cache write, reasoning, reported total. Tag each with its relation
  (subset of input, additive to input, subset of output) per dialect rule.
- Derive total input, fresh input and processed total only when every needed
  component is reported and valid; otherwise UNKNOWN with a lower bound.
- Keep a reported total and compare it with the derived one; record conflicts.
- Preserve unrecognized fields by path. Never coerce strings, negatives or
  nulls into counts.
- Write expected results by hand from the documentation BEFORE running, then
  replay real stored responses and a synthetic corpus, checking that source
  hashes are unchanged. No network calls. No billing.

Failure modes

  • Normalizing by field name. input_tokens has opposite relations to cached input in two major dialects.
  • Adding nested details to their parents. Cached and reasoning tokens counted twice.
  • Filling “not reported” with zero. A lower bound silently becomes a total.
  • Replacing a provider’s total with your derivation, or the reverse. Keep both and record the conflict.
  • Coercing bad counts. -5 and "100" became plausible numbers in the historical path.
  • Believing a provider’s breakdown. Zero reasoning tokens arrived twice beside visible reasoning.
  • Comparing token counts across models. Different tokenizers, different units, even after normalization.
  • Treating tokens as cost. Fresh input, cache reads, cache writes and output can carry different rates; a token total does not determine price.
  • Overwriting the old interpretation. Decisions made under it lose their basis.

What this chapter established

  • A field name is not a unit. The same name, input_tokens, counts cached input inside the number in OpenAI Responses and beside it in Anthropic Messages. Arithmetic across those fields is wrong even when every number is right, just as the Mars Climate Orbiter’s thruster data was right in the wrong unit.
  • Normalize by rule, quantity by quantity. Normalize the equivalent. Keep the related-but-not-equivalent separate, with its relation declared. Keep unknown semantics raw and UNKNOWN. Keep a reported total as an observation and record any conflict with the derived one. When a formula needs an unreported component, keep that derived quantity UNKNOWN; record a lower bound only where the known additive pieces justify one.
  • Keep the layers apart. Observation is preserved evidence. Usage interpretation is a versioned reading of that evidence. Accounting is a separate job with its own pricing assumptions. Reinterpreting usage must not rewrite the observation or retroactively move historical records and decisions.
  • Know the limit. A declared rule can make one route’s usage interpretation coherent without proving that the route conforms to that rule, and it does not make different routes’ or models’ token counts one unit. Compare chamber economics using outcome-level evidence and process cost, not cross-model token totals.

What CodeAI showed. usage-semantics-v2 was built as a pure, versioned projection beside the historical two-number v1 view. Across four real responses and twenty synthetic cases, all 24 outputs matched expectations written before the run; source hashes were unchanged, no network calls or ledger writes occurred, and recorded usage stayed untouched. Under the declared rules, v2 projected 38 fresh-input tokens for the Responses capture, 87 for Chat after separating 192 reported cached tokens from 279 reported input, and 73 for Messages while leaving total input UNKNOWN because cache components were not reported. Those matches establish conformance of the implementation to the frozen expectations; they do not establish that the gateway routes themselves conform to every documented dialect semantic.

v2 also surfaced what v1 could not express: reasoning counts of zero beside visible reasoning on two calls, reasoning content with no count, invalid counts silently coerced, unrecognized fields, and a gateway "cost": "0" string whose accounting meaning is unknown. This usage interpretation did not build a bill; richer accounting needs a separately justified pricing shape.

Evidence notes

The preserved Stage 13 run was made from uncommitted CodeAI code; its manifest pins the base version, the modified and untracked paths, and the exact source hashes used. The implementation was committed afterwards, and a clean rerun from that committed code is preserved separately in usage-semantics/offline-862d64d/, with no uncommitted paths, identical source hashes and identical results. The original bundle is kept as it was: a run from uncommitted code.

Next

Earlier in this part, Chapter 11’s offline trace ended on a line the book has been carrying ever since: task_status: not automatically completed.

Every chapter since has made the call more inspectable. It records what was intended, preserves what came back, survives a change of dialect, and can now interpret its usage numbers under named rules. None of that says whether the work is done. Even a normally completed generation with a coherent usage interpretation is still only a candidate. The producing call cannot establish its own task completion; Chapter 14 adds the criteria, check, acceptance and completion evidence needed to make that a process fact.

Continue with A Successful Call Is Not Finished Work.

References

Implementation sources: the evidence run and its clean rerun are described under Evidence notes; their manifests are experiments/applied-ai/evidence/usage-semantics/offline/manifest.json and .../offline-862d64d/. Relevant symbols are interpret_usage, USAGE_SEMANTICS_V1, USAGE_SEMANTICS_V2, Component, Derived, UsageInterpretation, and Runtime.interpret_usage_as; tests and harness are tests/test_usage_semantics.py and experiments/usage_semantics_demo.py. Evidence: experiments/applied-ai/evidence/usage-semantics/ (expected.json authored before the run, offline/ report and table, runtime-projection.json). Source observations: experiments/applied-ai/evidence/protocol-conformance/live/ and experiments/applied-ai/evidence/ch11-live-opencode/.