← Applied AI

The Price of Intelligence

Cost shape is an architectural property: a cheaper model changes the coefficient, while converting a model call into owned machinery changes the shape. This chapter works out where intelligence should stay a recurring runtime purchase and where the system should own it — explore, extract, own, escalate — priced per verified outcome and disciplined by a verifier that checks what you actually care about.

Part 1 — Where You Stand

The invoice nobody runs at prototype time

The paragraph review from Chapter 1 costs about 1,200 tokens in and 400 out. At prototype scale that is a rounding error, and it is why nobody computes the next four numbers.

A manuscript has 3,000 paragraphs. That is one full pass.

The review is not run once. Every edit invalidates the review of the paragraph it touched, and a careful author edits most paragraphs four or five times.

The product has users. Say a thousand, each with their own manuscript.

And it runs every month, forever, because the product is not a thing you finish — it is a thing you operate.

Nothing in that chain is exotic. Each step is a multiplication anyone could do, and almost nobody does it at the point where the architecture is still cheap to change. By the time the invoice is large enough to notice, the design that produces it has been load-bearing for a year.

So do it now, with round numbers and prices chosen only for illustration:

tokens_in, tokens_out = 1_200, 400        # one paragraph review
paragraphs, edits_per_paragraph = 3_000, 4.5
users, months = 1_000, 12
price_in, price_out = 1.00, 4.00          # $ per million tokens: illustrative, not a quote

reviews = paragraphs * edits_per_paragraph * users * months              # 162,000,000
annual = reviews * (tokens_in * price_in + tokens_out * price_out) / 1e6 # ≈ $453,600

A review that costs about a quarter of a cent becomes roughly $450,000 a year. That figure is illustrative arithmetic, not a measurement. Change the prices and the total moves proportionally.

Now notice what a price change can and cannot do. A model ten times cheaper divides the total by ten. It is still a product of five factors, four of which grow with success, billed every month for as long as the product runs. A cheaper model changes the coefficient. Only an architectural change alters the shape. A cheap recurring variable cost is still a recurring variable cost, and that distinction outlives every price in this chapter.

Where should intelligence remain a recurring runtime purchase, and where should it be converted into something the system owns?

Software used to get done

Here is the economic fact that made software one of the best businesses of the last forty years, stated carefully because we are about to break it.

Traditional software has a high fixed cost to produce a capability and a low marginal cost to exercise it. You pay engineers for six months to build a spell-checker. Checking the ten-thousandth document still costs something — compute, storage, bandwidth, support, security, operations. But it does not cost the development again. The expensive part, working out how to check spelling, was paid once and amortizes across every use.

That is what people mean, without quite saying it, when they say a game ships or an application is finished. Not that no further work will ever happen — that the cost of producing the capability has stopped, and everything from here is operation and improvement.

Now put a model call in the request path.

Every use of that capability now carries a metered inference charge: a per-token payment, usually to a third party, repeated on every execution and growing with your success. Servers were always a running cost. What is new is that part of the thinking — the part that used to be paid for once, in development — is bought again each time the capability runs.

You have moved part of the cost of producing a capability from bounded construction into recurring operation. That is not an implementation detail. It changes the economics of that capability.

This is the thing that does not add up, and it is worth being precise about where the mistake is. The mistake is not using AI. The mistake is leaving it in the finished product’s critical path, by default, without ever asking whether it still needs to be there.

So the working principle for the rest of this book:

A model call in your steady-state product is a recurring cost that should have to earn its place.

Read its scope before you apply it. Some runtime calls genuinely are the product: interpreting a novel document, handling arbitrary user language, or producing a proposal whose useful space cannot be specified economically in advance. The principle is aimed at operations that stayed in the request path after they became understood, bounded and checkable because nobody went back to reconsider them. Every steady-state call is a candidate for review. The ones that survive that review have earned their place.

Intelligence is a means of production

Turn that around and the strategy becomes obvious, and a bit surprising.

Model use during construction is usually a bounded project cost rather than a charge repeated on every execution. Strong models can justify premium rates here when they shorten exploration, design, or implementation, because the calls are attached to the work of building an increment rather than to every later use of it.

Model use at runtime is different. Every call is a usage-linked recurring cost, repeated for as long as that operation remains in the request path.

The purpose of strong intelligence during construction is often to discover structure that no longer needs the strong model once it has been found. That gives a pattern you can reuse on any feature:

    flowchart LR
    subgraph BUILD["construction — fixed cost, ends"]
        B1["EXPLORE<br/><i>strong model maps the space</i>"] --> B2["EXTRACT<br/><i>rules, schemas, checks, examples</i>"]
        B2 --> B3["OWN<br/><i>deterministic code or a smaller component</i>"]
    end
    subgraph RUN["runtime — recurring cost"]
        R1["ESCALATE<br/><i>only the irreducible long tail</i>"]
    end
    B3 -->|"baked into the artifact"| RUN
    style RUN stroke-dasharray: 4 4
  
  • Explore. Use strong intelligence to understand the space: what the inputs look like, what the categories are, where the hard cases live.
  • Extract. Turn the repeated structure into something that is not a model call — rules, a schema, checks, a labeled set of examples.
  • Own. Move the stable part into deterministic code, or into a smaller component whose cost and behavior you control.
  • Escalate. Leave only the irreducible long tail to runtime intelligence, and send it there on purpose.

Much of what looks like a runtime intelligence requirement is a construction task in disguise. You need to classify support tickets into eleven categories — is that a model call per ticket forever, or one afternoon with a frontier model working out what the eleven categories actually are (explore), written down as definitions and examples (extract), followed by a classifier you control (own), with the tickets it cannot place sent up to a model or a person (escalate)? You need to extract fields from invoices — every invoice forever, or a build phase that discovers the twenty layouts covering most of your volume, with the rest escalated?

Not every exploration compiles away. Sometimes the space has no stable structure worth extracting, and the long tail is most of the traffic. Then runtime intelligence has earned its place, and the job is to price it honestly.

Chapter 3 gave a useful heuristic for the stable part: if variance on a re-run would itself be a defect, look for a deterministic implementation first. It is a heuristic, not a proof that a model is unnecessary; a stochastic generator can still be the right mechanism when the solution is hard to produce but an adequate check is cheap. Now the distinction has a price attached. Every operation you can move from the stochastic column to the deterministic one turns a recurring inference charge into construction cost plus ordinary code upkeep.

The input space alone does not decide the boundary. Open-ended documents can contain deterministic sub-operations, and narrow inputs can still demand semantic judgment that is expensive to specify mechanically.

So the economic version of Chapter 3’s rule is: can the stable part of this operation be specified, implemented and verified more cheaply over its lifetime than continuing to rent general intelligence? If yes, conversion is a candidate. If not — because the specification is expensive, the semantics keep moving, or the long tail is the operation you actually care about — runtime intelligence may have earned its place.

There is a second consequence, and it runs the other way. If intelligence spent on construction is a fixed cost that ends, software that was never worth building for a small audience starts to be worth building for one. The old economics amortized a high fixed cost across many users, which is why applications were built for everyone and then configured by each. Lower the construction cost far enough and a tool shaped around one person’s work can pay for itself — provided its runtime stays mostly deterministic, because that person pays the running cost too. Chapter 30 follows that thread.

What are you actually buying?

Now the question the industry mostly avoids. Models get better. Better at what, and is that the thing you need?

Start with a property that has a cheap mechanical check. In code, compilation and tests often provide exactly that. A candidate either clears those particular checks or it does not, and once it clears an adequate verifier for the property you care about, a more expensive model cannot make that verified property more passed.

That does not make passing tests equivalent to correct or maintainable software. Those are different properties, and they need their own evidence. The ceiling exists only for what the verifier actually establishes.

Writing, images, and other judged artifacts can have mechanical properties too: factual claims can be sourced, constraints can be checked, required structure can be validated. What often lacks a cheap adequate verifier is the quality dimension people actually care about — style, usefulness, originality, taste. Without an operational criterion for that dimension, there is no crisp pass condition that tells you when further capability stopped buying value.

Put those together and you get the claim this chapter needs:

Where an adequate verifier establishes the property you care about, model spending has a ceiling for that property. Where no adequate verifier exists for the objective, spending has no natural stopping rule.

Read the first half precisely, because the ceiling applies only to what the verifier actually checks.

tests pass            ≠  maintainability established
compiler succeeds     ≠  correct behavior established
benchmark score high  ≠  production usefulness established

The exact form of the claim is: where an adequate verifier establishes the property you care about, intelligence beyond what clears it has sharply diminishing value for that property. For every other property, the ceiling says nothing. Adequacy is the hidden variable. Chapter 28 shows what happens when it is missing: a check built to verify grounding accepted a grounded, wrong answer, because whether a question had a single answer was never something it looked for.

This is Chapter 4’s checkability axis, seen from the finance side. Chapter 4 argued that a cheap verifier is what makes an AI process deployable. The same verifier is what gives it price discipline:

without a verifier                      with an adequate verifier
  better output is a preference           acceptable / unacceptable is observable
  preference has no stopping point        cheaper candidates can be compared on the same items
  so spend has no stopping point          so escalation has somewhere to stop

That produces a genuinely uncomfortable corollary:

You are most tempted to buy the expensive model precisely where you are least able to measure whether it helped.

For a code property covered by tests, you can run the experiment directly: does the cheaper model clear those tests at an acceptable rate? Writing can be evaluated too — with factual checks, blinded preference studies, downstream task outcomes, or another criterion chosen in advance — but only after you decide what improvement means. If you have not operationalized the quality you are paying for, buying the best available model does not establish that the extra spend helped. The absence of an adequate verifier is therefore not only a quality problem. It is a budget problem.

That is also the honest answer to “is a smarter model worth it.” Name the property first, then ask whether your evidence can detect improvement in it. If an adequate check does, find the cheapest mechanism that clears it. If it does not, you are making a preference purchase without a natural stopping rule, and you should at least know that is where you are standing.

Floor, escalate, route, distill

The strategy follows. The evidence below is a set of external benchmark results and one measured experiment from this book. None of it is a production guarantee; all of it shows a mechanism.

Floor. Start with the cheapest mechanism that could plausibly work — a rule, a small model, an inexpensive API — and measure its pass rate against your verifier. Not its vibe: its pass rate. This is only possible if you built the verifier, which is why Chapter 4 put that row in bold. Model prices can differ by orders of magnitude — FrugalGPT’s 2023 survey of twelve commercial APIs found fees two orders apart — and the ordering changes faster than your architecture does. The live exercise at the end uses whatever the rates are on the day you run it.

The floor can have a zero token price. OpenCode Zen’s documentation lists multiple free routes, and Chapter 28’s free rung accepted 23 of the 30 items that reached it, one of them wrongly. Treat “free model” as a changing category, not as an architectural dependency on one route. OpenCode’s own privacy documentation also attaches model-specific data-use or trial terms to some free routes, so zero price is not the same thing as zero constraint (OpenCode Zen).

A free model still needs the same adequate check as a paid one. Its availability and terms can change, and its data policy can rule it out for sensitive work before quality is considered. That is another reason the model slot stays replaceable (Chapter 10): price, capability, availability and data terms are properties of the current occupant, not of the operation.

Escalate. Send only what the floor could not settle upward:

cheapest plausible mechanism
            ↓
   accepted by the check?
      ┌─────┴─────┐
     yes          no
      │            │
    stop     escalate to the next rung

The saving does not come from having several models. It comes from not paying for a stronger mechanism on items a cheaper one already resolved adequately. Chen, Zaharia, and Zou’s FrugalGPT learned exactly this kind of cascade over twelve 2023-era commercial APIs, with a small DistilBERT scorer deciding when an answer was good enough to accept. Matching the accuracy of the best single API cost 98.3% less on a financial-headline classification task, 73.3% less on a legal classification task, and 59.2% less on a reading-comprehension task; at equal cost, accuracy improved by up to 4% over GPT-4 (Chen et al., 2023).

The “up to 98%” in the abstract is the best of those three datasets. And the authors name the dependency that matters most here: the scorer was trained on labeled examples, which have to come from the same or a similar distribution as the queries it will judge.

That is the whole risk in one sentence. A cascade with a weak acceptance check is a machine for confidently stopping too early.

This book eventually runs a cascade to find out. Chapter 28’s ladder sends forty extraction items up from a deterministic rule, through free, cheap and strong models, to a person. Against sending everything to the strong model first, it was cheaper per accepted outcome under its declared cost scenario and produced one more correct acceptance. It also accepted one wrong answer its check could not recognize as wrong, and because the rule written before the run did not allow any increase in wrong acceptances, the cheaper ladder was not adopted. The two paid rungs also resolved nothing the free one had declined. The ladder looked attractive on cost and failed on exactly the clause that protected against a weak check. Escalation saves money only where the thing deciding “good enough” is adequate to the question.

Route. A cascade and a router save money in different places:

cascade   try the cheaper rung first; escalate after observing failure or doubt
          → saves by stopping cheaply

router    choose the rung from the input, before any call
          → saves by never making the unnecessary first call

The router’s advantage is also its limit: it never sees the answer it is betting on, so its decisions are only as good as the data it learned from. Ong and colleagues trained routers on Chatbot Arena preference data to choose between GPT-4 and Mixtral-8x7B. Their best routers cut cost relative to always calling GPT-4 by 3.66× on MT Bench while keeping 95% of GPT-4’s score; on MMLU and GSM8K the savings were about 1.4× and 1.5×, at 92% and 87% of its score. Trained on Arena data alone, the routers did no better than random on those two benchmarks until the training data was augmented with labeled or judge-scored examples. The routers also kept their advantage when routing between model pairs they had not been trained on (Ong et al., 2024).

Hold onto that last property. A model-selection router that survives having its candidate models replaced is more durable than one tied to fixed model identities.

But keep two decisions separate. RouteLLM is choosing which model to call before generation. Chapter 1’s second bet, and Chapter 28’s implemented scheduler, concern which operation the process should perform next; that deterministic decision happens before model selection is relevant. Chapter 28 keeps a separate model-router challenge as an unrun experiment. Evidence for model-selection routing therefore does not count as evidence for the next-operation scheduler, or vice versa.

Distill. Once the task and the target behavior are narrow and measured, the expensive general model can become a teacher: its outputs become training data or supervision for a cheaper specialized component. Hsieh and colleagues used rationales from the 540B-parameter PaLM as extra supervision for small T5 models on four NLP benchmarks. A 770M-parameter model, over 700× smaller, beat few-shot-prompted PaLM on ANLI using 80% of the labeled training set; on e-SNLI a 220M-parameter model, over 2,000× smaller, beat it using 0.1% of the dataset (Hsieh et al., 2023).

Read that as the answer to “paying for intelligence forever,” within its limits. Distillation did not turn a frontier model into a permanent possession. It needed a narrow task, training data, an evaluation to show the student was good enough, and infrastructure to train and serve the student. The student is reliable on the distribution it learned and can degrade off it without signaling, so it needs the same verifier as the thing it replaced, plus monitoring for drift, plus someone to retrain it when the task moves. Provider terms may restrict training on outputs, so the right to do it has to be checked, not assumed.

What you gain is not ownership of intelligence. It is reduced dependence on the expensive general model for one well-measured job: per-call rent replaced by training, evaluation and serving costs you control.

The ratchet, ordered by how much of the capability you end up controlling:

MoveWhat it does to your cost curve
Prompt/context trimmingReduces the coefficient. Does not change the shape.
Cheaper model at the floorReduces the coefficient, sometimes by orders of magnitude.
Cascade / routePays for premium capability only on the items that need it — if the check or the routing data is adequate.
Distill into a smaller specialized componentReplaces per-call rent on the general model with training, evaluation and serving you control.
Move the operation to deterministicRemoves the inference cost; leaves ordinary code upkeep (Chapter 3).

Every row down that table is more permanent than the one above it. Row one is where most cost work starts, because it needs no architecture.

The table is not a staircase every feature must climb. A frontier prototype can go straight to deterministic code once exploration has revealed the rule. A feature can jump from a frontier model to a small one. And a feature whose inputs stay genuinely open-ended can legitimately keep a frontier model in its request path for its whole life. The ratchet orders the moves by permanence; it is not a maturity model.

The unit is the verified outcome

Every number so far has been a price per call. That is usually the wrong unit.

A model at a tenth of the price that needs twenty attempts per usable result is not cheaper. A stronger model that avoids sending an item to a person may be. But keep acceptance and correctness separate. The process can always measure cost per accepted outcome under its declared check. It can measure cost per correct accepted outcome only where independent evidence or frozen gold establishes which acceptances were actually correct.

Chapter 28 deliberately reports both. Its economic metric is cost per accepted outcome, while correct acceptances and accepted-but-wrong outcomes remain separate quality columns because the checker itself can be inadequate.

The arithmetic is simple enough to keep in your head:

And “total process cost” is not only tokens:

total process cost = model calls
                   + deterministic compute
                   + human review and escalation
                   + failure and rework
                   + engineering and maintenance

You do not need to price every term precisely to use this. You need to remember the terms exist, because the cheapest model path can be the most expensive process if it creates more review or rework.

Here is what that looks like in numbers. Everything below is invented for teaching — the prices, the acceptance rates and the cost of a wrong answer found later are illustrations, not measurements — but the structure is the one every cascade has. Ten thousand items go either straight to a frontier model or first to a small model behind a check; whatever the frontier model does not settle goes to a person:

N = 10_000                                   # items
SMALL, FRONTIER = 0.001, 0.01                # $ per attempt
PERSON, REWORK = 5.00, 20.00                 # $ per escalation, $ per wrong answer found later

def run(first_small, small_accept, small_wrong, frontier_accept=0.96, frontier_wrong=0.005):
    model = wrong = 0.0
    to_frontier = N
    if first_small:
        model += N * SMALL
        accepted_small = N * small_accept
        wrong += accepted_small * small_wrong
        to_frontier = N - accepted_small
    model += to_frontier * FRONTIER
    accepted_frontier = to_frontier * frontier_accept
    wrong += accepted_frontier * frontier_wrong
    people = (to_frontier - accepted_frontier) * PERSON
    total = model + people + wrong * REWORK
    return model, people, wrong, total, total / (N - wrong)

run(False, 0, 0)         # frontier first:          model $100, people $2,000, 48 wrong,  total $3,060, $0.307 per verified outcome
run(True, 0.80, 0.002)   # cascade, adequate check: model  $30, people   $400, ~26 wrong, total   $942, $0.094
run(True, 0.90, 0.030)   # cascade, weak check:     model  $20, people   $200, ~275 wrong, total $5,716, $0.588

Read the three rows slowly. The adequate cascade is cheaper than frontier-first on every line: the saving on models is $70, and the saving on people is $1,600, because the small model settled most items without anyone being asked. The weak cascade has the lowest model bill and the lowest people bill of the three — its lax check accepted 90% of the small model’s answers instead of 80% — and it is by far the most expensive process, because the wrong answers it let through came back as rework. On a dashboard that shows model spend, the weak cascade is the winner. On the only unit that matters, it is the loser.

The relationship is easier to see than to describe:

Stacked bars comparing total process cost across three strategies: the adequate-check cascade is cheapest overall, while the weak-check cascade is the most expensive because rework dominates

Illustrative totals from the scenario above: the adequate cascade is cheapest on every line, while the weak cascade pairs the lowest model bill with by far the highest process cost. Prices and rates are invented for teaching.

Two measurement rules fall out of that. Count accepted-but-wrong outcomes separately, because a check that accepts wrong answers makes every cost-per-outcome figure look better than it is. And never compare arms on model spend alone.

Chapter 28 measured a real, smaller instance. Under a declared scenario of $5 for every question put to a person, its ladder spent about half a cent on models and its top-first control about three cents — a sixfold difference — yet the costs per accepted outcome were $1.06 and $1.45, and almost the whole gap came from two fewer questions to people. Making the model work six times cheaper barely moved the total, because the model was never most of it. Chapter 8 makes the same point about time: optimizing a component that is not the constraint barely moves the whole.

Prices fall; architecture persists

One more force, and it argues against a decision most teams make early.

Epoch AI measured how fast the cheapest list price for reaching a fixed benchmark score fell, across six benchmarks, from 2022 to early 2025. Prices were per token, as a 3:1 blend of input and output rates, with reasoning models excluded. The rate varied enormously by milestone — from 9× to 900× per year, with a median of 50×. For GPT-4-level performance on PhD-level science questions (GPQA Diamond), the price fell about 40× per year (Epoch AI, 2025). Their own caveat matters: the fastest declines began after January 2024, so it is less clear they will persist.

Now assume the decline slows sharply. The argument does not depend on the rate. Capability keeps being commoditized, providers change order, new providers appear, and the model that is the right choice this quarter is unlikely to stay the right choice for the life of your product. Which means:

The durable engineering decision is not which model you chose. It is whether you can change your mind about it without a rewrite.

That is the entire justification for the adapter boundary built in Part 2, and it is why Chapters 12 and 13 contain provider differences and normalize what providers report. Bind the architecture to the operation, not to the current model. A codebase that can swap models can take a price decline as soon as the new model clears the same checks. A codebase welded to one provider’s response format pays for the privilege of not noticing.

And the necessary counterweight, because “prices are falling” is not the same as “your bill is falling”:

unit price falls
  → uses that were not worth paying for become worth it
  → usage rises
  → total spend can stay flat, or rise

That is the familiar rebound pattern of cheapening inputs, not a law of AI; it may or may not happen to your product. But it is common enough that falling prices are a reason to build for substitutability, not a reason to skip the arithmetic at the top of this chapter. A smaller coefficient on a growing shape is still a growing bill.

You cannot manage what you do not measure

You start with one model and one provider, and a set of costs made visible. That is deliberate, and the measurement has to be more careful than a dashboard total.

What has to be counted, per operation:

  • Input and output tokens, separately — they are usually priced differently.
  • Cached versus fresh input. A repeated system prompt and a novel document are not the same purchase.
  • Attempts, not logical operations. One review that silently retried twice cost three times what your code thinks it did. Chapter 11 shows a real client that retries empty output without telling the caller.
  • Failed calls that still bill. A response that arrived and was unusable is a cost with no output.
  • Reasoning tokens, where a provider charges for tokens you never see.
  • Which model, at which version, under which price, at which date.

That last one is the difference between a number and a measurement.

CodeAI records most of this at the attempt level today. Each attempt carries its own input and output counts with a provenance label — measured, estimated or unavailable, with unknown kept as unknown rather than zero — plus the resolved model, any reported revision, the pricing version, and a cost marked estimated or unknown. A model absent from the pricing table gets an unknown cost, not a guessed one. Cached and reasoning tokens are read out of the preserved responses by Chapter 13’s versioned usage semantics, but that interpreter deliberately produces no bill: the recorded cost estimate still prices input and output only. The ledger can tell you how much input was cached; it does not yet tell you what caching saved.

Chapters 12 and 13 are about making these counts survive the trip out of three different provider response formats without being silently double-counted or zero-filled — accounting that is easy to get wrong in ways that quietly understate your spend.

Where this argument is weakest

  • Cascades and routers cost engineering time. At low volume, a week of work to save $40 a month is a bad trade. The ratchet earns its keep at volume, and volume is exactly when the architecture is hardest to change. There is no comfortable answer to that; it is a judgment call about expected scale, made early, on poor information.
  • The strongest cost results are benchmark results. FrugalGPT’s 98% is its best of three datasets (59–98%), on 2023 APIs. RouteLLM’s larger savings are on MT Bench and nearer 1.5× on MMLU and GSM8K. The distillation ratios come from four NLP benchmarks with a 2022-era teacher. The mechanisms transfer; your numbers will be different.
  • Cost per verified outcome is only as honest as the check. A weak check inflates accepted outcomes and flatters the cheap path. Chapter 28’s ladder looked cheaper precisely where its check was weakest.
  • Human and rework costs are scenarios, not constants. The $5 per question in Chapter 28 is declared, not measured, and so are the prices in the cascade example above. Change them and the absolute numbers move, even where the ordering holds.
  • “Use it during construction” assumes construction ends. For genuinely open-ended products it may not, and then the runtime cost is real and must simply be priced into the business.

Do this now

Ten minutes, one table. Run the invoice.

For one AI feature you have built or are planning:

  1. Tokens in and out for a single attempt, priced separately at your provider’s current rates. Write the date beside the price. If the model is free or covered by a subscription, write that instead of zero, with its limits and data terms.
  2. × attempts per useful outcome — retries, re-runs, rejected outputs.
  3. × operations per user per month (be honest about edits).
  4. × users at the scale you are actually aiming for, not today’s. × 12.
  5. Then one row per operation:
OperationCurrent mechanismCost per attemptAttempts per useful outcomeHuman escalation?Verifier?Candidate conversion
deterministic · cheaper model · cascade · route · distill · leave as frontier runtime

The annual figure is the one nobody computes while the architecture is still cheap to change. The row with the largest annual cost and a verifier is your most expensive avoidable runtime dependency — start there. A row with no verifier cannot be safely converted yet, and the next chapter is about that.

Failure modes

  • Discovering the cost curve after the architecture has set. The arithmetic is trivial and almost nobody does it while the design is still cheap to change.
  • Mistaking a price cut for a cost strategy. A cheaper model changes the coefficient; the shape stays.
  • Treating the frontier model as the steady-state default. It is the right tool for exploration and the expensive choice for work that has become well understood.
  • Buying capability in a domain with no ceiling and calling it an investment. Without a verifier there is no evidence the upgrade helped, only a feeling that it did.
  • Trusting a verifier beyond what it checks. Tests passing is not maintainability; a grounded answer is not necessarily the right one.
  • A cascade with a weak acceptance check. It stops early, cheaply and confidently, and its savings show up on the model bill while its costs show up as rework.
  • Optimizing cost per call. The unit is cost per verified outcome, with people and rework in the total.
  • Measuring logical operations instead of attempts. Hidden retries make your cost model wrong in the direction of optimism.
  • Welding to one provider’s response shape. You pay full price through every price decline.
  • A router that spends a frontier call deciding whether to spend a frontier call. The learned routers and scorers above are small classifiers, far cheaper than the calls they gate. A general model in that seat adds cost to the decision about cost.

What this chapter established

  • Cost shape is an architectural property. A cheaper model changes the coefficient; converting a model call into owned machinery changes the shape. The $453,600 invoice is illustrative arithmetic; its five-factor shape is the point.
  • Traditional software amortizes the cost of producing a capability across its uses. A model in the request path puts a metered inference charge on every use of that capability again.
  • A steady-state model call is a recurring cost that should have to earn its place. Stable, well-specified operations are candidates for conversion into owned machinery; open-ended input alone does not prove that a model is required.
  • Intelligence is a means of production: explore with strong models, extract the structure, own the stable part, escalate only the irreducible long tail. Conversion is worth it when the stable part can be specified, built and verified more cheaply over its lifetime than renting general intelligence.
  • Where an adequate verifier establishes the property you care about, extra intelligence has sharply diminishing value for that property. Verifier adequacy is the hidden variable, and the same verifier that makes a process deployable (Chapter 4) gives it price discipline.
  • Floor, escalate, route, distill: cascades save by stopping cheaply, routers by avoiding the first call, distillation by reducing dependence on the general model for one measured job. External results: FrugalGPT 59–98% cheaper at matched accuracy on three datasets; RouteLLM 3.66× cheaper at 95% of GPT-4’s MT Bench score and about 1.5× on MMLU and GSM8K; distilled students 700–2,000× smaller beating a few-shot PaLM teacher on specific benchmarks.
  • A cascade with a weak acceptance check stops too early. This book’s measured ladder was cheaper per accepted outcome and was still not adopted, because it accepted a wrong answer its check could not see.
  • The ratchet orders moves by permanence, not by maturity; some features rightly stay at a frontier model.
  • Optimize the whole process, not cost per call: report cost per accepted outcome alongside accepted-but-wrong outcomes, and where independent correctness is available, compute cost per correct acceptance. Human review, escalation and rework belong in the total.
  • Epoch measured fixed-capability prices falling 9×–900× per year from 2022 to early 2025. The durable response does not depend on that rate: bind the architecture to the operation, keep models substitutable, and remember that cheaper units can mean more use rather than a smaller bill.
  • Cost is measured at the attempt level, including cached input, hidden retries, billed failures and unseen reasoning tokens. CodeAI records attempts, usage provenance and pricing versions today, and does not yet price cache or reasoning components.

Next

One thing is still missing, and without it everything in this chapter is unusable.

You cannot optimize price until you can measure quality. You cannot pay only what clears the verifier if you have not built the verifier. You cannot tell whether the cheaper model was good enough, whether the distilled model held up, whether the cascade stopped at the right place, or whether this month’s process is better than last month’s. Measurement comes before routing.

Vendors sell capability, which is a scalar. Direction comes from measurement, and measuring a stochastic process is harder than measuring a deterministic one — hard enough that most teams quietly skip it and buy capability instead.

The next chapter is about how you would know.

Continue with Intelligence in the Wrong Direction.

References

  • Lingjiao Chen, Matei Zaharia, and James Zou. FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance. arXiv:2305.05176, 2023. https://arxiv.org/abs/2305.05176 — cascade results on HEADLINES, OVERRULING and COQA (Table 3).
  • Isaac Ong, Amjad Almahairi, Vincent Wu, Wei-Lin Chiang, Tianhao Wu, Joseph E. Gonzalez, M Waleed Kadous, and Ion Stoica. RouteLLM: Learning to Route LLMs with Preference Data. International Conference on Learning Representations (ICLR), 2025. https://arxiv.org/abs/2406.18665 — cost-saving ratios over GPT-4 in Table 6.
  • Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alexander Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes. Findings of the Association for Computational Linguistics: ACL 2023, pp. 8003–8017. https://aclanthology.org/2023.findings-acl.507/
  • Ben Cottier, Ben Snodin, David Owen, and Tom Adamczewski (Epoch AI). LLM Inference Prices Have Fallen Rapidly but Unequally Across Tasks. Data Insight, March 12, 2025. https://epoch.ai/data-insights/llm-inference-price-trends