← Agents From First Principles

The Action Boundary

Turn probabilistic model output into explicit, inspectable action proposals by defining a finite action space, validating structure and semantics, separating authorization from execution, and treating failures as runtime evidence.

The previous chapter ended with a division of responsibility:

The model proposes. The runtime decides what may execute. The environment supplies observations about what happened.

This chapter takes the first clause seriously. What does it actually mean for a model to propose an action?

Suppose the model emits the sentence “Search the documentation for the latest PyTorch optimizer API.” A human reads that and understands the intention immediately. A runtime cannot act on it at all, because it needs answers to six questions the sentence does not contain: which action, which arguments, whether those arguments are well formed, whether their values are meaningful, whether this action is permitted in this run, and whether its execution preconditions currently hold.

Only when all six have answers should anything in the environment change. So the mechanism this chapter adds is a boundary that asks them in order:

    flowchart TD
    M[model output] --> P[parse representation]
    P --> S[validate schema]
    S --> V[validate semantics]
    V --> A[authorize]
    A --> C[check preconditions]
    C --> X[execute]
    X --> O[observation]

    P -.rejected.-> R[rejection<br/>tagged with stage]
    S -.rejected.-> R
    V -.rejected.-> R
    A -.rejected.-> R
    C -.rejected.-> R
    R --> O
  

Note the dotted edges. Every stage can refuse, and every refusal ends up in the same place as a success — as an observation the next decision can use. That is the whole design, and it rests on one idea:

A model output can propose an action without having authority to execute it.

That distinction is what turns tool use from a prompt trick into a software interface.


1. Start with the failure

Imagine an agent with two capabilities, search(query) and calculate(expression), and a naive implementation that asks for a command in natural language:

def decide(task: str) -> str:
    return llm(
        f"""
        Available tools:
        - search(query)
        - calculate(expression)

        Task:
        {task}

        Tell me which tool to call.
        """
    )

Sometimes this produces exactly what we hoped for. Often it produces something else, and the something-elses are not variations of one problem:

Model outputWhat went wrong
SearchWeb("PyTorch optimizer state")invented a capability that does not exist
calculate(expression="2 +")real action, unusable argument
I should probably search the docs first.described an intention rather than encoding an action
{"tool": "search", "query": 17}structured, but wrong shape and wrong type

These fail at four different places and want four different responses. Prompt engineering cannot collapse them into one category, because they are not one category. What the runtime needs is an explicit contract it can check.


2. Move from language to a finite action space

The fix begins by changing the question. Instead of what do you want to do?, ask: choose one action from this set and supply its required arguments.

With an action space of {search, calculate, final}, a proposal becomes a data structure:

{
  "action": "search",
  "arguments": {
    "query": "PyTorch optimizer state"
  }
}

JSON is convenient here, but JSON is not the principle. A native function-call object, a typed Python object, a protocol buffer, a JSON Schema or a grammar-constrained representation would all serve. What matters is the shift from free-form intention to a finite action vocabulary with explicit arguments — because that is what makes a proposal inspectable by software rather than only by a reader.

Tool-use research treats function selection and argument construction as distinct capabilities for exactly this reason. Models choose the wrong function, invent APIs, and supply incorrect arguments even when the surrounding natural-language answer sounds entirely plausible.[1][4]


3. Parsing is not validation, and the pre-execution ladder has five rungs

Suppose the model returns this, and it parses perfectly:

{ "action": "launch_missiles", "arguments": {} }

The parser’s success tells us precisely one thing: the representation can be decoded. It says nothing about whether the action belongs to our system. Decoding and validating are different jobs, and collapsing them is how launch_missiles reaches a dispatch table.

Treating validation as a single Boolean loses even more. A proposal can pass one check and fail the next, and which check it failed is the most useful debugging signal the boundary produces. So the boundary has five distinct rungs:

RungQuestionExample failure
1. RepresentationCan the output be decoded at all?truncated JSON, prose instead of an object
2. SchemaDoes it have the right structure?unknown action, missing field, wrong primitive type
3. SemanticsAre the values meaningful for this action?empty query, max_results of 10,000,000
4. AuthorizationIs this allowed for this run, user and context?valid publish in a read-only run
5. PreconditionsCan it sensibly execute right now?file missing, budget exhausted, no connection

We are going to make those five rungs a real type rather than a diagram, because the difference between naming a concept and implementing it is most of this book:

from enum import StrEnum

class Stage(StrEnum):
    REPRESENTATION = "representation"
    SCHEMA = "schema"
    SEMANTICS = "semantics"
    AUTHORIZATION = "authorization"
    PRECONDITION = "precondition"
    EXECUTION = "execution"

The first five values name failures that occur before a side effect is permitted. EXECUTION names the separate failure boundary after an action has crossed those checks and the environment is actually invoked. Together they give the chapter one vocabulary for the complete path from proposal to observation.


4. Validated data should become a different type

There is a design move here worth stating on its own, because it does more work than it looks like it does.

Do not let arbitrary dictionaries flow to execution. After validation, convert the proposal into a typed action:

from dataclasses import dataclass
from typing import Literal

@dataclass(frozen=True)
class SearchAction:
    kind: Literal["search"]
    query: str
    max_results: int

@dataclass(frozen=True)
class CalculateAction:
    kind: Literal["calculate"]
    expression: str

@dataclass(frozen=True)
class FinalAction:
    kind: Literal["final"]
    answer: str

Action = SearchAction | CalculateAction | FinalAction

The executor’s signature now says Action, not dict. That makes the boundary visible in the program’s type structure: static analysis can distinguish accepted actions from untrusted decoded data, and the runtime can be organized so that only the acceptance boundary constructs an Action. Python type hints do not enforce this boundary by themselves. The enforcement comes from the control flow: raw output is decoded and validated first, and the executor is only called with the accepted object produced by that path.


5. Validate the schema exactly

Let us build the conversion by hand so the mechanism stays visible. First, two exception types, so that schema failures and semantic failures stay distinguishable all the way up:

class SchemaError(ValueError):
    """Structure, keys or primitive types are wrong."""

class SemanticError(ValueError):
    """Structure is correct; the values are not usable."""

def require_exact_keys(data: dict, expected: set[str], *, where: str) -> None:
    actual = set(data)
    missing = expected - actual
    extra = actual - expected

    if missing:
        raise SchemaError(f"Missing fields in {where}: {sorted(missing)}")

    if extra:
        raise SchemaError(f"Unexpected fields in {where}: {sorted(extra)}")

Rejecting extra keys matters as much as requiring missing ones. An unexpected field usually means the model has invented a parameter, and silently discarding it hides a decision the model believed it was making.

Now the outer proposal, which dispatches to a per-action validator:

def validate_proposal(data: object) -> Action:
    if not isinstance(data, dict):
        raise SchemaError("Action proposal must be an object")

    require_exact_keys(data, {"action", "arguments"}, where="action proposal")

    name, arguments = data["action"], data["arguments"]

    if not isinstance(name, str):
        raise SchemaError("action must be a string")

    if not isinstance(arguments, dict):
        raise SchemaError("arguments must be an object")

    validators = {
        "search": validate_search,
        "calculate": validate_calculate,
        "final": validate_final,
    }

    if name not in validators:
        raise SchemaError(f"Unknown action: {name}")

    return validators[name](arguments)

An unknown action now fails here and never reaches a tool implementation.


6. Schema validity is not semantic validity

Consider a proposal whose types are all correct:

{
  "action": "search",
  "arguments": { "query": "", "max_results": 0 }
}

query is a string. max_results is an integer. A type-only schema could therefore accept the proposal even though the search is useless. Richer schemas can encode constraints such as minimums, maximums, string lengths, patterns and enumerated values. Where a stable invariant fits cleanly in the schema, encode it there. But not every semantic rule is merely structural. The application may still need to decide whether a value is meaningful in the current domain and context. So each validator does both jobs, and says which one failed:

def validate_search(args: dict) -> SearchAction:
    require_exact_keys(args, {"query", "max_results"}, where="search arguments")

    query, max_results = args["query"], args["max_results"]

    if not isinstance(query, str):
        raise SchemaError("query must be a string")

    if type(max_results) is not int:
        raise SchemaError("max_results must be an integer")

    query = query.strip()

    if not query:
        raise SemanticError("query cannot be empty")

    if not 1 <= max_results <= 20:
        raise SemanticError("max_results must be between 1 and 20")

    return SearchAction("search", query, max_results)

validate_calculate and validate_final follow the same shape: exact keys, then types as SchemaError, then values as SemanticError. The pattern is deliberately boring, and the boringness is the point — this is the layer where we want no cleverness at all.


7. Authorization is a third question again

Suppose the action space later grows to include read_file, write_file, send_message and publish. A publish proposal can be perfectly formed, entirely meaningful, and still not permitted in the current run.

That is not a schema failure and not a semantic one. It is an authorization decision, and it depends on context the proposal itself cannot contain:

@dataclass(frozen=True)
class RunPolicy:
    allowed_actions: frozenset[str]

def authorize(action: Action, policy: RunPolicy) -> None:
    if action.kind not in policy.allowed_actions:
        raise PermissionError(f"Action not allowed in this run: {action.kind}")

Validation asks does this describe a legitimate action? Authorization asks may this run perform it? Keeping them apart matters enormously once tools have side effects, because they change on different schedules and for different reasons — the action contract is a property of the system, while the policy is a property of a particular run.

The model can propose more than the runtime permits. The runtime owns authority.


8. Constrained generation helps, and does not replace the boundary

Many production systems no longer rely only on a prompt that says please return valid JSON. They may use native function calling, JSON Schema constrained output, formal grammars, or constrained decoding. This is a genuine improvement: it can eliminate entire classes of representation and schema failure and, depending on the constraint language, enforce some value restrictions.

It does not eliminate the runtime boundary. Authorization and execution preconditions remain external to constrained generation, and application-level semantic invariants may still depend on context the decoder does not own. Constrained decoding can produce a structurally flawless {"action": "search", "arguments": {"query": "", "max_results": 20}} that is semantically useless, or a perfectly formed publish action in a run with no publishing rights.

The four mechanisms answer four different questions:

MechanismQuestion it answersLives where
Constrained generationCan the output be shaped correctly?inside the model call
Runtime validationDo the structure and values satisfy application invariants?at the boundary
AuthorizationDoes this run have the authority?at the boundary, against policy
Precondition checkIs the environment currently in a state where this action may execute?at the boundary, against current runtime state

Use all four. Expect none of them to cover for another.


9. Repair representation, never invent intent

Suppose a model wraps otherwise valid JSON in a Markdown fence. Stripping a known fence is deterministic representation normalisation, and it is fine:

def strip_json_fence(text: str) -> str:
    text = text.strip()

    if text.startswith("``json"):
        text = text[len("``json"):]
    elif text.startswith("``"):
        text = text[len("``"):]

    if text.endswith("``"):
        text = text[:-3]

    return text.strip()

Now compare it with this, which looks superficially similar:

if "query" not in arguments:
    arguments["query"] = "something useful"

The first recovers a value the model unambiguously supplied. The second fabricates a decision the model never made, and then hides the fabrication behind a successful validation. Every downstream log will show a well-formed search that the agent never chose.

Repair representation when the intended value is already unambiguous. Never manufacture semantic intent merely to make validation pass.

This does not forbid defaults that belong to the action contract itself. If max_results is optional and the runtime defines its default as 5, applying that default is deterministic contract behaviour rather than reconstruction of model intent. The dangerous case is inventing a value only because a particular proposal failed validation.

When required semantic intent is missing or ambiguous, a visible rejection is better than silently inventing it. The rejection preserves the fact that the agent did not make the required decision; an ad hoc substitution erases it.


10. Rejections become observations

An invalid proposal does not have to end the run. It is information, and the previous chapter gave us somewhere to put information.

This is where the Stage enum earns its place. We need an outcome type that carries not just failed but failed where:

@dataclass(frozen=True)
class Accepted:
    action: Action

@dataclass(frozen=True)
class Rejected:
    stage: Stage
    message: str

Decision = Accepted | Rejected

A rejection converts straight into the observation format from the previous chapter:

observation = {
    "kind": "invalid_action",
    "stage": rejection.stage,
    "error": rejection.message,
}

And the policy can retry with the exact complaint rather than a vague instruction to try harder:

def correction_prompt(previous_output: str, error: str) -> str:
    return f"""
Your previous action proposal was rejected.

Previous proposal:
{previous_output}

Validation error:
{error}

Return one corrected action proposal.
"""

The causal structure — proposal, rejection, observation, new proposal — is the loop from the previous chapter becoming concrete for the first time. Notice what the validator did and did not do. It did not make the agent smarter. It made failure observable and attributable, which gives the next decision something specific to respond to.


11. Correction needs a budget

A correction loop must be bounded, and the bound must be visible:

MAX_REPAIRS = 2

REPAIRABLE_STAGES = {
    Stage.REPRESENTATION,
    Stage.SCHEMA,
    Stage.SEMANTICS,
}

def get_valid_action(prompt: str, llm, policy: RunPolicy, context) -> Decision:
    raw = llm(prompt)

    for attempt in range(MAX_REPAIRS + 1):
        decision = prepare_action(raw, policy, context)

        if isinstance(decision, Accepted):
            return decision

        if decision.stage not in REPAIRABLE_STAGES:
            return decision

        if attempt == MAX_REPAIRS:
            return decision

        raw = llm(correction_prompt(raw, decision.message))

    raise RuntimeError("unreachable")

Representation, schema and semantic failures may justify correction of the proposal itself. Authorization and precondition failures normally belong back in the main agent loop, because changing the wording of the same proposal should not change what the runtime permits or what the environment currently makes possible.

A model that fails the same contract three times in a row is telling us something worth knowing — most often that the action schema is ambiguous, or that the task cannot be expressed in the action space we defined. An unbounded retry loop destroys that signal by eventually succeeding through sheer resampling.

Loop budgets and termination get proper treatment later, when we build runtime state and progress. Here we need only one commitment:

A validation repair policy must have an explicit failure state.


12. Validate completely before any side effect

The boundary should be atomic from the validator’s point of view: parse the complete proposal, validate it completely, authorize it completely, check preconditions, and only then execute once.

This is ordinary software engineering, and stochastic model output raises the stakes. A half-interpreted proposal that has already begun writing files leaves the system in a state no rung of the ladder ever approved. The executor should only ever operate on something the runtime has already accepted in full:

def execute(action: Action, tools):
    if isinstance(action, SearchAction):
        return tools.search(query=action.query, max_results=action.max_results)

    if isinstance(action, CalculateAction):
        return tools.calculate(action.expression)

    if isinstance(action, FinalAction):
        return action.answer

    raise TypeError(f"Unsupported validated action: {type(action)}")

The model never calls a tool. It never has a reference to one.


13. Execution is another boundary

Passing validation does not guarantee that execution succeeds. A search service can be down, a referenced file can vanish, an API can reject a request that was legal when it was proposed. The environment is free to change between the decision and the act.

So execution returns an observation rather than silently becoming truth:

from typing import Any

class ToolExecutionError(RuntimeError):
    """Expected operational failure reported by a capability adapter."""

@dataclass(frozen=True)
class Observation:
    ok: bool
    kind: str
    data: Any
    failure_stage: Stage | None = None

def execute_as_observation(action: Action, tools) -> Observation:
    try:
        value = execute(action, tools)
    except ToolExecutionError as exc:
        return Observation(
            ok=False,
            kind="execution_error",
            data=repr(exc),
            failure_stage=Stage.EXECUTION,
        )

    return Observation(
        ok=True,
        kind="action_result",
        data=value,
    )

This is the sixth Stage, and it is the one the runtime cannot prevent — only observe and report.

One execution failure deserves special treatment: an ambiguous result is not evidence that no side effect occurred. If a network connection disappears after charge_card() is sent, blindly retrying may charge the card twice. Retry policy therefore depends on the capability contract. A genuinely idempotent operation may be repeated under its declared conditions; a non-idempotent mutation should first reconcile against the environment through an idempotency key, transaction identifier, receipt or state check. Later chapters make idempotency part of the capability contract. For now the rule is enough: a failed response does not prove a failed action.


14. Assemble the boundary

Everything above composes into one function, and it is worth reading closely because it is the chapter compressed:

import json

def prepare_action(raw: str, policy: RunPolicy, context) -> Decision:
    try:
        data = json.loads(strip_json_fence(raw))
    except json.JSONDecodeError as exc:
        return Rejected(Stage.REPRESENTATION, str(exc))

    try:
        action = validate_proposal(data)
    except SchemaError as exc:
        return Rejected(Stage.SCHEMA, str(exc))
    except SemanticError as exc:
        return Rejected(Stage.SEMANTICS, str(exc))

    try:
        authorize(action, policy)
    except PermissionError as exc:
        return Rejected(Stage.AUTHORIZATION, str(exc))

    try:
        check_preconditions(action, context)
    except PreconditionError as exc:
        return Rejected(Stage.PRECONDITION, str(exc))

    return Accepted(action)

The five pre-execution rungs from §3 are now something the program can execute rather than something the reader has to remember. Every rejection is labelled with the rung it fell from.

The property that matters here is architectural rather than syntactic: raw output cannot reach execute() without crossing this function. A framework may hide the same responsibility behind decorators, generated schemas or provider-native function calling. The responsibility does not go away; it only becomes someone else’s code.

The full implementation, including validate_calculate, validate_final and the precondition rung, is in the companion repository.


15. Test the boundary, then measure it

The action boundary is ordinary software and deserves ordinary tests. Because rejections now carry a stage, we can assert on where a proposal failed rather than merely that it did:

ProposalExpected stage
malformed representationREPRESENTATION
unknown actionSCHEMA
missing argumentSCHEMA
unexpected argumentSCHEMA
wrong primitive typeSCHEMA
empty or out-of-range valueSEMANTICS
valid but disallowed actionAUTHORIZATION
valid, authorized action with unmet runtime requirementPRECONDITION
valid, authorized action with satisfied preconditionsreaches the executor
expected operational failure during executionEXECUTION, returned as an observation
def test_unknown_action_is_rejected_at_the_schema_rung():
    raw = json.dumps({"action": "teleport", "arguments": {}})
    policy = RunPolicy(allowed_actions=frozenset({"search", "calculate", "final"}))

    decision = prepare_action(raw, policy, context=None)

    assert isinstance(decision, Rejected)
    assert decision.stage is Stage.SCHEMA

def test_forbidden_action_is_rejected_at_the_authorization_rung():
    raw = json.dumps({
        "action": "search",
        "arguments": {"query": "PyTorch", "max_results": 5},
    })
    policy = RunPolicy(allowed_actions=frozenset({"final"}))

    decision = prepare_action(raw, policy, context=None)

    assert isinstance(decision, Rejected)
    assert decision.stage is Stage.AUTHORIZATION

Both tests deliberately pass context=None because they are expected to terminate before the precondition rung ever reads runtime context. A precondition test must instead supply the concrete state that makes the requirement hold or fail.

Those two tests fail for different reasons, and the assertions say so. If every rejection collapses into bad tool call, we lose the diagnostic structure this mechanism produces.

The same structure makes measurement straightforward: count boundary failures by stage rather than reporting only a single success rate. Pre-execution failures come from Rejected.stage; expected execution failures come from Observation.failure_stage. A number like agent success rate = 72% tells us nothing about this chapter’s mechanism, and it can move in misleading directions. Suppose a schema change causes silent wrong executions to fall while visible validation failures rise. The aggregate looks worse. The system has become safer and considerably easier to debug.

This is also why function-calling benchmarks such as BFCL evaluate function selection, argument matching, relevance and multi-turn behaviour separately: “can call a tool” is not one indivisible capability.[4] Full observability comes later in the book. For now:

Record the rung at which a proposed action failed.


16. When the boundary earns its cost

Not every model call needs this boundary. The test is whether model output is about to acquire operational authority:

    flowchart TD
    Q{Does model output<br/>trigger external behaviour?} -->|no| T[plain output]
    Q -->|yes| A[define action space]
    A --> B[structured proposal]
    B --> C[validate]
    C --> D[authorize]
    D --> E[execute]
  

Rewrite this paragraph stays on the left branch. Choose one of these executable capabilities and supply its arguments belongs on the right.

There is a useful diagnostic hiding in the right-hand branch. If a capability cannot be described precisely enough to validate, that is a design signal rather than a validation problem. It often means the action interface is too broad. A capability such as run_anything(command) can still be constrained by sandboxing, allowlists, path policies and resource limits, but its semantic contract is much coarser than a capability such as read_file(path) or run_test(test_name). The harder it is to state what a valid invocation means, the more work the runtime must do to constrain and reason about that capability.


17. What we earned

The previous chapter gave us adaptive control. This chapter made the act step of that loop explicit, and the important addition was never JSON. It was the separation between probabilistic proposal and deterministic runtime authority, which lets us state the contract precisely:

The model may propose any output it can generate. Only the runtime can turn an accepted proposal into an environmental action.

That is the first mechanism added to the agent we began with, and it immediately exposes the next problem.

A proposal can be well formed, semantically valid, authorized and executable — and still be a poor choice. Our boundary is good at rejecting proposals that are malformed, invalid, unauthorized or currently unexecutable. It has nothing whatsoever to say about a proposal that passes every one of those checks and is still a poor choice. Validity and quality are different properties. So how do we improve the decision without changing the model underneath it? One answer spends additional inference-time compute to create alternatives: stop accepting the first proposal, generate several, then compare them.


Research roots

This chapter derives the action boundary from ordinary software-engineering principles, but several research threads make the motivation concrete.

  1. Patil et al. — Gorilla: Large Language Model Connected with Massive APIs (2023). API and tool use exposes failures such as incorrect API selection and arguments, motivating explicit interfaces and evaluation of tool-use behaviour.
    https://arxiv.org/abs/2305.15334

  2. Geng et al. — Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning (EMNLP 2023). Shows how formal grammars can constrain generation to valid structures rather than relying only on prompt compliance.
    https://arxiv.org/abs/2305.13971

  3. Geng et al. — Generating Structured Outputs from Language Models: Benchmark and Studies (2025). Introduces JSONSchemaBench and studies constrained decoding across real-world JSON schemas, including efficiency, coverage and quality.
    https://arxiv.org/abs/2501.10868

  4. Patil et al. — The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models (ICML 2025). Evaluates function selection, argument construction, relevance, multi-turn tool use and related failure modes across models.
    https://gorilla.cs.berkeley.edu/leaderboard

Next: Candidate Generation and Selection

We have a runtime that can reject invalid, unauthorized or currently unexecutable proposals and admit proposals that satisfy its boundary. It still accepts the first proposal that clears those checks, which means proposal quality remains the quality of one sample from one model call.

The next chapter keeps the boundary exactly as it is and changes what flows into it. Instead of one proposal we generate several, then select among them — which splits agent quality into two measurable halves: whether a good candidate was produced at all, and whether we managed to pick it.