You add a tool to your Agent, say "fetch a platform's public posts." It runs in production for two weeks with no trouble. One day you change the function's limit default from 25 to 20 and add a new value to the sort enum. Code changed, tests green, merged.
Three days later production starts throwing the occasional error. The model called that tool with an enum value you deleted last week, runtime validation rejected it, and the stack trace points at your dispatch layer. You stare at the dispatch code for half an hour. Not a line of it is wrong. The real fault is somewhere else entirely: you changed the function signature and left the tool description the model reads untouched. The model is still holding the old schema, so it generates calls against the old shape, and of course they no longer fit.
This is tool-description drift. It is the most common and the hardest-to-trace class of bug in Agent engineering, and the reason it's hard to trace is specific. The error and the root cause live in different places. The error surfaces in the execution layer. The cause sits in a JSON file nobody thinks to open. This piece is about deleting the source of that bug structurally. Not "remember to keep them in sync." Arrange things so there is no second copy left to drift.
How drift happens
Pull it apart and the root of drift is that you are maintaining two sources of truth.
One is the code that actually runs: the function signature, the argument validation, the defaults, the enum constraints. This one is hard. If it's wrong, it fails loudly.
The other is the tool description the model reads: the name, the description, and the parameters JSON schema. This one is soft. Get it wrong and nothing blows up right away. The model just generates a bad call, which then fails down in the execution layer.
As long as a human keeps those two in agreement, drift is only a question of when. You change an argument in the code and forget the description. Or you change the description and forget the code. Or you change both and the meanings still don't line up. None of these show themselves at the moment you make the change. They wait until the model happens to generate a call that touches the difference, by which point you've long forgotten what you edited two weeks ago. The fix has only one direction: turn the two copies into one.
Single source of truth: the declaration is the interface
The shift in thinking is this: you don't actually need that tool-description JSON.
A function's parser declaration, together with its docstring, already contains every field a tool description needs. Here is what a plain argparse declaration looks like:
subreddit = commands.add_parser("subreddit", help="Query a public board's feed")
subreddit.add_argument("subreddit")
subreddit.add_argument(
"--sort",
choices=("hot", "new", "top", "rising", "controversial"),
default="hot",
)
subreddit.add_argument("--limit", type=int, default=25)
The help is the one-line description of the command. The choices are the enum constraint. The default is the default value. The type is the parameter type, and the positional argument is the required field. Everything the model needs to call the tool is right there: what the command does, which parameters it takes, which are required, what the enum values are, what the defaults are. And this declaration is the same one the runtime uses to parse and validate, so it cannot drift away from the execution logic, because it is the execution logic.
So stop writing a second tool description. The right posture is that the document does not exist. There is only code, and when you need a tool description, you project one out of the code. The projection runs one way, from code to description, and never the reverse.
Deriving the whole catalog from declarations
Once you accept that the declaration is the interface, the tool description shouldn't be hand-written. A deriver should produce all of them.
What the deriver does is mechanical. It walks every platform context, imports its parser, and reads the argparse action list into three immutable structures: Platform, Command, and Parameter. Each Parameter carries its name, type, required flag, enum values, default, and help text. That is an interface read-model derived entirely from code.
With that read-model in hand, any output format is downstream of it. A describe --format json emits the full machine-readable interface to feed an Agent's tool selection. A render_skill() emits a capability catalog a human or a model can read. The command count in that catalog isn't a hand-typed constant. It's sum(len(platform.commands)) computed on the spot. Right now that comes to 22 platform contexts and 241 commands, and not one of them was typed into the catalog by hand.
That buys a comfortable property. Adding a platform means adding a platform context, and the catalog picks up all of its commands on its own. Changing a parameter means editing the parser declaration, and the matching enum and default in the catalog update themselves. You never hit "wrote a new command but forgot to register it" or "changed a parameter but the catalog is stale," because registration as an action does not exist. The catalog is computed, not maintained.
(This instinct to derive rather than maintain is the same one that shows up when you use set operations to decide what has actually been migrated across languages, which is the subject of the cross-language migration piece.)
CI catches drift: turn the ghost into a red X at commit time
Derivation handles "new command lands in the catalog automatically," but one gap remains. Someone edits a parser declaration, forgets to re-run the deriver, and doesn't commit the regenerated catalog. Now the copy in the repo is stale again, and drift slips back in through the side door.
The last gate goes in CI, and its core is a single assertion:
docs-check:
$(PYTHON) -c 'from pathlib import Path; from reverse.catalog import render_skill; \
path = Path("skill/SKILL.md"); \
assert path.read_text(encoding="utf-8") == render_skill(), \
"skill/SKILL.md is out of sync with the code; run make docs"'
It takes the catalog committed to the repo and byte-compares it against a catalog regenerated from the current code. One character of difference and CI fails, with a message that just says the catalog is stale and to run make docs.
The value of that one line is that it moves when drift gets found. Drift used to be a runtime ghost: it blew up in production two weeks later with a stack trace pointing at the wrong place. Now it's a red X at commit time. You get stopped in the pull request, the error tells you the catalog is out of date, and regenerating it fixes the problem. Drift drops from "the hardest bug to trace" to a compile error you clear with one command. That is the full loop of interface-as-code. The declaration is the source, the capability catalog is the build artifact, and CI is the type check. You wouldn't hand-write a build artifact or tolerate one that disagreed with its source, and a tool description deserves the same treatment.
What to freeze as code, what to leave to the model
Derivation and CI guarantee that the interface description is accurate. There is an earlier judgment that comes before that, though: should a given capability be written as fixed code, or left for the model to orchestrate on the spot? Get that division wrong and an accurate interface won't save you.
Look at capabilities in three layers.
A low-level primitive reads one kind of data or performs one clear action. Its input is stable, its output is structured, and it can be tested on its own. This layer is pure code and spends no reasoning at all. The large majority of the 241 commands sit here.
A deterministic workflow is a strongly ordered process inside one platform, with shared state and a clear success condition. Take a creative pipeline, creative-pipeline, that runs in sequence: find opportunities, then Top Ads, then creator matching, then a creative brief, then a generation preflight. The order and the dependencies between steps are fixed. This layer is also frozen into code, because if the order is already determined, having the model re-plan it every time is both slower and less stable. Marking it takes one line: give the command set_defaults(_command_level="workflow"). That is the only such line in the codebase, and the catalog shows workflows and primitives on two separate levels because of it.
Agent orchestration is the layer of cross-platform research, live tradeoffs, and rerouting after a failure. This is the layer you actually leave to the model, because what to query next depends on what the last query turned up, and you can't write that down in advance.
The test is fairly clear. If a capability needs stable stage status, shared context, or generation side effects, freeze it into code. If it involves query expansion, cross-platform verification, or rerouting after a failure, leave it to the model. Both mistakes cost you. Hardcoding a research hypothesis into the client is over-freezing, and the day the platform changes you're back editing code. Handing a fixed sequence to the model to reassemble each time is under-freezing, where you save one model decision and buy a pile of instability.
Letting the model decide when to degrade: six stage statuses
For the orchestration layer to make decisions, the things the lower layer returns have to be legible to the model. An opaque success-or-failure boolean isn't enough. Hand the model a success: false and all it can do is guess at the next step.
So each stage of a workflow returns a stage status rather than a boolean, and there are six of them: completed, empty, ready, skipped, unavailable, and blocked. The information lives in the distinctions among the ones that didn't proceed:
skippedmeans the operator turned this step off on purpose, for example by setting one collection path's cap to 0. This is not an error, and the model should not retry it.unavailablemeans something this step depends on is temporarily out, for example an interface throwing an error or a missing session. The model can skip it and continue, or prompt for a fresh session and come back.blockedmeans a precondition isn't met, for example the research evidence being empty or the preflight failing. The model should not force the next step through. It should go back and fill in the evidence.
Take that creative pipeline. It judges "platform preflight ready" and "research evidence ready" as two separate things, with a final ready = platform_ready and research_ready. If either one fails, the generation stage comes back blocked with a blockers list saying what it's stuck on, and when every commercial search result is empty it simply doesn't submit the generation job.
Why is this design for the model's benefit? An orchestration model that reads seedance_generation: blocked plus blockers: [research_evidence_empty] knows to go back for evidence instead of retrying the submission. Reading organic_discovery: skipped, it knows this is the user's intent and not a fault, so it leaves it alone. Reading a step marked unavailable, it knows it can degrade around it. Once you separate "turned off on purpose," "temporarily out," and "precondition not met," the model can pick the right degradation path. Flatten all three into false and even a strong model just spins in place.
Which model for which layer
The stack above asks very different things of a model at each layer. (The four-stage reverse-engineering piece lays out the same four-tier table in a reverse-engineering setting; here it moves onto the Agent stack.) Assign by layer and you stop wasting capacity:
Work in the Agent stack | Capability it needs | Pick | model id |
|---|---|---|---|
Load the | Long context, reads the whole catalog in one pass | Kimi K3 |
|
Orchestration: read stage status and blockers, decide to degrade, reroute, or continue | Strong reasoning, makes the right call against the status | Claude Opus 5 |
|
Generate model-friendly tool-description text from docstrings in bulk | Cheap, runs hundreds of calls at high concurrency | Claude Sonnet 5 |
|
Tool-call error attribution: read the error and the declaration, decide if it's drift or an upstream change | Mid reasoning, explains against specific fields | GPT-5.6 Sol |
|
The orchestration layer is the one worth dwelling on. Reading blocked and skipped to decide the next move is the only step in this flow where switching models visibly changes the result, because what it tests is exactly whether the model can make the right call against a piece of status. A weaker model treats a skipped as a failure and retries it, or sees a blocked and submits anyway. A strong reasoning model reads the blockers and reroutes precisely. It's the same kind of gap as whether the counter-evidence section is really arguing against itself in the fingerprinting piece: anyone can produce the candidate, the hard part is the judgment call.
You don't have to take my word for the gap. Test it:
Take a real return from one of your own workflows, with its
stagesandblockers, or fabricate a response that isblockedwithblockers: [research_evidence_empty].Feed that response, plus your capability catalog (the
describejson), plus an instruction to decide the next action, toclaude-opus-5andgpt-5.6-solseparately.Look at one thing: does the next action the model proposes correctly separate
blocked(go back for evidence),skipped(user intent, leave it), andunavailable(get a session or degrade around), or does it retry theskippedas if it failed?The share of correct degradation paths is your selection criterion. It decides whether your Agent spins in place under a real failure or routes around it on its own.
The switching cost is the real obstacle
The four models come from three vendors, and for function calling the switching cost is especially punishing. OpenAI's tools / tool_calls and Anthropic's tool_use / tool_result are two different formats. Swap in a model with a better judgment call at the orchestration layer and you have to rewrite your whole tool-dispatch and error-parsing path. That is the real reason most people end up locking one model into the orchestration layer, even when that model often misreads stage status.
AIReiter takes that layer away. One key, one OpenAI-compatible interface, all four models behind it, and switching is a matter of changing the model field in the request body.
# Orchestration decision: hand the reasoning tier the catalog plus one blocked workflow response, ask for the next action
curl https://aireiter.com/api/v1/chat/completions \
-H "Authorization: Bearer $AIREITER_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-opus-5",
"messages": [{"role": "user", "content": "<describe json> + <stages/blockers response> + decide the next action"}]
}'
# Generate tool-description text in bulk: change the model field, leave the rest
# "model": "claude-sonnet-5"
# Error attribution:
# "model": "gpt-5.6-sol"
Native function calling just adds a tools array, and the OpenAI tool protocol passes through this interface unchanged, so switching models is still a one-field edit. If you already use the OpenAI SDK, point base_url at https://aireiter.com/api/v1 and change nothing else. On the Anthropic SDK, hit POST /api/v1/messages with the same key.
On price, Claude models run at 30% off list, GPT models at half, and Kimi K3 is available on the same key. For this stack the discount lands where it counts. Every step the orchestration layer takes forward is another reasoning-tier call, which makes it the most frequent and most expensive tier in the whole Agent, and the Claude discount sits right on it. Generating tool descriptions from 241 docstrings in bulk is high-concurrency Sonnet work, discounted too. Those two are the bulk of the cost. The GPT-5.6 calls for error attribution are far fewer.
Try it without signing up: run a few rounds by hand, feed the same
blockedresponse to both models, and see for yourself which one degrades correctly before you wire one into the orchestration layer.
In closing
Tool-description drift is not something "remember to sync" can cure. That approach just compresses a structural defect into a matter of personal discipline. The real fix is to delete the two-sources structure: the parser declaration plus the docstring is the only source, the capability catalog is a build artifact derived from it, and one CI assertion acts as the type check. Drift goes from a runtime ghost to a red X at commit time.
But derivation only guarantees the description is accurate. It says nothing about whether the layering is right. Which capabilities you freeze into code and which you leave to the model to orchestrate, plus the six stage statuses that let the model read "retry or degrade," are the two things that decide whether your Agent can run on its own. The model does two concrete jobs in this stack: it makes the tradeoff at the orchestration layer, and it attributes the cause when a tool call fails. Whether to freeze a capability and which degradation path to take are settled by the stage statuses you design and the CI you write, not by the model.
That is the same stance as the pieces on set-based migration reconciliation and on not building a unified response Model: AI compresses the time of one step, and the verdict stays inside the constraints you hardcode. Once the whole thing runs smoothly, the only friction left is model switching, which is an infrastructure problem that one unified interface solves.