AIREITER

329 Commands, 128 Still Missing: Reconcile a Migration With Set Math, Not a Model

Last Updated: 2026-07-31 06:41:26

You hand a model a function in an old language and ask it to translate it into a new one. It gives you clean, idiomatic new code, naming conventions and all, and the test passes. You do this two hundred times, wrap up, and declare the migration finished.

This is where "translated" and "migrated" get mixed up. Whether a single function is translated correctly is the model's strength, and it does it well. But whether the whole thing migrated is a different kind of question. It's a set operation, and set operations are exactly what a model should not touch and is most likely to fool you on.

Here's the ledger from a real Go-to-Python migration, and why this ledger has to be settled by a script, with the model only explaining.

Start with three numbers

The old Go registry had 23 platforms and 329 commands. The command set the new Python side derives from its argparse declarations has an exact intersection with the old registry of 201. Which means 128 commands exist only on the old side, neither migrated into Python nor left as a stub.

329 = 201 + 128. That subtraction has no technical content, yet it's the only thing in the whole migration that answers "is it finished," and it's exactly what you never see while translating one function at a time. A missing item is an absence error: no error thrown, no exception, no failing test, just a name that should exist and doesn't. It never once entered the chat box, and you can stare at two hundred green checkmarks without seeing it.

Why no stubs, and no compatibility proxies

Halfway through a migration, the tempting move is to leave a placeholder for the commands you haven't done: a raise NotImplementedError stub, or a compatibility proxy that forwards to the old binary, so "the endpoint catalog looks complete." Don't. An empty shell costs more than a gap, for three reasons.

A stub breaks the reconciliation. The command name enters the new side's set, the diff comes out 0, and you think you're done. A gap is an honest red, a stub is a green lie that paints "128 to go" as "all present."

A compatibility proxy freezes an un-purified dependency permanently. The proxy forwards to the old Go binary, so the old runtime can never be deleted. The point of the migration is to shed the old stack, and a forwarding proxy lets the old stack move in under the name of "temporary compatibility" and stay for life.

A half-built endpoint deceives the caller. An Agent or a person reads the catalog, assumes it works, calls down, and hits runtime_unavailable, or worse, a false success that silently returns an empty result.

An honest gap is actually the cheapest: the diff flags red immediately, and everyone can see how much is left. This is the same principle as the evidence threshold in app reverse engineering: marking something "not usable yet" is always cheaper than shipping a half-built one.

The reconciliation script skeleton

The core of reconciliation is one line: both sides' command sets are derived from declarations, and nobody transcribes by hand. Transcribe a "migrated list" by hand and you've introduced a third source of truth that drifts from the code, and in two weeks it's the first thing to go wrong.

The new (Python) side's single source of truth is the argparse declarations in each platform's cli.py. A catalog module walks the subcommands and exports a {platform/command} set, emitted by python -m reverse describe --format json. Why a declaration can be the single source of truth and how the catalog derives fully automatically is the subject of the interface-as-code piece. The old (Go) side is already a platform -> command map (an immutable allowlist compiled into the binary), and exporting the same-shaped JSON is trivial.

With the two JSON files in hand, the rest is set operations:

# Both sides' command sets derive from declarations, not transcription.
# Transcribe by hand and you've added a third source of truth that will drift.
import json
from collections import Counter

def ids(path):
    doc = json.load(open(path))
    return {f"{p['name']}/{c['name']}"
            for p in doc["platforms"] for c in p["commands"]}

old = ids("go-registry.dump.json")      # old registry: immutable platform->command allowlist
new = ids("python-catalog.dump.json")   # python -m reverse describe --format json

missing = old - new     # old side only: each one needs a keep-or-drop verdict
added   = new - old     # new side only: new capability, logged separately
kept    = old & new     # intersection: migrated, but still check for semantic drift

assert missing | kept == old            # every old-side item classified, none dropped

by_platform = Counter(pc.split("/")[0] for pc in missing)  # goes straight into the README table

That runs in a few milliseconds, at zero cost, deterministic, 100% correct. missing is those 128 commands, and aggregated by platform it's this table:

Platform

Unmigrated commands

xiaohongshu

33

tiktok

30

hotspot

21

douyin

19

reddit

8

weibo

7

bilibili

5

zhihu

3

linkedin

1

netease_music

1

Total

128

There's no place for a model in this step.

How expensive, and how wrong, "line-by-line compare" by a model is

Skip the script, paste the two lists into the chat box, and ask "which of the 329 don't appear in these 201," and three things happen without fail.

It drops items: once the list is long, the model doesn't do an element-by-element set difference, it samples on "looks about right," the tail items get diluted, and you get an answer that looks complete and is short by a dozen. It invents: reports items present on both sides as missing, or counts truly missing ones as migrated, because it's imitating what a reconciliation report looks like, not doing the difference. It isn't reproducible: ask the same input twice and the missing list differs, and a "reconciliation" with a different result every time is not a reconciliation.

It doesn't pay off on cost either: the script is a few milliseconds, having the model compare is a few hundred thousand tokens plus multiple self-check rounds, expensive and slow and untrustworthy. Give the set operation to the thing that does set operations is the least controversial line in this piece.

What the model should do: explain the difference, not decide whether it migrated

The script hands you 128 "didn't migrate" facts, but a fact isn't a conclusion. Each one needs a keep-or-drop decision, and a decision needs a reason. This is the model's arena.

Explain "why it didn't migrate," one by one. Dead code? An upstream endpoint retired? Deferred? Or, the trickiest, it wasn't deleted but merged into another command, the name gone but the capability still there. That "merged, not deleted" hidden correspondence can't be spotted from the missing list alone, you have to read both registries at once to line it up.

The 201 in the intersection aren't safe either. Migrated doesn't mean the semantics held: a same-named command with a quietly changed default, a swapped pagination semantic, two error codes merged into one. This is semantic drift, sneakier than a gap, because the diff is green and it never enters missing at all. Checking for drift needs the model to read both implementations and judge "is this behavior equivalent," confirmed in the end by differential testing (the fixture comparison in stage three of the four-stage workflow). That ability to look at an already-successful translation and still say "behavior changed here" is exactly what the counter-evidence section in the fingerprinting piece is about, and a weak model only recites "successfully migrated."

So the division is clear: deciding "is it there" belongs to the script, deciding "should it stay, did it change" is what needs reasoning. This time 4 commands the diff flagged red turned out, on review, to belong, and were restored as first-class new commands. Script decides, model explains, human decides, three layers each in place.

Which model for which step

Note that all four tiers below land in the explanation layer. The decision layer (the diff) uses no model at all, and that's the line between this piece and other "AI migration" articles.

Step

Capability it needs

Pick

model id

Feed both registries at once, spot "not deleted but merged elsewhere" correspondences

Long context, reads both sides' full declarations at once

Kimi K3

kimi-k3

First-pass keep-or-drop on 128 missing items, structured draft

Cheap, hundreds of calls at high concurrency

Claude Sonnet 5

claude-sonnet-5

Semantic-drift judgment (migrated, did behavior change), reads both implementations

Strong reasoning, willing to say "this changed"

Claude Opus 5

claude-opus-5

Migrated but the fixture won't match, explain against parameter or response shape

Mid reasoning attribution

GPT-5.6 Sol

gpt-5.6-sol

The one most worth testing is the third tier. Semantic-drift judgment tests exactly "will you argue against an already-successful translation," and that's where switching models changes the result most. The protocol:

  1. Take a real two-language migration of your own, run the script for the missing set (zero model in this step).

  2. Hand-label 10 to 15 of them with a ground truth (drop, keep, merged elsewhere, deferred) as a control.

  3. Feed the same "explain keep-or-drop per item" prompt to claude-opus-5 and a cheap tier, and look at two things: does the keep-or-drop reason point to a specific code fact or hand you "possibly deprecated" mush, and how many merged-elsewhere correspondences each spots.

  4. The number of hidden correspondences it spots is your basis for whether you dare hand it the first pass.

The friction isn't picking a model, it's the switching cost

Four models from three vendors, three SDKs, three auth schemes, three error formats. Rewriting your client three times to switch tiers isn't worth it, so most people run one model the whole way, and at the semantic-drift review that most needs the reasoning tier, they use a cheap tier that only speaks mush and wave all the green drift right through.

AIReiter flattens that layer: one key, one OpenAI-compatible interface, all four tiers behind it, switching by changing the model field in the request body.

# Semantic-drift review / per-item keep-or-drop: the reasoning tier
curl https://aireiter.com/api/v1/chat/completions \
  -H "Authorization: Bearer $AIREITER_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-opus-5",
    "messages": [{"role": "user", "content": "<both implementations + this command migration status, ask if behavior is equivalent>"}]
  }'

# First pass on 128 missing items in bulk: change the model field, leave the rest
#   "model": "claude-sonnet-5"
# Diff attribution when a fixture won't match:
#   "model": "gpt-5.6-sol"

If you already use the OpenAI SDK, point base_url at https://aireiter.com/api/v1 and change nothing else. On the Anthropic SDK, hit POST /api/v1/messages with the same key.

Price lands on this flow: the first pass runs hundreds of items at once and reruns every migration round, where high-concurrency claude-sonnet-5 is cheapest; semantic-drift review is a dozen hard cases asked of claude-opus-5 over and over, the most expensive per item. Both are Claude tiers, and the 30% off lands right on the densest and the most expensive parts. gpt-5.6-sol does the diff attribution, GPT at half price.

In closing

"Translated" is a single-function illusion. "Migrated" is settled by the diff. Set operations go to the script, explanation goes to the model, decisions go to the human, and the order can't be scrambled, least of all by letting the model do the deciding.

There's one more step that's easiest to skip: the missing list has to go into the README, visible for the long haul. The 128 stays up until it becomes 0, or until every one carries a written "not migrating, because X." A reconciliation that lives only in some PR discussion is no reconciliation, because the next person to take over can't see it and will step on all 128 again. That's the shared ground between this piece, the interface-as-code piece, and the piece on not building a unified response Model: let the single source of truth speak for itself, and don't scatter conclusions across people's memories.