You want to pull public data from twenty-some platforms, and the first thing you want to do is almost certainly abstract: draw one unified Post, one unified User, and map every platform's response into them. A bilibili video, a tiktok video, a zhihu answer, a linkedin post, they all look like "a piece of content plus an author," and merging them into one model feels obvious. That abstraction feels great on the first three platforms and buries you alive by the twentieth.
I didn't build that unified model in the end. Across twenty-some platforms and two hundred-some commands, the seemingly dumber split of "every platform minds its own" is what survived.
How the unified model collapses
It collapses slowly. By the eighth platform, your Post has a dozen optional fields hanging off it: some platforms have a danmaku count, some don't; "published time" is a second-level timestamp on some and a string like "3 days ago" on others. By twenty-some it's fully collapsed, and the way it collapses isn't a compile error, it's that it stops saving you anything. Every piece of downstream code has to first decide "did this platform even fill this field," the decision logic runs longer than reading the raw response would, and the unified layer has become an obstacle you route around to get work done. The write side pays the same tax: every new platform you integrate sends you back to it, stuffing fields into boxes shaped to fit the old platforms.
Each platform is its own bounded context
The split that holds is the reverse: don't abstract a unified model, let every platform mind its own. In the catalog, each platform is a <platform>_reverse/ context that owns four things and lends out none:
Input validation. What this platform's ids look like, which parameter combinations are legal, only it knows.
Protocol. HTTP direct or a fragment of local JS signing, which domain, which headers, all platform-private.
Signing. The signing mechanisms differ wildly across platforms, and cramming them into one shared signer only builds a monster full of if-else.
Response normalization. Tidy the raw response into a structure it owns, not a global unified structure.
The fourth is the one most easily misread. "No unified model" is not "no normalization." Every platform of course normalizes, the target shape is just its own to define, not something it forces onto a shared model. Unification holds only where it really is the same thing: two endpoints inside one platform sharing a post structure is legitimate, because that genuinely is the same domain object. The mistake is dragging that within-platform unification out to span platforms.
The shared layer holds only what is genuinely cross-platform
What goes in the shared layer? Capabilities that are genuinely cross-platform and behave identically for every platform, not the ones that merely look alike. My shared layer has only three things:
The interface read-model: derive a unified capability catalog from each platform's argparse declarations. What it unifies is how commands are discovered and described, not what commands return. The former is genuinely cross-platform, the latter is platform-private. (This "declaration is the interface" idea is unpacked in the interface-as-code piece.)
Local loopback transport: authenticated requests pass through a local WebSocket session service that treats all platforms alike and touches none of any platform's business fields.
The dispatch entry: discover the platform, hand the command to its context, and nothing more.
The test is simple: to enter the shared layer, it has to genuinely behave the same for every platform. Transport, dispatch, and the way interface descriptions are generated are the same for every platform, whereas "a piece of content" behaves nothing alike on bilibili and linkedin, so it doesn't go in. "Looks alike" is abstraction's biggest trap: a video and a video look alike, so you want to unify them, but surface similarity isn't behavioral sameness, and treating it as a shareable domain model is the root of the unified model's collapse.
The command-count distribution tells you where abstraction pays
Still on the fence about a unified model? One look at the real command-count distribution is enough. 22 platforms, 241 commands, distributed extremely unevenly:
Platform | Commands |
|---|---|
tiktok | 34 |
bilibili | 26 |
18 | |
zhihu | 18 |
douyin | 17 |
xiaohongshu | 16 |
The other 16 platforms | 1 to 13 each |
The top six platforms total 129 commands, more than half of everything, and the remaining half is split across 16 long-tail platforms, many with only two or three commands, some with just one.
That distribution fixes the economics of abstraction: the cost of a unified model is fixed (every integrator has to fill fields, check for null, route around it), but the benefit is spread per platform. For a long-tail platform with two or three commands, the benefit of abstraction is negative, because the adapter code you write to fit it into the unified model runs longer than all of its business code.
Don't reserve abstraction for a single implementation
Following that distribution, one more rule: don't reserve abstraction for a single implementation. When a platform has only one implementation, don't add a repository, a factory, or an interface layer for "there might be another implementation later." Adding a platform is adding one <platform>_reverse/ context, with no need to touch a shared base class first.
An interface layer's value is making multiple implementations swappable, and with one implementation its value is zero while its maintenance cost is positive. Reserving for a second implementation that doesn't exist and reserving for a cross-platform commonality that doesn't exist are the same mistake. The cross-language migration proved it again: the old registry's few hundred commands explicitly left a batch unmigrated, and left no stubs or compatibility proxies either, because an empty shell costs more than a gap, it makes the next person think something's there. A reserved abstraction is the same.
For cross-platform normalization, feed the model per-platform context too
This "split by platform, no unified model" thinking holds just as well when you use a model to normalize. To tidy twenty-some platforms' raw responses into an analyzable structure, it's natural to hand it to a model, and here the easiest mistake is exactly the one from the code layer: define a unified schema, and feed each platform's raw JSON with "map this to the schema" tacked on. It fails, because the model doesn't know whether bilibili's play-count field and tiktok's play-count field are the same thing, and forcing it toward a lowest-common-denominator schema means it either drops a field the platform depends on or fills it half right.
The right way is to give context per platform: tell the model "this is bilibili, here's what these fields mean, I want this shape for this platform," normalize one platform at a time, and leave cross-platform merging to the analysis layer. This runs in a few steps, each asking something different of a model:
Step | Capability it needs | Pick | model id |
|---|---|---|---|
Read the shape of one platform's whole raw response | Long context, swallows the full response plus field notes at once | Kimi K3 |
|
Set the normalization boundary (which fields are genuinely cross-platform, which are platform-specific) | Strong reasoning, resists over-unifying | Claude Opus 5 |
|
Extract fields per platform in bulk, map item by item | Cheap, hundreds to thousands of calls at high concurrency | Claude Sonnet 5 |
|
Explain why two platforms' same-named fields don't line up | Mid reasoning, explains differences against the fields | GPT-5.6 Sol |
|
The second step is the only place where switching models visibly changes the result. What it tests is whether you'll admit two fields aren't actually the same thing, the same as the counter-evidence section in algorithm-family identification: a weak model follows your hint to "unify," a strong one points out the boundary.
The switching cost is the real obstacle
The four tiers come from three vendors, three SDKs, three auth schemes, three error formats. Rewriting your client three times to switch models between steps isn't worth it, so most people run one tier the whole way, and often use a tier that only flattens at the "set the boundary" step, building a schema that collapses again at twenty-some platforms.
AIReiter flattens that layer: one key, one OpenAI-compatible interface, all four tiers behind it, switching by changing the model field in the request body.
# Set the normalization boundary: the reasoning tier
curl https://aireiter.com/api/v1/chat/completions \
-H "Authorization: Bearer $AIREITER_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-opus-5",
"messages": [{"role": "user", "content": "<one platform response sample + have it mark which fields are platform-specific>"}]
}'
# Extract fields per platform in bulk: change the model field, leave the rest
# "model": "claude-sonnet-5"
# Field-difference attribution:
# "model": "gpt-5.6-sol"
If you already use the OpenAI SDK, point base_url at https://aireiter.com/api/v1 and change nothing else. On the Anthropic SDK, hit POST /api/v1/messages with the same key.
On price, Claude models run at 30% off list, GPT models at half, and Kimi K3 is callable on the same key. The discount lands right on the main cost: bulk per-platform field extraction is the most call-dense step, twenty-some platforms with hundreds to thousands of records each, one call per record, running on the cheapest Sonnet with 30% off on top. Reading a whole long response on Kimi K3 is a few hundred thousand tokens per input, another chunk of cost. The reasoning tier for setting boundaries is few calls, so it barely costs anything.
Try it without signing up: run one platform's response by hand first and see whether the model honestly flags the differences or rushes to flatten them, then decide whether to wire it in.
In closing
The first reaction in cross-platform collection, abstracting one unified Post/User, feels great at small scale and inevitably collapses at twenty-some platforms: fixed cost, per-platform benefit, and an extremely long-tail command distribution. The split that holds is one bounded context per platform, each owning its input validation, protocol, signing, and response normalization. The shared layer holds only what genuinely behaves the same for every platform (transport, dispatch, interface-description generation), not a domain model that merely looks alike, and it reserves no abstraction for a single implementation or a commonality that doesn't exist. On a model it's the same sentence: normalize with per-platform context, don't feed it a unified schema, and let cross-platform merging happen only at the analysis layer. The full four-stage workflow covers the four-tier split in more detail, and connected through one unified interface, the switching cost stops being a reason not to use them.