I located the native links on all three platforms. Where device registration enters, which layer the security SDK intercepts the request at, how the native interceptor rewrites the outgoing packet, which methods JNI dynamically registers onto the runtime with RegisterNatives. Frida script attached, logs scrolling, every hook landing reliably. Then I opened the command catalog and counted the callable native app endpoints for those three platforms: zero.
In the same project, another platform had 11. Field-validated, already migrated into the code and running as first-class commands.
The difference isn't hook technique. I got all the hooks landing, the link diagram is drawn out cleanly. The difference is a threshold a lot of people don't realize is there. Between a hook landing and a capability being usable sits an evidence tier. This piece is about how to set that threshold, why marking something "not usable yet" is cheaper than shipping a half-built endpoint, and what the model actually does for you along the way.
A hook landing is not a usable capability
People who reverse apps tend to treat "the hook landed" as the finish line. Your script attached to the target method, the log shows the arguments, the return value, the call stack. That "I'm inside" feeling is very real, and very misleading.
But the command catalog doesn't accept "I'm inside." What it accepts is: given a normal input, can this command reliably emit a non-empty, correctly structured payload that downstream can consume. Those two things are a long way apart.
The typical collapse goes like this. The hooks all land, the link is complete, you can even watch the device-registration request go out in the log, watch the security SDK compute its value, watch the native interceptor add the signature header. Everything looks right. Then you run it on a real device and device registration returns a zero-value device ID, or the detail endpoint comes back with an empty body. The link is open, the data is empty.
What do you have at that point? A set of probes that can observe one version of this app's internal behavior. What you don't have is a callable endpoint. Ship the first as if it were the second and you've buried a landmine for everyone who comes after you.
The control group: what separates 11 from 0
Put the two sets of platforms side by side and the threshold shows itself.
Control platform (a video-community app) | Three head content apps | |
|---|---|---|
Hook link | Located and field-validated | All located, all hooks landing |
Real install state | Got a non-zero device identity | Zero device ID / missing real device profile |
Non-empty response | Structured detail | Empty detail / empty body |
In the command catalog | 11 | 0 |
Those 11 on the control platform aren't "nicer hooks." Each one crossed the four evidence thresholds in the next section before it migrated to a first-class Python command, and it brought its structured validation evidence along with it.
The three platforms weren't idle. The security SDK's interception layer, the native request interceptor, the batch of methods JNI registered dynamically, all figured out, and the Frida hooks are still in place. But as long as real install state doesn't pass, as long as device registration can't get a non-zero identity, everything after it is empty. So their native app commands honestly stay at 0, and what's currently usable is the entirely separate path of Web or browser transport.
On the larger ledger: this batch of mobile capabilities totals 32, of which 23 actually landed as implementations, and the remaining 9 are stuck at "waiting for real device profile," none of them allowed into the catalog early. The number 9 isn't failure, it's discipline. It marks the exact size of "understood the link, not enough evidence."
Four evidence thresholds
Break the comparison apart and an app capability has to satisfy four conditions to enter the command catalog. One short and it doesn't go in.
One: real install state. The request has to come from a non-zero install identity that upstream will acknowledge. A hook landing in an emulated environment, or on a broken profile, doesn't count. Device registration returning a zero-value ID is this gate failing, and after that no amount of link completeness matters, the output is empty. This is where all three platforms got stuck together.
Two: non-empty response. Open isn't the same as has data. The request goes out, status code 200, body empty. This "successful empty response" is more dangerous than an error, because it slips past every check that only looks for whether an exception was thrown. The threshold wants structured, complete, non-empty data you can feed downstream directly.
Three: complete error classification. A mature capability, when it fails, has to say why it failed rather than throw a blanket error. Missing page, not logged in, runtime not ready, empty response: these fail in completely different ways and have to be distinguishable error codes, not mashed into one. A command that can only say "it failed" isn't ready for the catalog, because the caller can't decide from it whether to retry, to re-authenticate, or to skip.
Four: a closed, repeatable test. One lucky success is not a capability. The same input and the same flow have to run through repeatably, and you have to freeze that success into stored, structured validation evidence. Runs once and comes back empty the next time means you don't control the link yet, you just happened to line up with the runtime state that once.
All four closed, then it registers into the command catalog. Any one open and the capability is still a research asset, not an endpoint. The difference between those two words is the core of this whole piece.
When the Frida log explodes, split it across model tiers
For those four thresholds, the raw material for judging which one a link is stuck at is the trace Frida produces. And a Frida trace explodes. Attach a few dozen methods, run one full link, and tens of thousands to hundreds of thousands of lines is normal, most of it noise from polyfills, heartbeats, and unrelated business modules.
Dump that whole thing into one model and ask "why didn't this hook get data" and you get a blanket guess, and the bigger the context the more likely it stitches two unrelated call segments together. The right approach is to chunk it, then hand different segments to different model tiers by capability.
This step is the textbook case for combining tiers, not "pick the strongest one and use it for everything." The four jobs ask completely different things of a model:
Frida log-processing step | Capability it needs | Pick | model id |
|---|---|---|---|
Read a full call chain's trace in one pass | Long context, reads the whole chain at once | Kimi K3 |
|
Label tens of thousands of lines (device reg / network / crypto / noise) | Cheap, thousands of calls at high concurrency | Claude Sonnet 5 |
|
Judge which of the four thresholds a link is stuck at | Strong reasoning, willing to make the call | Claude Opus 5 |
|
Compare the divergence point of two traces, explain why a response is empty | Mid-reasoning attribution, explains against specific lines | GPT-5.6 Sol |
|
The labeling step is the one worth calling out. It's pure grunt work, marking off the noise lines in tens of thousands and keeping only the device-registration, crypto, and network classes, and the volume is too high to do by hand. This "simple judgment at very high frequency" is exactly the cheap tier's home turf, and running it on the reasoning tier is a straight waste of money. Once the log is compressed to a few hundred labeled lines, hand it to the reasoning tier for the "which threshold" judgment, and both the cost and the precision line up.
The effect of this split, don't take my word for it, run one round:
Cut a segment of trace from one link (a few thousand to tens of thousands of lines).
Label it in chunks with
claude-sonnet-5first, marking off the noise lines and keeping the device-registration, crypto, and network classes.Feed the labeled summary to
claude-opus-5and have it output "which of the four thresholds this link is currently stuck at, plus which lines the evidence points to."Control: dump the same raw trace whole into a single model and ask the same question.
Look at one thing: does it hand you a blanket guess, or land on a specific threshold and specific lines. That difference is your selection criterion.
A hook is a version-bound observation probe, not a signer
Why are the hooks on these three platforms retained yet firmly kept out of the command catalog? Because a hook is a probe, not a signer, and those two are fundamentally different in nature.
A hook is bound to one specific build of the app, say version 32.x. The symbols, offsets, and method layout it depends on all belong to that version. Upstream ships a new version, all of it shifts, and your hook dies on the spot. It's inherently perishable and version-bound. It answers "what is this version doing internally right now," an observation question. That's exactly what an instrumentation tool is for: the Frida documentation describes Interceptor as a means of observing and rewriting calls at runtime. It lets you see how a function is called and what the arguments are, but it is not itself a deliverable that "computes a signature from an input."
A signer that goes into the endpoint catalog is the reverse: it has to be stable, repeatable, standalone, and CI-able. It answers "give me an input and I compute the right signature," a reusable-capability question. Even if a hook lets you see the skeleton of the signature algorithm clearly, that's only the algorithm-fingerprinting step, still a whole purification and differential-verification flow away from a standalone signer.
This also explains the order in the purification ladder: a capability should first fight for pure Python direct connection, then fall back to running a minimal signature fragment on local Node/V8, and only grudgingly accept a passive browser bridge. A hook isn't even on that ladder yet. It sits upstream of it, in "research," not "implementation." Registering a version-bound observation probe as a signer hangs the "implementation" sign on a half-built piece of "research."
How to archive a research asset so it doesn't rot
Keeping a hook out of the endpoint catalog doesn't mean throwing it away. It's a research asset you spent real effort on, with link knowledge, samples, and validation evidence inside it, and deleting it is a net loss. The point is to archive it properly, or in two months even you won't know how far it actually got.
An app research asset records at least these four things:
Transport type. Whether the link runs on the native protocol, a local Node/V8 signature, or a passive browser bridge. This decides how far it can be purified later.
Evidence tier. How many of the four thresholds it passed. Whether it's "link located," or "got a non-zero install state but the response is empty," or "non-empty but not repeatable." This one tells whoever comes next exactly what stands between it and usable.
App version / build. Which version the hook is bound to. Without this, the next time upstream updates you have no way to tell whether you wrote it wrong or they changed the build.
Sample summary. What the input and output from this run looked like, kept in a desensitized copy. It's the fastest anchor for restarting this research.
Record those four and a link stuck at 0 commands is an asset you can push forward, not a heap of expired logs. It also blocks the two worst handlings at once: deleting the hook and pretending you never did the research, or forcing it into the catalog and pretending it's usable. The second is especially expensive. An endpoint that "looks callable, actually returns empty" spreads its cost across every downstream caller who trusts the catalog: they write their integration against it, write retries, get burned once by empty data, and end up distrusting the whole catalog. An honest gap marked "research asset, 0 commands" costs only you, once, in a single line of a README.
An empty shell costs more than a gap, and nowhere is that clearer than here.
One key smooths the log-splitting friction
Back to the log split from the fourth section: long context to read the whole trace, the cheap tier to label in bulk, the reasoning tier to classify, the mid-reasoning tier to attribute. Four tiers from several vendors, four SDKs, four auth schemes, four error formats. To classify a Frida log across four clients, most people run the numbers, decide it isn't worth it, and end up gnawing on hundreds of thousands of trace lines with one model, either burning money or unable to get through it.
AIReiter removes that friction: one key, one OpenAI-compatible interface, all four tiers behind it, switching by changing the model field in the request body.
# Label the log: the cheap tier, thousands of calls at high concurrency
curl https://aireiter.com/api/v1/chat/completions \
-H "Authorization: Bearer $AIREITER_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-5",
"messages": [{"role": "user", "content": "<one chunk of trace + labeling instruction>"}]
}'
# Classify the threshold: switch to the reasoning tier, leave the rest
# "model": "claude-opus-5"
# Read the whole call chain: the long-context tier
# "model": "kimi-k3"
# Attribute the response difference:
# "model": "gpt-5.6-sol"
If you already use the OpenAI SDK, point base_url at https://aireiter.com/api/v1 and change nothing else. On the Anthropic SDK, hit POST /api/v1/messages with the same key.
This piece has a lopsided cost structure, and the discount sits right on the lopsided part. One link's trace is tens of thousands to hundreds of thousands of lines, labeling runs chunk by chunk, and one platform comes to hundreds or thousands of claude-sonnet-5 calls. That is the bulk of the cost. The long-context calls that read a whole trace are a few hundred thousand tokens per input, another large chunk. Classification and attribution are few calls at a higher unit price. Claude at 30% off lands right on labeling and threshold classification, the two most expensive parts. GPT at half price lands on the attribution tier. Long-context reading runs on Kimi K3, callable on the same key.
Try it without signing up: hand a segment of trace to
claude-sonnet-5to label, then haveclaude-opus-5classify it, and see whether it can point straight at which threshold it's stuck at before you decide to wire it in.
In closing
In app reverse engineering, a hook landing gives you the observation capability of "I can see what this version is doing internally." The endpoint catalog wants the call capability of "give me an input and I reliably emit non-empty, correct data." Between those two sit four evidence thresholds: real install state, non-empty response, error classification, and a repeatable test.
Three platforms fully located, all hooks landing, commands still at 0 isn't failure, it's discipline. Real install state doesn't pass, so you honestly stop at 0, archive the hook as a research asset, and push it forward once the device profile is in hand.
The model's place in this flow is specific. It splits, labels, classifies, and attributes hundreds of thousands of trace lines for you, compressing "where is this link stuck" from an afternoon of reading logs to a few minutes. But the judgment of "did it pass the threshold," like the judgment of "is the hypothesis correct" in differential testing, doesn't fall to the model in the end. It falls to the four criteria you set and the evidence that runs through repeatably. How the whole reverse-engineering workflow puts the model where it belongs is in the four-stage overview.