You open the Network panel in DevTools and watch a GraphQL-driven page finish loading. You go through every request body looking for the query { ... } you want, and there isn't a word of it. All that's there is an operationName, a sixty-four-character hash, and a bag of variables.
You didn't miss it. This is a persisted operation: the client no longer sends the plaintext query, only a pre-registered hash. The server takes the hash, looks it up in its own registry, reconstructs the real query, and runs it. The packet-capture route stops working right here. You can see which operation was called and what variables went with it, but not which fields this operation selects or what the response structure looks like.
Most people's first reaction here is the wrong one: they think they have to "crack" that hash. A hash is one-way, you can't crack it, and there's no need to. The real problem is classification. Figure out which situation you're in first, then decide how to get it. The ways of getting it differ by an order of magnitude in cost, and pick the wrong one and you've wasted the effort.
What a persisted operation is, and why the capture shows no query
First, why this mechanism exists, or you can't judge how to get around it.
The trouble with plaintext GraphQL is direct: the query string is long and heavy, sending the whole field tree on every request is wasteful, and the server has to accept arbitrary queries, which opens the entire schema's attack surface. A persisted operation answers both. At build time, every query the client will use is extracted, hashed, and registered as a whitelist on the server. At runtime the client sends only the hash plus variables, the server only honors registered hashes, and any query not on the whitelist is rejected. This is a genuine performance design, not deliberate anti-scraping. Apollo's Automatic Persisted Queries docs present it as a recommended practice, using the query's SHA-256 in place of the plaintext, and the capture not showing the query is just a side effect.
For reverse engineering, that side effect is specific: the "what to fetch" has been pulled out of the request. You're left holding three things: an operation identifier (a hash, or a readable operationName), a bag of variables, and a response. The middle layer, "which fields this operation selected," isn't on the wire.
In real projects you hit both ends of this spectrum. At one end, the platform never adopted persisted anything, and the plaintext query is still sitting honestly in the request body. At the other, the request is compressed into an opaque envelope where you can't even read a field name. Each end calls for a different way of getting it, covered separately below.
Two ways to get it, an order of magnitude apart in cost
Way one: the plaintext or a mapping table is right there in the client build, and you find it. Way two: you don't chase the plaintext at all, you treat the whole operation as a black box and replay it verbatim.
The first sounds more thorough, so most people charge down that road by default. That's exactly where the wasted effort starts. It's only cheap when the plaintext really was shipped to the client, and that premise often doesn't hold.
Way one: go find the plaintext or mapping in the client build
The lowest-cost case is when the plaintext was never hidden.
One Chinese short-video platform's hot-rank works this way: a single /graphql endpoint, the request body is the standard {operationName, variables, query} trio, the query field holds the complete plaintext GraphQL, and operationName is a readable name like hotRankQuery. There's nothing to "get" here, one capture has it all in front of you. It didn't even adopt persisted operations, and it's the easiest end of the spectrum.
A bit more work is the case where persisted operations really are in use but the client still holds the mapping. To send a hash, the client has to know which operation maps to which hash, and that operationName-to-hash table (sometimes with the plaintext query alongside) is usually built into the frontend bundle. Build tools often generate it as a manifest file, or inline it into some module. Find it and you have the plaintext and the hash at once, free to add fields or change the selection set afterward.
The hard part isn't the act of finding, it's that the build is tens of thousands of lines, minified and obfuscated, and the mapping may be split and inlined. This step is where the model earns its place in this piece, covered on its own below. It's chunked retrieval, not reasoning.
Way one's test is simple: if the plaintext or the mapping is anywhere in the client, spend ten minutes finding it first. Found, this is the strongest way to get it, and you have full control over the endpoint.
Way two: don't chase the plaintext, replay the operation as a black box
The problem is that the plaintext often isn't in the client.
A properly done persisted operation leaves the client holding only the hash, with the plaintext query living only in the server's registry. You can search the bundle to shreds and not find it, because it was never shipped. Grinding on way one here is hunting for something that doesn't exist.
Way two is the most underrated solution here: you don't need the plaintext. What you want is the response, not the query's field tree. So record this operation's identifier (hash or operationName) and its variable envelope, send it verbatim, and swap only the input parameters you care about. You never know which fields it selected, but the server returns them all the same. For the vast majority of data-fetching and monitoring work, that's enough.
YouTube's innertube is the standard form of this way. It isn't even GraphQL: a fixed set of self-describing endpoints youtubei/v1/{player,search,next}, with a request body of a context envelope (client type, version) plus a bag of parameters. Nobody goes to "reconstruct" YouTube's internal query graph, which is neither possible nor worthwhile. The real move is to read the client version and context out of the current page resources once, then carry that envelope verbatim on every request, swapping only inputs like videoId or a search term, hitting that fixed endpoint. The operation's semantics stay a black box throughout. This case, where a key value isn't in the static code and has to be read out of runtime resources, is its own class of reverse-engineering problem, covered separately.
Way two's advantage is that it's independent of whether the plaintext is in the client. Hash or opaque envelope, you don't try to understand it, you just reproduce it faithfully. The cost is that you're locked to the requests the client already sends. Want a field the client never requests, and black-box replay can't give it to you.
There's one more case that breaks replay: the envelope carries a signature field computed fresh per request that expires. Black box ends there, and you have to work out that one field on its own, and recognizing which signature algorithm family it belongs to is the job of another piece.
Where each platform lands, and why
Put those two real cases in one table and the landing spot and the reason are clear:
Platform case | Request shape | Plaintext in the client? | Natural way | Why |
|---|---|---|---|---|
A short-video hot-rank GraphQL |
| Yes, plaintext right in the request body | Way one (near zero cost) | No persisted operations, readable operationName, plaintext query, one capture has it all |
YouTube innertube | Fixed endpoints + a | No plaintext query to speak of | Way two (black-box replay) | Not GraphQL, no query to reconstruct, read the context envelope once and carry it verbatim |
The contrast at the two ends says one thing: the way to get it isn't your choice, it's decided by the platform's API design. The first put its strength elsewhere (plaintext in the open is fine, it guards against abuse by other means), and you pick it up in passing. The second made "what to fetch" an opaque envelope, so there's no plaintext to chase and all you can do is replay.
The big stretch in the middle, real persisted GraphQL, is where the judgment gets tested. The plaintext might be in the client (the mapping built into the bundle, way one) or only on the server (the client has only the hash, way two). You have to classify it to one end before you start.
Pick wrong and it's wasted: first ask whether you need the plaintext
The order-of-magnitude cost difference rides entirely on one judgment.
If a platform really ships only the hash and locks the plaintext on the server, and you insist on way one to recover the plaintext, you might dig through the bundle for days and find in the end that the thing you're hunting was never shipped. That's not a difficulty problem, it's a direction problem, and no amount of effort buys a result.
The other way around, if you need to modify the query (request a field the client never requests), way two's black-box replay can't help, and you're back to way one for the plaintext, and stuck if you can't get it.
So the order shouldn't be "how do I get the query," it should start with one question: do I actually need the plaintext?
Just reproduce a request the client already makes and read its response: way two (black-box replay). Cheapest, most overlooked, works whether or not the plaintext is around. Default to this.
Need to change the selection set, assemble a query the client never sends: way one (recover the plaintext) is required. Whether it's cheap depends on whether the client shipped the mapping. If it didn't, that's the order-of-magnitude cost blowup, and you either accept it or reconsider whether you really need to change the query.
Putting this judgment first heads off most of the "charge down way one, three days stuck before realizing it should have been way two" waste. This kind of "rewrite or tolerate the black box" degradation tradeoff is the subject of the purification ladder piece, and here it just drives the choice of which way to get it.
Locating the mapping is chunked retrieval, not reasoning
The one technically demanding step in way one, finding the operation declaration and mapping in tens of thousands of lines of frontend build, is exactly where the model saves you real time. But first be clear about what kind of task it is.
It isn't a reasoning task. You don't need the model to understand what the code computes, you need it to locate, in a large body of text: which block declares the operationName-to-hash mapping, which module inlines the plaintext query, where the context envelope is assembled. This is chunked retrieval, and what it tests is whether you can fit enough context in at once and point precisely inside it, not whether you can argue against yourself.
Do the mechanical splitting step first: use a script to break the build into modules, index it, and filter out polyfills and unrelated business modules. Hand the model the result and the task becomes "find the declaration in these blocks," clean and simple.
The four tiers split cleanly by capability in this flow:
Step | Capability it needs | Pick | model id |
|---|---|---|---|
Locate the mapping / operation declaration in tens of thousands of lines | Long context, swallows a big chunk of build and points precisely | Kimi K3 |
|
Decide way one or way two, induce from a few samples which envelope fields vary | Strong reasoning, reads structure and makes the tradeoff | Claude Opus 5 |
|
Label hundreds of operations in bulk, generate replay stubs, fill in variable types | Cheap, high concurrency | Claude Sonnet 5 |
|
When replay won't connect, read the diff to attribute (missing context field? hash version changed?) | Mid reasoning, explains against the response difference | GPT-5.6 Sol |
|
The first tier is this piece's home ground. Switching models at the locating step visibly changes the result, because what it's bounded by is the context window. The build is tens of thousands of lines, a short-context model can't fit it and has to truncate, and one truncation may cut out the mapping, after which its "can't find it" isn't because it can't search, it's because it never saw it.
Don't take my word for the difference, test it:
Capture a real request and store its operation identifier (operationName or hash) and variables.
Chunk the frontend build with a script, feed it along with that identifier to
kimi-k3, and have it locate "where this operation is declared, and which block holds the matching plaintext query or hash mapping."Look at one thing: does its location jump straight to the line. Hit, miss, or an adjacent but wrong spot.
Control: feed the same input to a short-context model and see whether it misses because it can't fit. The hit rate is your selection criterion.
One round shows you that long context on this kind of task isn't "a bit better," it's the difference between can and can't.
The switching cost is the real obstacle
Four models from three vendors, three SDKs, three auth schemes, three error formats. To use a different model for retrieval, judgment, bulk, and attribution, the naive move is to wire all three clients, and most people run the numbers, decide it isn't worth it, and end up on one model the whole way, using a tier that can't fit the build at the long-context retrieval step, then blaming "the model can't find it."
AIReiter flattens that layer: one key, one OpenAI-compatible interface, all four tiers behind it, switching by changing the model field in the request body.
# Locate the mapping: the long-context tier, swallows a big chunked build at once
curl https://aireiter.com/api/v1/chat/completions \
-H "Authorization: Bearer $AIREITER_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "kimi-k3",
"messages": [{"role": "user", "content": "<chunked frontend build + the operation identifier to locate>"}]
}'
# Label operations / generate replay stubs in bulk: change the model field, leave the rest
# "model": "claude-sonnet-5"
# Replay attribution:
# "model": "gpt-5.6-sol"
If you already use the OpenAI SDK, point base_url at https://aireiter.com/api/v1 and change nothing else. On the Anthropic SDK, hit POST /api/v1/messages with the same key.
On price, Claude models run at 30% off list, GPT models at half, and Kimi K3 is callable on the same key. This flow's cost sits in two places: way one's locating, feeding the whole chunked build in at a few hundred thousand tokens per input, on kimi-k3; and labeling hundreds of operations and generating replay stubs in bulk, the most call-heavy step, on claude-sonnet-5 at 30% off. The discount lands right on the most call-dense batch.
Try it without signing up: feed a chunk of build by hand first and watch whether the long-context tier locates the mapping in one pass, then decide whether to wire it in.
In closing
A GraphQL API with no docs doesn't mean you can't integrate it. A persisted operation only pulled "what to fetch" out of the request, into one of two places: the client build (find it, way one), or the server only (don't chase the plaintext, replay the operation as a black box, way two).
Those two ways cost an order of magnitude apart, and what decides which one is never "which is more thorough," it's two earlier questions: is the plaintext in the client, and do you need to change the query. Answer those two before you start and you head off most of the wasted effort.
The model's place here is specific: the "locate the declaration in tens of thousands of lines" step in way one is pure retrieval, and a long-context tier swallows it in one pass and points precisely, turning days of digging into minutes. It doesn't decide which way to get it for you, that's the judgment you should have after reading this, it just does the grunt work of finding the mapping.