Between a keyword tool's output and the decision you have to make sits one conversion nobody does for you. The tool hands you a table: a search volume, a CPC, and a competition level per word. What you want isn't that table, it's one number, how much budget this keyword set should get this month. How those three isolated metrics collapse into one budget figure is what this piece is about, along with where a model can actually save you time and where it only makes things worse.
First, the boundary: all data here is from each platform's public ad library and creative center, logged in with your own account, no signatures, nothing bypassed. This is about turning public data anyone can see into a budget decision.
Three metrics, three traps
Most people take the keyword table and treat the three columns as three fixed values. None of them are.
Search volume is a window function. The same word can have search volume differing several times over across a 7-day, 30-day, or 180-day window, usually with month-over-month and year-over-year figures next to it. A word that's "low volume but doubled year over year" and one that's "high volume but halved year over year" mean opposite things, and if you copy only the volume column, those two words look identical.
CPC is a range, not a point. The planner returns a low-to-high range. Grabbing a median offhand erases the uncertainty the platform explicitly told you about. The right use is the high end for a conservative budget and the low end for an optimistic return. Taking a point value is pretending you know something you don't.
Competition is a relative tier, not absolute difficulty. That "low / medium / high" is a label applied relative to the platform's whole keyword pool at filter time. It only says "among these words it's on the hard side," not "it's this hard in absolute terms to win." Change the market or the category and the tier shifts. Treating it as an absolute score comparable across keyword pools is the sneakiest of the three traps.
Aggregate by set, not by word
Once you understand all three metrics are dishonest, the next question is: why can't you compute "volume x CPC" per word and sum it up?
Because near-synonyms' search volumes don't add. "Buy X" and "X price" are very likely the same people searching in the same purchase, and summing them directly counts one person twice, inflating the budget out of nowhere.
So a mature planner, beyond the per-word view, has a keyword-set aggregate view: submit a group of words as a set and get back the whole set's search-volume range and a budget-estimate range, the total after it removes the overlap between words for you, not a per-word sum. The platform itself gives budget by set. Those three per-word columns are for sorting and grouping, not for summing.
A topic word is not a search word
This is the most valuable trap in this piece, and the one the most people fall into: what you call your product and what the market actually types into the search box are often two different things.
A business topic word is your internal language, the product name, the feature name, the positioning word on the slide. A search word is the user's language, more colloquial, more problem-oriented. You do "AI video generation," your topic word is "AI video generation," and the people actually searching type "how to turn a photo into a video" or "text to video free." The two groups of words aren't in the same tray for volume, CPC, or competition.
This mismatch is especially fatal in cross-market launches (the topic word translated straight over, and the target market simply doesn't talk that way) and new categories (you invented a word, and the market only describes the new need with old words). The fix isn't mysterious: the topic word is only a seed, and the words that actually enter the budget table have to grow out of the market's search words. A mature flow treats "creative topic words" and "market search words" as two separate fields. The topic word decides what the creative talks about, the search word decides where the budget goes. Merge them and you'll buy a batch of words nobody searches, in your own language.
Model expansion: give it seeds and context, ask for intent alongside
Expanding from a few seeds to a keyword table that covers the market's real searches is the first place a model can genuinely save time, provided you don't ask the wrong question. Ask "expand related keywords for me" and you get a pile of structureless synonyms. The right prompt feeds in the business context and forces the model to hand back three things per word:
I'm doing search advertising for <one-sentence product description> in
<target market/country>.
Seeds: <3 to 10 words you already know>
Expand the keywords, and output each in the structure below, no prose:
1. Keyword
2. Intent stage it belongs to (awareness / comparison / decision / brand)
awareness: the user is describing a problem, doesn't know the solution yet
comparison: the user is weighing several solutions
decision: the user is ready to buy, looking for a channel/price/alternative
brand: the user is searching a specific brand name
3. Why it's adjacent to the seed (a different phrasing of the same need?
the downstream next step? an adjacent need in the same scenario?)
The second item is the hook that turns a word into budget. Without an intent stage, an expanded word is just a string. The third is for your review: words the model invented that are actually unrelated to your business tend to fall apart in that column. Two columns force the model to give a verifiable structure, not a plausible-reading association.
Feedback validation: the model's guesses have to be scored by real search volume
This step is the gate on the whole chain, and where the most people skip and then launch on a table of hallucinated words. A model's expanded words are guesses until you feed them back to the tool. It can expand a word that's grammatically clean, plausibly intent-tagged, and searched by zero people in a month. The model has no live search volume, it samples on "does this look like a real search word," not on "does anyone actually search this."
The method is mechanical: feed each word back to the planner, get real search volume, CPC, and competition, and keep only the words with non-zero volume whose intent tag is consistent with the real metrics. Consistent means: a word the model tagged "decision intent" usually has a higher real CPC and fiercer competition. If it's tagged decision intent but the CPC is absurdly low and competition is the lowest tier, that's not a find, it's a mislabeled intent, send it back.
This division of "model gives the hypothesis, real data makes the call" is the same skeleton as letting a model recognize algorithms in obfuscated code and then gating its hallucinations with differential testing in reverse engineering: the model is strong at generating candidates, weak at verifying facts, and you don't let it do both at once. "Feedback validation" for keywords is the differential test here, and the assertion volume > 0 is far more authoritative than the model saying "this is a good word."
Budget lands as numbers by intent stage
The validated keyword table now carries a usable structure: each word has a real search-volume range, CPC range, competition tier, plus an intent-stage tag. Only now can budget be computed. The key is not to spread it evenly over words, but to allocate by intent stage:
Decision-intent words ("X price," "X alternative," "buy X"): closest to conversion, most valuable per click, but low volume, fiercest competition, highest CPC ceiling. High priority, high price tolerance, but with a ceiling, because only so many people are ready to buy.
Comparison-intent words: middling on both conversion and volume, catching the budget that overflows from decision words.
Awareness-intent words ("how to do X," "what is X"): highest volume, lowest CPC, furthest from conversion. Low allocation for coverage, and don't hold them to a decision word's conversion expectation.
Within each group, sort further by "volume range x target share x CPC range." A keyword set collapses in the end not into a number pulled from the air, but into a budget structure split by intent stage, one range per stage.
Which model for which step
The jobs the model does on this chain differ widely in what they need, and one model for the whole thing either burns money or burns precision:
Step | Capability it needs | Pick | model id |
|---|---|---|---|
Keyword expansion + intent classification | Strong reasoning, understands business context, judges intent, argues "why adjacent" | Claude Opus 5 |
|
Fill the whole context then expand (keyword set + launch history + landing-page copy) | Long context, feed the full business context without dropping detail | Kimi K3 |
|
Tag and dedupe hundreds to thousands of expanded words per item | Cheap, high concurrency, runs at volume | Claude Sonnet 5 |
|
Post-feedback deviation attribution (why the model's intent doesn't match real metrics) | Mid reasoning, explains against numbers | GPT-5.6 Sol |
|
The first step is the one worth calling out: expansion and intent classification is the only step where switching models visibly changes the result. It tests "does it understand your business" plus "will it admit a word is actually irrelevant." A weaker model gives you a fully-expanded table with every intent tagged "decision," and one feedback pass shows a dismal hit rate. A strong reasoning tier flags the non-adjacent words on its own. Test the difference yourself, the protocol is simple:
Pick one seed keyword set (5 to 10 words whose business context you know).
Feed the same "seeds + business context" to
claude-opus-5andgpt-5.6-sol, ask each for 50 words, both output as "word + intent stage + why adjacent."Feed both back to the planner for real volume, CPC, competition.
Look at two things: how many expanded words have non-zero real volume (not invented strings), and whether the intent tags are consistent with real CPC and competition (decision-intent words should be pricier and more crowded).
That hit rate is your selection criterion, it decides directly how many hallucinated words downstream has to throw away. One round shows you the difference, more directly than any benchmark.
The switching cost is the real obstacle
Four models from three vendors, three SDKs, three auth schemes, three error formats. Rewriting your client three times to switch models between expansion, bulk tagging, and attribution isn't worth it, so most people end up on one model the whole way, use a tier that doesn't get intent tiering for expansion, and get a low feedback hit rate without knowing why.
AIReiter removes that layer: one key, one OpenAI-compatible interface, all four tiers behind it, switching by changing the model field in the request body.
# Keyword expansion + intent classification: the reasoning tier
curl https://aireiter.com/api/v1/chat/completions \
-H "Authorization: Bearer $AIREITER_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-opus-5",
"messages": [{"role": "user", "content": "<expansion prompt + seeds + business context>"}]
}'
# Tag hundreds of expanded words in bulk: change the model field, leave the rest
# "model": "claude-sonnet-5"
# Post-feedback deviation attribution:
# "model": "gpt-5.6-sol"
If you already use the OpenAI SDK, point base_url at https://aireiter.com/api/v1 and change nothing else. On the Anthropic SDK, hit POST /api/v1/messages with the same key.
On price, Claude models run at 30% off list and GPT models at half. This flow's main cost isn't expansion, which is a few reasoning-tier calls, since one keyword set is only a few rounds of asking. What actually burns volume is bulk tagging: hundreds to thousands of candidate words per planning pass, tagged for intent and deduped item by item, the cheap tier's high-concurrency bulk, and Claude at 30% off lands right on that densest call.
Try it without signing up: expand one keyword set by hand first and compare the feedback hit rate of the two tiers' words, then decide whether to wire it in.
In closing
A keyword tool gives you three dishonest metrics: search volume is a window function, CPC is a range, competition is a relative tier. A launch needs a budget structure split by intent stage. The skeleton of that middle conversion is: dedupe by set rather than summing per word, seed from market search words rather than business topic words, have the model expand words with intent attached but validate every word against real search volume, and allocate budget by intent stage rather than spreading it evenly over words. The model is an accelerator for expansion and intent tiering, not the decider of budget.
Once this budget figure is computed, it's the input to the commercial-validation stage of the keyword-to-finished-ad creative pipeline: only keyword sets whose budget clears the bar are worth carrying on to creator matching and asset generation. The parallel decision line is reading an ad's second-by-second retention curve: one governs which words you put money on, the other governs which creative you put money on.