Blog
Qwen3.8-27B: Local Setup, VRAM, and Reasoning
A hardware-first guide to Qwen3.8-27B: what officially shipped, what 12GB to 32GB setups can realistically do, and why the default reasoning mode matters.
How to Use DeepSeek in Codex: Setup, Limits, and Cost
Codex reaches models over the Responses API, which DeepSeek supports natively. Three setup routes, how to verify it took effect, the four limits that surprise people, and verified pricing.
Qwen3.8-Max API Pricing: $2/$6 per Million Tokens
Qwen3.8-Max went GA on August 3, 2026 at $2 input and $6 output per million tokens. All five cache tiers, the rate limits, why the xhigh reasoning default inflates the bill, and how the cost compares.
OpenRouter Alternatives: 7 Options With Real 2026 Prices
OpenRouter tops out when you need image/video generation, cheaper tokens, or production guardrails. Seven alternatives mapped to why you're leaving, with verified prices.
GPT-5.6 Price Cut: What Luna and Terra Really Cost Now
A verified breakdown of the July 30 GPT-5.6 price cut: exact new rates for Luna, Terra and Sol, why cache-write billing can eat the discount, and which tier each workload should move to.
Invalid API Key: Diagnose 401 and 403 Before You Fix
Three different causes produce the same 401. Reproduced responses show how to tell a rejected key from one that never arrived, why an env var outranks your login, and which 403s leave the key alone.
Fix OpenRouter 429: Provider Error or Rate Limit?
A field-by-field OpenRouter 429 diagnosis with current free-model limits, provider fallback checks, and retry code that respects Retry-After.
DeepSeek V4 Flash vs GLM-5.2: 0731 Update Tested
I ran DeepSeek V4 Flash (0731) and GLM-5.2 on four graded tasks. They tied on three, GLM-5.2 never returned an answer on the fourth, and Flash cost 19x less overall. Prices verified 2026-07-31.
B2B Ad Intelligence Only Comes From HTML: Writing a Parser That Survives Redesigns
Most B2B ad libraries have no JSON API, only server-rendered HTML. How to write a redesign-resistant streaming parser, and where a model helps: diffing old and new HTML for repairs.
Cross the Creator Library Against the Asset Library: Stop Picking by Follower Count
Picking creators by follower count is systematically wrong. Extract structure from your effective creative, reverse-look-up creators whose content matches, and make each match trace to a specific video.
From a Keyword Set to a Budget: How Volume, CPC, and Competition Collapse Into One Number
A keyword tool gives three isolated metrics; a launch needs one budget number. The traps in each, why you aggregate by set, and why expanded keywords are hypotheses until real volume validates them.
How to Read an Ad's Second-by-Second Retention Curve: Three Curves and One Percentile
Reading competitor creative by play count and likes is reading nothing. What guides your next script is where three per-second curves inflect and where CTR sits in its category, not the absolute number.
22 Platforms, 241 Commands, and Why I Didn't Build a Unified Response Model
The unified Post/User model is the first instinct in cross-platform collection, and it collapses at twenty-plus platforms. Make each platform its own bounded context and normalize per-platform.
329 Commands, 128 Still Missing: Reconcile a Migration With Set Math, Not a Model
A model translates one function well, but 'did the whole thing migrate' is set math. Derive both command sets from declarations, diff with a script, and let the model explain the misses, not find them.
A GraphQL API With No Query in the Request: Two Ways to Get a Persisted Operation
Persisted GraphQL leaves your capture an operationName and a hash, no query. Find the mapping in the client build, or black-box replay the hash. Two ways to get it, an order of magnitude apart in cost.
Can't Find the Signing Logic in the Code? It Might Not Be in the Code at All
Some signatures aren't in the static code at all. The server ships a one-time script at runtime, or the value is a build constant that rolls every release. Execute or read it, don't reverse it.
Every Frida Hook Landed. Why I Still Had Zero Usable Endpoints.
Every Frida hook landed and I got zero usable endpoints. A hook landing isn't a capability: four evidence thresholds sit between them, and an honest 'not usable yet' beats shipping an empty shell.
Feed 300 Competitor Ads to a Model and You Get Four Adjectives: Do This Instead
Ask a model to summarize 300 competitor ads and you get four useless adjectives. Extract structured fields one ad at a time first, then cluster the fields and make the model bring a counterexample.
From Keyword to Finished Ad: A Five-Stage Pipeline and the Six Statuses That Keep It Operable
Ad-creative automation isn't 'type a keyword, get a video.' A five-stage pipeline, the six statuses that make it operable, and the evidence gate that stops it generating off zero evidence.
Agent Tool Descriptions Drift: Make the Declaration the Single Source of Truth
Hand-maintained tool descriptions drift from the real function signatures, so the model calls tools with arguments you changed. Derive them from the declaration so nothing is left to drift.