A $0.15 input price is only the starting point for GLM-5.3-Flash. Cache hits, output tokens, forced reasoning, and a roughly 320B-parameter checkpoint determine whether the API or open-weight route fits a real deployment.
The price that is actually live
Z.AI’s official pricing page lists GLM-5.3-Flash at $0.15 per 1 million input tokens, $0.03 per 1 million cached-input tokens, and $0.50 per 1 million output tokens. It also shows a temporary 50% promotion, ending at 24:00 on September 9, 2026, in UTC+8 (Singapore time).
| GLM-5.3-Flash charge | List price / 1M tokens | Promotion / 1M tokens |
|---|---|---|
| New input | $0.15 | $0.075 |
| Cached input | $0.03 | $0.015 |
| Output | $0.50 | $0.25 |
| Cached-input storage | Limited-time free | Limited-time free |
Treat the 50% reduction as temporary promotional pricing, not a production-budget baseline; Z.AI’s launch announcement identifies ox-alpha as the earlier preview identity.
The same official table lists standard GLM-5.3 at $1.40 input, $0.26 cached input, and $4.40 output per 1 million tokens. On list prices, Flash is about 89% cheaper for both input and output. That is a price comparison, not a claim that the two models have identical quality, latency, or operational behavior.
What you are buying: API access, not a smaller local model
GLM-5.3-Flash is an MIT-licensed open-weight multimodal model with 320B total parameters and 18B active parameters per token. The official Hugging Face model card documents the public zai-org/GLM-5.3-Flash checkpoint and local-serving paths; Z.AI describes a 1-million-token context window and text, image, and video input in its release material.
“18B active” refers to selected expert computation, not total memory. Full weights, KV cache, runtime overhead, multimodal processing, and context length still determine local feasibility.
| Access route | What it provides | Main constraint |
|---|---|---|
| Z.AI API | Hosted inference billed by input, cached input, and output tokens | Account limits and reasoning latency |
| Official checkpoint | MIT-licensed weights for self-hosting or adaptation | High-memory infrastructure, depending on precision and quantization |
| Quantized community build | Smaller files for some local runtimes | Build quality, modality support, speed, and compatibility vary |
The low token price and low hosting cost are different claims: bursty users pay as API traffic arrives, while self-hosting reserves serving capacity even during idle periods.
API billing in real workloads
GLM-5.3-Flash cost depends on the mix of new input, recognized cached input, and generated output:
cost = (new_input_tokens / 1,000,000 × input_rate)
+ (cached_input_tokens / 1,000,000 × cached_rate)
+ (output_tokens / 1,000,000 × output_rate)
At list price, 10 million new input tokens plus 2 million output tokens costs $2.50: $1.50 for input and $1.00 for output. During the promotion, the same volume costs $1.25.
| Workload | List-price calculation | List-price total | Promotional total |
|---|---|---|---|
| 10M new input + 2M output | $1.50 + $1.00 | $2.50 | $1.25 |
| 2M new input + 8M cached input + 2M output | $0.30 + $0.24 + $1.00 | $1.54 | $0.77 |
| 100K new input + 10K output | $0.015 + $0.005 | $0.020 | $0.010 |
The $0.10-per-million blended figure shown by Artificial Analysis uses a defined 7:2:1 cache-hit, input, and output mix. It is useful for a comparable workload, not a universal GLM-5.3-Flash rate.
Cache savings require an actual cache hit. Z.AI’s Chat Completion response schema exposes usage.prompt_tokens_details.cached_tokens, so production accounting should record that field instead of assuming repeated-looking prompts will be discounted. One user in r/opencodeCLI put the launch economics this way:
“There's a 50% discount offer until September 9th, and its context consumption is much better than in dsv4f...” — u/CriteriumA, Reddit discussion
The discount date is confirmed by Z.AI’s rate card; the comparison and user experience are that commenter’s opinion. Stable repeated context can improve cache reuse, but it does not remove output, retry, tool, or account-limit costs.
A compact budget worksheet
Forecast four separate rows: new input, cached input, output, and paid tools. Z.AI lists Web Search at $0.01 per use, so repeated agent searches can exceed a token-only estimate.
| Monthly usage assumption | Formula at list price | Monthly cost |
|---|---|---|
| 100M new input + 20M output | 100 × $0.15 + 20 × $0.50 | $25.00 |
| 20M new input + 80M cached input + 20M output | 20 × $0.15 + 80 × $0.03 + 20 × $0.50 | $23.40 |
| 100 Web Search calls | 100 × $0.01 | $1.00 |
These are arithmetic examples, not a quota. Add retries and abandoned agent runs separately. If output volume is high, setting a realistic max_tokens value and using a lower reasoning_effort where suitable can matter more than optimizing a small prompt prefix.
Z.AI API constraints before migration
The official Z.AI API uses https://api.z.ai/api/paas/v4/ as its general base URL and https://api.z.ai/api/paas/v4/chat/completions for chat completions. Authentication uses a Bearer API key. Z.AI documents OpenAI-compatible Python, Node.js, and Java client patterns through a custom base URL in its API introduction.
The current Chat Completion reference mentions GLM-5.3-FLASH in its thinking and reasoning descriptions, but its rendered model enum does not display glm-5.3-flash; the official pricing page and launch materials do list the Flash SKU. Test the exact hosted model ID with a small request before switching production traffic.
| Integration detail | Verified constraint |
|---|---|
| Hosted model ID to test | glm-5.3-flash; confirm acceptance against the live account schema |
| Hosted endpoint | https://api.z.ai/api/paas/v4/chat/completions |
| Authentication | Authorization: Bearer YOUR_API_KEY |
| Thinking | GLM-5.3-FLASH uses enabled thinking; depth is controlled with reasoning_effort |
| Reasoning effort | low, high, or max |
Maximum max_tokens | 131,072 in the current parameter reference |
| Coding Plan | Separate key, endpoints, and credit system; it is not a general API balance. See Z.AI’s Coding Plan quick start |
Z.AI sets max as the default reasoning effort. The token rate does not change by effort level, but reasoning can change output usage and waiting time. For account limits, Z.AI directs users to an account-specific rate-limit dashboard; its public guidance does not provide one universal numeric ceiling for every API account.
Open weights versus hosted API deployment
The official model card provides local-serving recipes for Transformers, vLLM, SGLang, TokenSpeed, and KTransformers, but those recipes do not establish practical consumer-GPU deployment.
Use the hosted API for intermittent traffic or when high-memory inference hardware is unavailable. Consider distributed or local serving when utilization is sustained, or when data-handling control justifies the operational cost. A reduced-precision text build may not reproduce the official multimodal setup.
When self-hosting starts to make sense
A GetDeploying estimate puts a 4-bit deployment around 192 GB of GPU memory, about 384 GB for 8-bit, and 768 GB for BF16. It estimates a $2.90-per-hour MI300X configuration at roughly $2,088 per month when run continuously, with API/self-hosting parity near 14 million tokens per hour using a 5:1 input-to-output mix.
These are third-party estimates, not Z.AI hardware guarantees. The 192 GB 4-bit configuration is labeled tight; KV cache, runtime overhead, multimodal processing, batching, and idle time can move the break-even point. Compare infrastructure cost per hour, utilization, effective tokens per hour, cache-hit rate, and cost per completed task before choosing local inference.
GLM-5.3-Flash versus GLM-5.3: the decision boundary
Flash and standard GLM-5.3 have different rate cards and deployment positions.
| Question | GLM-5.3-Flash | GLM-5.3 |
|---|---|---|
| List input price | $0.15 / 1M | $1.40 / 1M |
| List cached input | $0.03 / 1M | $0.26 / 1M |
| List output price | $0.50 / 1M | $4.40 / 1M |
| Temporary promotion | 50% off through September 9, 2026 UTC+8 | No matching Flash promotion shown |
| Open-weight checkpoint | Official MIT checkpoint listed | Separate model listing; do not infer identical weights |
| Multimodal positioning | Native multimodal; image/video input described by Z.AI | Treat standard GLM-5.3 coverage separately |
| Local deployment | Large checkpoint and high-memory serving required | Use its own model and deployment documentation |
Choose Flash first when token cost, multimodal input, or an open-weight route matters and the application can tolerate forced reasoning and provider constraints. Choose standard GLM-5.3 only after measuring whether its higher rate buys a capability or reliability difference that matters to the application.
FAQ
What is the current GLM-5.3-Flash API price?
The official list rates are $0.15 input, $0.03 cached input, and $0.50 output per 1M tokens; the rate table above shows the temporary 50% prices.
Does cached input use the $0.03 rate?
Yes, when Z.AI recognizes the tokens as cached. Cached-input storage is listed separately as limited-time free, so storage status and cache-hit billing should not be combined.
Is GLM-5.3-Flash open source?
The precise description is an MIT-licensed open-weight checkpoint. The license permits broad use, but the 320B-total-parameter model still needs substantial memory, compatible software, and serving capacity.
Can I run GLM-5.3-Flash on a normal laptop?
Not as a practical full-precision deployment. Even third-party 4-bit estimates are around 192 GB of GPU memory, and 18B active parameters does not remove the storage and runtime cost of the full checkpoint.
Is GLM-5.3-Flash faster than GLM-5.3?
“Flash” is supported as a cost and active-compute position, not a universal latency guarantee. Artificial Analysis reports 48.7 output tokens per second through Z.AI and 42.57 seconds to the first answer token in its measured setup; reasoning time and workload shape can dominate perceived response time.
Does the Coding Plan provide the same API balance?
No. Z.AI’s Coding Plan quick start documents a separate Coding Plan key, separate endpoints, and access limited to officially supported tools and products. A Coding Plan subscription should not be converted directly into a dollar-per-token public API budget.
The API is the lower-risk route when traffic is intermittent and hardware is not already paid for; local serving is a different economic choice that depends on sustained utilization, memory, and data-control requirements.