AIREITER

GLM-5.3-Flash API Pricing: Rates, Cache, Hosting

Last Updated: 2026-08-26 20:06:25

A $0.15 input price is only the starting point for GLM-5.3-Flash. Cache hits, output tokens, forced reasoning, and a roughly 320B-parameter checkpoint determine whether the API or open-weight route fits a real deployment.

The price that is actually live

Z.AI’s official pricing page lists GLM-5.3-Flash at $0.15 per 1 million input tokens, $0.03 per 1 million cached-input tokens, and $0.50 per 1 million output tokens. It also shows a temporary 50% promotion, ending at 24:00 on September 9, 2026, in UTC+8 (Singapore time).

GLM-5.3-Flash chargeList price / 1M tokensPromotion / 1M tokens
New input$0.15$0.075
Cached input$0.03$0.015
Output$0.50$0.25
Cached-input storageLimited-time freeLimited-time free
GLM-5.3-Flash API price comparison for input, cached input, and output tokens

Treat the 50% reduction as temporary promotional pricing, not a production-budget baseline; Z.AI’s launch announcement identifies ox-alpha as the earlier preview identity.

The same official table lists standard GLM-5.3 at $1.40 input, $0.26 cached input, and $4.40 output per 1 million tokens. On list prices, Flash is about 89% cheaper for both input and output. That is a price comparison, not a claim that the two models have identical quality, latency, or operational behavior.

What you are buying: API access, not a smaller local model

GLM-5.3-Flash is an MIT-licensed open-weight multimodal model with 320B total parameters and 18B active parameters per token. The official Hugging Face model card documents the public zai-org/GLM-5.3-Flash checkpoint and local-serving paths; Z.AI describes a 1-million-token context window and text, image, and video input in its release material.

“18B active” refers to selected expert computation, not total memory. Full weights, KV cache, runtime overhead, multimodal processing, and context length still determine local feasibility.

Access routeWhat it providesMain constraint
Z.AI APIHosted inference billed by input, cached input, and output tokensAccount limits and reasoning latency
Official checkpointMIT-licensed weights for self-hosting or adaptationHigh-memory infrastructure, depending on precision and quantization
Quantized community buildSmaller files for some local runtimesBuild quality, modality support, speed, and compatibility vary

The low token price and low hosting cost are different claims: bursty users pay as API traffic arrives, while self-hosting reserves serving capacity even during idle periods.

API billing in real workloads

GLM-5.3-Flash cost depends on the mix of new input, recognized cached input, and generated output:

cost = (new_input_tokens / 1,000,000 × input_rate)
     + (cached_input_tokens / 1,000,000 × cached_rate)
     + (output_tokens / 1,000,000 × output_rate)

At list price, 10 million new input tokens plus 2 million output tokens costs $2.50: $1.50 for input and $1.00 for output. During the promotion, the same volume costs $1.25.

WorkloadList-price calculationList-price totalPromotional total
10M new input + 2M output$1.50 + $1.00$2.50$1.25
2M new input + 8M cached input + 2M output$0.30 + $0.24 + $1.00$1.54$0.77
100K new input + 10K output$0.015 + $0.005$0.020$0.010

The $0.10-per-million blended figure shown by Artificial Analysis uses a defined 7:2:1 cache-hit, input, and output mix. It is useful for a comparable workload, not a universal GLM-5.3-Flash rate.

Cache savings require an actual cache hit. Z.AI’s Chat Completion response schema exposes usage.prompt_tokens_details.cached_tokens, so production accounting should record that field instead of assuming repeated-looking prompts will be discounted. One user in r/opencodeCLI put the launch economics this way:

“There's a 50% discount offer until September 9th, and its context consumption is much better than in dsv4f...” — u/CriteriumA, Reddit discussion

The discount date is confirmed by Z.AI’s rate card; the comparison and user experience are that commenter’s opinion. Stable repeated context can improve cache reuse, but it does not remove output, retry, tool, or account-limit costs.

A compact budget worksheet

Forecast four separate rows: new input, cached input, output, and paid tools. Z.AI lists Web Search at $0.01 per use, so repeated agent searches can exceed a token-only estimate.

Monthly usage assumptionFormula at list priceMonthly cost
100M new input + 20M output100 × $0.15 + 20 × $0.50$25.00
20M new input + 80M cached input + 20M output20 × $0.15 + 80 × $0.03 + 20 × $0.50$23.40
100 Web Search calls100 × $0.01$1.00

These are arithmetic examples, not a quota. Add retries and abandoned agent runs separately. If output volume is high, setting a realistic max_tokens value and using a lower reasoning_effort where suitable can matter more than optimizing a small prompt prefix.

Z.AI API constraints before migration

The official Z.AI API uses https://api.z.ai/api/paas/v4/ as its general base URL and https://api.z.ai/api/paas/v4/chat/completions for chat completions. Authentication uses a Bearer API key. Z.AI documents OpenAI-compatible Python, Node.js, and Java client patterns through a custom base URL in its API introduction.

The current Chat Completion reference mentions GLM-5.3-FLASH in its thinking and reasoning descriptions, but its rendered model enum does not display glm-5.3-flash; the official pricing page and launch materials do list the Flash SKU. Test the exact hosted model ID with a small request before switching production traffic.

Integration detailVerified constraint
Hosted model ID to testglm-5.3-flash; confirm acceptance against the live account schema
Hosted endpointhttps://api.z.ai/api/paas/v4/chat/completions
AuthenticationAuthorization: Bearer YOUR_API_KEY
ThinkingGLM-5.3-FLASH uses enabled thinking; depth is controlled with reasoning_effort
Reasoning effortlow, high, or max
Maximum max_tokens131,072 in the current parameter reference
Coding PlanSeparate key, endpoints, and credit system; it is not a general API balance. See Z.AI’s Coding Plan quick start

Z.AI sets max as the default reasoning effort. The token rate does not change by effort level, but reasoning can change output usage and waiting time. For account limits, Z.AI directs users to an account-specific rate-limit dashboard; its public guidance does not provide one universal numeric ceiling for every API account.

Open weights versus hosted API deployment

The official model card provides local-serving recipes for Transformers, vLLM, SGLang, TokenSpeed, and KTransformers, but those recipes do not establish practical consumer-GPU deployment.

Use the hosted API for intermittent traffic or when high-memory inference hardware is unavailable. Consider distributed or local serving when utilization is sustained, or when data-handling control justifies the operational cost. A reduced-precision text build may not reproduce the official multimodal setup.

When self-hosting starts to make sense

A GetDeploying estimate puts a 4-bit deployment around 192 GB of GPU memory, about 384 GB for 8-bit, and 768 GB for BF16. It estimates a $2.90-per-hour MI300X configuration at roughly $2,088 per month when run continuously, with API/self-hosting parity near 14 million tokens per hour using a 5:1 input-to-output mix.

These are third-party estimates, not Z.AI hardware guarantees. The 192 GB 4-bit configuration is labeled tight; KV cache, runtime overhead, multimodal processing, batching, and idle time can move the break-even point. Compare infrastructure cost per hour, utilization, effective tokens per hour, cache-hit rate, and cost per completed task before choosing local inference.

GLM-5.3-Flash versus GLM-5.3: the decision boundary

Flash and standard GLM-5.3 have different rate cards and deployment positions.

QuestionGLM-5.3-FlashGLM-5.3
List input price$0.15 / 1M$1.40 / 1M
List cached input$0.03 / 1M$0.26 / 1M
List output price$0.50 / 1M$4.40 / 1M
Temporary promotion50% off through September 9, 2026 UTC+8No matching Flash promotion shown
Open-weight checkpointOfficial MIT checkpoint listedSeparate model listing; do not infer identical weights
Multimodal positioningNative multimodal; image/video input described by Z.AITreat standard GLM-5.3 coverage separately
Local deploymentLarge checkpoint and high-memory serving requiredUse its own model and deployment documentation

Choose Flash first when token cost, multimodal input, or an open-weight route matters and the application can tolerate forced reasoning and provider constraints. Choose standard GLM-5.3 only after measuring whether its higher rate buys a capability or reliability difference that matters to the application.

FAQ

What is the current GLM-5.3-Flash API price?

The official list rates are $0.15 input, $0.03 cached input, and $0.50 output per 1M tokens; the rate table above shows the temporary 50% prices.

Does cached input use the $0.03 rate?

Yes, when Z.AI recognizes the tokens as cached. Cached-input storage is listed separately as limited-time free, so storage status and cache-hit billing should not be combined.

Is GLM-5.3-Flash open source?

The precise description is an MIT-licensed open-weight checkpoint. The license permits broad use, but the 320B-total-parameter model still needs substantial memory, compatible software, and serving capacity.

Can I run GLM-5.3-Flash on a normal laptop?

Not as a practical full-precision deployment. Even third-party 4-bit estimates are around 192 GB of GPU memory, and 18B active parameters does not remove the storage and runtime cost of the full checkpoint.

Is GLM-5.3-Flash faster than GLM-5.3?

“Flash” is supported as a cost and active-compute position, not a universal latency guarantee. Artificial Analysis reports 48.7 output tokens per second through Z.AI and 42.57 seconds to the first answer token in its measured setup; reasoning time and workload shape can dominate perceived response time.

Does the Coding Plan provide the same API balance?

No. Z.AI’s Coding Plan quick start documents a separate Coding Plan key, separate endpoints, and access limited to officially supported tools and products. A Coding Plan subscription should not be converted directly into a dollar-per-token public API budget.

The API is the lower-risk route when traffic is intermittent and hardware is not already paid for; local serving is a different economic choice that depends on sustained utilization, memory, and data-control requirements.