Before the reveal, an anonymous model circulated as Ox Alpha. Z.ai has now attached a public name to that experiment: GLM-5.3-Flash. The reveal matters, but the “Flash” label needs a careful reading: the model activates 18B parameters per token, yet its published checkpoint is roughly 320B and still demands serious hardware locally.
The short answer: Ox Alpha was the GLM-5.3-Flash preview
Z.ai’s August 26, 2026 announcement identifies GLM-5.3-Flash as the model previously previewed anonymously as Ox Alpha on OpenCode and OpenRouter. The official Hugging Face model card now lists the public checkpoint, MIT license, multimodal inputs, and local-serving references.
Treat Ox Alpha reports as preview evidence, not shipped-model specifications. Ox Alpha had its own routing, traffic, and operational conditions, so early speed, quota, and policy reports may not describe every GLM-5.3-Flash endpoint.
“For me it has changed nothing. Both models are way too big for my 32 GB RAM system.” — u/dampflokfreund in r/LocalLLaMA
GLM-5.3-Flash at a glance
| Specification | GLM-5.3-Flash |
|---|---|
| Official release | August 26, 2026 |
| Earlier preview name | Ox Alpha / ox-alpha |
| Model shape | 320B total parameters, 18B active per token |
| Context window | 1,048,576 tokens (1M) |
| Input | Text, images, and video |
| Output | Text |
| License | MIT |
| Official checkpoint | zai-org/GLM-5.3-Flash |
| Listed API input | $0.15 per 1M tokens |
| Listed cached input | $0.03 per 1M tokens |
| Listed API output | $0.50 per 1M tokens |
The Hugging Face page displays a stored model size of about 321B parameters, while the model description uses the shorthand 320B-A18B. Those figures describe total stored capacity and active per-token computation respectively; 18B active does not turn the checkpoint into an 18B local model.
What changed from Ox Alpha to the public model
Z.ai says it ran GLM-5.3-Flash anonymously as Ox Alpha before the public release, using OpenCode and OpenRouter to collect real-world feedback. The OpenRouter listing showed the preview as a multimodal, long-context model with a temporary $0 price; the public release now makes the provider, checkpoint, and license attributable.
The preview and the shipped model remain separate operating environments. A tokenizer match or GLM-style API error can support lineage, but it cannot establish that every anonymous request used the final weights or the same routing layer. The public checkpoint is now the reference for deployment and reproducibility.
The architecture behind the “Flash” label
GLM-5.3-Flash combines a sparse mixture-of-experts design with attention changes intended to reduce the cost of very long contexts. Z.ai describes 320B total parameters, 18B active parameters, 45 layers, hybrid sparse and linear attention, IndexPool retrieval, and Manifold-Constrained Hyper-Connections (mHC). The model card also attributes training to a 30T-token multimodal pretraining corpus.
Z.ai reports 3x less attention compute and a 4.4x smaller KV cache at 1M context compared with GLM-5.3. Those are vendor architecture claims, not a hardware-neutral latency guarantee; actual speed still depends on context length, concurrency, visual preprocessing, and the serving stack.
The model’s native multimodality is more useful for agents than for a generic “vision chatbot” label. Z.ai’s intended loop is to let an agent inspect rendered interfaces, browser states, documents, or other visual artifacts and then revise code or an output. The official model card demonstrates image input, while the public model and deployment documentation also describe video input.
What the benchmark table supports—and what it does not
The official release and model card report strong results, especially for coding and agent tasks. The visible model-card entries include TerminalBench 2.1 at 84.3, DeepSWE at 63.4, and HLE at 55.3. The Z.ai release material also reports an Artificial Analysis Intelligence Index score of 57, AutomationBench at 48.8, and Z.ai Code Bench maximum-effort performance of 29.0.
| Benchmark | GLM-5.3-Flash | Comparator or baseline | Source and limitation |
|---|---|---|---|
| TerminalBench 2.1 | 84.3 | GLM-5.2: 81.0; Claude Opus 4.8: 85.0 | Z.ai/model-card result; harnesses and budgets matter |
| DeepSWE v1.1 | 63.4 | GLM-5.2: 46.2 | Reported evaluation; do not generalize to every repository |
| AutomationBench | 48.8 | GLM-5.2: 26.2 | Reported evaluation with a specified benchmark version |
| Z.ai Code Bench, maximum effort | 29.0 | Claude Opus 4.8: 29.5 | Near-parity claim under Z.ai’s setup |
| Artificial Analysis Intelligence Index v4.1.1 | 57 | Z.ai says comparable to Claude Opus 4.8 | External index; task mix is not a coding-only score |
Reported results use different harnesses, context limits, timeouts, and judges; compare them directionally rather than as one universal leaderboard. The fairest conclusion is that GLM-5.3-Flash has credible evidence of a major improvement over GLM-5.2 on coding-agent evaluations, not that it dominates every frontier model.
The same caution applies to early community tests. @prz_chojecki reported that Ox Alpha scored better than GLM-5.2 on all 226 ErdosBench problems. That observation is useful for forming test cases, not a substitute for a controlled evaluation of your own tools and repository.
The practical limits early users found
“Flash” describes active compute and cost positioning more reliably than interactive latency. Real-user discussions in r/opencodeCLI reported both strong coding results and frustrating waits, particularly during the anonymous preview. These are provider- and workload-specific observations, not official latency benchmarks.
“How can it be called ‘Flash’ if it’s much slower than the main model?” — u/ahriad in r/opencodeCLI
The most useful operational signal is long-running agent reliability under load. @solomonneas reported 30 successful builds from 49 attempts in a seven-day test, with 14 failures arriving as timeouts. That does not prove a fixed failure rate for the official API, but it does justify a fallback route for tasks that can run for hours.
The model also has a local-deployment limit that benchmark headlines hide: the official FP8 checkpoint is about 306 GiB before runtime and KV-cache overhead. A 320B total model therefore needs high-memory infrastructure even when only 18B parameters are active for each token.
API, hosted access, and local deployment choices
For hosted use, Z.ai’s pricing page lists GLM-5.3-Flash at $0.15 per 1M input tokens, $0.03 per 1M cached input tokens, and $0.50 per 1M output tokens. A temporary 50% promotion lists $0.075 input, $0.015 cached input, and $0.25 output, ending at 24:00 on September 9, 2026, UTC+8. Check the official Z.AI pricing table before budgeting because the promotion has an end date.
The public list price is separate from the temporary Ox Alpha preview price. For a simple example, 10M input tokens plus 2M output tokens at list price would cost $2.50 before any other provider fees: 10 × $0.15 + 2 × $0.50. Cached input can lower the input portion when the provider bills those tokens as cached.
Local serving is possible, but it is an infrastructure project rather than a laptop download. The official vLLM recipe lists approximately 306 GiB for FP8 weights and 386GB of VRAM for the default FP8 deployment; BF16 is listed at 772GB. It currently calls for NVIDIA Hopper or newer GPUs in the supported path, a recent FlashInfer version, and a dedicated vLLM image while integration matures.
The KTransformers tutorial documents CPU-GPU heterogeneous inference, direct use of the official FP8 weights, and at least 350GB of free system memory. Its single-GPU example is an offload configuration, not evidence that one consumer GPU can hold the model. The tutorial also uses a validated 501,025-token setting and limits each multimodal request to either eight images or one video.
| Deployment route | What the documentation supports | Main constraint |
|---|---|---|
| Z.ai API | Metered hosted access; list and promotional token rates published | Verify rate limits, regional terms, and current promotion |
| vLLM | FP8/BF16 serving, OpenAI-compatible endpoint, tensor parallelism, multimodal path | High-memory NVIDIA infrastructure; current recipe targets Hopper+ |
| KTransformers | CPU-GPU heterogeneous inference and FP8 loading | About 306 GiB of weights; at least 350GB system memory recommended |
| Ollama/LM Studio/llama.cpp ecosystem | Model card points to quantized distributions and compatible tools | Quantized support and speed depend on the specific build |
Who should use GLM-5.3-Flash now?
Use it for cost-sensitive coding agents and long-context engineering pilots. The low list price, 1M context, and reported improvement over GLM-5.2 make it a strong candidate for repository analysis, multi-file changes, test generation, and repeated agent calls. Keep acceptance tests and a fallback model in the loop until your own success rate is stable.
Use it for multimodal technical work when text-only GLM models are the bottleneck. Image and video input can help an agent inspect browser output, UI layouts, diagrams, documents, and rendered artifacts. The model is not a video generator; it reasons over visual input and returns text or code.
Do not choose it solely because “18B active” sounds like a desktop model. The total checkpoint is roughly 320B parameters, and local serving needs hundreds of gigabytes of storage or memory before context and runtime overhead. Hosted API access is usually the simpler economic choice unless you already operate high-memory inference hardware.
Do not send confidential code to an unverified preview endpoint. The public GLM-5.3-Flash release is attributable to Z.ai, but provider terms, retention, and routing still matter. For production traffic, record the endpoint’s current data policy and use the official API or a provider whose operating terms you can audit.
FAQ
Is Ox Alpha exactly the same as GLM-5.3-Flash?
Z.ai officially identifies GLM-5.3-Flash as the model previously previewed as Ox Alpha. Early Ox Alpha behavior remains useful context, but preview routing and limits should not be treated as public-model specifications.
Is GLM-5.3-Flash open source?
The official Hugging Face checkpoint is released under the MIT license, and Z.ai provides local-serving references. “Open weights” is the precise operational description: the weights are available, but running the full model still requires substantial infrastructure.
How much memory does GLM-5.3-Flash need locally?
The official FP8 weights occupy about 306 GiB before runtime and KV-cache overhead. The vLLM recipe lists 386GB of VRAM for its default FP8 deployment, while KTransformers recommends at least 350GB of system memory for heterogeneous inference.
Is the API really $0.15 per million input tokens?
Yes. Z.AI’s official pricing page lists $0.15 for input, $0.03 for cached input, and $0.50 for output; its 50% promotion is listed through September 9, 2026, UTC+8.
Is GLM-5.3-Flash faster than other Flash models?
The available evidence does not establish a universal latency ranking. Early users reported slow or timing-out agent sessions, so measure time to first token, completion time, retries, and accepted-task rate on your own provider route.
Does GLM-5.3-Flash support images and video?
Yes. Z.ai and the official model card describe text, image, and video input with text output. Local-serving documentation adds request boundaries such as up to eight images or one video in the documented KTransformers setup.
Use the hosted API when you want GLM-5.3-Flash’s cost and multimodality without buying a multi-hundred-gigabyte inference box; pilot local serving only if that infrastructure already exists. For long agent jobs, keep a tested fallback, because the public evidence is stronger on capability and price than on sustained interactive reliability.