If you are checking whether GLM-5.3 will run on a desktop GPU, the answer is no: the live flagship checkpoint needs server-class memory. GLM-5.3-Flash lowers the barrier, but it is still a large multi-GPU model; loading a quantized checkpoint is not the same as serving a responsive coding agent.
The hardware answer in one table
The full GLM-5.3 model has a documented 8-GPU FP8 topology, while the complete 1-million-token context target is documented for 8× B200. A 24 GB, 64 GB, 128 GB, or 192 GB machine is not a practical full-model target, even though sparse MoE routing activates only part of the parameters for each token.
| Target | Published evidence | Evidence type | Practical decision |
|---|---|---|---|
| GLM-5.3 native FP8 | 8× H200 or H20 in the official vLLM recipe | Official topology | Server-class or specialist workstation deployment. |
| GLM-5.3 BF16 | Separate BF16 checkpoint; multi-node serving in the vLLM recipe | Official deployment note | Evaluation or high-end production only. |
| GLM-5.3 NVFP4 | The official vLLM recipe lists Inferact/GLM-5.3-NVFP4, a roughly 465 GB Blackwell checkpoint | Community checkpoint listed by official recipe | Blackwell-specific experiment or service. |
| GLM-5.3 quantized | A community 2-bit GGUF report used a roughly 281 GB file | Community packaging, not Z.ai sizing | Possible offload experiment, not a normal desktop install. |
| GLM-5.3-Flash | A validated profile uses 2× RTX PRO 6000 Blackwell 96GB with a 175.6 GB 4-bpw checkpoint | Community validation | The realistic local GLM option, still multi-GPU. |
The live Hugging Face model card reports 753,329,940,480 total parameters, 751,226,191,872 FP8 parameters, and a total safetensors file size of 755,643,409,571 bytes. The vLLM recipe rounds the model to about 743B total and 39B active parameters. Those rounded figures differ slightly, but neither changes the hardware conclusion.
Raw weight-memory math
These figures are arithmetic estimates before KV cache, activations, runtime buffers, and allocator overhead. FP8 and BF16 use the live flagship’s roughly 753B-parameter scale; NVFP4 and 2-bit values come from specific implementations.
| Representation | Approximate raw weight memory | What it means |
|---|---|---|
| FP8 | ~753 GB | Matches the native-FP8, 8-GPU class deployment plan. |
| BF16 | ~1.5 TB | Requires multi-node-class memory before serving overhead. |
| NVFP4 | ~465 GB for the listed community checkpoint | Blackwell-only path in the official recipe; not the default Z.ai checkpoint. |
| 2-bit | Hundreds of GB for community builds | Offload or large-memory experimentation, not a 24 GB deployment. |
What changed from the pre-release hardware advice
The GLM-5.3 repository is now live with native FP8 files. Use the live model card for checkpoint metadata, the official vLLM recipe for topology and flags, and community reports for measured deployment experience; the current recipe advertises a 1,048,576-token window.
Choose by memory budget, not by parameter count
For GLM-5.3, weight memory is the first constraint and context memory is the second; active parameters reduce compute but do not remove the need to store routed experts and runtime buffers.
24–64 GB: do not plan a full local GLM-5.3 run
A single RTX 4090, RTX 5090, or other 24 GB-class card cannot hold the flagship’s native FP8 checkpoint, which has about 756 GB of safetensors files in the live Hugging Face repository. A 64 GB workstation card is still far below the weight footprint.
CPU offload can make a quantized experiment load, but it is a debugging path rather than a default for an interactive coding service. The agent still has to finish repeated tool calls at a tolerable speed.
128–192 GB: flagship no; Flash only in a specific configuration
A 128 GB or 192 GB unified-memory machine remains below the flagship’s native FP8 requirement. Flash has a concrete path: a public validated profile runs a pinned EXL3/TR3 4-bpw checkpoint across 2× RTX PRO 6000 Blackwell 96GB GPUs.
That profile uses 175.6 GB of checkpoint data, approximately 220 GB of free storage, PCIe peer-to-peer communication, and a pinned runtime. It reports a 262,144-token request ceiling, but it is a discrete two-GPU deployment, not 192 GB of ordinary system RAM. The profile also reports 171.7 tokens per second decode and 0.059 seconds median time to first token for that exact Flash setup.
2–4 high-memory GPUs: consider Flash, not the flagship
Two or four high-memory cards are the first range worth investigating for Flash because precision, runtime, context length, batching, and image inputs all change the memory budget. The validated two-GPU profile supports text, structured tools, and semantic image input; video is disabled, the endpoint has no built-in authentication, and the shipped vision template needs a reversible repair before multimodal verification. Those details apply to that pinned recipe, not to every Flash build or to the flagship.
8× H200 or H20: documented FP8 flagship topology
The official vLLM recipe names eight H200 or H20 GPUs as the standard native-FP8 topology. It uses eight-way tensor parallelism, FP8 KV cache, five-token multi-token prediction, automatic tool choice, and GLM-specific reasoning and tool-call parsers.
This is the documented topology, not a promise of a particular tokens-per-second result. Actual throughput depends on interconnect, context length, batch size, concurrent sequences, and the serving build; the vLLM page supplies configuration but no measured production throughput.
8× B200: use this when the 1M context target matters
The official recipe positions eight B200 GPUs for the full 1,048,576-token context configuration. That extra context is a VRAM decision: KV cache grows with active sequences and context, so a deployment that works at 32K or 128K tokens may not support one million tokens with the same concurrency.
Start with a smaller --max-model-len value and raise it only after measuring KV-cache usage. A large advertised context window is valuable for long repositories and documents, but it does not make every request economical or low-latency.
Host requirements the official recipe does not specify
The vLLM recipe gives GPU topology and launch flags, but it does not publish a universal system-RAM, power, cooling, storage-headroom, or network requirement. Those values vary by checkpoint, runtime, context target, and provider platform.
| Host item | What the collected evidence supports |
|---|---|
| Model storage | The live flagship repository reports 755.6 billion bytes of safetensors files; allocate additional space for caches and temporary shards. |
| GPU interconnect | The recipe requires eight-way tensor parallelism; verify the rental or server platform’s topology rather than assuming PCIe-only performance. |
| System RAM | No universal official number is published in the vLLM recipe. Do not substitute a system-RAM figure for required GPU memory. |
| Power and cooling | No universal official figure is published. Use the platform’s eight-GPU electrical and thermal specifications before buying hardware. |
| Software | vLLM 0.28.0 or newer and Transformers 5.15.0 or newer are shown in the official recipe; DeepGEMM is required for FP8 performance. |
Minimal official serving path
The documented deployment path uses vLLM 0.28.0 and an OpenAI-compatible endpoint. It is designed for a multi-GPU node; copying the command onto a smaller machine does not remove the model-memory requirement.
Install the documented runtime
uv venv
source .venv/bin/activate
uv pip install "vllm==0.28.0" --torch-backend=auto
uv pip install "transformers>=5.15.0"
The vLLM recipe also notes that DeepGEMM is required for FP8 performance. Check the current recipe against the target GPU image before provisioning a paid node.
Start native FP8 GLM-5.3
vllm serve zai-org/GLM-5.3 \
--kv-cache-dtype fp8 \
--tensor-parallel-size 8 \
--speculative-config.method mtp \
--speculative-config.num_speculative_tokens 5 \
--tool-call-parser glm47 \
--reasoning-parser glm45 \
--enable-auto-tool-choice \
--served-model-name glm-5.3
The flags have specific jobs:
--tensor-parallel-size 8spreads the checkpoint across eight GPUs.--kv-cache-dtype fp8reduces cache pressure compared with a higher-precision cache.- The five-token MTP setting enables speculative decoding from the documented recipe.
--tool-call-parser glm47and--reasoning-parser glm45format model output for tool use and reasoning.--enable-auto-tool-choicelets the server select tools when the client provides them.
Validate the endpoint before connecting an agent
curl -X POST "http://localhost:8000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "glm-5.3",
"messages": [
{"role": "user", "content": "Write a Python function that reverses a linked list."}
],
"max_tokens": 256
}'
A successful response verifies loading, not tool use or long-context behavior. The official model card documents SGLang as another OpenAI-compatible path, but this guide uses vLLM because its GPU topology and flags are documented in a dedicated recipe.
The hidden memory budget: KV cache and always-on reasoning
The GLM-5.3 model card and vLLM recipe treat thinking as always enabled. The supported reasoning-effort values are low, high, and max; max is the default when no supported lower setting is supplied.
Longer reasoning consumes more output tokens, while a coding agent can keep large repository prefixes in the KV cache across repeated tool calls. More concurrent sequences multiply the cache requirement even when the model weights remain unchanged.
Use the settings as a deployment control:
- Low: start here for interactive coding, short requests, and latency-sensitive tools.
- High: use when a task needs more planning but still has an interactive response budget.
- Max: reserve for difficult, long-horizon tasks where extra reasoning tokens are justified.
The official recipe recommends --max-num-seqs 32 for the full-context B200 configuration and uses FP8 cache settings. Treat that number as a starting point: reduce concurrency when the server runs out of memory, and do not claim one-million-token support until a real request reaches that range without cache truncation.
When the API is the rational hardware choice
Self-hosting reserves GPU capacity even when no developer is sending requests. For intermittent traffic, low concurrency, or a team still validating GLM-5.3, the API avoids buying or continuously renting an eight-GPU node; the trade-off is hosted data handling and provider dependence.
Z.ai’s current pricing page lists GLM-5.3 at $1.40 per 1M input tokens, $0.26 per 1M cached input tokens, and $4.40 per 1M output tokens. It lists GLM-5.3-Flash at $0.15 / $0.03 / $0.50 at list price, with a 50% promotion shown through September 9, 2026; verify the live billing page before budgeting.
| Workload | GLM-5.3 list-price calculation | GLM-5.3-Flash list-price calculation |
|---|---|---|
| 10M new input + 2M output | $14.00 + $8.80 = $22.80 | $1.50 + $1.00 = $2.50 |
| 2M new input + 8M cached input + 2M output | $2.80 + $2.08 + $8.80 = $13.68 | $0.30 + $0.24 + $1.00 = $1.54 |
These are token-cost examples, not a self-hosting break-even calculation. A valid local comparison needs the node’s hourly price, sustained tokens per second, utilization, electricity, storage, engineering time, and failed or retried agent runs. For a rate-by-rate explanation, see AIReiter’s GLM-5.3-Flash API pricing guide.
The decision is:
- Choose the hosted GLM-5.3 API when you need the flagship but have irregular or moderate demand.
- Choose GLM-5.3-Flash when lower token cost, multimodal input, or a smaller self-hosting target matters more than flagship capacity.
- Choose self-hosted flagship GLM-5.3 when privacy, control, or sustained utilization justifies an eight-GPU deployment.
GLM-5.3 hardware requirements FAQ
Can GLM-5.3 run on a single RTX 4090, RTX 5090, or 24 GB GPU?
No, not as the full model. The live native-FP8 repository contains roughly 756 GB of safetensors files, so a 24 GB card can only participate in an extreme offload or quantized experiment rather than hold the model for normal serving.
Is 128 GB or 192 GB of RAM enough?
It is not enough for the flagship’s native FP8 deployment. The validated Flash profile uses two discrete 96 GB GPUs, a pinned 4-bpw checkpoint, and extra storage; that is not equivalent to a 192 GB laptop or unified-memory workstation.
What is the difference between GLM-5.3 and GLM-5.3-Flash?
They are separate models: the flagship model card reports roughly 753B total parameters, while Z.ai describes Flash as 320B total and 18B active with native multimodal positioning. Flash is smaller and cheaper, but it is still a server-class model rather than an 18B desktop model.
Which precision should I choose?
Use native FP8 when the documented eight-GPU topology is available. Use BF16 for reference-quality or specialized evaluation when multi-node memory is acceptable; consider NVFP4 only on supported Blackwell hardware and remember that the listed NVFP4 checkpoint is a community re-quantization, not the original default package.
Can I connect GLM-5.3 to a coding agent?
Yes. The official vLLM recipe enables an OpenAI-compatible endpoint plus tool-call and reasoning parsers, so clients that support that interface can connect after you validate the server. Test the exact harness, tool schema, and long-session behavior rather than assuming a successful chat completion proves agent compatibility.
What to do next
For a desktop or sub-192 GB host, test the hosted API first. Rent the documented eight-GPU topology for a representative workload, or evaluate Flash on a high-memory multi-GPU machine; self-host only after measured utilization shows that control and privacy justify the infrastructure cost.