AIREITER

GLM-5.3: Benchmarks, Pricing, and How to Access Z.ai's New Model

Last Updated: 2026-08-17 07:40:30

GLM-5.3 turns GLM-5.2's 753B-parameter base into a near-frontier coding model through post-training alone - DeepSWE climbed from 46.2 to 66.9 without a new base model. The catch three days after launch: there is still no token-priced API, and the "open-weights" model shipped with its weights held back for a two-week safety review.

What GLM-5.3 actually is

GLM-5.3 is Z.ai's latest flagship model for long-horizon software engineering and agent tasks, released on August 14, 2026. The official model documentation confirms the most unusual design fact of this launch: GLM-5.3 uses the same base model as GLM-5.2, and every claimed improvement comes from post-training on realistic engineering workflows - full task cycles from problem identification through implementation, verification, and delivery.

That makes GLM-5.3 a sibling, not a successor architecture. GLM-5.2 shipped in mid-June 2026 as a 753B-parameter Mixture-of-Experts model, MIT-licensed with weights on Hugging Face since June 17, as TechTimes reported at release.

SpecGLM-5.3GLM-5.2
Base modelSame 753B MoE (per Z.ai)753B MoE, MIT license
Context window1M tokens1M tokens
Max output128K tokens128K tokens
Input / outputText / textText / text
Tool supportFunction calling, MCP, JSON output, context caching, thinking modesSame list (GLM-5.2 docs)
WeightsNot yet published - delayed for safety reviewOn Hugging Face since June 17
Access todayGLM Coding Plan only; API "coming soon"Coding Plan + token-priced API

What the benchmarks say (and what they don't prove)

The numbers below come straight from Z.ai's launch documentation; here is the full before/after table, with GLM-5.2 as the baseline.

BenchmarkGLM-5.2GLM-5.3Change
Terminal-Bench 3.04.628.3+23.7
DeepSWE v1.146.266.9+20.7
Agents' Last Exam23.828.5+4.7
ExploitBench24.4%54.4%+30.0 pts

A 20.7-point DeepSWE gain from post-training alone is the claim that made the rounds on launch day - one widely shared Reddit analysis framed it as proof of "how much capability may still be hiding inside today's largest base models." The same launch page reports 1,769 points on GDPval-AA v2 across 44 occupational categories, positioning GLM-5.3 beyond pure coding, and claims programming capability "comparable to Claude Fable 5."

Treat the direction as real and the precision as unproven. The launch-page figures are vendor-reported, published without harness details or confidence intervals, and the community benchmark thread tracking them notes that independent reproductions are still absent. The most useful external cross-check available: a user-compiled comparison in r/ZaiGLM puts GPT-5.6 Luna Max at 67.2 on DeepSWE v1.1 versus GLM-5.3's 66.9 - a 0.3-point gap that suggests GLM-5.3 is competing at the frontier's edge on agentic code, not lapping it. To bench that pairing today, GPT-5.6 Luna is callable now, and GLM-5.2 remains the family's API-available flagship until the 5.3 API lands - not a stand-in for GLM-5.3's post-trained behavior, but the same base model.

Why the open weights are delayed

Z.ai is holding GLM-5.3's weights back for roughly two weeks while it strengthens safety controls, according to Axios reporting from launch day. The reason is the model's cybersecurity performance, which Z.ai itself describes as having "exceeded expectations" as its long-horizon task environment expanded.

On CyberGym, a benchmark for discovering known security vulnerabilities, GLM-5.3 scored 84.5% - ahead of Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6% in Z.ai's comparison table. On ExploitBench, which tests reasoning through real exploit development, GLM-5.3's 54.4% more than doubled GLM-5.2's 24.4% but still ranked behind Claude Fable 5 and GPT-5.6 Sol in Axios's reading of the results. Z.ai's own documentation is blunt about the limit: the advantage concentrates in the front end of the exploitation chain, and deeper exploitation plus complete offensive/defensive workflows "need improvement."

Two supporting facts give the delay context:

  • According to Axios, Z.ai launched a disclosure site the same day crediting its GLM models with finding more than 2,400 security flaws - over 1,000 of them critical or high severity - in targets including the Linux kernel, VMware projects, and Apache projects. Open-source maintainers can submit repositories for GLM-powered bug scans.
  • Per Axios, in lieu of public weights, Z.ai is standing up a tiered program giving selected security partners controlled access during the review window - a separate track from the Coding Plan.

Once weights are public, Z.ai loses the ability to control downstream modification - its own announcement framed the stakes: "An open world cannot have only open attack surfaces."

How to access GLM-5.3 today

For the public, GLM-5.3 is available only through a GLM Coding Plan subscription right now. The official documentation states the API is "coming soon," and GLM-5.3 has no entry in Z.ai's token pricing table as of August 17, 2026 - GLM-5.2, GLM-5.1, and GLM-5 all have listed rates, which tells you the 5.3 row simply hasn't been switched on yet.

The subscription tiers from the official subscribe page:

TierList priceDisplayed priceAllowanceTarget workload
Lite$18/mo$12.60/mo10,000 credits/weekSmall repositories, light iteration
Pro$80/mo$56/mo6× Lite usageDay-to-day development, mid-sized repos
Max$168/mo$117.60/mo14× Lite usageMid-to-large repositories, heavy agents

The displayed prices sit alongside higher reference prices on the page without a labeled billing condition - treat $18/$80/$168 as the monthly anchors and the lower figures as the current discount. Setup is a single command, npx @z_ai/coding-helper, which wires the plan into your tool of choice: Z.ai lists 20+ supported agent tools including Claude Code, OpenClaw, and its own ZCode, with MCP tool access from the Pro tier up. The Max tier promises first access to new flagships, and all Coding Plan tiers include rolling access to the current one.

The pricing change hiding behind the launch

The Coding Plan was repriced around the GLM-5.3 launch, and heavy users noticed immediately. A subscriber who tracked the numbers in r/ZaiGLM posted this subscriber-reported comparison against the legacy V2 plans - treat it as community data, not an official Z.ai table, and check your own dashboard before acting on it:

PlanLegacy priceLegacy tokens/weekNew priceNew tokens/week
Pro$65~500M$80~526M
Max$144~2B$168~1.22B

"new Max gives you ~39% fewer tokens while costing ~17% more." - u/x-primez-x, r/ZaiGLM

That works out to roughly 1.9× worse cost per token at the Max tier by the poster's math - ~7.26M tokens per dollar per week, down from ~13.89M. Pro gets hit less: about 5% more allowance for about 23% more money, which is roughly 14% fewer tokens per dollar by the same math. The same post rates the model itself highly and pairs it with a warning:

"GLM-5.3 looks awesome… but this new coding plan probably pushes me away from Z.AI for good." - u/x-primez-x, r/ZaiGLM

For context on the API side: when GLM-5.3's token pricing does appear, the family anchor is GLM-5.2 at $1.40 per 1M input tokens, $0.26 cached, and $4.40 output on the official pricing page - with cached-input storage currently free. One commenter in the same thread already argues the subscription-vs-API tradeoff favors pay-as-you-go: "The flexibility is worth it when new models come out every week or two."

Should you switch to GLM-5.3?

The honest answer splits by what you're running today:

You are…Move now?Why
Coding Plan Pro subscriberYesSame $80 tier, immediate GLM-5.3 access; the DeepSWE jump is vendor-reported, so validate on your own repo first
Coding Plan Max heavy userDo the math first~39% fewer weekly tokens than legacy Max at a higher price; throughput per dollar may have fallen below your break-even
Pay-as-you-go API userWaitNo token price exists yet; GLM-5.2 at $1.40/$4.40 per 1M remains the only GLM flagship you can meter
Self-hosterWait for weightsNothing to run until the ~two-week review ends; same 753B MoE base means multi-GPU server hardware, and vLLM/SGLang support is an expectation from that shared base, not a Z.ai confirmation
Vision-dependent userNoGLM-5.3 is text-in/text-out, same as GLM-5.2; vision still lives in the separate, closed GLM-5V line

That last row deserves emphasis because it was the community's loudest pre-launch request: Jie Tang's June 29 poll asking what GLM-5.3 should include drew 466,000 views with vision the overwhelming answer, and the shipped model still has no visual encoder. If your workflow feeds screenshots, PDFs, or UI mockups, GLM-5V-Turbo at $1.20/$4.00 per 1M tokens on Z.ai's API is the workaround, not GLM-5.3.

The unresolved trade-off to watch: subscribing now gets you GLM-5.3 immediately but locks you into repriced quotas, while waiting costs you nothing except access. The weights carry Z.ai's stated review timeline; the API carries only a "coming soon" label with no pricing or date attached. If you're not bottlenecked on agentic coding throughput today, the two-week weight window is short enough to wait out.

FAQ

When will GLM-5.3's weights be released?

Z.ai delayed the weight release by about two weeks from the August 14 launch for a safety review, which points to late August if the timeline holds. Watch the zai-org organization on Hugging Face - that's where GLM-5.2's weights landed within days of its June release.

Is GLM-5.3 open source?

GLM-5.3 is positioned as an open-weights model, but the weights are not published yet, and the launch documentation does not state a license. GLM-5.2 shipped under MIT, which is the community's working assumption for GLM-5.3 - an assumption, not a commitment from Z.ai.

Does GLM-5.3 support images or vision?

No. The official documentation lists text input and text output only, with a 1M-token context and 128K maximum output. Native vision remains in the separate GLM-5V family (GLM-5V-Turbo, GLM-4.6V), which Z.ai serves through its API rather than open weights.

What does GLM-5.3 cost per token?

There is no per-token price yet - the API is listed as "coming soon" and GLM-5.3 is absent from Z.ai's pricing table. Today the only access is the GLM Coding Plan at $18/$80/$168 per month list (currently displayed at $12.60/$56/$117.60).

Can GLM-5.3 run locally?

Not until the weights are released. GLM-5.2 shipped with model-card support for vLLM, SGLang, and KTransformers (per TechTimes); the launch documentation says nothing about GLM-5.3's deployment stack yet, so plan for multi-GPU server hardware rather than consumer rigs.