the AI bench
VERIFIED AUGUST 2026
All fast takes

FAST TAKE · 2026-08-26 · Z.AI GLM-5.3 + GLM-5.3-FLASH

GLM-5.3 and GLM-5.3-Flash: same flagship price, and a genuinely cheap MIT multimodal under it

Two Z.AI drops we are late on, recorded together. GLM-5.3 replaced GLM-5.2 as the hosted flagship on August 14 at an identical $1.40 in / $4.40 out per 1M — a rename in our cost calculator, not a repricing. The more interesting release is August 26's GLM-5.3-Flash: the first natively multimodal model in the GLM-5 series, 320B total with 18B active, MIT-licensed with weights on HuggingFace, priced at roughly one-tenth of the flagship.

Verdict: Z.AI's flagship steps to 5.3 at the same price, and a 320B/18B-active MIT sibling ships at one-tenth the cost


The take

GLM-5.3 first, briefly. Z.AI's blog leads on stronger agentic coding at better token efficiency than GLM-5.2 — vendor benchmarks, unreproduced. The API price is unchanged (verified on docs.z.ai), the open weights were put on a HuggingFace countdown ending August 28, and the GLM Coding Plan quietly moved to a points-based quota with this release while still publishing no dollar figure we can cite. Our calculator row renames 5.2 → 5.3 and its price does not move.

GLM-5.3-Flash is the real story. The published card: $0.15 in / $0.50 out per 1M list, on a 50% promo ($0.075 / $0.25) through September 9 — we quote the list price, for the same reason we refused DeepSeek's clock-dependent rates a single number: a 36-month projection built on a two-week promo is the misleading option. Architecture: 320B total / 18B active, hybrid sparse + linear attention that Z.AI says cuts attention compute 3.0× and KV cache 4.4× versus GLM-5.3, 45 layers versus GLM-4.5's 92. License is a clean MIT on the card (`zai-org/GLM-5.3-Flash`), weights in BF16 and FP8, vLLM/SGLang supported at launch.

Why it matters: the cheap tier is now where the open-weight frontier competes hardest. GLM-5.3-Flash lands in the same $0.15-in band as Qwen3.8-Flash (released the same day) and undercuts everything American at that capability claim — Z.AI positions it as approaching Opus 4.8 on coding and agentic benchmarks, which we pass along as a vendor claim and nothing more. It is also, notably, being served at scale on Chinese AI chips, which is its own story about where inference economics are heading.

Our call: no local planner-pick change. At 320B total, even FP8 is ~320 GB — multi-GPU / big-iron territory regardless of the friendly license and the tiny active count; this sits in the hosted-frontier-comparator bucket with DeepSeek V4-Flash and Kimi K3. On the cloud side, the calculator note now carries GLM-5.3-Flash's list price alongside the flagship row. What we watch: whether the MIT weights + FP8 release makes this the default self-hosting choice for orgs with real hardware, and whether GLM-5.3's own weights land on schedule today.

Where this fits

Models: GLM-5.1 · DeepSeek V4-Flash · Kimi K3 (2.8T-A50B)

Hardware: NVIDIA RTX 5090 · NVIDIA RTX 4090

Sources

Next step

Try this in the planner