GLM-5.3-Flash: Multimodal GLM-5 at Flash Price
Z.ai GLM-5.3-Flash is the first natively multimodal GLM-5. 320B/18B, 1M context, MIT weights. Self-reported DeepSWE 63.4 vs GLM-5.2 46.2. List $0.15/$0.50 per 1M. Not text-only GLM-5.3.

Key Numbers
GLM-5.3-Flash · Aug 26, 2026
Price Rail
per 1M tokens
Z.ai did not ship a cheaper 5.3. GLM-5.3-Flash is a different model: 320B total / 18B active, native vision and video, a 1M window, and MIT weights. The text flagship stays glm-5.3 at $1.40 / $4.40. Flash list is $0.15 / $0.50, with a 50% promo that ends September 9. They ran it as ox-alpha on OpenCode and OpenRouter first.
The vs-5.2 table is the reason this is a post. Self-reported DeepSWE +17.2 (63.4 vs 46.2) and AutomationBench +22.6 (48.8 vs 26.2). Versus Opus 4.8 they approach on several agent rows and lose NL2Repo (56.3 vs 69.7). Versus DeepSeek-V4-Flash-Vision-Exp, DeepSWE 63.4 vs 59.3 and Chartography 78.0 vs 64.3. All vendor numbers.
At a Glance
- Catalog / API:
glm-5.3-flash(OpenRouterz-ai/glm-5.3-flash, Vercelzai/glm-5.3-flash) - Org: Z.ai (zai-org)
- Release: August 26, 2026
- Architecture: 320B / 18B MoE, 45 layers, hybrid linear + sparse attention (IndexPool, mHC)
- Context: 1M
- Modalities: text + image + video
- Weights: MIT on zai-org/GLM-5.3-Flash
- List: $0.15 / $0.03 cached / $0.50 per 1M; promo 50% through 2026-09-09 16:00 UTC
- Thinking: always on (
thinking.typeenabled only) - Stealth id: was
ox-alpha - Not
glm-5.3(753B text at $1.40 / $4.40)
What's New
First native multimodal GLM-5. Vision sits in the coding loop (self-visual judgment), not a bolted VL head. Text, image, and video in one stack.
Architecture: 18B active versus the 4.5-series 32B, 45 layers versus 92, hybrid linear + sparse attention with IndexPool and mHC, and a 30T multimodal pretrain. Compared with text GLM-5.3, Z.ai cites ~3.0× less attention compute and ~4.4× smaller KV cache.
Flash list versus 5.3 list is roughly 10× cheaper on input ($0.15 vs $1.40). The 50% promo is not the steady state.
Benchmarks
All scores below are self-reported on the Z.ai Aug 26 launch post. Not LLM Stats verified.
GLM-5.3-Flash vs GLM-5.2
DeepSWE and AutomationBench are the movement: +17.2 and +22.6 versus GLM-5.2. Terminal Bench +3.3 (84.3 vs 81.0) is small. HLE w/ Tools 55.3 vs 54.7 is a wash. Versus Opus 4.8 they approach on several agent rows and lose NL2Repo (56.3 vs 69.7). GDPval-AA v2 lands at 1773 Elo (5.2: 1504); one Elo number, not a percent bar.
Vision · Self-reported
OfficeQA-Pro
vs Vision-Exp 57.9 · Opus 48.9
CharXiv Reasoning
w/ tools · near Opus 89.9
Chartography
w/ tools · vs Vision-Exp 64.3
MVBench
video · vs Vision-Exp 69.4
GLM-5.2 has no vision column, so the vision story is versus Vision-Exp, Opus 4.8, and Gemini 3.7 Flash. Chartography 78.0 vs Vision-Exp 64.3 is a clear lead. BabyVision 53.4 trails Gemini 3.7 Flash 70.9; that is the miss.
In-house Z.ai Code Bench v1.0 at max effort: Flash 29.0 vs Opus 4.8 29.5. Labeled in-house, not a hero number.
Harness footnotes from the blog: DeepSWE on mini-swe-agent (temperature 0.95, top_p 1.0, 6h timeout, 400K context); Terminal Bench 2.1 on Claude Code 2.1.207; AutomationBench v1.0.6; Toolathlon Verified via the official service (pass@1, 3 runs). Read the full footnote block on Z.ai before treating any row as apples to apples.
Base-model table (MMLU and friends) is optional color: Flash-Base sits competitive with prior GLM bases at 18B active.
Pricing
From the official pricing docs, per 1M tokens:
| Detail | List | Promo 50% |
|---|---|---|
| Input | $0.15 | $0.075 |
| Cached input | $0.03 | $0.015 |
| Output | $0.50 | $0.25 |
Write the list as the story. Promo ends 2026-09-09 16:00 UTC (Z.ai writes 24:00 September 9 UTC+8). Put a September 9 price check on the calendar.
Coding Plan: Flash gives 3× the usable quota versus GLM-5.3. Off-peak hours and all-day weekends consume 50% of standard points on the Coding Plan. Do not confuse Flash list with text GLM-5.3 at $1.40 / $0.26 / $4.40.
Serving
API id glm-5.3-flash. thinking.type supports enabled only; thinking cannot be turned off. Recommended: temperature 1, top_p 0.95, reasoning_effort: max, thinking.clear_thinking: false. Images via image_url content blocks. See the VLM guide.
Local weights support SGLang, vLLM, and TokenSpeed. The Chinese-chip serving story (SGLang stack, EPD disaggregation, ox-alpha traffic on domestic accelerators) is infrastructure color, not the thesis.
When to Use It
Good fit
- Multimodal coding / CUA loops at Flash list
- ox-alpha users graduating to a named id
- Shops that want MIT weights at 18B active and a 1M window
Prefer glm-5.3
When you want the 753B text stack and will pay $1.40 / $4.40.
Prefer Vision-Exp
If you are already on DeepSeek and only need image-in at Flash price.
Put a September 9 price check on the calendar. After the promo, list is the rate that matters.
Outlook
This post ages when the promo ends, when LLM Stats verifies DeepSWE and AutomationBench, or when a larger GLM-5.3 multimodal lands. The through-line holds either way: multimodal GLM-5 at Flash dollars.
Sources: Z.ai blog, VLM docs, pricing, Hugging Face, GitHub.
Questions
Frequently Asked Questions
- August 26, 2026. Z.ai stealth-tested it first as
ox-alphaon OpenCode and OpenRouter. - List is $0.15 / $0.03 cached / $0.50 per 1M. A 50% promo runs at $0.075 / $0.015 / $0.25 until 2026-09-09 16:00 UTC. Budget the list.
- 1M tokens.
- Different model. GLM-5.3 is text-only, 753B, list $1.40 / $4.40. Flash is 320B / 18B active, natively multimodal (text, image, video), list $0.15 / $0.50.
- Self-reported: DeepSWE 63.4 vs 46.2, AutomationBench 48.8 vs 26.2. Not LLM Stats verified.
- Yes. MIT license on zai-org/GLM-5.3-Flash.
Continue Reading
