The AI arena is free today

Open Superagent
Back to blog
Model Release·Vision·Multimodal·API

GLM-5.3-Flash: Multimodal GLM-5 at Flash Price

Z.ai GLM-5.3-Flash is the first natively multimodal GLM-5. 320B/18B, 1M context, MIT weights. Self-reported DeepSWE 63.4 vs GLM-5.2 46.2. List $0.15/$0.50 per 1M. Not text-only GLM-5.3.

Sebastian Crossa
Sebastian Crossa
Co-Founder @ LLM Stats
·9 min read
GLM-5.3-Flash: Multimodal GLM-5 at Flash Price

Key Numbers

GLM-5.3-Flash · Aug 26, 2026

0.0%
DeepSWE
0.0%
Terminal Bench
0.0%
Toolathlon
0.0%
OfficeQA-Pro
0B
Active params
$0.00
List input

Price Rail

per 1M tokens

$0.15 / $0.03 / $0.50
List
in / cached / out
$0.075 / $0.015 / $0.25
Promo 50%
in / cached / out
Promo ends 2026-09-09 16:00 UTC. Budget the list. Promo is a clock, not the permanent rate.

Z.ai did not ship a cheaper 5.3. GLM-5.3-Flash is a different model: 320B total / 18B active, native vision and video, a 1M window, and MIT weights. The text flagship stays glm-5.3 at $1.40 / $4.40. Flash list is $0.15 / $0.50, with a 50% promo that ends September 9. They ran it as ox-alpha on OpenCode and OpenRouter first.

The vs-5.2 table is the reason this is a post. Self-reported DeepSWE +17.2 (63.4 vs 46.2) and AutomationBench +22.6 (48.8 vs 26.2). Versus Opus 4.8 they approach on several agent rows and lose NL2Repo (56.3 vs 69.7). Versus DeepSeek-V4-Flash-Vision-Exp, DeepSWE 63.4 vs 59.3 and Chartography 78.0 vs 64.3. All vendor numbers.


At a Glance

  • Catalog / API: glm-5.3-flash (OpenRouter z-ai/glm-5.3-flash, Vercel zai/glm-5.3-flash)
  • Org: Z.ai (zai-org)
  • Release: August 26, 2026
  • Architecture: 320B / 18B MoE, 45 layers, hybrid linear + sparse attention (IndexPool, mHC)
  • Context: 1M
  • Modalities: text + image + video
  • Weights: MIT on zai-org/GLM-5.3-Flash
  • List: $0.15 / $0.03 cached / $0.50 per 1M; promo 50% through 2026-09-09 16:00 UTC
  • Thinking: always on (thinking.type enabled only)
  • Stealth id: was ox-alpha
  • Not glm-5.3 (753B text at $1.40 / $4.40)

What's New

First native multimodal GLM-5. Vision sits in the coding loop (self-visual judgment), not a bolted VL head. Text, image, and video in one stack.

Architecture: 18B active versus the 4.5-series 32B, 45 layers versus 92, hybrid linear + sparse attention with IndexPool and mHC, and a 30T multimodal pretrain. Compared with text GLM-5.3, Z.ai cites ~3.0× less attention compute and ~4.4× smaller KV cache.

Flash list versus 5.3 list is roughly 10× cheaper on input ($0.15 vs $1.40). The 50% promo is not the steady state.


Benchmarks

All scores below are self-reported on the Z.ai Aug 26 launch post. Not LLM Stats verified.

GLM-5.3-Flash vs GLM-5.2

5.3-Flash5.2
DeepSWE
63.446.2+17.2
AutomationBench
48.826.2+22.6
Toolathlon
78.459.9+18.5
Terminal Bench 2.1
84.381.0+3.3
NL2Repo
56.348.9+7.4
Agents' Last Exam
26.320.4+5.9
Self-reported by Z.ai, Aug 26 launch post. DeepSWE and AutomationBench are the movement. Not LLM Stats verified. Percent benches only; GDPval-AA v2 (Elo) omitted.

DeepSWE and AutomationBench are the movement: +17.2 and +22.6 versus GLM-5.2. Terminal Bench +3.3 (84.3 vs 81.0) is small. HLE w/ Tools 55.3 vs 54.7 is a wash. Versus Opus 4.8 they approach on several agent rows and lose NL2Repo (56.3 vs 69.7). GDPval-AA v2 lands at 1773 Elo (5.2: 1504); one Elo number, not a percent bar.

Vision · Self-reported

62.4%

OfficeQA-Pro

vs Vision-Exp 57.9 · Opus 48.9

89.4%

CharXiv Reasoning

w/ tools · near Opus 89.9

78.0%

Chartography

w/ tools · vs Vision-Exp 64.3

77.8%

MVBench

video · vs Vision-Exp 69.4

First native multimodal GLM-5. GLM-5.2 has no vision column. BabyVision 53.4 trails Gemini 3.7 Flash 70.9; that is the miss. Scores are self-reported by Z.ai.

GLM-5.2 has no vision column, so the vision story is versus Vision-Exp, Opus 4.8, and Gemini 3.7 Flash. Chartography 78.0 vs Vision-Exp 64.3 is a clear lead. BabyVision 53.4 trails Gemini 3.7 Flash 70.9; that is the miss.

In-house Z.ai Code Bench v1.0 at max effort: Flash 29.0 vs Opus 4.8 29.5. Labeled in-house, not a hero number.

Harness footnotes from the blog: DeepSWE on mini-swe-agent (temperature 0.95, top_p 1.0, 6h timeout, 400K context); Terminal Bench 2.1 on Claude Code 2.1.207; AutomationBench v1.0.6; Toolathlon Verified via the official service (pass@1, 3 runs). Read the full footnote block on Z.ai before treating any row as apples to apples.

Base-model table (MMLU and friends) is optional color: Flash-Base sits competitive with prior GLM bases at 18B active.


Pricing

From the official pricing docs, per 1M tokens:

DetailListPromo 50%
Input$0.15$0.075
Cached input$0.03$0.015
Output$0.50$0.25

Write the list as the story. Promo ends 2026-09-09 16:00 UTC (Z.ai writes 24:00 September 9 UTC+8). Put a September 9 price check on the calendar.

Coding Plan: Flash gives the usable quota versus GLM-5.3. Off-peak hours and all-day weekends consume 50% of standard points on the Coding Plan. Do not confuse Flash list with text GLM-5.3 at $1.40 / $0.26 / $4.40.


Serving

API id glm-5.3-flash. thinking.type supports enabled only; thinking cannot be turned off. Recommended: temperature 1, top_p 0.95, reasoning_effort: max, thinking.clear_thinking: false. Images via image_url content blocks. See the VLM guide.

Local weights support SGLang, vLLM, and TokenSpeed. The Chinese-chip serving story (SGLang stack, EPD disaggregation, ox-alpha traffic on domestic accelerators) is infrastructure color, not the thesis.


When to Use It

Good fit

  • Multimodal coding / CUA loops at Flash list
  • ox-alpha users graduating to a named id
  • Shops that want MIT weights at 18B active and a 1M window

Prefer glm-5.3

When you want the 753B text stack and will pay $1.40 / $4.40.

Prefer Vision-Exp

If you are already on DeepSeek and only need image-in at Flash price.

Put a September 9 price check on the calendar. After the promo, list is the rate that matters.


Outlook

This post ages when the promo ends, when LLM Stats verifies DeepSWE and AutomationBench, or when a larger GLM-5.3 multimodal lands. The through-line holds either way: multimodal GLM-5 at Flash dollars.

Sources: Z.ai blog, VLM docs, pricing, Hugging Face, GitHub.

Questions

Frequently Asked Questions

  • August 26, 2026. Z.ai stealth-tested it first as ox-alpha on OpenCode and OpenRouter.
  • List is $0.15 / $0.03 cached / $0.50 per 1M. A 50% promo runs at $0.075 / $0.015 / $0.25 until 2026-09-09 16:00 UTC. Budget the list.
  • 1M tokens.
  • Different model. GLM-5.3 is text-only, 753B, list $1.40 / $4.40. Flash is 320B / 18B active, natively multimodal (text, image, video), list $0.15 / $0.50.
  • Self-reported: DeepSWE 63.4 vs 46.2, AutomationBench 48.8 vs 26.2. Not LLM Stats verified.
  • Yes. MIT license on zai-org/GLM-5.3-Flash.

Continue Reading