Coding Model · Z.ai

GLM-5.3 Flash

GLM-5.3 Flash is Z.ai's efficient tier under GLM-5.3, released August 26, 2026 under the MIT licence: a 320B-parameter hybrid linear/sparse-attention model with 18B active parameters, native vision (text, image, and video input), and a 1M-token context window — the first natively multimodal GLM-5 model. TeamDay runs it through the Pi harness via OpenRouter as the budget option below GLM-5.3.

At a glance (verified 2026-09-05)

Vendor
Z.ai
Type
Coding Model
Context window
1M tokens
Pricing
OpenRouter lists $0.075 input / $0.015 cached / $0.25 output per 1M tokens (50% launch promo; $0.15 / $0.03 / $0.50 list).
OpenRouter ID
z-ai/glm-5.3-flash

Best for

High-volume codingLong-horizon agent tasksScreenshot and UI understandingLong-context (1M) work

How to use GLM-5.3 Flash in TeamDay

Pick GLM-5.3 Flash on TeamDay's Pi harness, routed through OpenRouter (OPENROUTER_API_KEY).

Frequently asked questions

How much does GLM-5.3 Flash cost?

OpenRouter currently bills $0.075 per 1M input tokens, $0.015 cached, and $0.25 per 1M output tokens under a 50% launch promotion; the list price is $0.15 / $0.03 / $0.50. That is roughly 5% of GLM-5.3's $1.40 / $4.40.

What is the difference between GLM-5.3 and GLM-5.3 Flash?

GLM-5.3 is the flagship reasoning-capable coding model. GLM-5.3 Flash is a leaner 320B / 18B-active hybrid-attention model that adds native vision, ships under the MIT licence, and costs a fraction as much — Z.ai positions it for efficient coding and long-horizon agent tasks.

Does GLM-5.3 Flash support vision?

Yes — it is the first natively multimodal model in the GLM-5 series, accepting text, image, and video input so it can observe interfaces and rendered results.

How do I run GLM-5.3 Flash in TeamDay?

Select it on the Pi harness in the chat model picker, routed via OpenRouter with your own API key.

Related

GLM-5.3 Flashis one of the models TeamDay's AI employees already run — no API keys, no infrastructure.

Hire your first AI employee →