Z.ai: GLM 5.3 Flash
Vendor: z-ai|
glm-5.3-flashModalities
TextText
In / out price70% off70% cheaper than the provider's official price per 1M tokens.
$0.045 / $0.045 per 1M
Context
1.05M
Released
Aug 18, 2026
Overview
A fast, low-cost version of GLM 5.3: the same 1M-token context and up to 131K output, but significantly cheaper and with lower latency. Built for high-throughput agentic runs, code completion and bulk text processing.
Context
1.05M
Max output
131 072
Released
Aug 18, 2026
Step-by-step reasoning
Yes
FastCoding
Price per 1M tokens
Input and output cost the same, billed per actual usage — no subscriptions, no minimums.
Input
$0.045/M
$0.150/M70% off70% cheaper than the provider's official price per 1M tokens.
Output
$0.045/M
$0.500/M91% off91% cheaper than the provider's official price per 1M tokens.
Official list price: $0.150 / $0.500 per 1M
Performance
Gateway metrics for the last 24 hours.
Loading metrics…
FAQ
Are there any rate limits?
The default limit is request rate per key. Each key can be restricted by models, spend and IP in your dashboard.
How is billing calculated?
Pay as you go: one price per 1M tokens, input and output cost the same, charged proportionally to real usage.
Is streaming supported?
Yes, stream: true behaves exactly like in the OpenAI API — tokens arrive as they are generated.