Umans GLM 5.3 Flash
also served as umans-coder
93.1tok/s
throughput · p50 · last 5 min
620ms
TTFT · p50 · last 5 min
100.00%
uptime · 24h
GLM-5.3-Flash is Z.ai's fast open-weights model for coding and agentic work: a 320B mixture-of-experts with 18B active parameters per token, the first natively multimodal release in the GLM-5 series with native image and video understanding, on a 1M-token context window. It always thinks at max effort by default; dial reasoning to low or high (thinking cannot be turned off). Billed per token ($0.15 / $0.50 / $0.03 per 1M; input / output / cache read).
Trends
Speed over the last 90 days
90 days agopre-release before Sep 9, 2026today
90 days agopre-release before Sep 9, 2026today
Changelog
Events for Umans GLM 5.3 Flash
Sep 92026
Released pay-per-token: Umans GLM 5.3 Flash Released
umans-glm-5.3-flash joins the lineup: GLM-5.3-Flash, the checkpoint that served here as the seat-gated pre-release lab since August 26: a 320B mixture-of-experts (18B active per token), the first natively multimodal release in the GLM-5 series with native image and video understanding, on a 1M context window. It always thinks at max effort by default; dial reasoning to low or high (thinking cannot be turned off). Billed per token: $0.15 / $0.50 / $0.03 per 1M (input / output / cache read). The pre-release window's metrics stay on the model's status page as its pre-release period (before Sep 9).