GLM-5.3-Flash
First natively multimodal GLM-5 model, a 320B-total (18B active) MIT-licensed MoE on a hybrid sparse and linear attention design.
A newly trained base, said to outperform GLM-5.2 at one-tenth the price and approach Claude Opus 4.8 on coding and agentic benchmarks. Wikipedia lists a closed GLM-5.3-FlashX variant in September 2026.
- Date
- Wednesday, 26 August 2026
- Lab
- Zhipu AI / Z.ai
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Total / active parameters | 320B / 18B | company |
| DeepSWE | 63.4 | company |
| Terminal Bench 2.1 | 84.3 | company |
Date from Z.ai release notes. Benchmarks and the price claim are company-reported; HF card links the GLM-5 arXiv paper rather than a dedicated report.
Sources
- huggingface.co/zai-org/GLM-5.3-Flash
- docs.z.ai/release-notes/new-released
- en.wikipedia.org/wiki/ChatGLM
This record was checked against its sources on 6 October 2026. How we check