Zhipu has released and open-sourced GLM-5.3-Flash, a 320B-A18B mixture-of-experts model that is the first natively multimodal entry in its GLM-5 family, handling text, images, video and interleaved visual documents in one system. Weights are published under a permissive MIT licence on Hugging Face, with the API opened in sync.
The model carries 320 billion total parameters with 18 billion active per token, designed for low inference cost. Zhipu says GLM-5.3-Flash scores 57 on the Artificial Analysis Intelligence Index - level with Anthropic's Claude Opus 4.8 - while costing about one-tenth as much to run as GLM-5.2 and roughly one-fortieth of Opus 4.8.
Crucially, Zhipu disclosed that during a large stealth trial - when the model appeared anonymously as 'Ox Alpha' on OpenRouter and OpenCode - all request traffic was served by a cluster of more than 100,000 domestically produced chips, linked by a self-developed high-bandwidth interconnect, with no Nvidia hardware involved. The trial processed about 62 trillion tokens, and in its first three days on OpenRouter alone handled more than 11 trillion tokens.
The architecture combines sparse and linear attention, cutting attention computation by about 3x and context-cache memory by more than 4x versus GLM-5.3 while supporting a 1-million-token window. Zhipu says end-to-end service performance on the same hardware improved about 3x versus the initial baseline, with hardware efficiency and per-token cost approaching those of mainstream Nvidia GPUs - evidence that Chinese chips can economically support frontier-scale inference at global usage volumes.
Image: Efficiency of AI-related computer chips. Source: Wikimedia Commons (CC BY-SA 4.0).




