On September 18, Zhipu's BigModel platform opened GLM-5.3-FlashX, with a peak inference speed of 200 tokens per second—about 5x faster than GLM-5.3-Flash—while keeping pricing in the same tier. The upgrade did not change the base model but focused on speed and service capacity.
The model carries 320B total parameters with only 18B activated, using a sparse plus linear-attention hybrid architecture, and runs entirely on a domestic AI accelerator cluster of over 100,000 cards. Its predecessor GLM-5.3-Flash famously went from first run on domestic accelerators to carrying full production traffic in just two weeks, with end-to-end throughput up 3.2x; an Infra Agent driven by GLM-5.3 participated in optimizing the inference system.
Zhipu also disclosed that Claude at Anthropic now leads 26% of internal R&D and writes 80% of the company's code—a sign that "AI building AI" is moving from news to daily routine. GLM-5.3-FlashX's 100,000-card domestic backing underscores China's push to run frontier models on homegrown compute.
Source: Zhipu / National Business Daily / AIBase




