At the "Computing Power Premiere" session of the 2026 China Computing Power Conference, China Mobile Cloud, together with partners including the South Lake Research Institute of China Electronics Technology Group, Beijing Lingxi Technology, Iluvatar CoreX, Jiangsu Mobile, Tsinghua University and Peking University, released China's first domestic GPU + brain-inspired chip heterogeneous hybrid inference system for large models.
As the AI industry shifts from model training to large-scale inference services, inference compute has become the new digital infrastructure behind Token factories, AI agents and text-to-video. Traditional single-GPU inference architectures struggle with the "memory wall" and "power wall," while domestic advanced-process supply is constrained.
The team's breakthrough splits Transformer inference into Attention-FeedForward (AF) separation: Attention computation runs on domestic GPUs while the feed-forward modules run on brain-inspired chips that exploit in-memory computing and large on-chip SRAM — overcoming the bottlenecks of conventional GPU inference.
Self-developed model compilers, high-speed interconnect protocols and a unified inference engine coordinate the two architectures. Validated on the DeepSeek V4-Flash model, the system more than doubles inference throughput and energy efficiency versus comparable domestic GPU clusters and cuts operating costs by over 40%, offering a replicable path to localized large-model inference.




