ByteDance is internally discussing the training of a large language model with over 5 trillion parameters, which would surpass Alibaba's Qwen 3.8-Max (2.4 trillion) and Moonshot's K3 (2.8 trillion) to become China's largest known model by parameter count.
The project is led by Xiang Liang, head of ByteDance's Seed Foundation team, in collaboration with pre-training data head Shen Ke. It remains in early discussion stages and may not ultimately be released.
According to insiders, the team hopes to achieve technological catch-up through dramatically increasing parameter scale, rather than incremental iteration. This approach aligns with CEO Liang Rubo's statement at the company's mid-year all-hands meeting that ByteDance will persist in self-developed research and accept short-term setbacks.
Founder Zhang Yiming also expressed opposition to model distillation, advocating for deep investment in long-term technological breakthroughs.
If realized, this model would represent a significant leap in China's AI capabilities, potentially closing the gap with leading international models in terms of raw scale and capability.
Source: LatePost, August 6, 2026



