Alibaba's Qwen team released Qwen3.8-Flash-Next on August 26, 2026 — a multimodal large language model with 125 billion total parameters but only 6 billion active per token, designed to preview architecture choices planned for the full Qwen4 model fa
Alibaba's Qwen team released Qwen3.8-Flash-Next on August 26, 2026 — a multimodal large language model with 125 billion total parameters but only 6 billion active per token, designed to preview architecture choices planned for the full Qwen4 model family. The release includes a novel 51-billion-parameter embedding layer the team says is one of several innovations slated for Qwen4. By giving developers an early open preview of next-generation architecture months before the full model ships, Alibaba is pursuing a distinct release strategy that signals confidence in the underlying design while allowing the research community to validate key components ahead of the full release. Qwen3.8-Flash-Next scores competitively against models two to three times its active parameter count on key benchmarks. The open-source release under a permissive license reinforces Alibaba's commitment to democratizing access to frontier AI technology. (Original source: AI Impact Hub / MarkTechPost, September 2, 2026)