On September 18, Alibaba's Qwen released Qwen3.8-Omni-Flash, a native omni-modal model that takes text, image, audio and video as direct input (no transcription step) with a 1-million-token context and support for function calling, web search, deep thinking and caching. The model covers audio/video understanding, meeting minutes and content analysis, folding multimodal understanding and agent tool use into one model.
Official figures show an average gain of over 26% across 30 benchmarks, with audio-video capability close to Gemini 3.8 Flash. The headline is price: audio input cost fell more than 98% and audio-video input more than 93% per hour—pushing scenarios like transcription, meeting records and call-quality inspection past their break-even line at once. On the same day Qwen3.8-27B topped Hugging Face's trending open-source model list.
Analysts said the second half of 2026 has seen domestic large-model competition shift decisively from "who scores highest" to "who has the lowest unit cost," with omni-modal pricing the new battleground.
Source: Alibaba Qwen / IT Home




