HANGZHOU, July 31, 2026 - DeepSeek officially launched the V4-Flash API via its documentation log, marking a significant milestone for the Chinese AI ecosystem. The V4-Flash official version maintains the MoE architecture with 284B total parameters and 13B activated parameters, identical to the preview version, with all capability gains derived from newly conducted post-training.
Five major innovations stand out: first, smaller parameters outperforming larger ones, breaking the ceiling that AI model performance relies solely on stacking parameters; second, post-training directional enhancement, achieving qualitative leaps in Agent capability; third, native support for 1M context length; fourth, dual-mode support for thinking/non-thinking; and fifth, day-one compatibility with the Codex development ecosystem.
In terms of pricing, V4-Flash sets a new industry benchmark, with cached input at only $0.0028 per million tokens, uncached input at $0.14, and output at $0.28. A 20 million-token Agent task costs only about $3.5, delivering performance close to flagship models at a fraction of the cost.
DeepSeek founder Liang Wenfeng net worth has soared to $36 billion, making him the richest person among dedicated AI companies globally, as the V4-Flash release reshapes global developer choices.
Source: DeepSeek Official, Sina Finance





