On September 14, Alibaba Cloud announced that DeepSeek-V4.1-Flash had officially gone live on the Qianwen AI platform, with related API services and Token Plans opened simultaneously. The model uses a 552B-parameter MoE architecture, natively supports a 1-million-Token context, and through a new caching-compression technique cuts HBM demand to one-quarter and SSD storage to one-eighth of the previous generation. Off-peak pricing is 1 yuan per million input Tokens and 4 yuan per million output Tokens, doubling at peak.
Separately, multiple sources told Phoenix Tech that DeepSeek's CFO is about to be appointed, with the closest final candidate being Hillhouse venture partner Yan Wentao, who is said to be a post-90s and is already leaving his current role. The hire would professionalize DeepSeek's finance as the company prepares its STAR Market IPO.
The V4.1 Flash release extends DeepSeek's aggressive price-performance push and reinforces China's leading position in cost-efficient large-model inference.
Source: NetEase / Phoenix Tech




