Days before the National Day holiday, DeepSeek open-sourced its complete low-level toolchain targeting Huawei's Ascend chips, with all components mapping one-to-one to their open-source counterparts on the Nvidia platform.
The significance is not simply "supporting domestic chips" but the exactness of the mapping. Computational programs that previously ran only on Nvidia GPUs now have high-performance Ascend implementations, meaning the same code runs on both platforms. Developers use the same Python interface, and programs automatically select whether to run on an Nvidia GPU or a Huawei NPU — switching chips no longer requires rewriting code.
The release included five core components: DeepGEMM for accelerated matrix operations, DeepEP for cross-device communication, TileKernels for vector computation, FlashMLA for optimized sparse attention, and DeepSelect for data selection, alongside the TileLang compiler toolchain. TileLang is a high-performance programming language developed by a team led by Peking University's School of Computer Science, whose core logic is to let developers write in a high-level language while automatically adapting to different chips at the lower level without sacrificing performance.
Testing data cited in Chinese media shows FlashMLA's measured performance on Ascend 950 reaching 95% of theoretical hardware peak during the prefill phase and 83% during the decode phase for sparse attention. Multiple components have computation and communication performance approaching hardware theoretical limits.
Market reaction was swift. Reuters reported that ByteDance, Tencent and Alibaba quickly approached Huawei after the DeepSeek V4 release to procure Ascend 950 chips at scale, with Alibaba Cloud's Bailian platform launching service on the day of the V4 release and Tencent Cloud following shortly after.
Huawei has also been closing the software gap. The Ascend CANN software stack continues to open up, the PTO virtual instruction set now provides more than 120 instructions, and Huawei has made 10,000-card-scale compute clusters available to developers. Xu Zhijun has said at least three to four leading Chinese large-model companies are widely recognized as using Ascend compute.
The economics remain real. DeepSeek founder Liang Wenfeng has noted that training a top-tier model on Ascend 950 would require roughly 200,000 chips, versus perhaps 50,000 Nvidia GB300 — a fourfold difference. Efficiency gaps exist, but high-end Nvidia chips are not always purchasable, and tighter US export controls mean Chinese companies cannot place all hopes on a single supply chain. Liang has described training large models on domestic chips as "a direction that must succeed," a statement about compute security and industrial choice rather than romance.
Nvidia's own quarterly report acknowledged the shift, stating that U.S. export restrictions effectively kept the company out of China's data center market while helping competitors build larger developer ecosystems.




