DeepSeek open-sourced a software toolkit for Huawei's Ascend accelerators on Wednesday, giving Chinese developers a domestic alternative to the Nvidia software layer that has defined AI development for more than a decade. The company announced the release on its official WeChat account, saying every component maps one-to-one onto what it had already published for Nvidia hardware.
The centerpiece is an Ascend build of TileLang, DeepSeek's high-level language and compiler for writing model operators — the small programs that carry out matrix math, data movement and attention inside a model. DeepSeek wraps Ascend C, the chip's low-level instruction set, behind TileLang's higher-level programming model and says the abstraction costs no hardware performance. Alongside it came DeepGEMM for general matrix operations, DeepEP for large-scale cross-device communication, TileKernels for vector compute and memory access, FlashMLA for sparse attention on long contexts, and DeepSelect for data filtering.
TileLang is not experimental at DeepSeek. The company says the language already carries most of the operators used to train its V4-series models on Nvidia hardware, and that every TileLang operator it uses in training now has a corresponding high-performance implementation on Ascend. A community-maintained Ascend adapter, tilelang-ascend, has existed on GitHub since September 2025; DeepSeek's release makes the official version available to anyone.
The performance figures are the company's own and have not been independently benchmarked. DeepSeek says its components approach the hardware's theoretical limits in several key tests. Project documentation cited by Chinese media puts FlashMLA's sparse attention operator at up to 95% of theoretical peak during the prefill stage on the Ascend 950, and up to 83% during decoding. Huawei's involvement was not nominal: DeepSeek says the Huawei team gave "unreserved and vigorous support," and that the two companies worked together on a 128-chip supernode design built on the Ascend 950, deeply optimizing both compute and communication paths. Reports differ on how close that supernode is to production.
The economics of the release are straightforward. What keeps customers on Nvidia is software, not silicon: switching accelerators has traditionally meant months of rewriting kernels, debugging and clawing back lost performance. Handing over a working compiler and operator library erases much of that cost for any Chinese lab willing to move to Ascend. Bloomberg reported in August that DeepSeek planned to install at least 160,000 of Huawei's most powerful accelerators at a data center it is building in Inner Mongolia, and that the startup was finalizing a funding round valuing it at close to 500 billion yuan, or about $74 billion, ahead of a possible listing.
The technical groundwork has been laid in public. Research DeepSeek published earlier put Huawei's Ascend 910C at roughly 60% of Nvidia's H100 on inference, with manual kernel tuning pushing that higher, and DeepSeek's PyTorch library now converts CUDA code into Huawei's CUNN equivalent. The company also gave Chinese chipmakers — not Nvidia or AMD — early access to test its V4 model, after which ByteDance, Tencent and Alibaba moved to lock in orders for Ascend 950 parts.
What to watch now is adoption outside the two companies that built this. A compiler that mirrors someone else's stack only matters if other developers port their code onto it, and DeepSeek's numbers will mean little until independent labs reproduce them on their own workloads. If that happens, the moat that has protected Nvidia's software franchise in China gets noticeably shallower — and an open-source license is a lot harder to re-fence than a chip.
Comments (0)
Log in to join the discussion
Log InNo comments yet