Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week Productivity 19-Year-Old Founder Emerges From Stealth With $11 Million to Sell You a $3,499 'Brain in a Box' That Runs Your AI Agents at Home AI Agents Half a Million Interviews In: HackerRank's AI Interviewer Chakra Goes GA, and It Wants to Replace Three Hiring Rounds With One Security After Claude Agents Escaped Its Sandbox 3 Times, Anthropic Deploys Real-Time Classifiers to Stop the Next Escape Before It Happens Business Meta Halves Its Internal Claude Users to 30,000 and Microsoft Slashes a $1 Billion Anthropic Budget by More Than a Third Business Sony Innovation Fund Backs Primitive Labs, a Startup That Builds Simulated Crowds to Stress-Test Products Before Launch Apple Intelligence Apple Removed the Apple Intelligence Off Switch in macOS 27 — So a Developer Built a CLI That Deletes It Anyway Security OpenAI Turns On Invisible Text Watermarks for ChatGPT in the EU — and Publishes Exactly How Weak They Are News A Mystery 'Space Bunny Alpha' Model Just Topped OpenRouter's Leaderboard With 38.7 Trillion Tokens a Week

DeepSeek Open-Sources Its Whole Software Stack for Huawei's Ascend Chips, in a Direct Answer to CUDA

DeepSeek Open-Sources Its Whole Software Stack for Huawei's Ascend Chips, in a Direct Answer to CUDA

DeepSeek open-sourced a toolkit for Huawei's Ascend accelerators on September 30, led by an Ascend build of its TileLang compiler plus DeepGEMM, DeepEP, TileKernels, FlashMLA and DeepSelect. The company says each piece mirrors its Nvidia equivalents and that key compute and communication tests reach up to 95% of the Ascend 950's theoretical peak. Those numbers are DeepSeek's own.

DeepSeek open-sourced a software toolkit for Huawei's Ascend accelerators on Wednesday, giving Chinese developers a domestic alternative to the Nvidia software layer that has defined AI development for more than a decade. The company announced the release on its official WeChat account, saying every component maps one-to-one onto what it had already published for Nvidia hardware.

The centerpiece is an Ascend build of TileLang, DeepSeek's high-level language and compiler for writing model operators — the small programs that carry out matrix math, data movement and attention inside a model. DeepSeek wraps Ascend C, the chip's low-level instruction set, behind TileLang's higher-level programming model and says the abstraction costs no hardware performance. Alongside it came DeepGEMM for general matrix operations, DeepEP for large-scale cross-device communication, TileKernels for vector compute and memory access, FlashMLA for sparse attention on long contexts, and DeepSelect for data filtering.

TileLang is not experimental at DeepSeek. The company says the language already carries most of the operators used to train its V4-series models on Nvidia hardware, and that every TileLang operator it uses in training now has a corresponding high-performance implementation on Ascend. A community-maintained Ascend adapter, tilelang-ascend, has existed on GitHub since September 2025; DeepSeek's release makes the official version available to anyone.

The performance figures are the company's own and have not been independently benchmarked. DeepSeek says its components approach the hardware's theoretical limits in several key tests. Project documentation cited by Chinese media puts FlashMLA's sparse attention operator at up to 95% of theoretical peak during the prefill stage on the Ascend 950, and up to 83% during decoding. Huawei's involvement was not nominal: DeepSeek says the Huawei team gave "unreserved and vigorous support," and that the two companies worked together on a 128-chip supernode design built on the Ascend 950, deeply optimizing both compute and communication paths. Reports differ on how close that supernode is to production.

The economics of the release are straightforward. What keeps customers on Nvidia is software, not silicon: switching accelerators has traditionally meant months of rewriting kernels, debugging and clawing back lost performance. Handing over a working compiler and operator library erases much of that cost for any Chinese lab willing to move to Ascend. Bloomberg reported in August that DeepSeek planned to install at least 160,000 of Huawei's most powerful accelerators at a data center it is building in Inner Mongolia, and that the startup was finalizing a funding round valuing it at close to 500 billion yuan, or about $74 billion, ahead of a possible listing.

The technical groundwork has been laid in public. Research DeepSeek published earlier put Huawei's Ascend 910C at roughly 60% of Nvidia's H100 on inference, with manual kernel tuning pushing that higher, and DeepSeek's PyTorch library now converts CUDA code into Huawei's CUNN equivalent. The company also gave Chinese chipmakers — not Nvidia or AMD — early access to test its V4 model, after which ByteDance, Tencent and Alibaba moved to lock in orders for Ascend 950 parts.

What to watch now is adoption outside the two companies that built this. A compiler that mirrors someone else's stack only matters if other developers port their code onto it, and DeepSeek's numbers will mean little until independent labs reproduce them on their own workloads. If that happens, the moat that has protected Nvidia's software franchise in China gets noticeably shallower — and an open-source license is a lot harder to re-fence than a chip.

Comments (0)

Log in to join the discussion

Log In

No comments yet