DeepSeek ports its core AI software to Huawei's Ascend chips
Ascend versions of six open-source libraries, including TileLang and DeepGEMM, target the part of Nvidia's lead that chips alone cannot close: the software developers write against.

Key takeaways
- DeepSeek released Huawei Ascend versions of six libraries, including TileLang and DeepGEMM.
- DeepGEMM-Ascend keeps DeepGEMM's API and claims up to 99.8% of Ascend 950's hardware limit.
- DeepSeek says it is jointly developing a 128-chip Ascend 950 supernode with Huawei.
DeepSeek has released versions of six of its open-source AI infrastructure libraries that run on Huawei's Ascend chips as well as Nvidia's. The Chinese lab announced the release on its official WeChat account on Wednesday, September 30, saying Huawei supported the work and that the two companies are jointly developing a supernode built from 128 Ascend 950 chips.
The six are TileLang, a language for writing AI kernels (the small, heavily tuned routines that do the maths on a chip); DeepGEMM for matrix multiplication; DeepEP for communication between chips; TileKernels; FlashMLA for DeepSeek's attention design; and DeepSelect for fast data selection.
What developers actually get
The clearest example is DeepGEMM-Ascend, now on GitHub under the MIT licence. Its README calls it "a port of DeepGEMM to the HUAWEI Ascend platform" and says it "is fully API-compatible with DeepGEMM and supports BF16, FP8, FP4 GEMM, MQA logits, and MegaMoE". In plain terms, code written for DeepSeek's Nvidia library should be able to call the Ascend version the same way.
DeepSeek claims dense matrix multiplication "reaches up to 99.8% of the hardware limit" on the Ascend 950 series, where the library was developed and validated, and thanks Huawei "for its technical support and engineering expertise". These are the developer's own figures on Huawei's newest chip. Nobody outside has reproduced them yet.
According to DeepSeek's announcement, as translated by the Geopolitechs newsletter, "Every TileLang operator currently used in DeepSeek training now has a corresponding high-performance implementation on Ascend." That is the line that matters for DeepSeek's own roadmap. It means the company's training code, not only its inference code, has a working path onto Huawei hardware.
Why the software matters more than the chip
Nvidia's lead has never rested on silicon alone. It rests on CUDA, the programming platform most AI code is written for, and on the thick layer of tuned libraries built on top of it. A rival chip that is fast on paper is of little use if moving a model onto it means rewriting and re-tuning thousands of kernels.
DeepSeek is going after that problem head on. "To build a new generation of independent, self-controlled GPU software ecosystems, the first priority is establishing a high-level language that is universal, easy to program, and still capable of reaching the hardware's full performance potential," the company said, according to Reuters as quoted by Dataquest. Its announcement argues that TileLang offers a simpler programming model than CUDA.
The work did not start this week. The open TileLang-Ascend project lists the release of "DeepSeek V4 kernels" in April. The project reports that its code runs at about 0.98 times the speed of hand-written Ascend C on typical matrix workloads and about 0.95 times on attention-style workloads, Dataquest reported. What is new is DeepSeek putting its own production libraries behind the effort, with API compatibility as the selling point.
Who gains, and the catch
Huawei gains most. A respected lab has done the tedious porting work that developers usually demand before they will try a new chip. Chinese AI companies that cannot buy Nvidia's best hardware because of US export controls get a clearer route to domestic chips. The likelier effect on Nvidia is a slow loss of lock-in inside China rather than any sudden shift in its business.
The catch is supply. Software cannot conjure chips, and The Neuron put the question in its headline: DeepSeek has opened the code, but can Huawei deliver the compute? It noted that Huawei has said demand for its AI computing equipment exceeded what it could currently supply inside China.
What to watch: independent benchmarks of these libraries on shipping Ascend 950 systems, whether other Chinese labs adopt them, and whether DeepSeek's next model is trained from start to finish on Huawei hardware.
- Open Source
- AI chips
- DeepSeek
- Huawei
- Ascend
- CUDA
Sources
- DeepSeek Builds for Huawei Ascend — Geopolitechs, Sep 30, 2026
- DeepSeek expands Huawei Ascend push with six open-source AI tools — Dataquest, Sep 30, 2026
- DeepGEMM-Ascend: clean and efficient matrix multiplication kernel library for Huawei Ascend NPUs — DeepSeek (GitHub), Sep 30, 2026
- tilelang-ascend: Ascend TileLang adapter — tile-ai (GitHub)
- DeepSeek Opened the Code. Can Huawei Deliver the Compute? — The Neuron, Sep 30, 2026
Related stories

Nvidia launches an agent safety platform with a hardware kill switch
OpenShell, an open-source runtime that fences in AI agents, is available now. Sentry, a watchdog that runs on Nvidia's BlueField-4 chips, is a reference design with no price or date. Anthropic and SpaceXAI signed up; OpenAI is not named.
4 min read

Cognition says Devin passed a $1bn annual revenue run rate
The coding-agent maker roughly doubled its run rate in about four months. A run rate is a snapshot rather than a year of sales, and the market around Devin is getting cheaper and more crowded.
3 min read

GitHub open-sources an AI agent that writes and runs fuzz tests
The Security Lab's taskflow sets up AFL++, writes harnesses and triages crashes on its own. It found two real out-of-bounds reads in oniguruma, and GitHub says to run it only on a throwaway machine.
3 min read
Comments
No comments yet. Start the conversation.