Skip to content
nanoai

DeepSeek ports its core AI software to Huawei's Ascend chips

Ascend versions of six open-source libraries, including TileLang and DeepGEMM, target the part of Nvidia's lead that chips alone cannot close: the software developers write against.

By The Nano AI Staff3 min read

AI developer news graphic showing DeepSeek and Huawei Ascend chips in a futuristic data center, with six open-source libraries, code demonstrating Nvidia-compatible matrix operations, and callouts for Ascend support.
Image: AI Generated

Key takeaways

  • DeepSeek released Huawei Ascend versions of six libraries, including TileLang and DeepGEMM.
  • DeepGEMM-Ascend keeps DeepGEMM's API and claims up to 99.8% of Ascend 950's hardware limit.
  • DeepSeek says it is jointly developing a 128-chip Ascend 950 supernode with Huawei.

DeepSeek has released versions of six of its open-source AI infrastructure libraries that run on Huawei's Ascend chips as well as Nvidia's. The Chinese lab announced the release on its official WeChat account on Wednesday, September 30, saying Huawei supported the work and that the two companies are jointly developing a supernode built from 128 Ascend 950 chips.

The six are TileLang, a language for writing AI kernels (the small, heavily tuned routines that do the maths on a chip); DeepGEMM for matrix multiplication; DeepEP for communication between chips; TileKernels; FlashMLA for DeepSeek's attention design; and DeepSelect for fast data selection.

What developers actually get

The clearest example is DeepGEMM-Ascend, now on GitHub under the MIT licence. Its README calls it "a port of DeepGEMM to the HUAWEI Ascend platform" and says it "is fully API-compatible with DeepGEMM and supports BF16, FP8, FP4 GEMM, MQA logits, and MegaMoE". In plain terms, code written for DeepSeek's Nvidia library should be able to call the Ascend version the same way.

DeepSeek claims dense matrix multiplication "reaches up to 99.8% of the hardware limit" on the Ascend 950 series, where the library was developed and validated, and thanks Huawei "for its technical support and engineering expertise". These are the developer's own figures on Huawei's newest chip. Nobody outside has reproduced them yet.

According to DeepSeek's announcement, as translated by the Geopolitechs newsletter, "Every TileLang operator currently used in DeepSeek training now has a corresponding high-performance implementation on Ascend." That is the line that matters for DeepSeek's own roadmap. It means the company's training code, not only its inference code, has a working path onto Huawei hardware.

Why the software matters more than the chip

Nvidia's lead has never rested on silicon alone. It rests on CUDA, the programming platform most AI code is written for, and on the thick layer of tuned libraries built on top of it. A rival chip that is fast on paper is of little use if moving a model onto it means rewriting and re-tuning thousands of kernels.

DeepSeek is going after that problem head on. "To build a new generation of independent, self-controlled GPU software ecosystems, the first priority is establishing a high-level language that is universal, easy to program, and still capable of reaching the hardware's full performance potential," the company said, according to Reuters as quoted by Dataquest. Its announcement argues that TileLang offers a simpler programming model than CUDA.

The work did not start this week. The open TileLang-Ascend project lists the release of "DeepSeek V4 kernels" in April. The project reports that its code runs at about 0.98 times the speed of hand-written Ascend C on typical matrix workloads and about 0.95 times on attention-style workloads, Dataquest reported. What is new is DeepSeek putting its own production libraries behind the effort, with API compatibility as the selling point.

Who gains, and the catch

Huawei gains most. A respected lab has done the tedious porting work that developers usually demand before they will try a new chip. Chinese AI companies that cannot buy Nvidia's best hardware because of US export controls get a clearer route to domestic chips. The likelier effect on Nvidia is a slow loss of lock-in inside China rather than any sudden shift in its business.

The catch is supply. Software cannot conjure chips, and The Neuron put the question in its headline: DeepSeek has opened the code, but can Huawei deliver the compute? It noted that Huawei has said demand for its AI computing equipment exceeded what it could currently supply inside China.

What to watch: independent benchmarks of these libraries on shipping Ascend 950 systems, whether other Chinese labs adopt them, and whether DeepSeek's next model is trained from start to finish on Huawei hardware.

  • Open Source
  • AI chips
  • DeepSeek
  • Huawei
  • Ascend
  • CUDA

Sources

  1. DeepSeek Builds for Huawei Ascend — Geopolitechs, Sep 30, 2026
  2. DeepSeek expands Huawei Ascend push with six open-source AI tools — Dataquest, Sep 30, 2026
  3. DeepGEMM-Ascend: clean and efficient matrix multiplication kernel library for Huawei Ascend NPUs — DeepSeek (GitHub), Sep 30, 2026
  4. tilelang-ascend: Ascend TileLang adapter — tile-ai (GitHub)
  5. DeepSeek Opened the Code. Can Huawei Deliver the Compute? — The Neuron, Sep 30, 2026

Follow The Nano AI: Instagram · X · LinkedIn · YouTube

Was this article helpful?

Comments

No comments yet. Start the conversation.

Be respectful. Comments are moderated.

Related stories

The AI briefing, without the noise.

The stories that matter in AI, sourced and explained. Free, and you can unsubscribe at any time.