OossaAI is evolving fast. We explain it simply.

DeepSeek releases DeepGEMM-Ascend for Huawei Ascend NPU

The new repo ports DeepGEMM to Ascend, supporting BF16, FP8, FP4 GEMM and MegaMoE with near‑peak hardware performance.

NoteOossa1 min read

DeepSeek AI opened a new GitHub repo called DeepGEMM-Ascend. It brings the DeepGEMM matrix‑multiply library to Huawei’s Ascend NPUs. The package uses the same API as the original DeepGEMM, so code that runs on NVIDIA GPUs can be switched to Ascend with a simple pip install. It supports BF16, FP8, FP4 GEMM, MQA logits and MegaMoE kernels and claims to reach up to 99.8% of the hardware limits on Ascend 950 devices.

Why it matters

It lets developers reuse existing DeepGEMM code on Huawei’s AI chips, widening hardware options for large‑language‑model training and inference.

Was this article useful?
Share

Read next

Oossa · Newsletter

The week in AI, explained

Every Monday: the stories worth knowing, in plain language. Free, no spam.

Sources & references

#SourceOutletDateKey takeaway
1New repository deepseek-ai/DeepGEMM-Ascend: DeepGEMM-Ascend: clean and efficient matrix multiplication kernel library for Huawei Ascend NPUs ↗DeepSeekSep 29, 2026DeepGEMM-Ascend: clean and efficient matrix multiplication kernel library for Huawei Ascend NPUs

1 sources

Last updated: ·Markdown·llms.txt