# DeepSeek releases DeepGEMM-Ascend for Huawei Ascend NPU

> The new repo ports DeepGEMM to Ascend, supporting BF16, FP8, FP4 GEMM and MegaMoE with near‑peak hardware performance.

Oossa · 2026-09-30 · https://oossa.com/en/deepseek-releases-deepgemm-ascend-for-huawei-ascend-npu

DeepSeek AI opened a new GitHub repo called DeepGEMM-Ascend. It brings the DeepGEMM matrix‑multiply library to Huawei’s Ascend NPUs. The package uses the same API as the original DeepGEMM, so code that runs on NVIDIA GPUs can be switched to Ascend with a simple pip install. It supports BF16, FP8, FP4 GEMM, MQA logits and MegaMoE kernels and claims to reach up to 99.8% of the hardware limits on Ascend 950 devices.

## The facts

- Initial release announced 2026‑09‑30
- Supports Ascend 950 series, CANN 9.20 toolkit and torch_npu

## Why it matters

It lets developers reuse existing DeepGEMM code on Huawei’s AI chips, widening hardware options for large‑language‑model training and inference.

## Sources & references

1. [New repository deepseek-ai/DeepGEMM-Ascend: DeepGEMM-Ascend: clean and efficient matrix multiplication kernel library for Huawei Ascend NPUs](https://github.com/deepseek-ai/DeepGEMM-Ascend) – DeepSeek, 2026-09-29

Last updated: 2026-09-30
