DeepSeek AI opened a new GitHub repo called DeepGEMM-Ascend. It brings the DeepGEMM matrix‑multiply library to Huawei’s Ascend NPUs. The package uses the same API as the original DeepGEMM, so code that runs on NVIDIA GPUs can be switched to Ascend with a simple pip install. It supports BF16, FP8, FP4 GEMM, MQA logits and MegaMoE kernels and claims to reach up to 99.8% of the hardware limits on Ascend 950 devices.
Why it matters
It lets developers reuse existing DeepGEMM code on Huawei’s AI chips, widening hardware options for large‑language‑model training and inference.