# DeepSeek releases DeepEP-Ascend library for Huawei NPU communication

> DeepSeek open‑sources a high‑performance communication library that speeds up MoE training on Ascend NPUs, with early benchmarks showing up to 95% bandwidth utilization.

Oossa · 2026-09-30 · https://oossa.com/en/deepseek-releases-deepep-ascend-library-for-huawei-npu-communication

DeepSeek has published a new GitHub repo called DeepEP-Ascend. It provides a communication library for training and inference on Huawei Ascend NPUs, supporting expert‑parallel all‑to‑all, pipeline parallelism and remote memory access. The library matches the NVIDIA DeepEP API and can be installed with a pip command. Early tests on Ascend 950DT show dispatch bandwidth close to the hardware limit for expert‑parallel sizes up to 32.

## The facts

- Repo created Sep 30 2026; supports FP8 dispatch and BF16 combine
- Benchmarks on Ascend 950DT report 373‑375 GB/s dispatch bandwidth for EP size 8

## Why it matters

It lets developers use Huawei’s Ascend chips for large‑scale mixture‑of‑experts models with near‑hardware communication speeds.

## Sources & references

1. [New repository deepseek-ai/DeepEP-Ascend: A high-performance communication library for machine learning training and inference on Huawei Ascend NPUs.](https://github.com/deepseek-ai/DeepEP-Ascend) – DeepSeek, 2026-09-30

Last updated: 2026-09-30
