DeepSeek has published a new GitHub repo called DeepEP-Ascend. It provides a communication library for training and inference on Huawei Ascend NPUs, supporting expert‑parallel all‑to‑all, pipeline parallelism and remote memory access. The library matches the NVIDIA DeepEP API and can be installed with a pip command. Early tests on Ascend 950DT show dispatch bandwidth close to the hardware limit for expert‑parallel sizes up to 32.
Why it matters
It lets developers use Huawei’s Ascend chips for large‑scale mixture‑of‑experts models with near‑hardware communication speeds.