DeepSeek brings its matrix kernels to Huawei Ascend
DeepSeek has open-sourced a Huawei Ascend version of DeepGEMM, its library for matrix operations used in AI models. The repository documents Ascend 950 support and compatibility with the earlier DeepGEMM interface. Developers gain another hardware path for these operations, although the published speed figures are DeepSeek’s own measurements and do not establish how much a complete application would accelerate.
Artificial Intelligence··Midday
Code in the open
DeepSeek published an open-source Huawei Ascend version of its DeepGEMM matrix-computation library on September 30. DeepGEMM handles repeated operations, including matrix multiplication, during the training and use of large language models. This release is software beneath those models, rather than a new chatbot or model. The project repository documents support for Ascend 950 and compatibility with the earlier DeepGEMM interface. Developers can therefore keep the same software calls across supported hardware paths. Independent reporting also confirms that DeepSeek released a broader set of tools for Huawei Ascend chips, including DeepGEMM.[1], [2]
DeepGEMM covers three matrix formats
The repository lists matrix operations in BF16, FP8 and FP4 formats, scoring for multi-query attention, and kernels for mixture-of-experts models. Its Ascend implementation puts low-level details such as matrix layout, alignment and memory addresses behind a software layer. The project says it uses hardware-specific methods including sparse data loading and pipeline scheduling. Installation requires a Huawei Ascend NPU, the CANN 9.20 toolkit, Torch NPU and C++20 support. Downloading public code alone will therefore not make it run on an arbitrary system. The supported path is aimed at a particular accelerator family and its associated software stack.[1]
Speed figures cover kernels, not complete models
DeepSeek says development and validation took place on Ascend 950 series hardware. Tables in the repository show latency and throughput for different data types and matrix shapes, with utilization close to hardware limits reported for some dense operations. Those are measurements from the developer. The repository includes tests intended to reproduce them, but no independent comparison of complete model performance accompanies the release. A faster matrix kernel does not imply that an entire application accelerates by the same proportion: other model steps, memory movement and system configuration affect the result. The concrete change is an open implementation path for existing model operations on Huawei hardware. Compatibility with the earlier DeepGEMM interface offers a way to change hardware paths without rewriting every call. The CANN toolkit and Torch NPU dependencies still require preparation for the Ascend environment. Because the repository does not publish an independent whole-model result, its kernel figures for different matrix shapes and data types cannot be reduced to one system-wide speed number. Developers must examine which operations the library covers and how their own configuration meets its requirements.[1]