Community maintained hardware plugin for vLLM on Ascend
See the code
| About Ascend | Documentation | Support Matrix | #SIG-Ascend | Users Forum | Weekly Meeting |
Latest News 🔥
vLLM Ascend (vllm-ascend) is a community maintained hardware plugin for running vLLM seamlessly on the Ascend NPU.
It is the recommended approach for supporting the Ascend backend within the vLLM community. It adheres to the principles outlined in the [RFC]: Hardware pluggable, providing a hardware-pluggable interface that decouples the integration of the Ascend NPU with vLLM.
By using vLLM Ascend plugin, popular open-source models, including Transformer-like, Mixture-of-Experts (MoE), Embedding, Multi-modal LLMs can run seamlessly on the Ascend NPU.
For detailed information on supported models and features, please refer to the support matrix.
DeepWiki is a dynamic knowledge base collaboratively maintained by the community and AI, designed to provide you with deeper technical insights beyond conventional documentation. If official documentation serves as a "quick start" guide to help you get going, the DeepWiki is your technical companion for "deep understanding". Here, you can explore: core architecture and design principles, key source code interpretations, technical context, and decision-making logic.
Whether you are a performance tuner looking to deploy vLLM-Ascend in production or a contributor aiming to build upon it for secondary development, DeepWiki offers invaluable references for you. Welcome to explore now and join us in delving deep into the technical core of vLLM-Ascend.
If you need to access Ascend NPU computing resources for development or testing, please visit the HiDevLab - Online Development page on the Huawei HiDevLab platform to apply for and use them.
Please use the following recommended versions to get started quickly:
| Version | Release type | Doc |
|---|---|---|
| v0.26.0rc1 | Release candidate | See QuickStart and Installation for more details |
| v0.23.0 | Latest stable version | See QuickStart and Installation for more details |
vllm-ascend has a main branch and a dev branch.
releases/v0.13.0 is the dev branch for vLLM v0.13.0 version.Below are the maintained branches:
| Branch | Status | Note |
|---|---|---|
| main | Maintained | CI commitment for vLLM main branch and vLLM v0.30.0 tag |
| releases/v0.13.0 | Maintained | Only bug fixes are allowed, and no new release tags anymore. |
| releases/v0.18.0 | Maintained | CI commitment for vLLM 0.18.0 version |
| releases/v0.23.0 | Maintained | CI commitment for vLLM 0.23.0 version |
| rfc/ | Maintained | Feature branches for collaboration |
Please refer to Versioning policy for more details.
See CONTRIBUTING for more details, which is a step-by-step guide to help you set up the development environment, build and test.
We welcome and value any contributions and collaborations:
Apache License 2.0, as found in the LICENSE file.
Python
53.8%
C++
42.6%
Shell
1.5%
CMake
1.4%
Community maintained hardware plugin for vLLM on Ascend
See the code
| About Ascend | Documentation | Support Matrix | #SIG-Ascend | Users Forum | Weekly Meeting |
Latest News 🔥
vLLM Ascend (vllm-ascend) is a community maintained hardware plugin for running vLLM seamlessly on the Ascend NPU.
It is the recommended approach for supporting the Ascend backend within the vLLM community. It adheres to the principles outlined in the [RFC]: Hardware pluggable, providing a hardware-pluggable interface that decouples the integration of the Ascend NPU with vLLM.
By using vLLM Ascend plugin, popular open-source models, including Transformer-like, Mixture-of-Experts (MoE), Embedding, Multi-modal LLMs can run seamlessly on the Ascend NPU.
For detailed information on supported models and features, please refer to the support matrix.
DeepWiki is a dynamic knowledge base collaboratively maintained by the community and AI, designed to provide you with deeper technical insights beyond conventional documentation. If official documentation serves as a "quick start" guide to help you get going, the DeepWiki is your technical companion for "deep understanding". Here, you can explore: core architecture and design principles, key source code interpretations, technical context, and decision-making logic.
Whether you are a performance tuner looking to deploy vLLM-Ascend in production or a contributor aiming to build upon it for secondary development, DeepWiki offers invaluable references for you. Welcome to explore now and join us in delving deep into the technical core of vLLM-Ascend.
If you need to access Ascend NPU computing resources for development or testing, please visit the HiDevLab - Online Development page on the Huawei HiDevLab platform to apply for and use them.
Please use the following recommended versions to get started quickly:
| Version | Release type | Doc |
|---|---|---|
| v0.26.0rc1 | Release candidate | See QuickStart and Installation for more details |
| v0.23.0 | Latest stable version | See QuickStart and Installation for more details |
vllm-ascend has a main branch and a dev branch.
releases/v0.13.0 is the dev branch for vLLM v0.13.0 version.Below are the maintained branches:
| Branch | Status | Note |
|---|---|---|
| main | Maintained | CI commitment for vLLM main branch and vLLM v0.30.0 tag |
| releases/v0.13.0 | Maintained | Only bug fixes are allowed, and no new release tags anymore. |
| releases/v0.18.0 | Maintained | CI commitment for vLLM 0.18.0 version |
| releases/v0.23.0 | Maintained | CI commitment for vLLM 0.23.0 version |
| rfc/ | Maintained | Feature branches for collaboration |
Please refer to Versioning policy for more details.
See CONTRIBUTING for more details, which is a step-by-step guide to help you set up the development environment, build and test.
We welcome and value any contributions and collaborations:
Apache License 2.0, as found in the LICENSE file.
Python
53.8%
C++
42.6%
Shell
1.5%
CMake
1.4%