A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
1,559
stars
1,371
commits
C++
primary language
Sep 6, 2026
updated
xLLM is an efficient LLM inference framework, specifically optimized for Chinese AI accelerators, enabling enterprise-grade deployment with enhanced efficiency and reduced cost.
| Hardware | Abbreviation | Example | Remark |
|---|---|---|---|
| Ascend NPU | NPU | A2, A3 | HDK Driver 25.2.0 + |
| Cambricon MLU | MLU | MLU | |
| Moore Threads GPU | MUSA | S5000 | |
| Hygon DCU | DCU | BW1000 | |
| MetaX MACA | MACA | MXC500 | |
| Iluvatar CoreX GPU | ILU | BI150 |
This project was made possible thanks to the following open-source projects:
Thanks to the following collaborating university laboratories:
Thanks to all the following developers who have contributed to xLLM.
If you think this repository is helpful to you, welcome to cite us:
@article{liu2025xllm,
title={xLLM Technical Report},
author={xLLM team},
journal={arXiv preprint arXiv:2510.14686},
year={2025}
}
(top 30 of 89)
C++
87.3%
Python
6.8%
Cuda
3.9%
CMake
1.1%
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
1,559
stars
1,371
commits
C++
primary language
Sep 6, 2026
updated
xLLM is an efficient LLM inference framework, specifically optimized for Chinese AI accelerators, enabling enterprise-grade deployment with enhanced efficiency and reduced cost.
| Hardware | Abbreviation | Example | Remark |
|---|---|---|---|
| Ascend NPU | NPU | A2, A3 | HDK Driver 25.2.0 + |
| Cambricon MLU | MLU | MLU | |
| Moore Threads GPU | MUSA | S5000 | |
| Hygon DCU | DCU | BW1000 | |
| MetaX MACA | MACA | MXC500 | |
| Iluvatar CoreX GPU | ILU | BI150 |
This project was made possible thanks to the following open-source projects:
Thanks to the following collaborating university laboratories:
Thanks to all the following developers who have contributed to xLLM.
If you think this repository is helpful to you, welcome to cite us:
@article{liu2025xllm,
title={xLLM Technical Report},
author={xLLM team},
journal={arXiv preprint arXiv:2510.14686},
year={2025}
}
(top 30 of 89)
C++
87.3%
Python
6.8%
Cuda
3.9%
CMake
1.1%