A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
See the codexLLM is an efficient LLM inference framework, specifically optimized for Chinese AI accelerators, enabling enterprise-grade deployment with enhanced efficiency and reduced cost.
| Hardware | Abbreviation | Example | Remark |
|---|---|---|---|
| Ascend NPU | NPU | A2, A3 | HDK Driver 25.2.0 + |
| Cambricon MLU | MLU | MLU | |
| Moore Threads GPU | MUSA | S5000 | |
| Hygon DCU | DCU | BW1000 | |
| MetaX MACA | MACA | MXC500 | |
| Iluvatar CoreX GPU | ILU | BI150 |
This project was made possible thanks to the following open-source projects:
Thanks to the following collaborating university laboratories:
Thanks to all the following developers who have contributed to xLLM.
If you think this repository is helpful to you, welcome to cite us:
@article{liu2025xllm,
title={xLLM Technical Report},
author={xLLM team},
journal={arXiv preprint arXiv:2510.14686},
year={2025}
}
(top 30 of 97)
C++
83.8%
Python
10.9%
Cuda
3.5%
CMake
1.0%
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
See the codexLLM is an efficient LLM inference framework, specifically optimized for Chinese AI accelerators, enabling enterprise-grade deployment with enhanced efficiency and reduced cost.
| Hardware | Abbreviation | Example | Remark |
|---|---|---|---|
| Ascend NPU | NPU | A2, A3 | HDK Driver 25.2.0 + |
| Cambricon MLU | MLU | MLU | |
| Moore Threads GPU | MUSA | S5000 | |
| Hygon DCU | DCU | BW1000 | |
| MetaX MACA | MACA | MXC500 | |
| Iluvatar CoreX GPU | ILU | BI150 |
This project was made possible thanks to the following open-source projects:
Thanks to the following collaborating university laboratories:
Thanks to all the following developers who have contributed to xLLM.
If you think this repository is helpful to you, welcome to cite us:
@article{liu2025xllm,
title={xLLM Technical Report},
author={xLLM team},
journal={arXiv preprint arXiv:2510.14686},
year={2025}
}
(top 30 of 97)
C++
83.8%
Python
10.9%
Cuda
3.5%
CMake
1.0%