4
stars
1,519
commits
Python
primary language
Jul 29, 2026
updated
| Documentation | User Forum | Developer Slack | WeChat | Paper | Slides |
Latest News 🔥
vLLM was originally designed to support large language models for text-based autoregressive generation tasks. vLLM-Omni is a framework that extends its support for omni-modality model inference and serving:
vLLM-Omni is fast with:
vLLM-Omni is flexible and easy to use with:
vLLM-Omni seamlessly supports most popular open-source models on HuggingFace, including:
Visit our documentation to learn more.
We welcome and value any contributions and collaborations. Please check out Contributing to vLLM-Omni for how to get involved.
If you use vLLM-Omni for your research, please cite our paper:
@article{yin2026vllmomni,
title={vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models},
author={Peiqi Yin, Jiangyun Zhu, Han Gao, Chenguang Zheng, Yongxiang Huang, Taichang Zhou, Ruirui Yang, Weizhi Liu, Weiqing Chen, Canlin Guo, Didan Deng, Zifeng Mo, Cong Wang, James Cheng, Roger Wang, Hongsheng Liu},
journal={arXiv preprint arXiv:2602.02204},
year={2026}
}
Feel free to ask questions, provide feedbacks and discuss with fellow users of vLLM-Omni in #sig-omni slack channel at slack.vllm.ai or vLLM user forum at discuss.vllm.ai.
Apache License 2.0, as found in the LICENSE file.
(top 30 of 191)
Python
99.5%
4
stars
1,519
commits
Python
primary language
Jul 29, 2026
updated
| Documentation | User Forum | Developer Slack | WeChat | Paper | Slides |
Latest News 🔥
vLLM was originally designed to support large language models for text-based autoregressive generation tasks. vLLM-Omni is a framework that extends its support for omni-modality model inference and serving:
vLLM-Omni is fast with:
vLLM-Omni is flexible and easy to use with:
vLLM-Omni seamlessly supports most popular open-source models on HuggingFace, including:
Visit our documentation to learn more.
We welcome and value any contributions and collaborations. Please check out Contributing to vLLM-Omni for how to get involved.
If you use vLLM-Omni for your research, please cite our paper:
@article{yin2026vllmomni,
title={vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models},
author={Peiqi Yin, Jiangyun Zhu, Han Gao, Chenguang Zheng, Yongxiang Huang, Taichang Zhou, Ruirui Yang, Weizhi Liu, Weiqing Chen, Canlin Guo, Didan Deng, Zifeng Mo, Cong Wang, James Cheng, Roger Wang, Hongsheng Liu},
journal={arXiv preprint arXiv:2602.02204},
year={2026}
}
Feel free to ask questions, provide feedbacks and discuss with fellow users of vLLM-Omni in #sig-omni slack channel at slack.vllm.ai or vLLM user forum at discuss.vllm.ai.
Apache License 2.0, as found in the LICENSE file.
(top 30 of 191)
Python
99.5%