📖Paper |🤗Model | 🤗Training Data (Coming Soon)
1
9 commits
1 linked in READMEs
updated Oct 20, 2025
Che Liu
,
Yingji Zhang
,
Dong Zhang
,
Weijie Zhang
,
Chenggong Gong
,
Yu Lu
,
Shilin Zhou
,
Ziliang Gan
,
Ziao Wang,
Haipang Wu,
Ji Liu,
Andre Freitas,
Qifan Wang,
Zenglin Xu,
Rongjunchen Zhang♠,
Yong Dai♠
NEXUS-O is an industry-scale omni-modal large language model (LLM) that unifies audio, vision, and language understanding into a single modular framework. Human perception integrates sight, sound, and language — NEXUS-O aims to replicate this ability for intelligent agents across real-world scenarios such as ASR, Speech-to-Speech Chat, and Multimodal Reasoning.
Architecture of NEXUS-O
Training Stages
@article{liu2025nexus,
title={Nexus: An Omni-Perceptive And-Interactive Model for Language, Audio, And Vision},
author={Liu, Che and Zhang, Yingji and Zhang, Dong and Zhang, Weijie and Gong, Chenggong and Li, Haohan and Lu, Yu and Zhou, Shilin and Lu, Yue and Gan, Ziliang and others},
journal={arXiv preprint arXiv:2503.01879},
year={2025}
}
Usage and License Notices: The data and code are intended and licensed for research use only.
License: Attribution-NonCommercial 4.0 International It should abide by the policy of OpenAI: https://openai.com/policies/terms-of-use
7 commits
2 commits
📖Paper |🤗Model | 🤗Training Data (Coming Soon)
1
9 commits
1 linked in READMEs
updated Oct 20, 2025
Che Liu
,
Yingji Zhang
,
Dong Zhang
,
Weijie Zhang
,
Chenggong Gong
,
Yu Lu
,
Shilin Zhou
,
Ziliang Gan
,
Ziao Wang,
Haipang Wu,
Ji Liu,
Andre Freitas,
Qifan Wang,
Zenglin Xu,
Rongjunchen Zhang♠,
Yong Dai♠
NEXUS-O is an industry-scale omni-modal large language model (LLM) that unifies audio, vision, and language understanding into a single modular framework. Human perception integrates sight, sound, and language — NEXUS-O aims to replicate this ability for intelligent agents across real-world scenarios such as ASR, Speech-to-Speech Chat, and Multimodal Reasoning.
Architecture of NEXUS-O
Training Stages
@article{liu2025nexus,
title={Nexus: An Omni-Perceptive And-Interactive Model for Language, Audio, And Vision},
author={Liu, Che and Zhang, Yingji and Zhang, Dong and Zhang, Weijie and Gong, Chenggong and Li, Haohan and Lu, Yu and Zhou, Shilin and Lu, Yue and Gan, Ziliang and others},
journal={arXiv preprint arXiv:2503.01879},
year={2025}
}
Usage and License Notices: The data and code are intended and licensed for research use only.
License: Attribution-NonCommercial 4.0 International It should abide by the policy of OpenAI: https://openai.com/policies/terms-of-use
7 commits
2 commits