Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
50
8 commits
2 linked in READMEs
updated Jun 17, 2025
Shaolei Zhang, Shoutao Guo, Qingkai Fang, Yan Zhou, Yang Feng*
The introduction and usage of Stream-Omni refer to https://github.com/ictnlp/Stream-Omni.
Stream-Omni is an end-to-end language-vision-speech chatbot that simultaneously supports interaction across various modality combinations, with the following features💡:
| Microphone Input | File Input |
|---|---|
[!NOTE]
Stream-Omni can produce intermediate textual results (ASR transcription and text response) during speech interaction, offering users a seamless "see-while-hear" experience.
5 commits
3 commits
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
50
8 commits
2 linked in READMEs
updated Jun 17, 2025
Shaolei Zhang, Shoutao Guo, Qingkai Fang, Yan Zhou, Yang Feng*
The introduction and usage of Stream-Omni refer to https://github.com/ictnlp/Stream-Omni.
Stream-Omni is an end-to-end language-vision-speech chatbot that simultaneously supports interaction across various modality combinations, with the following features💡:
| Microphone Input | File Input |
|---|---|
[!NOTE]
Stream-Omni can produce intermediate textual results (ASR transcription and text response) during speech interaction, offering users a seamless "see-while-hear" experience.
5 commits
3 commits