50
stars
8
commits
3
repos using this model
2
linked in READMEs
Jan 16, 2025
updated
This repository contains the model of the paper VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.
Code: https://github.com/VITA-MLLM/VITA
4 commits
3 commits
1 commits
fushh7/ViSpeak-s3
1
fushh7/ViSpeak-s2
0
Vision-CAIR/MiniGPT4-Video
31
YaohuiW/LaVie
20
OpenGVLab/InternVideo2_chat_8B_HD
18
videocon/videocon-model
aurateam/AURA
13
Ricky06662/TaskRouter-1.5B
Tencent/VITA
The official implement of VITA, VITA15, LongVITA, VITA-Audio, VITA-VLA, and VITA-E.
166
HumanMLLM/ViSpeak
(ICCV2025) Official repository of paper "ViSpeak: Visual Instruction Feedback in Streaming Videos"
54
anisha0325/MUStReason
[LREC 2026] MUStReason: A Benchmark for Diagnosing Pragmatic Reasoning in Video-LMs for Multimodal…
VITA-MLLM/VITA
✨✨[NeurIPS 2025] VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
2,535