7
stars
5
commits
10
repos using this model
2
linked in READMEs
Nov 8, 2023
updated
5 commits
omni-research/Tarsier2-Recap-7b
39
omni-research/Tarsier2-7b-0115
9
chenjoya/videollm-online-8b-v1plus
30
MAGAer13/mplug-owl2-llama2-7b
26
Vision-CAIR/vicuna-7b
25
kuleshov/llama-7b-4bit
12
allenai/llama2-7b-WildJailbreak
0
kuleshov/llama-7b-3bit
EvolvingLMMs-Lab/lmms-eval
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
4,399
EvolvingLMMs-Lab/Otter
🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo),…
3,438
orca-wm/Orca
Orca: The World is in Your Mind
1,020
facebookresearch/tuna-2
Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding…
755
vision-x-nyu/thinking-in-space
Official repo and evaluation implementation of VSI-Bench
740
wenhaochai/MovieChat
[CVPR 2024] MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
706
VITA-Group/VLM-3R
[CVPR 2026] VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
445
XiaomiMiMo/MiMo-Embodied
MiMo-Embodied
405
dongyh20/Octopus
[ECCV2024] 🐙Octopus, an embodied vision-language model trained with RLEF, emerging superior in…
301
jacklishufan/LaViDa
Official Implementation of LaViDa: :A Large Diffusion Language Model for Multimodal Understanding
229
Hritikbansal/videophy
Video Generation, Physical Commonsense, Semantic Adherence, VideoCon-Physics
207
Hritikbansal/videocon
58