VITA-MLLM/Long-VITA

✨✨Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuracy

306

stars

50

commits

Python

primary language

May 14, 2025

updated

long-context
mllm
vision-language-model
Browse cluster: Long-Context Vision-Language Models

README

Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuracy

:fire: News

  • 2025.02.27 🌟 We have an Oneline Demo now.
  • 2025.02.27 🌟 VLMEvalKit of OpenCompass has supported our Long-VITA.
  • 2025.02.17 🌟 We support training on DeepSpeed and inference on Transformer.
  • 2025.02.09 🌟 We support training and inference on Megatron.
  • 2025.02.05 🌟 We release training code, training log, deployment code, and model weights, which support MindSpeed.
  • 2024.02.05 🌟 We are proud to launch Long-VITA, a strong long-context visual language model supporting over one million tokens.

Contents

✨ Highlights

  • Long Context. Long-VITA can process more than 4K frames or over 1M visual tokens. It achieves state-of-the-art performance on Video-MME under 20B models.
  • Open Source. Long-VITA is trained on open-source data only, consisting of a mix of 17M samples that are publicly available.
  • Strong Performance. Long-VITA achieves competitive results on image and video understanding benchmarks among cutting-edge models under 20B parameters.

📈 Experimental Results

  • Comparison of image understanding.

image image

  • Comparison of video understanding.

image

image

  • Effectiveness of Logits-Masked LM Head.

image

🐍 Models

⭐ Training, Inference and Evaluation

We implemented Long-VITA on three frameworks.

Contributors

shenyunhang

49 commits

BradyFU

1 commits

VITA-MLLM/Long-VITA

✨✨Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuracy

306

stars

50

commits

Python

primary language

May 14, 2025

updated

long-context
mllm
vision-language-model
Browse cluster: Long-Context Vision-Language Models

README

Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuracy

:fire: News

  • 2025.02.27 🌟 We have an Oneline Demo now.
  • 2025.02.27 🌟 VLMEvalKit of OpenCompass has supported our Long-VITA.
  • 2025.02.17 🌟 We support training on DeepSpeed and inference on Transformer.
  • 2025.02.09 🌟 We support training and inference on Megatron.
  • 2025.02.05 🌟 We release training code, training log, deployment code, and model weights, which support MindSpeed.
  • 2024.02.05 🌟 We are proud to launch Long-VITA, a strong long-context visual language model supporting over one million tokens.

Contents

✨ Highlights

  • Long Context. Long-VITA can process more than 4K frames or over 1M visual tokens. It achieves state-of-the-art performance on Video-MME under 20B models.
  • Open Source. Long-VITA is trained on open-source data only, consisting of a mix of 17M samples that are publicly available.
  • Strong Performance. Long-VITA achieves competitive results on image and video understanding benchmarks among cutting-edge models under 20B parameters.

📈 Experimental Results

  • Comparison of image understanding.

image image

  • Comparison of video understanding.

image

image

  • Effectiveness of Logits-Masked LM Head.

image

🐍 Models

⭐ Training, Inference and Evaluation

We implemented Long-VITA on three frameworks.

Contributors

shenyunhang

49 commits

BradyFU

1 commits

Languages

Python

91.7%

Shell

8.3%