OpenGVLab/VideoChat-Flash-Qwen2_5-7B-1M_res224

Model

2

stars

7

commits

2

repos using this model

2

linked in READMEs

May 16, 2025

updated

custom_code
endpoints_compatible
feature-extraction
model-index
multimodal
qwen2
safetensors
text-generation-inference
transformers
video-text-to-text
Browse cluster: Multimodal Vision-Language Models โ†’

README

๐ŸฆœVideoChat-Flash-Qwen2_5-7B-1M_res224โšก

[๐Ÿ“ฐ Blog] [๐Ÿ“‚ GitHub] [๐Ÿ“œ Tech Report] [๐Ÿ—จ๏ธ Chat Demo]

VideoChat-Flash-Qwen2_5-7B_InternVideo2-1B is constructed upon UMT-L (300M) and Qwen2.5-7B-1M, employing only 16 tokens per frame. By leveraging Yarn to extend the context window to 1M (Qwen2.5-7B-1M's native context window is 128k), our model supports input sequences of up to approximately 50,000 frames.

Note: Due to a predominantly English training corpus, the model only exhibits basic Chinese comprehension, to ensure optimal performance, using English for interaction is recommended.

๐Ÿ“ˆ Performance

ModelMVBenchLongVideoBenchVideoMME(w/o sub)Max input frames
VideoChat-Flash-Qwen2_5-2B@44870.058.357.010000
VideoChat-Flash-Qwen2-7B@22473.264.264.010000
VideoChat-Flash-Qwen2_5-7B-1M@22473.466.563.550000
VideoChat-Flash-Qwen2_5-7B_InternVideo2-1B@22474.364.565.110000
VideoChat-Flash-Qwen2-7B@44874.064.765.310000

๐Ÿš€ How to use the model

First, you need to install flash attention2 and some other modules. We provide a simple installation example below:

pip install transformers==4.40.1
pip install av
pip install imageio
pip install decord
pip install opencv-python
# optional
pip install flash-attn --no-build-isolation

Then you could use our model:

from transformers import AutoModel, AutoTokenizer
import torch

# model setting
model_path = 'OpenGVLab/VideoChat-Flash-Qwen2_5-7B-1M_res224'

tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
model = AutoModel.from_pretrained(model_path, trust_remote_code=True).to(torch.bfloat16).cuda()
image_processor = model.get_vision_tower().image_processor

mm_llm_compress = False # use the global compress or not
if mm_llm_compress:
    model.config.mm_llm_compress = True
    model.config.llm_compress_type = "uniform0_attention"
    model.config.llm_compress_layer_list = [4, 18]
    model.config.llm_image_token_ratio_list = [1, 0.75, 0.25]
else:
    model.config.mm_llm_compress = False

# evaluation setting
max_num_frames = 512
generation_config = dict(
    do_sample=False,
    temperature=0.0,
    max_new_tokens=1024,
    top_p=0.1,
    num_beams=1
)

video_path = "your_video.mp4"

# single-turn conversation
question1 = "Describe this video in detail."
output1, chat_history = model.chat(video_path=video_path, tokenizer=tokenizer, user_prompt=question1, return_history=True, max_num_frames=max_num_frames, generation_config=generation_config)

print(output1)

# multi-turn conversation
question2 = "How many people appear in the video?"
output2, chat_history = model.chat(video_path=video_path, tokenizer=tokenizer, user_prompt=question2, chat_history=chat_history, return_history=True, max_num_frames=max_num_frames, generation_config=generation_config)

print(output2)

โœ๏ธ Citation


@article{li2024videochatflash,
  title={VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling},
  author={Li, Xinhao and Wang, Yi and Yu, Jiashuo and Zeng, Xiangyu and Zhu, Yuhan and Huang, Haian and Gao, Jianfei and Li, Kunchang and He, Yinan and Wang, Chenting and others},
  journal={arXiv preprint arXiv:2501.00574},
  year={2024}
}

Contributors

lixinhao

7 commits

OpenGVLab/VideoChat-Flash-Qwen2_5-7B-1M_res224

Model

2

stars

7

commits

2

repos using this model

2

linked in READMEs

May 16, 2025

updated

custom_code
endpoints_compatible
feature-extraction
model-index
multimodal
qwen2
safetensors
text-generation-inference
transformers
video-text-to-text
Browse cluster: Multimodal Vision-Language Models โ†’

README

๐ŸฆœVideoChat-Flash-Qwen2_5-7B-1M_res224โšก

[๐Ÿ“ฐ Blog] [๐Ÿ“‚ GitHub] [๐Ÿ“œ Tech Report] [๐Ÿ—จ๏ธ Chat Demo]

VideoChat-Flash-Qwen2_5-7B_InternVideo2-1B is constructed upon UMT-L (300M) and Qwen2.5-7B-1M, employing only 16 tokens per frame. By leveraging Yarn to extend the context window to 1M (Qwen2.5-7B-1M's native context window is 128k), our model supports input sequences of up to approximately 50,000 frames.

Note: Due to a predominantly English training corpus, the model only exhibits basic Chinese comprehension, to ensure optimal performance, using English for interaction is recommended.

๐Ÿ“ˆ Performance

ModelMVBenchLongVideoBenchVideoMME(w/o sub)Max input frames
VideoChat-Flash-Qwen2_5-2B@44870.058.357.010000
VideoChat-Flash-Qwen2-7B@22473.264.264.010000
VideoChat-Flash-Qwen2_5-7B-1M@22473.466.563.550000
VideoChat-Flash-Qwen2_5-7B_InternVideo2-1B@22474.364.565.110000
VideoChat-Flash-Qwen2-7B@44874.064.765.310000

๐Ÿš€ How to use the model

First, you need to install flash attention2 and some other modules. We provide a simple installation example below:

pip install transformers==4.40.1
pip install av
pip install imageio
pip install decord
pip install opencv-python
# optional
pip install flash-attn --no-build-isolation

Then you could use our model:

from transformers import AutoModel, AutoTokenizer
import torch

# model setting
model_path = 'OpenGVLab/VideoChat-Flash-Qwen2_5-7B-1M_res224'

tokenizer = AutoTokenizer.from_pretrained(model_path, trust_remote_code=True)
model = AutoModel.from_pretrained(model_path, trust_remote_code=True).to(torch.bfloat16).cuda()
image_processor = model.get_vision_tower().image_processor

mm_llm_compress = False # use the global compress or not
if mm_llm_compress:
    model.config.mm_llm_compress = True
    model.config.llm_compress_type = "uniform0_attention"
    model.config.llm_compress_layer_list = [4, 18]
    model.config.llm_image_token_ratio_list = [1, 0.75, 0.25]
else:
    model.config.mm_llm_compress = False

# evaluation setting
max_num_frames = 512
generation_config = dict(
    do_sample=False,
    temperature=0.0,
    max_new_tokens=1024,
    top_p=0.1,
    num_beams=1
)

video_path = "your_video.mp4"

# single-turn conversation
question1 = "Describe this video in detail."
output1, chat_history = model.chat(video_path=video_path, tokenizer=tokenizer, user_prompt=question1, return_history=True, max_num_frames=max_num_frames, generation_config=generation_config)

print(output1)

# multi-turn conversation
question2 = "How many people appear in the video?"
output2, chat_history = model.chat(video_path=video_path, tokenizer=tokenizer, user_prompt=question2, chat_history=chat_history, return_history=True, max_num_frames=max_num_frames, generation_config=generation_config)

print(output2)

โœ๏ธ Citation


@article{li2024videochatflash,
  title={VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling},
  author={Li, Xinhao and Wang, Yi and Yu, Jiashuo and Zeng, Xiangyu and Zhu, Yuhan and Huang, Haian and Gao, Jianfei and Li, Kunchang and He, Yinan and Wang, Chenting and others},
  journal={arXiv preprint arXiv:2501.00574},
  year={2024}
}

Contributors

lixinhao

7 commits