DAMO-NLP-SG/Video-LLaMA-Series

Model

46

stars

8

commits

5

linked in READMEs

Jun 10, 2023

updated

visual-question-answering

README

Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

This is the Hugging Face repo for storing pre-trained & fine-tuned checkpoints of our Video-LLaMA, which is a multi-modal conversational large language model with video understanding capability.

Vision-Language Branch

CheckpointLinkNote
pretrain-vicuna7blinkPre-trained on WebVid (2.5M video-caption pairs) and LLaVA-CC3M (595k image-caption pairs)
finetune-vicuna7b-v2linkFine-tuned on the instruction-tuning data from MiniGPT-4, LLaVA and VideoChat
pretrain-vicuna13blinkPre-trained on WebVid (2.5M video-caption pairs) and LLaVA-CC3M (595k image-caption pairs)
finetune-vicuna13b-v2linkFine-tuned on the instruction-tuning data from MiniGPT-4, LLaVA and VideoChat
pretrain-ziya13b-zhlinkPre-trained with Chinese LLM Ziya-13B
finetune-ziya13b-zhlinkFine-tuned on machine-translated VideoChat instruction-following dataset (in Chinese)
pretrain-billa7b-zhlinkPre-trained with Chinese LLM BiLLA-7B
finetune-billa7b-zhlinkFine-tuned on machine-translated VideoChat instruction-following dataset (in Chinese)

Audio-Language Branch

CheckpointLinkNote
pretrain-vicuna7blinkPre-trained on WebVid (2.5M video-caption pairs) and LLaVA-CC3M (595k image-caption pairs)
finetune-vicuna7b-v2linkFine-tuned on the instruction-tuning data from MiniGPT-4, LLaVA and VideoChat

Usage

For launching the pre-trained Video-LLaMA on your own machine, please refer to our github repo.

Contributors

hangzhang-nlp

5 commits

lixin4ever

3 commits

DAMO-NLP-SG/Video-LLaMA-Series

Model

46

stars

8

commits

5

linked in READMEs

Jun 10, 2023

updated

visual-question-answering

README

Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

This is the Hugging Face repo for storing pre-trained & fine-tuned checkpoints of our Video-LLaMA, which is a multi-modal conversational large language model with video understanding capability.

Vision-Language Branch

CheckpointLinkNote
pretrain-vicuna7blinkPre-trained on WebVid (2.5M video-caption pairs) and LLaVA-CC3M (595k image-caption pairs)
finetune-vicuna7b-v2linkFine-tuned on the instruction-tuning data from MiniGPT-4, LLaVA and VideoChat
pretrain-vicuna13blinkPre-trained on WebVid (2.5M video-caption pairs) and LLaVA-CC3M (595k image-caption pairs)
finetune-vicuna13b-v2linkFine-tuned on the instruction-tuning data from MiniGPT-4, LLaVA and VideoChat
pretrain-ziya13b-zhlinkPre-trained with Chinese LLM Ziya-13B
finetune-ziya13b-zhlinkFine-tuned on machine-translated VideoChat instruction-following dataset (in Chinese)
pretrain-billa7b-zhlinkPre-trained with Chinese LLM BiLLA-7B
finetune-billa7b-zhlinkFine-tuned on machine-translated VideoChat instruction-following dataset (in Chinese)

Audio-Language Branch

CheckpointLinkNote
pretrain-vicuna7blinkPre-trained on WebVid (2.5M video-caption pairs) and LLaVA-CC3M (595k image-caption pairs)
finetune-vicuna7b-v2linkFine-tuned on the instruction-tuning data from MiniGPT-4, LLaVA and VideoChat

Usage

For launching the pre-trained Video-LLaMA on your own machine, please refer to our github repo.

Contributors

hangzhang-nlp

5 commits

lixin4ever

3 commits