[ECCV2024] Video Foundation Models & Data for Multimodal Understanding
2,382
stars
261
commits
Python
primary language
Jul 2, 2026
updated
This repo contains InternVideo series and related works in video foundation models.
2026.06: InternVideo3 is released with the technical report, 8B instruct model, long-video SFT dataset, evaluation scripts, and an initial video-agent implementation in Vidify.2025.12: InternVideo-Next is released with the technical report, pretrained model weights in the Hugging Face collection, and pretraining code.2025.01: InternVideo2.5 is now released! Check out the technical report for detailed insights, and access the model on HuggingFace.2024.08.12: We provide smaller models, InternVideo2-S/B/L, which are distilled from InternVideo2-1B. We also build smaller VideoCLIP with MobileCLIP.2024.08: InternVideo2-Stage3-8B and InternVideo2-Stage3-8B-HD are released. 8B indicates the use of InternVideo2-1B and the 7B LLM.2024.07: The video annotation for InternVid2 (HuggingFace) is released.2024.06: The full version of the video annotation (230M video-text pairs) for InternVid (OpenDataLab | HuggingFace) is released.2024.04: The Checkpoints and scripts for InternVideo2 are released.2024.03: The technical report of InternVideo2 is released.2024.01: InternVid (a video-text dataset for video understanding and generation) has been accepted for spotlight presentation of ICLR 2024.2023.07: A video-text dataset InternVid is released at here for facilitating multimodal understanding and generation.2023.05: Video instruction data are released at here for tuning end-to-end video-centric multimodal dialogue systems like VideoChat.2023.01: The code & models of InternVideo are released.2022.12: The technical report of InternVideo is released.2022.09: Press releases of InternVideo (official | 163 news | qq news).Python
94.5%
Shell
4.6%
[ECCV2024] Video Foundation Models & Data for Multimodal Understanding
2,382
stars
261
commits
Python
primary language
Jul 2, 2026
updated
This repo contains InternVideo series and related works in video foundation models.
2026.06: InternVideo3 is released with the technical report, 8B instruct model, long-video SFT dataset, evaluation scripts, and an initial video-agent implementation in Vidify.2025.12: InternVideo-Next is released with the technical report, pretrained model weights in the Hugging Face collection, and pretraining code.2025.01: InternVideo2.5 is now released! Check out the technical report for detailed insights, and access the model on HuggingFace.2024.08.12: We provide smaller models, InternVideo2-S/B/L, which are distilled from InternVideo2-1B. We also build smaller VideoCLIP with MobileCLIP.2024.08: InternVideo2-Stage3-8B and InternVideo2-Stage3-8B-HD are released. 8B indicates the use of InternVideo2-1B and the 7B LLM.2024.07: The video annotation for InternVid2 (HuggingFace) is released.2024.06: The full version of the video annotation (230M video-text pairs) for InternVid (OpenDataLab | HuggingFace) is released.2024.04: The Checkpoints and scripts for InternVideo2 are released.2024.03: The technical report of InternVideo2 is released.2024.01: InternVid (a video-text dataset for video understanding and generation) has been accepted for spotlight presentation of ICLR 2024.2023.07: A video-text dataset InternVid is released at here for facilitating multimodal understanding and generation.2023.05: Video instruction data are released at here for tuning end-to-end video-centric multimodal dialogue systems like VideoChat.2023.01: The code & models of InternVideo are released.2022.12: The technical report of InternVideo is released.2022.09: Press releases of InternVideo (official | 163 news | qq news).Python
94.5%
Shell
4.6%