chakravarthi589/Video-Question-Answering_Resources

Video Question Answering | Video QA | VQA

See the code

README

Video-Question-Answering (VideoQA) Resources

The Video-Question-Answering-Resources repository is a curated guide for beginners and researchers interested in the Video Question Answering (VQA) field. It provides an organized collection of the most relevant papers, models, datasets, and additional resources to help users understand and contribute to this evolving area. The repository focuses on the intersection of computer vision and natural language processing, particularly how video data can be used to answer complex questions, offering a range of materials from introductory guides to advanced research. (Last Update on 09/22/2025)

Keywords:

Video question answering (VideoQA), LLMs, Long video understanding, Spatial Reasoning, Temporal Reasoning, Multi-Choice QA, Open-Ended QA;

Curators:

Bharatesh Chakravarthi, Ph.D
Joseph Raj Vishal



Beginners Guide to Video Question Answering

  1. Answering Questions from YouTube Videos with OpenAI Whisper and GPT-4 (Medium article)

  2. Try a quick example on how to use LLMs for Video Question Answering here (Check Additional Resources for API key)

  3. Community Computer Vision Course (Unit 4) MultiModal Models


Publications

Survey/Review Papers

Conference/Journal Papers

2026

2025

2024

2023

2022

2021

2020

2019

2018

2017

2016

2015


Datasets

YearNameKey Features
2025RoadSocialRoadSocial is a large-scale, diverse VideoQA resource for road events, derived from social media videos spanning 14M frames and 414K social comments, resulting in a dataset with 13.2K videos, 674 unique tags, and 260K high-quality QA pairs.
2025CrossVideoQACrossVideoQA is a person-centric cross-video QA benchmark combining EOSD (surveillance) and HACS (web actions). EOSD: 20 videos across 3 indoor locations over 12 dates (~450K frames), suited for multi-day behavior analysis. HACS: 50K web videos with 1.55M action clips, offering high visual and semantic diversity.
2025LVSQALVSQA is a long-video, scene-level QA dataset with 100 ≥30-minute videos (from LVBench) and 500 human-refined QA pairs designed from a purely visual perspective (minimal subtitle reliance). It targets detailed understanding—scene localization and fine-grained visual reasoning in long videos—using an MLLM-assisted, expert-edited creation pipeline.
2025DeVE-QAThe DeVE-QA is a dataset featuring 78𝐾 questions about 26𝐾 events on 10.6𝐾 long videos.
2025CogStreamCogStream features a collection of 6,361 videos from six public sources: MovieChat (40.2%), MECD (16.8%), QVhighlights (9.8%), VideoMME (6.5%), COIN (18.0%), and YouCook2 (8.6%). Scale: The final dataset comprises 1,088 high-quality videos and 59,032 QA pairs, formally split into a training set (852 videos) and a testing set (236 videos).
2024NExT-GQAThe NExT-GQA dataset augments the NExT-QA dataset with temporal labels for Causal (“why/how”), Temporal (“before/when/after”) type questions. The annotations are done in a weakly supervised setup by labeling validation and test sets. 8,911 QA pairs from 1,557 videos are annotated with 10,531 valid temporal segments.
2024MVBenchThe MVBench dataset focuses on evaluating multi-modal video understanding by covering 20 complex video tasks that emphasize temporal reasoning, from perception to cognition.The MVBench dataset includes over 566,747 video clips from diverse sources, such as COCO, WebVid, YouCook2, and more. The dataset also covers a wide variety of task types, such as question-answering, captioning, and conversation tasks, with more than 200 multiple-choice questions generated for each temporal understanding task
2024LVBenchThe LVBench dataset consists of 103 videos, each with a minimum duration of 30 minutes. There are a total of 1549 question-answer pairs associated with these videos, with an average of 24 questions per hour of video content
2024FunQAFunQA is a video question-answering dataset featuring 4.3K counter-intuitive and humorous video clips with 312K free-text QA pairs, an average answer length of 34.2 words, and subsets like HumorQA, CreativeQA, and MagicQA highlighting humor, creativity, and magic-themed reasoning
2024MedVidQAMedVidQA dataset comprises 3,010 human-annotated instructional questions and visual answers from 900 health-related videos.This dataset forms apart of the challenge of two tasks,medical instructions question generation and Video Corpus Visual Answer Localization (VCVAL).
2024Video-MMEThe Video-MME is a comprehensive benchmark designed to evaluate Multi-Modal Large Language Models (MLLMs) in video analysis.Covers short (< 2min), medium (4-15min), and long (30-60min) videos to test MLLMs' ability to process varying time frames. This includes 6 primary domains, such as Knowledge, Film and TV, Sports, Life Records, and Multilingualism, with 30 subfields, ensuring broad generalizability. Integrates video frames, subtitles, and audio.
2024CinePileThe CinePile dataset consists of 9,396 movie clips sourced from the Movieclips YouTube channel, divided into training and testing splits of 9,248 and 148 videos, respectively. Through a question-answer generation and filtering pipeline, the dataset produced 298,888 training points and 4,940 test-set points, averaging 32 questions per video scene.
2024LongVideoBenchLongVideoBench is a long-context video–language QA benchmark with interleaved inputs up to 1 hour, comprising 3,763 web-collected videos (with subtitles) and 6,678 human-annotated multiple-choice questions across 17 categories; it introduces referring reasoning, which requires retrieving and reasoning over detailed, temporally grounded contexts from lengthy inputs.
2023TextVRThe TextVR dataset is a large-scale cross-modal video retrieval dataset, containing 42,200 sentence queries for 10,500 videos across eight scenario domains, including Street View, Game, Sports, Driving, Activity, TV Show, and Cooking.
2023Social-IQ-2.0This dataset is from the Social IQ challenge, consisting of 1000 videos,6000 questions and 24,000 answers. This challenge was co-hosted with the Artificial Social Intelligence Workshop at ICCV'23
2023VideoChatVideoChat is a video-centric multimodal instruction data based on WebVid-10M. The project features a 100K video-instruction dataset created using human-assisted and semi-automatic annotation techniques.
2022Ego4DEgo4D is a comprehensive egocentric video dataset comprising 3,670 hours of daily-life activities recorded by 931 camera wearers across 74 locations in 9 countries, covering various scenarios like household, outdoor, and workplace settings.
2022NEWSKVQANEWSKVQA is a new dataset of 12K news videos spanning across 156 hours with 1M multiple-choice question-answer pairs covering 8263 unique entities.
2022MedVidQACLThis dataset consists of Medical Instructional Videos and Questions based on those videos consists of 899 Videos each of 4 mins and 3K Questions manually annotated
2022FIBERThe FIBER dataset consists of 28,000 videos and description.The dataset consists of MCQ-type questions as well as video captioning data.Consists of 28K videos and 28K questions each of 10 seconds duration
2022Causal-VidQAThis dataset consists of 26K Videos with 107K questions.Manually annotated.
2022MUSIC-AVQAThis dataset consists of 9.3K Music Video each 60s long with 45K Manually annotated.
2022VQuADThis dataset consists of 7K videos with 1.3Million Questions offering spatial and temporal properties.It consist of Synthetic Videos
2022CRIPP-VQACRIPP-VQA is a VideoQA dataset for counterfactual reasoning about implicit physical properties, containing 4,000 training, 500 validation, and 500 test videos, plus ≈2,000 videos for out-of-distribution evaluation. The training split includes 41,761 descriptive questions, 41,761 counterfactual questions, and 10,440 planning-based questions.
2022STARThe STAR is a dataset for Situated Reasoning, which provides challenging question-answering tasks, symbolic situation descriptions and logic-grounded diagnosis via real-world video situations. It consists of 4 Question Types, 60K Situated Questions , 23K Situation Video Clips and 140K Situation Hypergraphs
2022In-the-WildThis consists of dataset with videos recorded outdoors(survival , agriculture,natural disaster and military),Consists of 369 videos with 916 questions each about a minute and 10 seconds long.
2022AGQA 2.0AGQA 2.0 is the succeeding dataset of AGQA.With this dataset, there exists a benchmark of 96.85M question-answer pairs and a balanced subset of 2.27M question-answer pairs
2022WebVidVQA3M DataConsists of Web Videos 2M with 3M Question and each video 4 mins long.This consists of automatically tagged videos
2021HowToVQA69M DataConsists of 69M videos, with 69M questions each video being 2 minutes long,This consists of manually tagged videos
2021iVQAThis dataset consists of 10K videos with 10K questions each of 8 minutes long.
2021PanoAVQAPanoAVQA dataset consists of 360 degree panoramic videos.Consists of a total of 5.4K videos and 20K spatial and 31.7K audio-video QAs.
2021AGQAAction Genome Question Answering (AGQA) is a benchmark for compositional spatiotemporal reasoning. AGQA contains 192M unbalanced question-answer pairs for 9.6K videos. It also contains a balanced subset of 3.9M question-answer pairs
2021Video-QAPVideoQAP dataset consists of Web videos consisting of 35K Videos and 162K Questions each 36.2 seconds long
2021KnowIT-X-VQAAn extension of KnowIT dataset.this dataset consists of TV videos (12.1K) and 21.4K questions.
2021Charades-SRL-QAThese consists of Charades , HomeMade videos with 9.5K videos with 71K questions each 29 seconds long.
2021NExTQAThe NExT-QA dataset comprises 5,440 videos, split into 3,870 for training, 570 for validation, and 1,000 for testing. It features around 52,044 question-answer pairs, with approximately 47,692 for multiple-choice QA and 52,044 for open-ended QA. The questions are divided into three main types: causal questions (48% of the dataset), temporal questions (29%), and descriptive questions (23%).
2021LSMDC-QA (Requires request access)LSMDC-QA (Large Scale Movie Description Challenge) contains 118,081 short video clips extracted from 202 movies. It consists of 7408 clips, and evaluation is performed on a test set of 1000 videos from movies disjoint.
2021Env-QAEnv-QA consists of 23.3K videos collected in AI2-THOR simulator and 85.1K questions
2021SUTD-TrafficQASUTD-TrafficQA takes the form of videoQA based dataset consists of 10080 in the wild videos and annotated 62535 QA pairs for complex traffic based scenarios.
2020CLEVRERCLEVRER focuses on temporal reasoning and inferencing of synthetic videos. Consists of 10K videos 305K Questions each of 5 seconds duration
2020LifeQALifeQA consists of videos of data-to-data activities.It consists 275 video clips and over 2.3k multiple-choice questions.
2020How2R-and-How2QAThe How2R and How2QA datasets contain 9,371 and 9,035 episodes, with 24,328 and 21,509 clips averaging around 17 seconds each, divided into training, validation, and testing sets.
2021TGIF-QA-RThis dataset consists of 71K GIFs each of 3 seconds long and 165K Questions.Extended version of the TGIF-QA dataset.
2020DramaQADramaQA dataset is built upon the TV drama "Another Miss Oh" and it contains 17,983 QA pairs from 23,928 various-length video clips, with each QA pair belonging to one of four difficulty levels.
2020KnowITVQAKnowITVQA is a video dataset with 24,282 human-generated question-answer pairs about The Big Bang Theory consisting of 207 videos each 20 minutes long
2020V2C-QAVideo2CommonSense dataset consists of Web Videos for image captioning and VideoQA consists of 1.5K videos and 37K questions.
2020PsTuts-VQAThe PsTuts dataset includes the following resources: 76 videos (5.6 hours in total), 17,768 question-answer pairs, and a domain knowledge-base with 1,236 entities and 2,196 options.It focuses on video tutorials
2019Social-QAThis Kaggle repository consists of the Social-QA dataset. Social-IQ contains 1,250 natural in-the-wild social situations, 7,500 questions and 52,500 correct and incorrect answers.
2019TutorialVQADTutorialVQAD consists of tutorial pertaining to image editing software. Total number of videos 76 and total number of questions 6195
2019AVSD (Audio-Visual Scene-Aware Dialog)AVSD is a dialog dataset grounded in the Charades human-activity videos, comprising dialogs about 11,816 short indoor videos (avg. length ~30 s; at least 2 actions per video). Each dialog discusses the video’s events and objects across multiple turns.
2019Moments in Time DatasetThe Moments in Time dataset consists of one million videos, each 3 seconds long, with 339 different classes.
2018TVQATVQA is a large-scale video question-answering dataset built from six popular TV shows, including Friends, The Big Bang Theory, and How I Met Your Mother. It contains 152.5K QA pairs sourced from 21.8K video clips, covering over 460 hours of content.
2018SVQASVQA dataset consists of Attribute comparison, count, integer comparison, exist and query type questions.This consists of synthetic videos almost 12K and 118K Questions
2018YouCook2YouCook2 is one of the largest instructional video datasets focused on task-oriented cooking, featuring 2,000 untrimmed videos from 89 recipes, with an average of 22 videos per recipe. Each video, averaging 5.26 minutes and totalling 176 hours, includes annotated procedure steps with their corresponding temporal boundaries.
2018TVQA+TVQA+ includes 29.4K multiple-choice questions grounded in both temporal and spatial domains. A set of visual concept words—objects and people—are identified to collect spatial groundings, and corresponding object regions in individual frames are annotated with bounding boxes.
2017TGIF-QATGIF-QA, a large-scale dataset, contains 165K question-answer pairs based on animated GIFs, testing video-based Visual Question Answering (VQA) across four question types: Repetition Count, Repeating Action, State Transition, and Frame QA.
2017MarioQAMarioQA is a dataset specifically designed for video-based question-answering in the context of Super Mario Bros. gameplay, containing over 70,000 question-answer pairs linked to gameplay footage.
2017VideoQAVideoQA dataset of 18100 automatically crawled user-generated videos and titles.Videos collected from web videos with 174k questions each of 90s each
2017Something-Something v1 & v2Something-Something is a collection of 220,847 labelled video clips of humans performing predefined basic actions with everyday objects. The dataset comprises 220,847 videos divided into a training set of 168,913, a validation set of 24,777, and a test set of 27,157 (without labels), totalling 174 unique labels.
2016MSVD-QAThe MSVD-QA dataset is a Video Question Answering (VideoQA) dataset derived from the Microsoft Research Video Description (MSVD) dataset, which includes around 120K sentences describing over 2,000 videos snippets. The dataset includes 1,970 video clips and approximately 50.5K QA pairs.
2016MSRVTT-QAMSRVTT-QA consists of 10K web video clips with a total duration of 41.2 hours. It spans 200k clip-sentence pairs. Each video clip is annotated with about 20 natural sentences.
2016MovieQAThe MovieQA dataset is designed for movie question answering, aimed at evaluating automatic story comprehension through both video and text. It contains nearly 15,000 multiple-choice questions derived from over 400 movies.
2016PororoQAThe Pororo dataset based on children's cartoons features a simple story structure with episodes averaging 7.2 minutes, where similar events are frequently repeated. The dataset comprises 8,834 QA pairs, with an average of 51.66 questions per episode, excluding ambiguous or unrelated questions.
2015VideoQA(FIB)This dataset consists of VideoQA, from multiple sources with videos 109K video clips and duration of over 1000 hours with 390744 questions.
2014Activity NetActivityNet is a large-scale video benchmark for human activity understanding. ActivityNet aims to cover a wide range of complex human activities. ActivityNet provides samples from 203 activity classes with an average of 137 untrimmed videos per class and 1.41 activity instances per video, for a total of 849 video hours.
2013YouTube2Text-QAYouTube2Text data consists of 1987 videos with 122708 descriptions.These include short descriptions of videos.

Models

Open Source Models

Model NameLinks
InternVLHugging Face , GitHub
LLaVaHugging Face , GitHub
LITAGitHub
End2End ChatBotHugging Face , GitHub
VideoLLAMA2Hugging Face, GitHub
FrozenBiLMGitHub
PercieverIOHugging Face,GitHub
InstructBlipVideoHugging Face , GitHub
VideoGPTHugging Face,GitHub
Qwen2-VLHugging Face,GitHub
ViLAGitHub
LongVLMGitHub

Closed Source Models

Model NameAPI Link
ChatGPTHere
GeminiHere
Llama 3.2Here

Additional-Resources

  1. OpenAI Docs
  2. Gemini Docs
  3. LLAMA Docs
  4. Azure Samples

:arrow_heading_up: Back to Top

aaai
acm
arxiv-papers
cvpr
eccv
iccv
ieee
neurips
video-query
video-question-answering
video-question-answering-dataset
video-questions
vqa
vqa-dataset
wacv

Contributors

joe-rabbit

126 commits

chakravarthi589

98 commits

monamavani

4 commits

chakravarthi589/Video-Question-Answering_Resources

Video Question Answering | Video QA | VQA

See the code

README

Video-Question-Answering (VideoQA) Resources

The Video-Question-Answering-Resources repository is a curated guide for beginners and researchers interested in the Video Question Answering (VQA) field. It provides an organized collection of the most relevant papers, models, datasets, and additional resources to help users understand and contribute to this evolving area. The repository focuses on the intersection of computer vision and natural language processing, particularly how video data can be used to answer complex questions, offering a range of materials from introductory guides to advanced research. (Last Update on 09/22/2025)

Keywords:

Video question answering (VideoQA), LLMs, Long video understanding, Spatial Reasoning, Temporal Reasoning, Multi-Choice QA, Open-Ended QA;

Curators:

Bharatesh Chakravarthi, Ph.D
Joseph Raj Vishal



Beginners Guide to Video Question Answering

  1. Answering Questions from YouTube Videos with OpenAI Whisper and GPT-4 (Medium article)

  2. Try a quick example on how to use LLMs for Video Question Answering here (Check Additional Resources for API key)

  3. Community Computer Vision Course (Unit 4) MultiModal Models


Publications

Survey/Review Papers

Conference/Journal Papers

2026

2025

2024

2023

2022

2021

2020

2019

2018

2017

2016

2015


Datasets

YearNameKey Features
2025RoadSocialRoadSocial is a large-scale, diverse VideoQA resource for road events, derived from social media videos spanning 14M frames and 414K social comments, resulting in a dataset with 13.2K videos, 674 unique tags, and 260K high-quality QA pairs.
2025CrossVideoQACrossVideoQA is a person-centric cross-video QA benchmark combining EOSD (surveillance) and HACS (web actions). EOSD: 20 videos across 3 indoor locations over 12 dates (~450K frames), suited for multi-day behavior analysis. HACS: 50K web videos with 1.55M action clips, offering high visual and semantic diversity.
2025LVSQALVSQA is a long-video, scene-level QA dataset with 100 ≥30-minute videos (from LVBench) and 500 human-refined QA pairs designed from a purely visual perspective (minimal subtitle reliance). It targets detailed understanding—scene localization and fine-grained visual reasoning in long videos—using an MLLM-assisted, expert-edited creation pipeline.
2025DeVE-QAThe DeVE-QA is a dataset featuring 78𝐾 questions about 26𝐾 events on 10.6𝐾 long videos.
2025CogStreamCogStream features a collection of 6,361 videos from six public sources: MovieChat (40.2%), MECD (16.8%), QVhighlights (9.8%), VideoMME (6.5%), COIN (18.0%), and YouCook2 (8.6%). Scale: The final dataset comprises 1,088 high-quality videos and 59,032 QA pairs, formally split into a training set (852 videos) and a testing set (236 videos).
2024NExT-GQAThe NExT-GQA dataset augments the NExT-QA dataset with temporal labels for Causal (“why/how”), Temporal (“before/when/after”) type questions. The annotations are done in a weakly supervised setup by labeling validation and test sets. 8,911 QA pairs from 1,557 videos are annotated with 10,531 valid temporal segments.
2024MVBenchThe MVBench dataset focuses on evaluating multi-modal video understanding by covering 20 complex video tasks that emphasize temporal reasoning, from perception to cognition.The MVBench dataset includes over 566,747 video clips from diverse sources, such as COCO, WebVid, YouCook2, and more. The dataset also covers a wide variety of task types, such as question-answering, captioning, and conversation tasks, with more than 200 multiple-choice questions generated for each temporal understanding task
2024LVBenchThe LVBench dataset consists of 103 videos, each with a minimum duration of 30 minutes. There are a total of 1549 question-answer pairs associated with these videos, with an average of 24 questions per hour of video content
2024FunQAFunQA is a video question-answering dataset featuring 4.3K counter-intuitive and humorous video clips with 312K free-text QA pairs, an average answer length of 34.2 words, and subsets like HumorQA, CreativeQA, and MagicQA highlighting humor, creativity, and magic-themed reasoning
2024MedVidQAMedVidQA dataset comprises 3,010 human-annotated instructional questions and visual answers from 900 health-related videos.This dataset forms apart of the challenge of two tasks,medical instructions question generation and Video Corpus Visual Answer Localization (VCVAL).
2024Video-MMEThe Video-MME is a comprehensive benchmark designed to evaluate Multi-Modal Large Language Models (MLLMs) in video analysis.Covers short (< 2min), medium (4-15min), and long (30-60min) videos to test MLLMs' ability to process varying time frames. This includes 6 primary domains, such as Knowledge, Film and TV, Sports, Life Records, and Multilingualism, with 30 subfields, ensuring broad generalizability. Integrates video frames, subtitles, and audio.
2024CinePileThe CinePile dataset consists of 9,396 movie clips sourced from the Movieclips YouTube channel, divided into training and testing splits of 9,248 and 148 videos, respectively. Through a question-answer generation and filtering pipeline, the dataset produced 298,888 training points and 4,940 test-set points, averaging 32 questions per video scene.
2024LongVideoBenchLongVideoBench is a long-context video–language QA benchmark with interleaved inputs up to 1 hour, comprising 3,763 web-collected videos (with subtitles) and 6,678 human-annotated multiple-choice questions across 17 categories; it introduces referring reasoning, which requires retrieving and reasoning over detailed, temporally grounded contexts from lengthy inputs.
2023TextVRThe TextVR dataset is a large-scale cross-modal video retrieval dataset, containing 42,200 sentence queries for 10,500 videos across eight scenario domains, including Street View, Game, Sports, Driving, Activity, TV Show, and Cooking.
2023Social-IQ-2.0This dataset is from the Social IQ challenge, consisting of 1000 videos,6000 questions and 24,000 answers. This challenge was co-hosted with the Artificial Social Intelligence Workshop at ICCV'23
2023VideoChatVideoChat is a video-centric multimodal instruction data based on WebVid-10M. The project features a 100K video-instruction dataset created using human-assisted and semi-automatic annotation techniques.
2022Ego4DEgo4D is a comprehensive egocentric video dataset comprising 3,670 hours of daily-life activities recorded by 931 camera wearers across 74 locations in 9 countries, covering various scenarios like household, outdoor, and workplace settings.
2022NEWSKVQANEWSKVQA is a new dataset of 12K news videos spanning across 156 hours with 1M multiple-choice question-answer pairs covering 8263 unique entities.
2022MedVidQACLThis dataset consists of Medical Instructional Videos and Questions based on those videos consists of 899 Videos each of 4 mins and 3K Questions manually annotated
2022FIBERThe FIBER dataset consists of 28,000 videos and description.The dataset consists of MCQ-type questions as well as video captioning data.Consists of 28K videos and 28K questions each of 10 seconds duration
2022Causal-VidQAThis dataset consists of 26K Videos with 107K questions.Manually annotated.
2022MUSIC-AVQAThis dataset consists of 9.3K Music Video each 60s long with 45K Manually annotated.
2022VQuADThis dataset consists of 7K videos with 1.3Million Questions offering spatial and temporal properties.It consist of Synthetic Videos
2022CRIPP-VQACRIPP-VQA is a VideoQA dataset for counterfactual reasoning about implicit physical properties, containing 4,000 training, 500 validation, and 500 test videos, plus ≈2,000 videos for out-of-distribution evaluation. The training split includes 41,761 descriptive questions, 41,761 counterfactual questions, and 10,440 planning-based questions.
2022STARThe STAR is a dataset for Situated Reasoning, which provides challenging question-answering tasks, symbolic situation descriptions and logic-grounded diagnosis via real-world video situations. It consists of 4 Question Types, 60K Situated Questions , 23K Situation Video Clips and 140K Situation Hypergraphs
2022In-the-WildThis consists of dataset with videos recorded outdoors(survival , agriculture,natural disaster and military),Consists of 369 videos with 916 questions each about a minute and 10 seconds long.
2022AGQA 2.0AGQA 2.0 is the succeeding dataset of AGQA.With this dataset, there exists a benchmark of 96.85M question-answer pairs and a balanced subset of 2.27M question-answer pairs
2022WebVidVQA3M DataConsists of Web Videos 2M with 3M Question and each video 4 mins long.This consists of automatically tagged videos
2021HowToVQA69M DataConsists of 69M videos, with 69M questions each video being 2 minutes long,This consists of manually tagged videos
2021iVQAThis dataset consists of 10K videos with 10K questions each of 8 minutes long.
2021PanoAVQAPanoAVQA dataset consists of 360 degree panoramic videos.Consists of a total of 5.4K videos and 20K spatial and 31.7K audio-video QAs.
2021AGQAAction Genome Question Answering (AGQA) is a benchmark for compositional spatiotemporal reasoning. AGQA contains 192M unbalanced question-answer pairs for 9.6K videos. It also contains a balanced subset of 3.9M question-answer pairs
2021Video-QAPVideoQAP dataset consists of Web videos consisting of 35K Videos and 162K Questions each 36.2 seconds long
2021KnowIT-X-VQAAn extension of KnowIT dataset.this dataset consists of TV videos (12.1K) and 21.4K questions.
2021Charades-SRL-QAThese consists of Charades , HomeMade videos with 9.5K videos with 71K questions each 29 seconds long.
2021NExTQAThe NExT-QA dataset comprises 5,440 videos, split into 3,870 for training, 570 for validation, and 1,000 for testing. It features around 52,044 question-answer pairs, with approximately 47,692 for multiple-choice QA and 52,044 for open-ended QA. The questions are divided into three main types: causal questions (48% of the dataset), temporal questions (29%), and descriptive questions (23%).
2021LSMDC-QA (Requires request access)LSMDC-QA (Large Scale Movie Description Challenge) contains 118,081 short video clips extracted from 202 movies. It consists of 7408 clips, and evaluation is performed on a test set of 1000 videos from movies disjoint.
2021Env-QAEnv-QA consists of 23.3K videos collected in AI2-THOR simulator and 85.1K questions
2021SUTD-TrafficQASUTD-TrafficQA takes the form of videoQA based dataset consists of 10080 in the wild videos and annotated 62535 QA pairs for complex traffic based scenarios.
2020CLEVRERCLEVRER focuses on temporal reasoning and inferencing of synthetic videos. Consists of 10K videos 305K Questions each of 5 seconds duration
2020LifeQALifeQA consists of videos of data-to-data activities.It consists 275 video clips and over 2.3k multiple-choice questions.
2020How2R-and-How2QAThe How2R and How2QA datasets contain 9,371 and 9,035 episodes, with 24,328 and 21,509 clips averaging around 17 seconds each, divided into training, validation, and testing sets.
2021TGIF-QA-RThis dataset consists of 71K GIFs each of 3 seconds long and 165K Questions.Extended version of the TGIF-QA dataset.
2020DramaQADramaQA dataset is built upon the TV drama "Another Miss Oh" and it contains 17,983 QA pairs from 23,928 various-length video clips, with each QA pair belonging to one of four difficulty levels.
2020KnowITVQAKnowITVQA is a video dataset with 24,282 human-generated question-answer pairs about The Big Bang Theory consisting of 207 videos each 20 minutes long
2020V2C-QAVideo2CommonSense dataset consists of Web Videos for image captioning and VideoQA consists of 1.5K videos and 37K questions.
2020PsTuts-VQAThe PsTuts dataset includes the following resources: 76 videos (5.6 hours in total), 17,768 question-answer pairs, and a domain knowledge-base with 1,236 entities and 2,196 options.It focuses on video tutorials
2019Social-QAThis Kaggle repository consists of the Social-QA dataset. Social-IQ contains 1,250 natural in-the-wild social situations, 7,500 questions and 52,500 correct and incorrect answers.
2019TutorialVQADTutorialVQAD consists of tutorial pertaining to image editing software. Total number of videos 76 and total number of questions 6195
2019AVSD (Audio-Visual Scene-Aware Dialog)AVSD is a dialog dataset grounded in the Charades human-activity videos, comprising dialogs about 11,816 short indoor videos (avg. length ~30 s; at least 2 actions per video). Each dialog discusses the video’s events and objects across multiple turns.
2019Moments in Time DatasetThe Moments in Time dataset consists of one million videos, each 3 seconds long, with 339 different classes.
2018TVQATVQA is a large-scale video question-answering dataset built from six popular TV shows, including Friends, The Big Bang Theory, and How I Met Your Mother. It contains 152.5K QA pairs sourced from 21.8K video clips, covering over 460 hours of content.
2018SVQASVQA dataset consists of Attribute comparison, count, integer comparison, exist and query type questions.This consists of synthetic videos almost 12K and 118K Questions
2018YouCook2YouCook2 is one of the largest instructional video datasets focused on task-oriented cooking, featuring 2,000 untrimmed videos from 89 recipes, with an average of 22 videos per recipe. Each video, averaging 5.26 minutes and totalling 176 hours, includes annotated procedure steps with their corresponding temporal boundaries.
2018TVQA+TVQA+ includes 29.4K multiple-choice questions grounded in both temporal and spatial domains. A set of visual concept words—objects and people—are identified to collect spatial groundings, and corresponding object regions in individual frames are annotated with bounding boxes.
2017TGIF-QATGIF-QA, a large-scale dataset, contains 165K question-answer pairs based on animated GIFs, testing video-based Visual Question Answering (VQA) across four question types: Repetition Count, Repeating Action, State Transition, and Frame QA.
2017MarioQAMarioQA is a dataset specifically designed for video-based question-answering in the context of Super Mario Bros. gameplay, containing over 70,000 question-answer pairs linked to gameplay footage.
2017VideoQAVideoQA dataset of 18100 automatically crawled user-generated videos and titles.Videos collected from web videos with 174k questions each of 90s each
2017Something-Something v1 & v2Something-Something is a collection of 220,847 labelled video clips of humans performing predefined basic actions with everyday objects. The dataset comprises 220,847 videos divided into a training set of 168,913, a validation set of 24,777, and a test set of 27,157 (without labels), totalling 174 unique labels.
2016MSVD-QAThe MSVD-QA dataset is a Video Question Answering (VideoQA) dataset derived from the Microsoft Research Video Description (MSVD) dataset, which includes around 120K sentences describing over 2,000 videos snippets. The dataset includes 1,970 video clips and approximately 50.5K QA pairs.
2016MSRVTT-QAMSRVTT-QA consists of 10K web video clips with a total duration of 41.2 hours. It spans 200k clip-sentence pairs. Each video clip is annotated with about 20 natural sentences.
2016MovieQAThe MovieQA dataset is designed for movie question answering, aimed at evaluating automatic story comprehension through both video and text. It contains nearly 15,000 multiple-choice questions derived from over 400 movies.
2016PororoQAThe Pororo dataset based on children's cartoons features a simple story structure with episodes averaging 7.2 minutes, where similar events are frequently repeated. The dataset comprises 8,834 QA pairs, with an average of 51.66 questions per episode, excluding ambiguous or unrelated questions.
2015VideoQA(FIB)This dataset consists of VideoQA, from multiple sources with videos 109K video clips and duration of over 1000 hours with 390744 questions.
2014Activity NetActivityNet is a large-scale video benchmark for human activity understanding. ActivityNet aims to cover a wide range of complex human activities. ActivityNet provides samples from 203 activity classes with an average of 137 untrimmed videos per class and 1.41 activity instances per video, for a total of 849 video hours.
2013YouTube2Text-QAYouTube2Text data consists of 1987 videos with 122708 descriptions.These include short descriptions of videos.

Models

Open Source Models

Model NameLinks
InternVLHugging Face , GitHub
LLaVaHugging Face , GitHub
LITAGitHub
End2End ChatBotHugging Face , GitHub
VideoLLAMA2Hugging Face, GitHub
FrozenBiLMGitHub
PercieverIOHugging Face,GitHub
InstructBlipVideoHugging Face , GitHub
VideoGPTHugging Face,GitHub
Qwen2-VLHugging Face,GitHub
ViLAGitHub
LongVLMGitHub

Closed Source Models

Model NameAPI Link
ChatGPTHere
GeminiHere
Llama 3.2Here

Additional-Resources

  1. OpenAI Docs
  2. Gemini Docs
  3. LLAMA Docs
  4. Azure Samples

:arrow_heading_up: Back to Top

aaai
acm
arxiv-papers
cvpr
eccv
iccv
ieee
neurips
video-query
video-question-answering
video-question-answering-dataset
video-questions
vqa
vqa-dataset
wacv

Contributors

joe-rabbit

126 commits

chakravarthi589

98 commits

monamavani

4 commits