NExTQA: nextqa_test.json
ID provided in the "image" field
Image:
Flickr30k: flickr30k_captions.json
(this is the standard 1k test set). ID provided in the "image" field.
TextVQA: textvqa.json
ID provided in the "image" field
GQA: testdev_balanced_questions_with_images.json
ID provided in the "image" field
Audio-visual:
How2: how2_test.json
ID provided in "image". Format: <video_id><start_second><end_second>.mp4 or .wav.
Audio-Visual Sound Source Detection (AVSSD): testdata_formatted.json
ID provided in the "image" field. The first one is image and the second one is the corresponding audio.
Audio Visual Matching (AVM): audiovisualmatching_combined.json
ID provided in the "image" field as a list of two values. The first one is the image and the second one is the audio/speech
Whether it is from VGGSS or is from SpokenCOCO is indicated in the ID as well
Audio-visual question answering (AVQA) Ego4D-QA: ego4d_qa.json
Video ID indicates the frame index
NExTQA: nextqa_test.json
ID provided in the "image" field
Image:
Flickr30k: flickr30k_captions.json
(this is the standard 1k test set). ID provided in the "image" field.
TextVQA: textvqa.json
ID provided in the "image" field
GQA: testdev_balanced_questions_with_images.json
ID provided in the "image" field
Audio-visual:
How2: how2_test.json
ID provided in "image". Format: <video_id><start_second><end_second>.mp4 or .wav.
Audio-Visual Sound Source Detection (AVSSD): testdata_formatted.json
ID provided in the "image" field. The first one is image and the second one is the corresponding audio.
Audio Visual Matching (AVM): audiovisualmatching_combined.json
ID provided in the "image" field as a list of two values. The first one is the image and the second one is the audio/speech
Whether it is from VGGSS or is from SpokenCOCO is indicated in the ID as well
Audio-visual question answering (AVQA) Ego4D-QA: ego4d_qa.json
Video ID indicates the frame index