VCR Visual Reasoning Datasets

16 repos

Visual Commonsense Reasoning (VCR) benchmark datasets for training and evaluating multimodal AI models on vision-language tasks. The cluster contains multiple language variants (English and Chinese) and difficulty levels (easy and hard) of the VCR dataset, which combines image understanding with reasoning about objects, relationships, and commonsense knowledge. These repositories represent standardized evaluation resources for researchers working on visual question answering and commonsense reasoning in computer vision.