14 repos
Benchmark datasets and evaluation frameworks for testing vision-language models across diverse tasks including visual question answering, document understanding, and real-world image comprehension. The cluster centers on standardized datasets like TextVQA, DocVQA, MME, MMStar, and SEED-Bench that measure multimodal AI performance on reading text in images, understanding documents, and reasoning about visual content. These resources enable systematic evaluation of how well models combine visual perception with language understanding across different domains and difficulty levels.