12 repos
Embeddings and benchmarks for tasks that combine document analysis, visual question-answering, and scene understanding. This cluster focuses on systems that interpret and reason over documents, images, and text together—including VQA (Visual Question Answering) over documents, scene graph extraction, and multimodal reasoning across diverse visual inputs. The central repos (DocVQA, TextVQA, LLaVA-Bench) represent key benchmarks for evaluating how well models understand structured visual content like forms, receipts, and documents alongside natural images.