10 repos
Benchmarking frameworks and evaluation tools for vision-language models and multimodal large language models. This cluster focuses on assessing VLM/MLLM capabilities across diverse tasks, with several repos developing comprehensive benchmark suites (PhysBench for physics reasoning, MOSSBench for open-source model evaluation) and instruction-tuning datasets. Includes related infrastructure for retrieval-augmented generation and multimodal dataset construction to support rigorous model evaluation.