Multimodal LLM Evaluation & Benchmarking

10 repos

Benchmarking frameworks and evaluation tools for vision-language models and multimodal large language models. This cluster focuses on assessing VLM/MLLM capabilities across diverse tasks, with several repos developing comprehensive benchmark suites (PhysBench for physics reasoning, MOSSBench for open-source model evaluation) and instruction-tuning datasets. Includes related infrastructure for retrieval-augmented generation and multimodal dataset construction to support rigorous model evaluation.

Python · 8
JavaScript · 2
vlm ·2,214
mllm ·1,522
multimodal ·938
benchmark ·880
rag ·698
embedding ·684
image-retrieval ·684
contrastive-learning ·684
mmeb ·684
representation-learning ·684