Vision-Language Model Evaluation

10 repos

Evaluation frameworks and benchmarking tools for assessing the performance of vision-language models (VLMs) and multimodal large language models (LLMs). These Python-based repositories provide standardized evaluation metrics, datasets, and testing harnesses for measuring how well models understand and reason about visual and textual information together. Central projects like VLMEvalKit and lmms-eval-mmllm offer comprehensive evaluation pipelines, while supporting repositories investigate model architecture redundancy and specialized benchmarking scenarios.

Python · 8
Jupyter Notebook · 1