Vision-Language Model Benchmarking

8 repos

Benchmarking frameworks and evaluation datasets for vision-language models (VLMs), including retrieval-augmented generation (RAG) systems, visual document understanding, and multimodal video analysis. The cluster focuses on standardized evaluation methodologies and performance measurement across different VLM architectures and retrieval scenarios, with tools for assessing model capabilities on real-world document and video understanding tasks.

Python · 4
JavaScript · 2
mllm ·35
vlm ·35
oversensitivity ·21
alignment ·18
attack ·18
multimodal ·14
llm ·14
rag ·14
croissant ·6
safety-alignment ·3