LLM and Vision-Language Model Evaluation

18 repos

Benchmark suites and evaluation frameworks for assessing the performance and capabilities of large language models and vision-language models across diverse domains. These repositories provide standardized evaluation datasets, metrics, and testing protocols—spanning medical imaging, code generation, multimodal understanding, and cyber security—enabling systematic comparison of model behavior and safety properties. Researchers and practitioners use these benchmarks to measure model improvements, identify failure modes, and validate alignment with real-world requirements.

Python · 6
Jupyter Notebook · 2
HTML · 1
TypeScript · 1
benchmark ·6,433
evaluation ·6,433
large-language-models ·4,671
llm-evaluation ·4,653
vision-language-model ·4,488
multimodal ·4,403
multimodal-evaluation ·4,399
agi ·4,399
audio-evaluation ·4,399
video-understanding ·4,399