LLM Evaluation & Benchmarking

27 repos

Datasets, benchmarks, and evaluation frameworks for assessing large language models across diverse tasks and domains. This cluster covers standardized evaluation suites, reward modeling datasets, and multi-domain benchmarking efforts that help measure LLM capabilities in areas ranging from code generation to medical reasoning. The repositories provide both the infrastructure and data needed to systematically evaluate and compare LLM performance.

Python · 8
Jupyter Notebook · 2
HTML · 1
TypeScript · 1
evaluation ·7,303
benchmark ·7,303
llm-evaluation ·5,491
large-language-models ·4,681
vision-language-model ·4,498
vlm ·4,459
multimodal ·4,418
audio-evaluation ·4,409
agi ·4,409
multimodal-evaluation ·4,409