27 repos
Datasets, benchmarks, and evaluation frameworks for assessing large language models across diverse tasks and domains. This cluster covers standardized evaluation suites, reward modeling datasets, and multi-domain benchmarking efforts that help measure LLM capabilities in areas ranging from code generation to medical reasoning. The repositories provide both the infrastructure and data needed to systematically evaluate and compare LLM performance.