MMLU Benchmarking & Evaluation

11 repos

Evaluation frameworks and datasets for assessing large language models across diverse knowledge domains, centered on the MMLU (Massive Multitask Language Understanding) benchmark and related multi-task testing methodologies. Repositories in this cluster provide implementations of standardized evaluation protocols, dataset curation tools (via Argilla), and domain-specific test suites for measuring model performance across scientific, professional, and general knowledge areas. These resources support researchers and practitioners in quantifying LLM capabilities across broad task distributions.

argilla ·86
hendrycks_test ·36
mmlu ·36
multi-task ·36
multitask ·36