11 repos
Evaluation frameworks and datasets for assessing large language models across diverse knowledge domains, centered on the MMLU (Massive Multitask Language Understanding) benchmark and related multi-task testing methodologies. Repositories in this cluster provide implementations of standardized evaluation protocols, dataset curation tools (via Argilla), and domain-specific test suites for measuring model performance across scientific, professional, and general knowledge areas. These resources support researchers and practitioners in quantifying LLM capabilities across broad task distributions.