14 repos
Benchmarking frameworks and leaderboard systems for evaluating large language models across diverse capabilities and domains. This area encompasses tools for systematic performance comparison, dataset curation for evaluation (like UltraChat and UltraFeedback), and infrastructure for hosting competitive leaderboards—both general-purpose and domain-specific (finance, Turkish language). Repositories here focus on creating standardized ways to measure LLM quality, track model progress over time, and enable reproducible comparison across the rapidly evolving landscape of language models.