LLM Evaluation and Leaderboards

14 repos

Benchmarking frameworks and leaderboard systems for evaluating large language models across diverse capabilities and domains. This area encompasses tools for systematic performance comparison, dataset curation for evaluation (like UltraChat and UltraFeedback), and infrastructure for hosting competitive leaderboards—both general-purpose and domain-specific (finance, Turkish language). Repositories here focus on creating standardized ways to measure LLM quality, track model progress over time, and enable reproducible comparison across the rapidly evolving landscape of language models.

Python · 6
Jupyter Notebook · 1
transformers ·6,028
rlhf ·5,676
llm ·5,676
generated_from_trainer ·352
llama ·352
conversational ·352
endpoints_compatible ·352
alignment-handbook ·352
safetensors ·352
text-generation ·352