4 repos
prometheus-eval/BiGGen-Bench
BIGGEN-Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models
17
15 commits
prometheus-eval/BiGGen-Bench-Results
BIGGEN-Bench Evaluation Results
12
28 commits
Quehry/HelloBench
HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models
60
9 commits
allenai/WildBench
🦁 WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
40
34 commits