5 repos
yuchenlin/ZeroEval
A simple unified framework for evaluating LLMs
274
258 commits
WildEval/ZeroEval
allenai/WildBench
Benchmarking LLMs with Challenging Tasks from Real Users
256
118 commits
zai-org/glm-simple-evals
GLM-SIMPLE-EVALS: The evaluation repository for the GLM-4.5 series of models by Z.ai.
43
13 commits
jankinf/llm_evals
A comprehensive evaluation framework for Large Language Models (LLMs), providing extensive…
5
96 commits