4 repos
microsoft/eureka-ml-insights
A framework for standardizing evaluations of large foundation models, beyond single-score reporting…
184
141 commits
evaleval/every_eval_ever
Every Eval Ever is a shared schema and crowdsourced eval database. It defines a standardized…
112
635 commits
stanford-crfm/helm
Holistic Evaluation of Language Models (HELM) is an open source Python framework created by the…
2,906
6,212 commits
microsoft/Eureka-Bench-Logs
No description
7
42 commits