8 repos
ashishsinha1602/dataset-integrity-audit
Measurements of public ML infrastructure: a census of all 1.46M Hugging Face Spaces, clone…
1
25 commits
TIGER-Lab/ClawBench
ClawBench — Leaderboard
0
17 commits
canyuchen/clinicalbench-results
No description
0 commits
vyang472/five-bugs
A 10-minute smoke test for AI coding agents. Five seeded Python bugs, one checker the agent sees…
5 commits
cy0307/ESGenius
ESGenius
4 commits
nihatgaribli/AZ-Eval
An open parallel Azerbaijani-English LLM benchmark. Kazakh fine-tuning produces negative transfer…
2
Virtue-Research/guard-eval-harness
One command to benchmark AI guardrails and coding agents across safety, security, jailbreak,…
19
108 commits
HeraFox-ai/Mental-Health-Safety-Eval
Dataset Overview
10