Mosescreates/arabic-agent-eval-leaderboard

Space

0

stars

10

commits

2

linked in READMEs

May 30, 2026

updated

static

README

Arabic Agent Eval — Leaderboard

Leaderboard for the arabic-agent-eval benchmark: an open, installable, dialect-split Arabic function-calling benchmark.

Static site (no runtime). index.html is generated by build_static.py from leaderboard.json, which is built from the dated, provenance-frozen result bundles in the repo (real OpenRouter runs, pinned by git SHA). Adding more models — including Hermes via a native endpoint — is open work.

Dataset: https://huggingface.co/datasets/Mosescreates/arabic-agent-eval

Contributors

Mosescreates

10 commits

Mosescreates/arabic-agent-eval-leaderboard

Space

0

stars

10

commits

2

linked in READMEs

May 30, 2026

updated

static

README

Arabic Agent Eval — Leaderboard

Leaderboard for the arabic-agent-eval benchmark: an open, installable, dialect-split Arabic function-calling benchmark.

Static site (no runtime). index.html is generated by build_static.py from leaderboard.json, which is built from the dated, provenance-frozen result bundles in the repo (real OpenRouter runs, pinned by git SHA). Adding more models — including Hermes via a native endpoint — is open work.

Dataset: https://huggingface.co/datasets/Mosescreates/arabic-agent-eval

Contributors

Mosescreates

10 commits