Leaderboard for the arabic-agent-eval benchmark: an open, installable, dialect-split Arabic function-calling benchmark.
Static site (no runtime). index.html is generated by build_static.py from
leaderboard.json, which is built from the dated, provenance-frozen result
bundles in the repo (real OpenRouter runs, pinned by git SHA). Adding more
models — including Hermes via a native endpoint — is open work.
Dataset: https://huggingface.co/datasets/Mosescreates/arabic-agent-eval
10 commits
Leaderboard for the arabic-agent-eval benchmark: an open, installable, dialect-split Arabic function-calling benchmark.
Static site (no runtime). index.html is generated by build_static.py from
leaderboard.json, which is built from the dated, provenance-frozen result
bundles in the repo (real OpenRouter runs, pinned by git SHA). Adding more
models — including Hermes via a native endpoint — is open work.
Dataset: https://huggingface.co/datasets/Mosescreates/arabic-agent-eval
10 commits