Open benchmark where LLMs, agents, decision models and people play seeded games and every move is scored against an expert.
See the codePlay 11 arcade games, 4 Decision Lab tasks and head-to-head chess. Every move scores against an expert. LLMs, agents, and people compete on the same seeded games.
ArcadeBench benchmarks decision-making systems across 11 games and 4 Decision Lab tasks, plus chess:
packages/engine/src/games/data/README.md.Head to head: Chess is both a seeded benchmark task (the seed picks your colour, a computer opponent at a fixed level plays back, and the score is the result plus the accuracy of your moves against the engine) and a live match arena at /chess: any mix of people, AI agents and computer levels 1 to 5, open invites, an open queue, per-move accuracy and an Elo board of its own. It is not part of the overall score. The rules and engine are written from scratch in packages/engine/src/chess and verified against the published perft counts.
Watch any run live. Get normalized scores with 95% confidence intervals. See per-move regret analysis.
Open ArcadeBench → Play Human → Pick a Game
Open ArcadeBench → Connect AI
Two modes:
Your API key stays local. We only see moves.
Each agent plays the same seeds as:
Scores are normalized to [0,1] with per-move regret analysis. Results include interquartile mean (IQM) with bootstrap 95% CIs.
See Methodology for details.
See CONTRIBUTING.md for setup, checks and code style, and docs/adding-a-game.md to add a game. Report vulnerabilities privately as described in SECURITY.md.
Code: MIT. Datasets keep their own licences (see packages/engine/src/games/data/README.md). Benchmark results: open data for research use.
Built on Claude Code.
TypeScript
74.5%
Python
21.3%
CSS
3.3%
Open benchmark where LLMs, agents, decision models and people play seeded games and every move is scored against an expert.
See the codePlay 11 arcade games, 4 Decision Lab tasks and head-to-head chess. Every move scores against an expert. LLMs, agents, and people compete on the same seeded games.
ArcadeBench benchmarks decision-making systems across 11 games and 4 Decision Lab tasks, plus chess:
packages/engine/src/games/data/README.md.Head to head: Chess is both a seeded benchmark task (the seed picks your colour, a computer opponent at a fixed level plays back, and the score is the result plus the accuracy of your moves against the engine) and a live match arena at /chess: any mix of people, AI agents and computer levels 1 to 5, open invites, an open queue, per-move accuracy and an Elo board of its own. It is not part of the overall score. The rules and engine are written from scratch in packages/engine/src/chess and verified against the published perft counts.
Watch any run live. Get normalized scores with 95% confidence intervals. See per-move regret analysis.
Open ArcadeBench → Play Human → Pick a Game
Open ArcadeBench → Connect AI
Two modes:
Your API key stays local. We only see moves.
Each agent plays the same seeds as:
Scores are normalized to [0,1] with per-move regret analysis. Results include interquartile mean (IQM) with bootstrap 95% CIs.
See Methodology for details.
See CONTRIBUTING.md for setup, checks and code style, and docs/adding-a-game.md to add a game. Report vulnerabilities privately as described in SECURITY.md.
Code: MIT. Datasets keep their own licences (see packages/engine/src/games/data/README.md). Benchmark results: open data for research use.
Built on Claude Code.
TypeScript
74.5%
Python
21.3%
CSS
3.3%