Ranking LLMs on agentic tasks
See the codeAn evaluation framework for AI agents across real-world business scenarios.
Two versions available:
Klarna's decision to replace 700 customer-service reps with AI backfired so much that they're now rehiring humans to patch the gaps. They saved money, but customer experience degraded. What if there were a way to catch those failures before flipping the switch?
That's precisely the problem Agent Leaderboard v2 is built to solve. Rather than simply testing whether an agent can call the right tools, we put AIs through real enterprise scenarios spanning five industries with multi-turn dialogues and complex decision-making.
Synthetic Dataset Generation & Agent-User Conversation Simulation
Each domain features 100 synthetic scenarios crafted to reflect the ambiguity, context-dependence, and unpredictability of real-world conversations. Every scenario includes:
Example Banking Scenario: "I need to report my Platinum credit card as lost, verify my mortgage payment on the 15th, set up automatic bill payments, find a branch near my Paris hotel, get EUR exchange rates, and configure travel alerts—all before I leave for Europe Thursday."
Action Completion (AC): Did the agent fully accomplish every user goal, providing clear answers or confirmations for every ask? This measures real-world effectiveness—can the agent actually get the job done?
Tool Selection Quality (TSQ): How accurately does an AI agent choose and use external tools? Perfect TSQ means picking the right tool with all required parameters correctly, while avoiding unnecessary or erroneous calls.
Step 1: Domain-Specific Tools
Step 2: Synthetic Personas
Step 3: Challenging Scenarios
Enterprise applications require AI agents tuned to specific needs, regulations, and workflows. Each sector brings unique challenges:
Our evaluation provides actionable insights into model suitability for specific business domains, answering questions like: Will this agent excel at healthcare scheduling but struggle with insurance claims?
Evaluates language models using standardized benchmarks and the Tool Selection Quality (TSQ) metric.
Features:
@misc{agent-leaderboard,
author = {Pratik Bhavsar},
title = {Agent Leaderboard},
year = {2025},
publisher = {Galileo.ai},
howpublished = "\url{https://huggingface.co/spaces/galileo-ai/agent-leaderboard}"
}
What we measure shapes AI. With our Agent Leaderboard initiative, we’re bringing focus back to the ground where real work happens. The Humanity’s Last Exam benchmark is cool, but we must still cover the basics.
225 followers · starred Jul 2025
Ranking LLMs on agentic tasks
See the codeAn evaluation framework for AI agents across real-world business scenarios.
Two versions available:
Klarna's decision to replace 700 customer-service reps with AI backfired so much that they're now rehiring humans to patch the gaps. They saved money, but customer experience degraded. What if there were a way to catch those failures before flipping the switch?
That's precisely the problem Agent Leaderboard v2 is built to solve. Rather than simply testing whether an agent can call the right tools, we put AIs through real enterprise scenarios spanning five industries with multi-turn dialogues and complex decision-making.
Synthetic Dataset Generation & Agent-User Conversation Simulation
Each domain features 100 synthetic scenarios crafted to reflect the ambiguity, context-dependence, and unpredictability of real-world conversations. Every scenario includes:
Example Banking Scenario: "I need to report my Platinum credit card as lost, verify my mortgage payment on the 15th, set up automatic bill payments, find a branch near my Paris hotel, get EUR exchange rates, and configure travel alerts—all before I leave for Europe Thursday."
Action Completion (AC): Did the agent fully accomplish every user goal, providing clear answers or confirmations for every ask? This measures real-world effectiveness—can the agent actually get the job done?
Tool Selection Quality (TSQ): How accurately does an AI agent choose and use external tools? Perfect TSQ means picking the right tool with all required parameters correctly, while avoiding unnecessary or erroneous calls.
Step 1: Domain-Specific Tools
Step 2: Synthetic Personas
Step 3: Challenging Scenarios
Enterprise applications require AI agents tuned to specific needs, regulations, and workflows. Each sector brings unique challenges:
Our evaluation provides actionable insights into model suitability for specific business domains, answering questions like: Will this agent excel at healthcare scheduling but struggle with insurance claims?
Evaluates language models using standardized benchmarks and the Tool Selection Quality (TSQ) metric.
Features:
@misc{agent-leaderboard,
author = {Pratik Bhavsar},
title = {Agent Leaderboard},
year = {2025},
publisher = {Galileo.ai},
howpublished = "\url{https://huggingface.co/spaces/galileo-ai/agent-leaderboard}"
}
What we measure shapes AI. With our Agent Leaderboard initiative, we’re bringing focus back to the ground where real work happens. The Humanity’s Last Exam benchmark is cool, but we must still cover the basics.
225 followers · starred Jul 2025