The Agent Leaderboard evaluates language models' ability to effectively utilize tools in complex scenarios. With major tech CEOs predicting 2025 as a pivotal year for AI agents, we built this leaderboard to answer: "How do AI agents perform in real-world business scenarios?"
Get latest update of the leaderboard on Hugging Face Spaces. For more info, checkout the blog post for a detailed overview of our evaluation methodology.
Our evaluation process follows a systematic approach:
Model Selection: Curated diverse set of leading language models (12 private, 5 open-source)
Agent Configuration: Standardized system prompt and consistent tool access
Metric Definition: Tool Selection Quality (TSQ) as primary metric
Dataset Curation: Strategic sampling from established benchmarks
Scoring System: Equally weighted average across datasets
Current standings across different models:
Comprehensive evaluation across multiple domains and interaction types by leveraging diverse datasets:
BFCL: Mathematics, Entertainment, Education, and Academic Domains
τ-bench: Retail and Airline Industry Scenarios
xLAM: Cross-domain Data Generation (21 Domains)
ToolACE: API Interactions across 390 Domains
Our evaluation metric Tool Selection Quality (TSQ) assesses how well models select and use tools based on real-world requirements:
We extend our sincere gratitude to the creators of the benchmark datasets that made this evaluation framework possible:
BFCL: Thanks to the Berkeley AI Research team for their comprehensive dataset evaluating function calling capabilities.
τ-bench: Thanks to the Sierra Research team for developing this benchmark focusing on real-world tool use scenarios.
xLAM: Thanks to the Salesforce AI Research team for their extensive Large Action Model dataset covering 21 domains.
ToolACE: Thanks to the team for their comprehensive API interaction dataset spanning 390 domains.
These datasets have been instrumental in creating a comprehensive evaluation framework for tool-calling capabilities in language models.
@misc{agent-leaderboard,
author = {Pratik Bhavsar},
title = {Agent Leaderboard},
year = {2025},
publisher = {Galileo.ai},
howpublished = "\url{https://huggingface.co/datasets/galileo-ai/agent-leaderboard}"
}
17 commits
3 commits
The Agent Leaderboard evaluates language models' ability to effectively utilize tools in complex scenarios. With major tech CEOs predicting 2025 as a pivotal year for AI agents, we built this leaderboard to answer: "How do AI agents perform in real-world business scenarios?"
Get latest update of the leaderboard on Hugging Face Spaces. For more info, checkout the blog post for a detailed overview of our evaluation methodology.
Our evaluation process follows a systematic approach:
Model Selection: Curated diverse set of leading language models (12 private, 5 open-source)
Agent Configuration: Standardized system prompt and consistent tool access
Metric Definition: Tool Selection Quality (TSQ) as primary metric
Dataset Curation: Strategic sampling from established benchmarks
Scoring System: Equally weighted average across datasets
Current standings across different models:
Comprehensive evaluation across multiple domains and interaction types by leveraging diverse datasets:
BFCL: Mathematics, Entertainment, Education, and Academic Domains
τ-bench: Retail and Airline Industry Scenarios
xLAM: Cross-domain Data Generation (21 Domains)
ToolACE: API Interactions across 390 Domains
Our evaluation metric Tool Selection Quality (TSQ) assesses how well models select and use tools based on real-world requirements:
We extend our sincere gratitude to the creators of the benchmark datasets that made this evaluation framework possible:
BFCL: Thanks to the Berkeley AI Research team for their comprehensive dataset evaluating function calling capabilities.
τ-bench: Thanks to the Sierra Research team for developing this benchmark focusing on real-world tool use scenarios.
xLAM: Thanks to the Salesforce AI Research team for their extensive Large Action Model dataset covering 21 domains.
ToolACE: Thanks to the team for their comprehensive API interaction dataset spanning 390 domains.
These datasets have been instrumental in creating a comprehensive evaluation framework for tool-calling capabilities in language models.
@misc{agent-leaderboard,
author = {Pratik Bhavsar},
title = {Agent Leaderboard},
year = {2025},
publisher = {Galileo.ai},
howpublished = "\url{https://huggingface.co/datasets/galileo-ai/agent-leaderboard}"
}
17 commits
3 commits