galileo-ai/agent-leaderboard

Dataset

Agent Leaderboard

33

20 commits

2 linked in READMEs

updated Jul 16, 2025

See the code

README

Agent Leaderboard

Leaderboard Blog Dataset

Overview

The Agent Leaderboard evaluates language models' ability to effectively utilize tools in complex scenarios. With major tech CEOs predicting 2025 as a pivotal year for AI agents, we built this leaderboard to answer: "How do AI agents perform in real-world business scenarios?"

Get latest update of the leaderboard on Hugging Face Spaces. For more info, checkout the blog post for a detailed overview of our evaluation methodology.

Methodology

Our evaluation process follows a systematic approach:

Model Selection: Curated diverse set of leading language models (12 private, 5 open-source)
Agent Configuration: Standardized system prompt and consistent tool access
Metric Definition: Tool Selection Quality (TSQ) as primary metric
Dataset Curation: Strategic sampling from established benchmarks
Scoring System: Equally weighted average across datasets

Model Rankings

Current standings across different models:

Dataset Structure

Comprehensive evaluation across multiple domains and interaction types by leveraging diverse datasets:

BFCL: Mathematics, Entertainment, Education, and Academic Domains
τ-bench: Retail and Airline Industry Scenarios
xLAM: Cross-domain Data Generation (21 Domains)
ToolACE: API Interactions across 390 Domains

Evaluation

Our evaluation metric Tool Selection Quality (TSQ) assesses how well models select and use tools based on real-world requirements:

Acknowledgements

We extend our sincere gratitude to the creators of the benchmark datasets that made this evaluation framework possible:

  • BFCL: Thanks to the Berkeley AI Research team for their comprehensive dataset evaluating function calling capabilities.

  • τ-bench: Thanks to the Sierra Research team for developing this benchmark focusing on real-world tool use scenarios.

  • xLAM: Thanks to the Salesforce AI Research team for their extensive Large Action Model dataset covering 21 domains.

  • ToolACE: Thanks to the team for their comprehensive API interaction dataset spanning 390 domains.

These datasets have been instrumental in creating a comprehensive evaluation framework for tool-calling capabilities in language models.

Citation

@misc{agent-leaderboard,
    author = {Pratik Bhavsar},
    title = {Agent Leaderboard},
    year = {2025},
    publisher = {Galileo.ai},
    howpublished = "\url{https://huggingface.co/datasets/galileo-ai/agent-leaderboard}"
}
agent
function-calling
LLM Agent
tools

Contributors

pratikbhavsar

17 commits

PB

galileo-ai/agent-leaderboard

Dataset

Agent Leaderboard

33

20 commits

2 linked in READMEs

updated Jul 16, 2025

See the code

README

Agent Leaderboard

Leaderboard Blog Dataset

Overview

The Agent Leaderboard evaluates language models' ability to effectively utilize tools in complex scenarios. With major tech CEOs predicting 2025 as a pivotal year for AI agents, we built this leaderboard to answer: "How do AI agents perform in real-world business scenarios?"

Get latest update of the leaderboard on Hugging Face Spaces. For more info, checkout the blog post for a detailed overview of our evaluation methodology.

Methodology

Our evaluation process follows a systematic approach:

Model Selection: Curated diverse set of leading language models (12 private, 5 open-source)
Agent Configuration: Standardized system prompt and consistent tool access
Metric Definition: Tool Selection Quality (TSQ) as primary metric
Dataset Curation: Strategic sampling from established benchmarks
Scoring System: Equally weighted average across datasets

Model Rankings

Current standings across different models:

Dataset Structure

Comprehensive evaluation across multiple domains and interaction types by leveraging diverse datasets:

BFCL: Mathematics, Entertainment, Education, and Academic Domains
τ-bench: Retail and Airline Industry Scenarios
xLAM: Cross-domain Data Generation (21 Domains)
ToolACE: API Interactions across 390 Domains

Evaluation

Our evaluation metric Tool Selection Quality (TSQ) assesses how well models select and use tools based on real-world requirements:

Acknowledgements

We extend our sincere gratitude to the creators of the benchmark datasets that made this evaluation framework possible:

  • BFCL: Thanks to the Berkeley AI Research team for their comprehensive dataset evaluating function calling capabilities.

  • τ-bench: Thanks to the Sierra Research team for developing this benchmark focusing on real-world tool use scenarios.

  • xLAM: Thanks to the Salesforce AI Research team for their extensive Large Action Model dataset covering 21 domains.

  • ToolACE: Thanks to the team for their comprehensive API interaction dataset spanning 390 domains.

These datasets have been instrumental in creating a comprehensive evaluation framework for tool-calling capabilities in language models.

Citation

@misc{agent-leaderboard,
    author = {Pratik Bhavsar},
    title = {Agent Leaderboard},
    year = {2025},
    publisher = {Galileo.ai},
    howpublished = "\url{https://huggingface.co/datasets/galileo-ai/agent-leaderboard}"
}
agent
function-calling
LLM Agent
tools

Contributors

pratikbhavsar

17 commits

PB