galileo-ai/agent-leaderboard-v2

Dataset

πŸ† Agent Leaderboard v2

13

18 commits

1 linked in READMEs

updated Jul 16, 2025

See the code

README

πŸ† Agent Leaderboard v2

Agent Leaderboard v2 is an enterprise-grade benchmark for evaluating AI agents in realistic customer support scenarios. This dataset simulates multi-turn conversations across five critical industries: 🏦 banking, πŸ₯ healthcare, πŸ›‘οΈ insurance, πŸ“ˆ investment, and πŸ“± telecom.

✨ Key Features

  • πŸ”„ Multi-turn dialogues with 5-8 interconnected user goals per conversation
  • πŸ”§ Domain-specific tools reflecting actual enterprise APIs
  • πŸ‘₯ Synthetic personas with varying communication styles and expertise levels
  • 🧩 Complex scenarios featuring context dependencies, ambiguous requests, and real-world edge cases
  • πŸ“Š Two evaluation metrics: Action Completion (AC) and Tool Selection Quality (TSQ)

πŸ“¦ Dataset Components

  1. πŸ”§ Tools: Domain-specific function definitions with JSON schemas
  2. πŸ‘€ Personas: Diverse user profiles with personality traits, communication preferences, and backgrounds
  3. 🎯 Adaptive Tool Use: Complete scenarios combining personas with multi-goal conversations

πŸ†• What's New in v2

Agent Leaderboard v2 addresses key limitations of v1:

  • πŸ“ˆ Beyond score saturation: v1 saw models clustering above 90%, making differentiation difficult
  • πŸ”„ Dynamic scenarios: Multi-turn conversations replace static, one-shot evaluations
  • 🏒 Domain isolation: Industry-specific datasets for targeted enterprise evaluation
  • 🌍 Real-world complexity: Ambiguous requests, context dependencies, and interdependent goals

πŸ“ Evaluation Metrics

βœ… Action Completion (AC)

Measures whether the agent fully accomplished every user goal, providing clear answers or confirmations for every request. This goes beyond correct tool calls to assess actual problem-solving effectiveness.

🎯 Tool Selection Quality (TSQ)

Evaluates how accurately an AI agent chooses and uses external tools, including:

  • βœ”οΈ Correct tool selection for the given context
  • βš™οΈ Proper parameter handling and formatting
  • 🚫 Avoiding unnecessary or erroneous calls
  • πŸ”— Sequential decision-making across multi-step tasks

πŸ”¬ Methodology

The benchmark uses a synthetic data approach with three key components:

  1. πŸ”§ Tool Generation: Domain-specific APIs created with structured JSON schemas
  2. πŸ‘₯ Persona Design: Diverse user profiles with varying communication styles and expertise
  3. πŸ“ Scenario Crafting: Complex, multi-goal conversations that challenge agent capabilities

Each scenario is evaluated through a simulation pipeline that recreates realistic customer support interactions, measuring both tool usage accuracy and goal completion effectiveness.

πŸš€ How to use it

Each domain contains 100 scenarios designed to test agents' ability to coordinate actions, maintain context, and handle the complexity of enterprise customer support interactions.

πŸ“„ Loading the Dataset

import json
import os
from datasets import load_dataset

# Choose domain (banking, healthcare, insurance, investment, or telecom)
domain = "banking"

# Load all configurations for the chosen domain
tools = load_dataset("galileo-ai/agent-leaderboard-v2", "tools", split=domain)
personas = load_dataset("galileo-ai/agent-leaderboard-v2", "personas", split=domain)
scenarios = load_dataset("galileo-ai/agent-leaderboard-v2", "adaptive_tool_use", split=domain)

# Required conversion to convert tool JSON strings to proper dictionaries
def convert_tool_json_strings(tool_record):
    tool = dict(tool_record)

    # Convert 'properties' from JSON string to dict
    if 'properties' in tool and isinstance(tool['properties'], str):
        tool['properties'] = json.loads(tool['properties'])

    # Convert 'response_schema' from JSON string to dict  
    if 'response_schema' in tool and isinstance(tool['response_schema'], str):
        tool['response_schema'] = json.loads(tool['response_schema'])

    return tool

# Apply conversion to tools dataset
converted_tools = [convert_tool_json_strings(tool) for tool in tools]

# Create directory structure
output_dir = f"v2/data/{domain}"
os.makedirs(output_dir, exist_ok=True)

# Save datasets as JSON files
with open(f'{output_dir}/tools.json', 'w') as f:
    json.dump(converted_tools, f, indent=2)

with open(f'{output_dir}/personas.json', 'w') as f:
    json.dump([dict(persona) for persona in personas], f, indent=2)

with open(f'{output_dir}/adaptive_tool_use.json', 'w') as f:
    json.dump([dict(scenario) for scenario in scenarios], f, indent=2)

Checkout our blog for more information on the methodology.

πŸ“š Citation

@misc{agent-leaderboard,
  author = {Pratik Bhavsar},
  title = {Agent Leaderboard},
  year = {2025},
  publisher = {Galileo.ai},
  howpublished = "\url{https://huggingface.co/spaces/galileo-ai/agent-leaderboard}"
}

πŸ“§ Contact

For inquiries about the dataset or benchmark:

agent
banking
finance
function-calling
healthcare
insurance
investment
LLM Agent
synthetic
telecom
tools

galileo-ai/agent-leaderboard-v2

Dataset

πŸ† Agent Leaderboard v2

13

18 commits

1 linked in READMEs

updated Jul 16, 2025

See the code

README

πŸ† Agent Leaderboard v2

Agent Leaderboard v2 is an enterprise-grade benchmark for evaluating AI agents in realistic customer support scenarios. This dataset simulates multi-turn conversations across five critical industries: 🏦 banking, πŸ₯ healthcare, πŸ›‘οΈ insurance, πŸ“ˆ investment, and πŸ“± telecom.

✨ Key Features

  • πŸ”„ Multi-turn dialogues with 5-8 interconnected user goals per conversation
  • πŸ”§ Domain-specific tools reflecting actual enterprise APIs
  • πŸ‘₯ Synthetic personas with varying communication styles and expertise levels
  • 🧩 Complex scenarios featuring context dependencies, ambiguous requests, and real-world edge cases
  • πŸ“Š Two evaluation metrics: Action Completion (AC) and Tool Selection Quality (TSQ)

πŸ“¦ Dataset Components

  1. πŸ”§ Tools: Domain-specific function definitions with JSON schemas
  2. πŸ‘€ Personas: Diverse user profiles with personality traits, communication preferences, and backgrounds
  3. 🎯 Adaptive Tool Use: Complete scenarios combining personas with multi-goal conversations

πŸ†• What's New in v2

Agent Leaderboard v2 addresses key limitations of v1:

  • πŸ“ˆ Beyond score saturation: v1 saw models clustering above 90%, making differentiation difficult
  • πŸ”„ Dynamic scenarios: Multi-turn conversations replace static, one-shot evaluations
  • 🏒 Domain isolation: Industry-specific datasets for targeted enterprise evaluation
  • 🌍 Real-world complexity: Ambiguous requests, context dependencies, and interdependent goals

πŸ“ Evaluation Metrics

βœ… Action Completion (AC)

Measures whether the agent fully accomplished every user goal, providing clear answers or confirmations for every request. This goes beyond correct tool calls to assess actual problem-solving effectiveness.

🎯 Tool Selection Quality (TSQ)

Evaluates how accurately an AI agent chooses and uses external tools, including:

  • βœ”οΈ Correct tool selection for the given context
  • βš™οΈ Proper parameter handling and formatting
  • 🚫 Avoiding unnecessary or erroneous calls
  • πŸ”— Sequential decision-making across multi-step tasks

πŸ”¬ Methodology

The benchmark uses a synthetic data approach with three key components:

  1. πŸ”§ Tool Generation: Domain-specific APIs created with structured JSON schemas
  2. πŸ‘₯ Persona Design: Diverse user profiles with varying communication styles and expertise
  3. πŸ“ Scenario Crafting: Complex, multi-goal conversations that challenge agent capabilities

Each scenario is evaluated through a simulation pipeline that recreates realistic customer support interactions, measuring both tool usage accuracy and goal completion effectiveness.

πŸš€ How to use it

Each domain contains 100 scenarios designed to test agents' ability to coordinate actions, maintain context, and handle the complexity of enterprise customer support interactions.

πŸ“„ Loading the Dataset

import json
import os
from datasets import load_dataset

# Choose domain (banking, healthcare, insurance, investment, or telecom)
domain = "banking"

# Load all configurations for the chosen domain
tools = load_dataset("galileo-ai/agent-leaderboard-v2", "tools", split=domain)
personas = load_dataset("galileo-ai/agent-leaderboard-v2", "personas", split=domain)
scenarios = load_dataset("galileo-ai/agent-leaderboard-v2", "adaptive_tool_use", split=domain)

# Required conversion to convert tool JSON strings to proper dictionaries
def convert_tool_json_strings(tool_record):
    tool = dict(tool_record)

    # Convert 'properties' from JSON string to dict
    if 'properties' in tool and isinstance(tool['properties'], str):
        tool['properties'] = json.loads(tool['properties'])

    # Convert 'response_schema' from JSON string to dict  
    if 'response_schema' in tool and isinstance(tool['response_schema'], str):
        tool['response_schema'] = json.loads(tool['response_schema'])

    return tool

# Apply conversion to tools dataset
converted_tools = [convert_tool_json_strings(tool) for tool in tools]

# Create directory structure
output_dir = f"v2/data/{domain}"
os.makedirs(output_dir, exist_ok=True)

# Save datasets as JSON files
with open(f'{output_dir}/tools.json', 'w') as f:
    json.dump(converted_tools, f, indent=2)

with open(f'{output_dir}/personas.json', 'w') as f:
    json.dump([dict(persona) for persona in personas], f, indent=2)

with open(f'{output_dir}/adaptive_tool_use.json', 'w') as f:
    json.dump([dict(scenario) for scenario in scenarios], f, indent=2)

Checkout our blog for more information on the methodology.

πŸ“š Citation

@misc{agent-leaderboard,
  author = {Pratik Bhavsar},
  title = {Agent Leaderboard},
  year = {2025},
  publisher = {Galileo.ai},
  howpublished = "\url{https://huggingface.co/spaces/galileo-ai/agent-leaderboard}"
}

πŸ“§ Contact

For inquiries about the dataset or benchmark:

agent
banking
finance
function-calling
healthcare
insurance
investment
LLM Agent
synthetic
telecom
tools