Agent Leaderboard v2 is an enterprise-grade benchmark for evaluating AI agents in realistic customer support scenarios. This dataset simulates multi-turn conversations across five critical industries: π¦ banking, π₯ healthcare, π‘οΈ insurance, π investment, and π± telecom.
Agent Leaderboard v2 addresses key limitations of v1:
Measures whether the agent fully accomplished every user goal, providing clear answers or confirmations for every request. This goes beyond correct tool calls to assess actual problem-solving effectiveness.
Evaluates how accurately an AI agent chooses and uses external tools, including:
The benchmark uses a synthetic data approach with three key components:
Each scenario is evaluated through a simulation pipeline that recreates realistic customer support interactions, measuring both tool usage accuracy and goal completion effectiveness.
Each domain contains 100 scenarios designed to test agents' ability to coordinate actions, maintain context, and handle the complexity of enterprise customer support interactions.
import json
import os
from datasets import load_dataset
# Choose domain (banking, healthcare, insurance, investment, or telecom)
domain = "banking"
# Load all configurations for the chosen domain
tools = load_dataset("galileo-ai/agent-leaderboard-v2", "tools", split=domain)
personas = load_dataset("galileo-ai/agent-leaderboard-v2", "personas", split=domain)
scenarios = load_dataset("galileo-ai/agent-leaderboard-v2", "adaptive_tool_use", split=domain)
# Required conversion to convert tool JSON strings to proper dictionaries
def convert_tool_json_strings(tool_record):
tool = dict(tool_record)
# Convert 'properties' from JSON string to dict
if 'properties' in tool and isinstance(tool['properties'], str):
tool['properties'] = json.loads(tool['properties'])
# Convert 'response_schema' from JSON string to dict
if 'response_schema' in tool and isinstance(tool['response_schema'], str):
tool['response_schema'] = json.loads(tool['response_schema'])
return tool
# Apply conversion to tools dataset
converted_tools = [convert_tool_json_strings(tool) for tool in tools]
# Create directory structure
output_dir = f"v2/data/{domain}"
os.makedirs(output_dir, exist_ok=True)
# Save datasets as JSON files
with open(f'{output_dir}/tools.json', 'w') as f:
json.dump(converted_tools, f, indent=2)
with open(f'{output_dir}/personas.json', 'w') as f:
json.dump([dict(persona) for persona in personas], f, indent=2)
with open(f'{output_dir}/adaptive_tool_use.json', 'w') as f:
json.dump([dict(scenario) for scenario in scenarios], f, indent=2)
Checkout our blog for more information on the methodology.
@misc{agent-leaderboard,
author = {Pratik Bhavsar},
title = {Agent Leaderboard},
year = {2025},
publisher = {Galileo.ai},
howpublished = "\url{https://huggingface.co/spaces/galileo-ai/agent-leaderboard}"
}
For inquiries about the dataset or benchmark:
Agent Leaderboard v2 is an enterprise-grade benchmark for evaluating AI agents in realistic customer support scenarios. This dataset simulates multi-turn conversations across five critical industries: π¦ banking, π₯ healthcare, π‘οΈ insurance, π investment, and π± telecom.
Agent Leaderboard v2 addresses key limitations of v1:
Measures whether the agent fully accomplished every user goal, providing clear answers or confirmations for every request. This goes beyond correct tool calls to assess actual problem-solving effectiveness.
Evaluates how accurately an AI agent chooses and uses external tools, including:
The benchmark uses a synthetic data approach with three key components:
Each scenario is evaluated through a simulation pipeline that recreates realistic customer support interactions, measuring both tool usage accuracy and goal completion effectiveness.
Each domain contains 100 scenarios designed to test agents' ability to coordinate actions, maintain context, and handle the complexity of enterprise customer support interactions.
import json
import os
from datasets import load_dataset
# Choose domain (banking, healthcare, insurance, investment, or telecom)
domain = "banking"
# Load all configurations for the chosen domain
tools = load_dataset("galileo-ai/agent-leaderboard-v2", "tools", split=domain)
personas = load_dataset("galileo-ai/agent-leaderboard-v2", "personas", split=domain)
scenarios = load_dataset("galileo-ai/agent-leaderboard-v2", "adaptive_tool_use", split=domain)
# Required conversion to convert tool JSON strings to proper dictionaries
def convert_tool_json_strings(tool_record):
tool = dict(tool_record)
# Convert 'properties' from JSON string to dict
if 'properties' in tool and isinstance(tool['properties'], str):
tool['properties'] = json.loads(tool['properties'])
# Convert 'response_schema' from JSON string to dict
if 'response_schema' in tool and isinstance(tool['response_schema'], str):
tool['response_schema'] = json.loads(tool['response_schema'])
return tool
# Apply conversion to tools dataset
converted_tools = [convert_tool_json_strings(tool) for tool in tools]
# Create directory structure
output_dir = f"v2/data/{domain}"
os.makedirs(output_dir, exist_ok=True)
# Save datasets as JSON files
with open(f'{output_dir}/tools.json', 'w') as f:
json.dump(converted_tools, f, indent=2)
with open(f'{output_dir}/personas.json', 'w') as f:
json.dump([dict(persona) for persona in personas], f, indent=2)
with open(f'{output_dir}/adaptive_tool_use.json', 'w') as f:
json.dump([dict(scenario) for scenario in scenarios], f, indent=2)
Checkout our blog for more information on the methodology.
@misc{agent-leaderboard,
author = {Pratik Bhavsar},
title = {Agent Leaderboard},
year = {2025},
publisher = {Galileo.ai},
howpublished = "\url{https://huggingface.co/spaces/galileo-ai/agent-leaderboard}"
}
For inquiries about the dataset or benchmark: