lixiaoxi45/DeepAgent-Datasets

Dataset

DeepAgent Datasets

4

18 commits

1 linked in READMEs

updated Feb 5, 2026

See the code

README

DeepAgent Datasets

Paper | GitHub

This repository contains the pre-processed evaluation datasets for DeepAgent, an end-to-end deep reasoning agent that performs autonomous thinking, tool discovery, and action execution within a single, coherent reasoning process.

Dataset Summary

The repository includes curated data for several major benchmarks used to evaluate reasoning agents across different domains:

General Tool Use

  • ToolBench: Features 16,000+ real-world RapidAPIs requiring multi-step, multi-tool reasoning.
  • RestBench (TMDB & Spotify): Simulates REST API applications with scenarios for TMDB (movie tools) and Spotify (music tools).
  • ToolHop: Tests multi-hop reasoning across thousands of locally executable tools.
  • API-Bank: Evaluates planning, retrieval, and calling across human-annotated dialogues.

Embodied AI and Web Navigation

  • ALFWorld: Text-based embodied AI environment where agents complete household tasks.
  • WebShop: Online shopping simulation requiring agents to search and navigate products.

Deep Research

  • GAIA: Complex tasks requiring web search, browsing, VQA, code execution, and file processing.
  • Humanity's Last Exam (HLE): Extremely challenging reasoning problems; this repository contains a sampled set of 500 questions.

Sample Usage

To run the DeepAgent system on these benchmarks, you can use the implementation from the official repository. For example, to run on the ToolBench dataset:

python src/run_deep_agent.py \
    --config_path ./config/base_config.yaml \
    --dataset_name toolbench \
    --enable_tool_search \
    --eval

Citation

If you find this work helpful, please cite:

@misc{deepagent,
      title={DeepAgent: A General Reasoning Agent with Scalable Toolsets}, 
      author={Xiaoxi Li and Wenxiang Jiao and Jiarui Jin and Guanting Dong and Jiajie Jin and Yinuo Wang and Hao Wang and Yutao Zhu and Ji-Rong Wen and Yuan Lu and Zhicheng Dou},
      year={2025},
      eprint={2510.21618},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2510.21618}, 
}
autonomous-agents
reasoning
tool-use

lixiaoxi45/DeepAgent-Datasets

Dataset

DeepAgent Datasets

4

18 commits

1 linked in READMEs

updated Feb 5, 2026

See the code

README

DeepAgent Datasets

Paper | GitHub

This repository contains the pre-processed evaluation datasets for DeepAgent, an end-to-end deep reasoning agent that performs autonomous thinking, tool discovery, and action execution within a single, coherent reasoning process.

Dataset Summary

The repository includes curated data for several major benchmarks used to evaluate reasoning agents across different domains:

General Tool Use

  • ToolBench: Features 16,000+ real-world RapidAPIs requiring multi-step, multi-tool reasoning.
  • RestBench (TMDB & Spotify): Simulates REST API applications with scenarios for TMDB (movie tools) and Spotify (music tools).
  • ToolHop: Tests multi-hop reasoning across thousands of locally executable tools.
  • API-Bank: Evaluates planning, retrieval, and calling across human-annotated dialogues.

Embodied AI and Web Navigation

  • ALFWorld: Text-based embodied AI environment where agents complete household tasks.
  • WebShop: Online shopping simulation requiring agents to search and navigate products.

Deep Research

  • GAIA: Complex tasks requiring web search, browsing, VQA, code execution, and file processing.
  • Humanity's Last Exam (HLE): Extremely challenging reasoning problems; this repository contains a sampled set of 500 questions.

Sample Usage

To run the DeepAgent system on these benchmarks, you can use the implementation from the official repository. For example, to run on the ToolBench dataset:

python src/run_deep_agent.py \
    --config_path ./config/base_config.yaml \
    --dataset_name toolbench \
    --enable_tool_search \
    --eval

Citation

If you find this work helpful, please cite:

@misc{deepagent,
      title={DeepAgent: A General Reasoning Agent with Scalable Toolsets}, 
      author={Xiaoxi Li and Wenxiang Jiao and Jiarui Jin and Guanting Dong and Jiajie Jin and Yinuo Wang and Hao Wang and Yutao Zhu and Ji-Rong Wen and Yuan Lu and Zhicheng Dou},
      year={2025},
      eprint={2510.21618},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2510.21618}, 
}
autonomous-agents
reasoning
tool-use