EverMind-AI/EvoAgentBench

Dataset

EvoAgentBench

17

14 commits

2 linked in READMEs

updated Jul 16, 2026

See the code

README

EvoAgentBench

EvoAgentBench is a benchmark for evaluating AI agent self-evolution β€” the ability of agents to improve their performance by learning from past experiences. It provides standardized train/test splits across five diverse task domains, enabling reproducible comparison of skill extraction and experience reuse methods.

Benchmark Overview

DomainBase DatasetTrainTestTask Format
Information RetrievalBrowseComp-Plus15465Multi-constraint entity identification via web search
Reasoning & Problem DecompositionOmniMath478100Competition-level mathematical reasoning
Software EngineeringSWE-Bench8756Real-world GitHub issue resolution
Code ImplementationLiveCodeBench18286Competitive programming problems
Knowledge WorkGDPVal10560Document-grounded question answering

Total: 1006 train + 367 test tasks

Dataset Structure

EvoAgentBench/
β”œβ”€β”€ Information Retrieval/
β”‚   └── task_split.json
β”œβ”€β”€ Reasoning & Problem Decomposition/
β”‚   └── test_set_100/
β”‚       β”œβ”€β”€ train.jsonl          # 478 OmniMath problems (train)
β”‚       β”œβ”€β”€ test.jsonl           # 100 OmniMath problems (test)
β”‚       β”œβ”€β”€ index_map.json       # test index -> problem metadata
β”‚       └── skip.json            # problems to skip during evaluation
β”œβ”€β”€ Software Engineering/
β”‚   └── task_split.json
β”œβ”€β”€ Code Implementation/
β”‚   └── task_split.json
└── Knowledge Work/
    β”œβ”€β”€ task_split.json
    β”œβ”€β”€ meta_prompts/
    └── reference_files/

Each task_split.json contains train/test task ID lists that reference the original benchmark datasets. For OmniMath the actual problems are included directly (test_set_100/*.jsonl). For Knowledge Work, task IDs reference the openai/gdpval dataset; per-occupation meta prompts are included under meta_prompts/, and per-task reference files are downloaded from GDPVal on first use (see reference_files/README.md).

Evaluation Protocol

EvoAgentBench follows a three-phase self-evolution protocol:

  1. Train: Run the agent on train tasks to collect interaction trajectories (sessions).
  2. Extract: Apply a self-evolution method to extract reusable knowledge (skills, cases, memories) from train trajectories.
  3. Evaluate: Run the agent on test tasks with extracted knowledge injected, and compare against the no-knowledge baseline.

The train/test splits are designed so that:

  • Train and test tasks have no overlap
  • Test tasks require similar capabilities to train tasks but are distinct problems
  • Performance improvement on test tasks demonstrates genuine generalization, not memorization

Usage

Download

# Option 1: git clone
git clone https://huggingface.co/datasets/EverMind-AI/EvoAgentBench

# Option 2: huggingface_hub
python -c "
from huggingface_hub import snapshot_download
snapshot_download('EverMind-AI/EvoAgentBench', repo_type='dataset', local_dir='data/')
"

See the paper for the full evaluation protocol and agent setup.

Loading Splits Directly

import json
from huggingface_hub import hf_hub_download

# Download a specific task split
path = hf_hub_download(
    "EverMind-AI/EvoAgentBench",
    "Information Retrieval/task_split.json",
    repo_type="dataset"
)
splits = json.loads(open(path).read())
train_ids = splits["train"]  # 154 task IDs
test_ids = splits["test"]    # 65 task IDs

Paper

EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer

Citation

@misc{gao2026evoagentbench,
  title={EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer},
  author={Xingze Gao and Chuanrui Hu and Hongda Chen and Pengfei Yao and Zhao Wang and Yi Bai and Zhengwei Wu and Yunyun Han and Xiaofeng Cong and Jie Gui and Yafeng Deng and Teng Li},
  year={2026},
  eprint={2607.05202},
  archivePrefix={arXiv},
  url={https://arxiv.org/abs/2607.05202}
}

License

Apache 2.0

agent
benchmark
evaluation
self-evolution

Contributors

EverMind-AI

14 commits

EverMind-AI/EvoAgentBench

Dataset

EvoAgentBench

17

14 commits

2 linked in READMEs

updated Jul 16, 2026

See the code

README

EvoAgentBench

EvoAgentBench is a benchmark for evaluating AI agent self-evolution β€” the ability of agents to improve their performance by learning from past experiences. It provides standardized train/test splits across five diverse task domains, enabling reproducible comparison of skill extraction and experience reuse methods.

Benchmark Overview

DomainBase DatasetTrainTestTask Format
Information RetrievalBrowseComp-Plus15465Multi-constraint entity identification via web search
Reasoning & Problem DecompositionOmniMath478100Competition-level mathematical reasoning
Software EngineeringSWE-Bench8756Real-world GitHub issue resolution
Code ImplementationLiveCodeBench18286Competitive programming problems
Knowledge WorkGDPVal10560Document-grounded question answering

Total: 1006 train + 367 test tasks

Dataset Structure

EvoAgentBench/
β”œβ”€β”€ Information Retrieval/
β”‚   └── task_split.json
β”œβ”€β”€ Reasoning & Problem Decomposition/
β”‚   └── test_set_100/
β”‚       β”œβ”€β”€ train.jsonl          # 478 OmniMath problems (train)
β”‚       β”œβ”€β”€ test.jsonl           # 100 OmniMath problems (test)
β”‚       β”œβ”€β”€ index_map.json       # test index -> problem metadata
β”‚       └── skip.json            # problems to skip during evaluation
β”œβ”€β”€ Software Engineering/
β”‚   └── task_split.json
β”œβ”€β”€ Code Implementation/
β”‚   └── task_split.json
└── Knowledge Work/
    β”œβ”€β”€ task_split.json
    β”œβ”€β”€ meta_prompts/
    └── reference_files/

Each task_split.json contains train/test task ID lists that reference the original benchmark datasets. For OmniMath the actual problems are included directly (test_set_100/*.jsonl). For Knowledge Work, task IDs reference the openai/gdpval dataset; per-occupation meta prompts are included under meta_prompts/, and per-task reference files are downloaded from GDPVal on first use (see reference_files/README.md).

Evaluation Protocol

EvoAgentBench follows a three-phase self-evolution protocol:

  1. Train: Run the agent on train tasks to collect interaction trajectories (sessions).
  2. Extract: Apply a self-evolution method to extract reusable knowledge (skills, cases, memories) from train trajectories.
  3. Evaluate: Run the agent on test tasks with extracted knowledge injected, and compare against the no-knowledge baseline.

The train/test splits are designed so that:

  • Train and test tasks have no overlap
  • Test tasks require similar capabilities to train tasks but are distinct problems
  • Performance improvement on test tasks demonstrates genuine generalization, not memorization

Usage

Download

# Option 1: git clone
git clone https://huggingface.co/datasets/EverMind-AI/EvoAgentBench

# Option 2: huggingface_hub
python -c "
from huggingface_hub import snapshot_download
snapshot_download('EverMind-AI/EvoAgentBench', repo_type='dataset', local_dir='data/')
"

See the paper for the full evaluation protocol and agent setup.

Loading Splits Directly

import json
from huggingface_hub import hf_hub_download

# Download a specific task split
path = hf_hub_download(
    "EverMind-AI/EvoAgentBench",
    "Information Retrieval/task_split.json",
    repo_type="dataset"
)
splits = json.loads(open(path).read())
train_ids = splits["train"]  # 154 task IDs
test_ids = splits["test"]    # 65 task IDs

Paper

EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer

Citation

@misc{gao2026evoagentbench,
  title={EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer},
  author={Xingze Gao and Chuanrui Hu and Hongda Chen and Pengfei Yao and Zhao Wang and Yi Bai and Zhengwei Wu and Yunyun Han and Xiaofeng Cong and Jie Gui and Yafeng Deng and Teng Li},
  year={2026},
  eprint={2607.05202},
  archivePrefix={arXiv},
  url={https://arxiv.org/abs/2607.05202}
}

License

Apache 2.0

agent
benchmark
evaluation
self-evolution

Contributors

EverMind-AI

14 commits