DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints
213
8 commits
1 linked in READMEs
updated Mar 3, 2026
DeepPlanningBench is a challenging benchmark for evaluating long-horizon agentic planning capabilities of large language models (LLMs) with verifiable constraints. It features realistic multi-day travel planning and multi-product shopping tasks that require proactive information acquisition, local constrained reasoning, and global constrained optimization.
🌐 Website: https://qwenlm.github.io/Qwen-Agent/en/benchmarks/deepplanning/
📄 Paper: https://arxiv.org/abs/2601.18137
While agent evaluation has shifted toward long-horizon tasks, most benchmarks still emphasize local, step-level reasoning rather than the global constrained optimization (e.g., time and financial budgets) that demands genuine planning ability. DeepPlanning addresses this gap by introducing practical long-horizon agent planning scenarios that require:
The benchmark includes two main domains:
If you find our work useful, please consider citing:
@article{deepplanning,
title={DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints},
author={
Yinger Zhang and Shutong Jiang and Renhao Li and Jianhong Tu and Yang Su and
Lianghao Deng and Xudong Guo and Chenxu Lv and Junyang Lin
},
journal={arXiv preprint arXiv:2601.18137},
year={2026}
}
DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints
213
8 commits
1 linked in READMEs
updated Mar 3, 2026
DeepPlanningBench is a challenging benchmark for evaluating long-horizon agentic planning capabilities of large language models (LLMs) with verifiable constraints. It features realistic multi-day travel planning and multi-product shopping tasks that require proactive information acquisition, local constrained reasoning, and global constrained optimization.
🌐 Website: https://qwenlm.github.io/Qwen-Agent/en/benchmarks/deepplanning/
📄 Paper: https://arxiv.org/abs/2601.18137
While agent evaluation has shifted toward long-horizon tasks, most benchmarks still emphasize local, step-level reasoning rather than the global constrained optimization (e.g., time and financial budgets) that demands genuine planning ability. DeepPlanning addresses this gap by introducing practical long-horizon agent planning scenarios that require:
The benchmark includes two main domains:
If you find our work useful, please consider citing:
@article{deepplanning,
title={DeepPlanning: Benchmarking Long-Horizon Agentic Planning with Verifiable Constraints},
author={
Yinger Zhang and Shutong Jiang and Renhao Li and Jianhong Tu and Yang Su and
Lianghao Deng and Xudong Guo and Chenxu Lv and Junyang Lin
},
journal={arXiv preprint arXiv:2601.18137},
year={2026}
}