This is only a demonstration result obtained by directly deploying and sampling HardGen in the BFCL environment, not the dataset directly used in the HardGen paper.
[!IMPORTANT] Important Hint
- This is an extension of the technical report FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use
- To allow the model to learn from errors, we specifically construct erroneous environmental responses. If you wish to delete this data, please delete the trajectories where
error_tool_responseis true.- This is an initial version of our data, we will release the up-to-date version after the acceptance of our paper.
The training set comprises 16,000 high-quality multi-turn samples. This dataset was generated using the three-phase HardGen (FunReason-MT) data synthesis framework, which focuses on generating complex trajectories.
The model built upon HardGen is rigorously evaluated on the Berkeley Function-Calling Leaderboard (BFCL).
| Model (4B - 235B) | Multi-Turn (Overall) | Single-Turn (Overall) |
|---|---|---|
| Qwen3-4B-Instruct (Base) | 22.13 | 82.14 |
| Qwen3-4B + HardGen (RL) | 63.13 | 87.14 |
| Gemini-3-Pro-Preview | 60.75 | 86.89 |
| DeepSeek-V3.2-Exp | 44.88 | 80.77 |
| GPT-5.2-2025-12-11 | 28.13 | 76.12 |
The performance of models trained upon Llama-3.1-8B-Instruct on agentic tasks (Web Search and Memory).
| Model | BFCLv4 Overall Score |
|---|---|
| HardGen-8B (RL) | 20.42 |
| CoALM-8B | 1.40 |
| ToolACE-2-8B | 13.50 |
| BitAgent-8B | 8.24 |
| xLAM-2-8b-fc-r | 10.24 |
This work is part of the open-source project AWorld, InclusionAI.
If you use this dataset in your research, please cite the related two papers:
@article{hao2026failure,
title={From Failure to Mastery: Generating Hard Samples for Tool-use Agents},
author={Hao, Bingguang and Xu, Zengzhuang and Wen, Yuntao and Xu, Xinyi and Liu, Yang and Zhao, Tong and Wang, Maolin and Chen, Long and Wang, Dong and Chen, Yicheng and others},
journal={arXiv preprint arXiv:2601.01498},
year={2026}
}
@article{xu2025funreason,
title={FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use},
author={Zengzhuang Xu, Bingguang Hao, Zechuan Wang, Yuntao Wen, Xinyi Xu, Yang Liu, Long Chen, Dong Wang, Maolin Wang, Tong Zhao, Yicheng Chen, Cunyin Peng, Jinjie Gu, Leilei Gan, Xiangyu Zhao, Chenyi Zhuang, Shi Gu},
journal={arXiv preprint arXiv:2510.24645},
year={2025}
}
For inquiries, please contact:
bingguanghao7@gmail.com29 commits
This is only a demonstration result obtained by directly deploying and sampling HardGen in the BFCL environment, not the dataset directly used in the HardGen paper.
[!IMPORTANT] Important Hint
- This is an extension of the technical report FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use
- To allow the model to learn from errors, we specifically construct erroneous environmental responses. If you wish to delete this data, please delete the trajectories where
error_tool_responseis true.- This is an initial version of our data, we will release the up-to-date version after the acceptance of our paper.
The training set comprises 16,000 high-quality multi-turn samples. This dataset was generated using the three-phase HardGen (FunReason-MT) data synthesis framework, which focuses on generating complex trajectories.
The model built upon HardGen is rigorously evaluated on the Berkeley Function-Calling Leaderboard (BFCL).
| Model (4B - 235B) | Multi-Turn (Overall) | Single-Turn (Overall) |
|---|---|---|
| Qwen3-4B-Instruct (Base) | 22.13 | 82.14 |
| Qwen3-4B + HardGen (RL) | 63.13 | 87.14 |
| Gemini-3-Pro-Preview | 60.75 | 86.89 |
| DeepSeek-V3.2-Exp | 44.88 | 80.77 |
| GPT-5.2-2025-12-11 | 28.13 | 76.12 |
The performance of models trained upon Llama-3.1-8B-Instruct on agentic tasks (Web Search and Memory).
| Model | BFCLv4 Overall Score |
|---|---|
| HardGen-8B (RL) | 20.42 |
| CoALM-8B | 1.40 |
| ToolACE-2-8B | 13.50 |
| BitAgent-8B | 8.24 |
| xLAM-2-8b-fc-r | 10.24 |
This work is part of the open-source project AWorld, InclusionAI.
If you use this dataset in your research, please cite the related two papers:
@article{hao2026failure,
title={From Failure to Mastery: Generating Hard Samples for Tool-use Agents},
author={Hao, Bingguang and Xu, Zengzhuang and Wen, Yuntao and Xu, Xinyi and Liu, Yang and Zhao, Tong and Wang, Maolin and Chen, Long and Wang, Dong and Chen, Yicheng and others},
journal={arXiv preprint arXiv:2601.01498},
year={2026}
}
@article{xu2025funreason,
title={FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use},
author={Zengzhuang Xu, Bingguang Hao, Zechuan Wang, Yuntao Wen, Xinyi Xu, Yang Liu, Long Chen, Dong Wang, Maolin Wang, Tong Zhao, Yicheng Chen, Cunyin Peng, Jinjie Gu, Leilei Gan, Xiangyu Zhao, Chenyi Zhuang, Shi Gu},
journal={arXiv preprint arXiv:2510.24645},
year={2025}
}
For inquiries, please contact:
bingguanghao7@gmail.com29 commits