Bingguang/FunReason-MT

Dataset

80

stars

29

commits

2

linked in READMEs

Mar 23, 2026

updated

agent
Agentic Learning
BFCL
tool use

README

From Failure to Mastery: Generating Hard Samples for Tool-use Agents

arXiv arXiv Model GitHub Project Page


This is only a demonstration result obtained by directly deploying and sampling HardGen in the BFCL environment, not the dataset directly used in the HardGen paper.

[!IMPORTANT] Important Hint

  • This is an extension of the technical report FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use
  • To allow the model to learn from errors, we specifically construct erroneous environmental responses. If you wish to delete this data, please delete the trajectories where error_tool_response is true.
  • This is an initial version of our data, we will release the up-to-date version after the acceptance of our paper.

Dataset

The training set comprises 16,000 high-quality multi-turn samples. This dataset was generated using the three-phase HardGen (FunReason-MT) data synthesis framework, which focuses on generating complex trajectories.

πŸ“Š HardGen Evaluation Results

The model built upon HardGen is rigorously evaluated on the Berkeley Function-Calling Leaderboard (BFCL).

BFCLv3 Multi-Turn and Single-Turn Performance

Model (4B - 235B)Multi-Turn (Overall)Single-Turn (Overall)
Qwen3-4B-Instruct (Base)22.1382.14
Qwen3-4B + HardGen (RL)63.1387.14
Gemini-3-Pro-Preview60.7586.89
DeepSeek-V3.2-Exp44.8880.77
GPT-5.2-2025-12-1128.1376.12

BFCL Agentic Evaluation (BFCLv4 OOD)

The performance of models trained upon Llama-3.1-8B-Instruct on agentic tasks (Web Search and Memory).

ModelBFCLv4 Overall Score
HardGen-8B (RL)20.42
CoALM-8B1.40
ToolACE-2-8B13.50
BitAgent-8B8.24
xLAM-2-8b-fc-r10.24

Training Details

  • Training Libraries: LLama-Factory and Verl.
  • Methodology: Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL).
  • Hardware: Conducted on 8 NVIDIA A100 GPUs.

This work is part of the open-source project AWorld, InclusionAI.

If you use this dataset in your research, please cite the related two papers:

@article{hao2026failure,
  title={From Failure to Mastery: Generating Hard Samples for Tool-use Agents},
  author={Hao, Bingguang and Xu, Zengzhuang and Wen, Yuntao and Xu, Xinyi and Liu, Yang and Zhao, Tong and Wang, Maolin and Chen, Long and Wang, Dong and Chen, Yicheng and others},
  journal={arXiv preprint arXiv:2601.01498},
  year={2026}
}
@article{xu2025funreason,
  title={FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use},
  author={Zengzhuang Xu, Bingguang Hao, Zechuan Wang, Yuntao Wen, Xinyi Xu, Yang Liu, Long Chen, Dong Wang, Maolin Wang, Tong Zhao, Yicheng Chen, Cunyin Peng, Jinjie Gu, Leilei Gan, Xiangyu Zhao, Chenyi Zhuang, Shi Gu},
  journal={arXiv preprint arXiv:2510.24645},
  year={2025}
}

Contact

For inquiries, please contact:

  • bingguanghao7@gmail.com

Contributors

Bingguang

29 commits

Bingguang/FunReason-MT

Dataset

80

stars

29

commits

2

linked in READMEs

Mar 23, 2026

updated

agent
Agentic Learning
BFCL
tool use

README

From Failure to Mastery: Generating Hard Samples for Tool-use Agents

arXiv arXiv Model GitHub Project Page


This is only a demonstration result obtained by directly deploying and sampling HardGen in the BFCL environment, not the dataset directly used in the HardGen paper.

[!IMPORTANT] Important Hint

  • This is an extension of the technical report FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use
  • To allow the model to learn from errors, we specifically construct erroneous environmental responses. If you wish to delete this data, please delete the trajectories where error_tool_response is true.
  • This is an initial version of our data, we will release the up-to-date version after the acceptance of our paper.

Dataset

The training set comprises 16,000 high-quality multi-turn samples. This dataset was generated using the three-phase HardGen (FunReason-MT) data synthesis framework, which focuses on generating complex trajectories.

πŸ“Š HardGen Evaluation Results

The model built upon HardGen is rigorously evaluated on the Berkeley Function-Calling Leaderboard (BFCL).

BFCLv3 Multi-Turn and Single-Turn Performance

Model (4B - 235B)Multi-Turn (Overall)Single-Turn (Overall)
Qwen3-4B-Instruct (Base)22.1382.14
Qwen3-4B + HardGen (RL)63.1387.14
Gemini-3-Pro-Preview60.7586.89
DeepSeek-V3.2-Exp44.8880.77
GPT-5.2-2025-12-1128.1376.12

BFCL Agentic Evaluation (BFCLv4 OOD)

The performance of models trained upon Llama-3.1-8B-Instruct on agentic tasks (Web Search and Memory).

ModelBFCLv4 Overall Score
HardGen-8B (RL)20.42
CoALM-8B1.40
ToolACE-2-8B13.50
BitAgent-8B8.24
xLAM-2-8b-fc-r10.24

Training Details

  • Training Libraries: LLama-Factory and Verl.
  • Methodology: Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL).
  • Hardware: Conducted on 8 NVIDIA A100 GPUs.

This work is part of the open-source project AWorld, InclusionAI.

If you use this dataset in your research, please cite the related two papers:

@article{hao2026failure,
  title={From Failure to Mastery: Generating Hard Samples for Tool-use Agents},
  author={Hao, Bingguang and Xu, Zengzhuang and Wen, Yuntao and Xu, Xinyi and Liu, Yang and Zhao, Tong and Wang, Maolin and Chen, Long and Wang, Dong and Chen, Yicheng and others},
  journal={arXiv preprint arXiv:2601.01498},
  year={2026}
}
@article{xu2025funreason,
  title={FunReason-MT Technical Report: Advanced Data Synthesis Solution for Real-world Multi-Turn Tool-use},
  author={Zengzhuang Xu, Bingguang Hao, Zechuan Wang, Yuntao Wen, Xinyi Xu, Yang Liu, Long Chen, Dong Wang, Maolin Wang, Tong Zhao, Yicheng Chen, Cunyin Peng, Jinjie Gu, Leilei Gan, Xiangyu Zhao, Chenyi Zhuang, Shi Gu},
  journal={arXiv preprint arXiv:2510.24645},
  year={2025}
}

Contact

For inquiries, please contact:

  • bingguanghao7@gmail.com

Contributors

Bingguang

29 commits