BingguangHao/FunReason

This is the official repository of the paper "BalanceSFT: Improving LLM Function Calling with Balanced Training Signals and Data Hardness"

57

stars

13

commits

Jul 2, 2026

updated

README

BalanceSFT: Improving LLM Function Calling with Balanced Training Signals and Data Hardness

  🤗 Dataset   |   🤗 Model   |    📑 Paper    |   📖 Github

[!IMPORTANT]

  • Our paper is accepted by ACL 2026.

  • We have released our training dataset!

  • Please give a ⭐️ to follow the update which is also an incentive for us.

Abstract

While Supervised Fine-Tuning (SFT) is the prevailing method for equipping Large Language Models (LLMs) with function calling capabilities, its effectiveness is often compromised by two critical challenges: 1) Imbalanced Training Signals, where lengthy Chain-of-Thought (CoT) reasoning tokens dominate the training signals over concise function calls in the learning objective, and 2) Imbalanced Data Hardness, characterized by a scarcity of hard training examples. To overcome these limitations, we propose Balanced Supervised Fine-tuning (BalanceSFT), a novel framework incorporates two key components: a Self-adjusted Signal Balancing (SSB) loss that employs a learnable hyperparameter to dynamically adjust the token contributions of CoT reasoning and function calls, together with a Hard Data Re-sampling (HDR) strategy that establishes a feedback loop to selectively generate new, high-quality complex data guided by model errors. Extensive experiments demonstrate the effectiveness of our proposed BalanceSFT framework. With BalanceSFT, a 7B model achieves function calling performance on par with state-of-the-art giants like GPT-4o. Our code, models, and dataset are open-sourced.

BalanceSFT

thinking_template

Overview of BalanceSFT's data refinement pipeline. It starts with a standard function call dataset, which is refined through a Base Quality Check and Answer Check to create initial training data and identify hard data. The model is first initialized via a Cold Start using the Self-adjusted Signal Balancing (SSB) Loss. Subsequently, the Hard Data Re-sampling (HDR) strategy creates a Self-evolving Loop where the model iteratively reasons on hard cases, generates new solutions, and undergoes quality-gated retraining.

We have released all the data that refined in this process, whcih contains high quality CoT data for function call.

Main Result

thinking_template

thinking_template

Citation

@inproceedings{hao-etal-2026-balancesft,
title = "{B}alance{SFT}: Improving {LLM} Function Calling with Balanced Training Signals and Data Hardness",
author = "Hao, Bingguang  and Xu, Zengzhuang  and Wang, Maolin  and Wen, Yuntao  and Chen, Yicheng  and Peng, Cunyin  and Chen, Long  and Zhao, Xiangyu  and Gu, Jinjie  and Zhuang, Chenyi  andZhang, Ji",
booktitle = "Findings of the {A}ssociation for {C}omputational {L}inguistics: {ACL} 2026",
year = "2026",
publisher = "Association for Computational Linguistics",
pages = "18094--18112",
ISBN = "979-8-89176-395-1"
}

Contributors

BingguangHao

13 commits

BingguangHao/FunReason

This is the official repository of the paper "BalanceSFT: Improving LLM Function Calling with Balanced Training Signals and Data Hardness"

57

stars

13

commits

Jul 2, 2026

updated

README

BalanceSFT: Improving LLM Function Calling with Balanced Training Signals and Data Hardness

  🤗 Dataset   |   🤗 Model   |    📑 Paper    |   📖 Github

[!IMPORTANT]

  • Our paper is accepted by ACL 2026.

  • We have released our training dataset!

  • Please give a ⭐️ to follow the update which is also an incentive for us.

Abstract

While Supervised Fine-Tuning (SFT) is the prevailing method for equipping Large Language Models (LLMs) with function calling capabilities, its effectiveness is often compromised by two critical challenges: 1) Imbalanced Training Signals, where lengthy Chain-of-Thought (CoT) reasoning tokens dominate the training signals over concise function calls in the learning objective, and 2) Imbalanced Data Hardness, characterized by a scarcity of hard training examples. To overcome these limitations, we propose Balanced Supervised Fine-tuning (BalanceSFT), a novel framework incorporates two key components: a Self-adjusted Signal Balancing (SSB) loss that employs a learnable hyperparameter to dynamically adjust the token contributions of CoT reasoning and function calls, together with a Hard Data Re-sampling (HDR) strategy that establishes a feedback loop to selectively generate new, high-quality complex data guided by model errors. Extensive experiments demonstrate the effectiveness of our proposed BalanceSFT framework. With BalanceSFT, a 7B model achieves function calling performance on par with state-of-the-art giants like GPT-4o. Our code, models, and dataset are open-sourced.

BalanceSFT

thinking_template

Overview of BalanceSFT's data refinement pipeline. It starts with a standard function call dataset, which is refined through a Base Quality Check and Answer Check to create initial training data and identify hard data. The model is first initialized via a Cold Start using the Self-adjusted Signal Balancing (SSB) Loss. Subsequently, the Hard Data Re-sampling (HDR) strategy creates a Self-evolving Loop where the model iteratively reasons on hard cases, generates new solutions, and undergoes quality-gated retraining.

We have released all the data that refined in this process, whcih contains high quality CoT data for function call.

Main Result

thinking_template

thinking_template

Citation

@inproceedings{hao-etal-2026-balancesft,
title = "{B}alance{SFT}: Improving {LLM} Function Calling with Balanced Training Signals and Data Hardness",
author = "Hao, Bingguang  and Xu, Zengzhuang  and Wang, Maolin  and Wen, Yuntao  and Chen, Yicheng  and Peng, Cunyin  and Chen, Long  and Zhao, Xiangyu  and Gu, Jinjie  and Zhuang, Chenyi  andZhang, Ji",
booktitle = "Findings of the {A}ssociation for {C}omputational {L}inguistics: {ACL} 2026",
year = "2026",
publisher = "Association for Computational Linguistics",
pages = "18094--18112",
ISBN = "979-8-89176-395-1"
}

Contributors

BingguangHao

13 commits