THUDM/ComplexFuncBench

Dataset

15

stars

7

commits

2

linked in READMEs

Jan 22, 2025

updated

README

Introduction

Complex Function Calling Benchmark (ComplexFuncBench) is specillly designed for complex function calling evaluation. The ComplexFuncBench dataset encompass 1,000 complex function calling samples from five aspects: (1) Function calling with multi-step in single turn; (2) Function calling with user-provided constraints; (3) Function calling that requires parameter value reasoning from implicit information; (4) Function calling with long parameter values that exceed 500 tokens; and (5) Function calling with 128k long-context length.

If you wish to use this dataset for automated evaluation, please refer to our github.

Paper: https://huggingface.co/papers/2501.10132

Leaderboard

ModelOverall Success RateOverall Call Acc.CompletenessCorrectness
Claude-3.5-Sonnet (20241022)61.0079.271.841.85
GPT-4o (2024-08-06)60.5080.551.661.75
GLM-4-Long57.1076.351.721.74
GPT-4-Turbo (2024-04-09)49.5071.381.721.81
Claude-3.5-Haiku (20241022)45.8069.501.791.71
Qwen2.5-72B40.1058.321.801.75
Mistral Large 220.1048.780.941.0
GLM-4-9B9.4027.971.151.03
Qwen2.5-7B5.018.191.51.47
Llama-3.1-405B4.0011.870.430.30
Llama-3.1-70B2.708.170.670.36
Llama-3.1-8B0.101.340.180.09

Dataset Statistics

HotelsFlightsCar RentalAttractionCrossTotal
Num Samples150150150150400600
Avg. Steps3.333.332.872.863.53.26
Avg. Calls4.295.334.573.66.05.07

Citation

If you find our work helpful for your research, please consider citing our work.

@misc{zhong2025complexfuncbench,
      title={ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario}, 
      author={Lucen Zhong and Zhengxiao Du and Xiaohan Zhang and Haiyi Hu and Jie Tang},
      year={2025},
      eprint={2501.10132},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2501.10132}, 
}

Contributors

anchorzhong

5 commits

nielsr

1 commits

ZR
zR

1 commits

THUDM/ComplexFuncBench

Dataset

15

stars

7

commits

2

linked in READMEs

Jan 22, 2025

updated

README

Introduction

Complex Function Calling Benchmark (ComplexFuncBench) is specillly designed for complex function calling evaluation. The ComplexFuncBench dataset encompass 1,000 complex function calling samples from five aspects: (1) Function calling with multi-step in single turn; (2) Function calling with user-provided constraints; (3) Function calling that requires parameter value reasoning from implicit information; (4) Function calling with long parameter values that exceed 500 tokens; and (5) Function calling with 128k long-context length.

If you wish to use this dataset for automated evaluation, please refer to our github.

Paper: https://huggingface.co/papers/2501.10132

Leaderboard

ModelOverall Success RateOverall Call Acc.CompletenessCorrectness
Claude-3.5-Sonnet (20241022)61.0079.271.841.85
GPT-4o (2024-08-06)60.5080.551.661.75
GLM-4-Long57.1076.351.721.74
GPT-4-Turbo (2024-04-09)49.5071.381.721.81
Claude-3.5-Haiku (20241022)45.8069.501.791.71
Qwen2.5-72B40.1058.321.801.75
Mistral Large 220.1048.780.941.0
GLM-4-9B9.4027.971.151.03
Qwen2.5-7B5.018.191.51.47
Llama-3.1-405B4.0011.870.430.30
Llama-3.1-70B2.708.170.670.36
Llama-3.1-8B0.101.340.180.09

Dataset Statistics

HotelsFlightsCar RentalAttractionCrossTotal
Num Samples150150150150400600
Avg. Steps3.333.332.872.863.53.26
Avg. Calls4.295.334.573.66.05.07

Citation

If you find our work helpful for your research, please consider citing our work.

@misc{zhong2025complexfuncbench,
      title={ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario}, 
      author={Lucen Zhong and Zhengxiao Du and Xiaohan Zhang and Haiyi Hu and Jie Tang},
      year={2025},
      eprint={2501.10132},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2501.10132}, 
}

Contributors

anchorzhong

5 commits

nielsr

1 commits

ZR
zR

1 commits