Code and data for "Test-time scaling in reasoning models is not effective for knowledge-intensive tasks yet" (COLM 2026)
Python
9
7 commits
updated Aug 4, 2026
This repository contains the data and code for the paper: "Test-time scaling in reasoning models is not effective for knowledge-intensive tasks yet" (COLM 2026)
This repository contains the benchmarks, runnable scripts, and model outputs for all experiments. It includes:
benchmarks/: Datasets used in our experiments: SimpleQA, FACTS Parametric, and FRAMES.scripts/: Generation and evaluation scripts, plus their shared utilities.outputs/: Evaluated model outputs from 14 reasoning models with different thinking levels on 3 benchmarks.figures/: Main result figures from the paper.First, clone the repository:
git clone https://github.com/XuZhao0/tts-knowledge.git
cd tts-knowledge
Create and activate a new Python 3.10+ environment using conda. For example:
conda create -n tts_knowledge python=3.10
conda activate tts_knowledge
Install package requirements with the following command:
pip install -r requirements.txt
Run the commands below from the repository root so the default benchmark and output paths resolve correctly. All executable Python files are kept in scripts/; scripts/tool.py contains shared JSONL and API-client helpers. See scripts/README.md for a compact entry-point reference.
To run experiments on proprietary models, you first need API keys for: OpenAI, Anthropic Claude, Google Gemini, XAI Grok
Run experiments on OpenAI closed-source models, use the following command. You can change the --model and --effort arguments to test different models and thinking levels. Replace 'your_openai_api_key' with your actual OpenAI API key. More details can be found in scripts/openai_close_generate.py.
python scripts/openai_close_generate.py --model gpt-5-mini --effort low --api_key 'your_openai_api_key'
Run experiments on XAI Grok models:
python scripts/grok_generate.py --effort low --api_key 'your_xai_api_key'
Run experiments on Gemini models. You can adjust the --thinking_budget parameter to set different thinking levels. Disable the thinking mode by setting --thinking_budget to 0. Note that only Gemini 2.5 Flash supports a thinking_budget of 0; the minimum budget for Gemini 2.5 Pro is 128. See Gemini API documentation for more details.
python scripts/gemini_generate.py --thinking_budget 512 --api_key 'your_google_api_key'
Run experiments on Anthropic Claude models. You can disable the thinking mode by setting --thinking_budget to 0.
python scripts/claude_generate.py --thinking_budget 1024 --api_key 'your_anthropic_api_key'
Run experiments on gpt-oss models, using the following command:
python scripts/gpt_oss_generate.py --effort low
Run experiments on DeepSeek-R1-Distill models with budget forcing, using the following command. You can adjust the --extend_times parameter to set different thinking levels. If it is set to 0, it means natural output without budget forcing.
python scripts/ds_r1_distill_exd.py --extend_times 2
Run experiments on Qwen3 models (thinking mode) with budget forcing:
python scripts/qwen3_exd.py --extend_times 2
Run experiments on Qwen3 models on non-thinking mode:
python scripts/qwen3_no_think.py
Evaluate model responses with the following command. We use gpt-4o-mini as the grader model. Replace your_openai_api_key with your OpenAI API key.
python scripts/evaluation.py --input_path '<model_response_file_path.jsonl>' --api_key 'your_openai_api_key'
We release evaluated outputs from 14 reasoning models with varying thinking levels (including non-thinking for several models). Each output includes: model response, reasoning trace (if applicable), and evaluation label (correct [A], incorrect [B], or not attempted [C]), and metadata such as token count.
If you find our code or data useful, please cite:
@article{zhao2025testtimescalingreasoningmodels,
title={Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet},
author={James Xu Zhao and Bryan Hooi and See-Kiong Ng},
year={2025},
eprint={2509.06861},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2509.06861},
}
For questions or suggestions, feel free to contact: xu.zhao@u.nus.edu
Our budget forcing implementation is adapted from r1-overthinker. We thank the developers for open-sourcing their code.
7 commits
Python
100.0%
Code and data for "Test-time scaling in reasoning models is not effective for knowledge-intensive tasks yet" (COLM 2026)
Python
9
7 commits
updated Aug 4, 2026
This repository contains the data and code for the paper: "Test-time scaling in reasoning models is not effective for knowledge-intensive tasks yet" (COLM 2026)
This repository contains the benchmarks, runnable scripts, and model outputs for all experiments. It includes:
benchmarks/: Datasets used in our experiments: SimpleQA, FACTS Parametric, and FRAMES.scripts/: Generation and evaluation scripts, plus their shared utilities.outputs/: Evaluated model outputs from 14 reasoning models with different thinking levels on 3 benchmarks.figures/: Main result figures from the paper.First, clone the repository:
git clone https://github.com/XuZhao0/tts-knowledge.git
cd tts-knowledge
Create and activate a new Python 3.10+ environment using conda. For example:
conda create -n tts_knowledge python=3.10
conda activate tts_knowledge
Install package requirements with the following command:
pip install -r requirements.txt
Run the commands below from the repository root so the default benchmark and output paths resolve correctly. All executable Python files are kept in scripts/; scripts/tool.py contains shared JSONL and API-client helpers. See scripts/README.md for a compact entry-point reference.
To run experiments on proprietary models, you first need API keys for: OpenAI, Anthropic Claude, Google Gemini, XAI Grok
Run experiments on OpenAI closed-source models, use the following command. You can change the --model and --effort arguments to test different models and thinking levels. Replace 'your_openai_api_key' with your actual OpenAI API key. More details can be found in scripts/openai_close_generate.py.
python scripts/openai_close_generate.py --model gpt-5-mini --effort low --api_key 'your_openai_api_key'
Run experiments on XAI Grok models:
python scripts/grok_generate.py --effort low --api_key 'your_xai_api_key'
Run experiments on Gemini models. You can adjust the --thinking_budget parameter to set different thinking levels. Disable the thinking mode by setting --thinking_budget to 0. Note that only Gemini 2.5 Flash supports a thinking_budget of 0; the minimum budget for Gemini 2.5 Pro is 128. See Gemini API documentation for more details.
python scripts/gemini_generate.py --thinking_budget 512 --api_key 'your_google_api_key'
Run experiments on Anthropic Claude models. You can disable the thinking mode by setting --thinking_budget to 0.
python scripts/claude_generate.py --thinking_budget 1024 --api_key 'your_anthropic_api_key'
Run experiments on gpt-oss models, using the following command:
python scripts/gpt_oss_generate.py --effort low
Run experiments on DeepSeek-R1-Distill models with budget forcing, using the following command. You can adjust the --extend_times parameter to set different thinking levels. If it is set to 0, it means natural output without budget forcing.
python scripts/ds_r1_distill_exd.py --extend_times 2
Run experiments on Qwen3 models (thinking mode) with budget forcing:
python scripts/qwen3_exd.py --extend_times 2
Run experiments on Qwen3 models on non-thinking mode:
python scripts/qwen3_no_think.py
Evaluate model responses with the following command. We use gpt-4o-mini as the grader model. Replace your_openai_api_key with your OpenAI API key.
python scripts/evaluation.py --input_path '<model_response_file_path.jsonl>' --api_key 'your_openai_api_key'
We release evaluated outputs from 14 reasoning models with varying thinking levels (including non-thinking for several models). Each output includes: model response, reasoning trace (if applicable), and evaluation label (correct [A], incorrect [B], or not attempted [C]), and metadata such as token count.
If you find our code or data useful, please cite:
@article{zhao2025testtimescalingreasoningmodels,
title={Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet},
author={James Xu Zhao and Bryan Hooi and See-Kiong Ng},
year={2025},
eprint={2509.06861},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2509.06861},
}
For questions or suggestions, feel free to contact: xu.zhao@u.nus.edu
Our budget forcing implementation is adapted from r1-overthinker. We thank the developers for open-sourcing their code.
7 commits
Python
100.0%