"what, how, where, and how well? a survey on test-time scaling in large language models" repository
See the code
Our repository, Awesome Test-time-Scaling in LLMs, gathers available papers on test-time scaling, to our current knowledge. Unlike other repositories that categorize papers, we decompose each paper's contributions based on the taxonomy provided by "What, How, Where, and How Well? A Survey on Test-Time Scaling in Large Language Models" facilitating easier understand and comparison for readers.
Figure 1: A Visual Map and Comparison: From What to Scale to How to Scale..
[13/Apr/2025] π The Second Version is released:
[9/Apr/2025] π Our repository is created.
[31/Mar/2025] π Our initial survey is on Arxiv!
As enthusiasm for scaling computation (data and parameters) in the pertaining era gradually diminished, test-time scaling (TTS)βalso referred to as βtest-time computingββhas emerged as a prominent research focus. Recent studies demonstrate that TTS can further elicit the problem-solving capabilities of large language models (LLMs), enabling significant breakthroughs not only in reasoning-intensive tasks, such as mathematics and coding, but also in general tasks like open-ended Q&A. However, despite the explosion of recent efforts in this area, there remains an urgent need for a comprehensive survey offering systemic understanding. To fill this gap, we propose a unified, hierarchical framework structured along four orthogonal dimensions of TTS research: what to scale, how to scale, where to scale, and how well to scale. Building upon this taxonomy, we conduct a holistic review of methods, application scenarios, and assessment aspects, and present an organized decomposition that highlights the unique contributions of individual methods within the broader TTS landscape.
Figure 2: omparison of Scaling Paradigms in Pre-training and Test-time Phases..
``What to scale'' refers to the specific form of TTS that is expanded or adjusted to enhance an LLMβs performance during inference.
HTML
87.2%
JavaScript
9.6%
CSS
3.2%
"what, how, where, and how well? a survey on test-time scaling in large language models" repository
See the code
Our repository, Awesome Test-time-Scaling in LLMs, gathers available papers on test-time scaling, to our current knowledge. Unlike other repositories that categorize papers, we decompose each paper's contributions based on the taxonomy provided by "What, How, Where, and How Well? A Survey on Test-Time Scaling in Large Language Models" facilitating easier understand and comparison for readers.
Figure 1: A Visual Map and Comparison: From What to Scale to How to Scale..
[13/Apr/2025] π The Second Version is released:
[9/Apr/2025] π Our repository is created.
[31/Mar/2025] π Our initial survey is on Arxiv!
As enthusiasm for scaling computation (data and parameters) in the pertaining era gradually diminished, test-time scaling (TTS)βalso referred to as βtest-time computingββhas emerged as a prominent research focus. Recent studies demonstrate that TTS can further elicit the problem-solving capabilities of large language models (LLMs), enabling significant breakthroughs not only in reasoning-intensive tasks, such as mathematics and coding, but also in general tasks like open-ended Q&A. However, despite the explosion of recent efforts in this area, there remains an urgent need for a comprehensive survey offering systemic understanding. To fill this gap, we propose a unified, hierarchical framework structured along four orthogonal dimensions of TTS research: what to scale, how to scale, where to scale, and how well to scale. Building upon this taxonomy, we conduct a holistic review of methods, application scenarios, and assessment aspects, and present an organized decomposition that highlights the unique contributions of individual methods within the broader TTS landscape.
Figure 2: omparison of Scaling Paradigms in Pre-training and Test-time Phases..
``What to scale'' refers to the specific form of TTS that is expanded or adjusted to enhance an LLMβs performance during inference.
HTML
87.2%
JavaScript
9.6%
CSS
3.2%