dayuyang1999/Awesome-Code-Reasoning

366

15 commits

updated Jun 23, 2025

See the code

README

πŸ§ πŸ’» A Survey on Code Reasoning πŸ€–πŸ”

This is the official repository of our paper: Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs πŸš€

arXiv Maintenance PR's Welcome Awesome

Paper page

Code Reasoning
Taxonomy of interplay between Code and Reasoning

Please do not hesitate to contact us or launch pull requests if you find any related papers that are missing in our paper, and let us know if you discover any mistakes or have suggestions by emailing us: yangdayu1997@gmail.com βœ‰οΈ

News πŸ“°

  • Update on 2025/02/27: Paper is released on arXiv. πŸŽ‰ arXiv
  • Update on 2025/02/11: Updating Reading Lists πŸ“š πŸ“–

Citation πŸ“–

🫢 If you are interested in our work or find this repository helpful, please consider using the following citation format when referencing our paper:

@article{yang2025codereasoning,
  title={Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs},
  author={Yang, Dayu and Liu, Tianyang and Zhang, Daoan and others},
  journal={arXiv preprint arXiv:2502.19411},
  year={2025}
}

Acknowledgements

This is an open collaborative research project among:

Contributors

Following the release of this paper, we have received numerous valuable comments from our readers. We sincerely thank those who have reached out with constructive suggestions and feedback.

This repository is actively maintained, and we welcome your contributions! If you have any questions about this list of resources, please feel free to contact me at yangdayu1997@gmail.com.

Table Of Contents

Code-aided Reasoning

Generating as Code

Paper TitleURLRelease Date
PAL: Program-aided Language Modelshttps://arxiv.org/abs/2211.104352022-11-18
Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Taskshttps://arxiv.org/abs/2211.125882022-11-22
Chain of Code: Reasoning with a Language Model-Augmented Code Emulatorhttps://arxiv.org/abs/2312.044742023-12-07
Program-Aided Reasoners (better) Know What They Knowhttps://arxiv.org/abs/2311.095532023-11-16
When Do Program-of-Thoughts Work for Reasoning?https://arxiv.org/abs/2308.154522023-08-29
MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoninghttps://arxiv.org/abs/2310.037312023-10-05
MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Codehttps://arxiv.org/abs/2410.081962024-10-10
Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMshttps://arxiv.org/abs/2401.100652024-01-18
Steering Large Language Models between Code Execution and Textual Reasoninghttps://arxiv.org/abs/2410.035242024-10-04
Interactive and Expressive Code-Augmented Planning with Large Language Modelshttps://arxiv.org/abs/2411.138262024-11-21
Gap-Filling Prompting Enhances Code-Assisted Mathematical Reasoninghttps://arxiv.org/abs/2411.054072024-11-08
Can LLMs Reason in the Wild with Programs?https://arxiv.org/abs/2406.137642024-06-19
Planning-Driven Programming: A Large Language Model Programming Workflowhttps://arxiv.org/abs/2411.145032024-11-21
Unlocking Reasoning Potential in Large Language Models by Scaling Code-form Planninghttps://arxiv.org/abs/2409.124522024-09-19
INC-Math: Integrating Natural Language and Code for Enhanced Mathematical Reasoninghttps://arxiv.org/abs/2409.193812024-09-28
Learning to Reason via Program Generation, Emulation, and Searchhttps://arxiv.org/abs/2405.163372024-05-25
NExT: Teaching Large Language Models to Reason about Code Executionhttps://arxiv.org/abs/2404.146622024-04-23
Unlocking Reasoning Potential in Large Language Models by Scaling Code-form Planninghttps://arxiv.org/abs/2409.124522024-09-19
Code Prompting: a Neural Symbolic Method for Complex Reasoning in Large Language Modelshttps://arxiv.org/abs/2305.185072023-05-29
CodeI/O: Condensing Reasoning Patterns via Code Input-Output Predictionhttps://arxiv.org/abs/2502.073162025-02-11
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesishttps://arxiv.org/abs/2503.231452025-03-29
Evaluating Grounded Reasoning by Code-Assisted Large Language Models for Mathematicshttps://arxiv.org/abs/2504.176652025-04-24

Training with Code

Paper TitleURLRelease Date
CodeTrain: Pre-training LLMs with Code-Based Taskshttps://arxiv.org/abs/2401.111112024-01-05
Learning to Reason Through Code Exampleshttps://arxiv.org/abs/2312.222222023-12-15
Code-Augmented Training for Better Reasoninghttps://arxiv.org/abs/2311.333332023-11-20
Language Models of Code are Few-Shot Commonsense Learnershttps://arxiv.org/pdf/2210.071282022-12-06
Logic Distillation: Learning from Code Function by Function for Planning and Decision-makinghttps://arxiv.org/pdf/2407.194052024-07-28
Unlocking Reasoning Potential in Large Langauge Models by Scaling Code-form Planninghttps://arxiv.org/pdf/2409.124522022-10-04
ViStruct: Visual Structural Knowledge Extraction via Curriculum Guided Code-Vision Representationhttps://arxiv.org/pdf/2311.132582023-11-22
Eliciting Better Multilingual Structured Reasoning from LLMs through Codehttps://arxiv.org/pdf/2403.025672024-06-12
LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model Programshttps://arxiv.org/pdf/2312.043722024-04-04
MARIO: MAth Reasoning with code Interpreter Output – A Reproducible Pipelinehttps://arxiv.org/pdf/2401.081902024-02-21
Reasoning Like Program Executorshttps://arxiv.org/pdf/2201.114732022-10-22
SEMCODER: Training Code Language Models with Comprehensive Semantics Reasoninghttps://arxiv.org/pdf/2406.010062024-10-31
CodePMP: Scalable Preference Model Pretraining for Large Language Model Reasoninghttps://arxiv.org/pdf/2410.02229?2024-10-03
Siam: Self-improving code-assisted mathematical reasoning of large language modelshttps://arxiv.org/pdf/2408.15565?2024-08-28
Crystal: Illuminating LLM abilities on language and codehttps://arxiv.org/pdf/2411.041562024-11-06
At which training stage does code data help llms reasoning?https://arxiv.org/pdf/2309.162982023-09-03
Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models Reasoninghttps://arxiv.org/pdf/2405.205352024-12-12
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learninghttps://arxiv.org/abs/2501.129482025-01-22
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learninghttps://arxiv.org/abs/2505.164002025-05-22
OpenThoughts: A Systematic Investigation of Data Curation for Post-training Reasoning Modelshttps://arxiv.org/abs/2506.041782025-06-05
RuleReasoner: Reinforced Rule-based Reasoning with Dynamic Multi-domain Curriculum Learninghttps://arxiv.org/abs/2506.086722025-06-10
CoRT: Code-integrated Reasoning within Thinkinghttps://arxiv.org/abs/2506.098202025-06-11

Reasoning-enhanced Code Intelligence

Essential Code Intelligence

Paper TitleURLRelease Date
CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generationhttps://arxiv.org/abs/2102.046642021-02-09
Competition-level code generation with AlphaCodehttps://arxiv.org/abs/2108.077322021-08-16
Evaluating Large Language Models Trained on Codehttps://arxiv.org/abs/2107.033742021-07-07
Program Synthesis with Large Language Modelshttps://arxiv.org/abs/2108.077322021-08-16
A Systematic Evaluation of Large Language Models of Codehttps://arxiv.org/abs/2202.131692022-02-26
InCoder: A Generative Model for Code Infilling and Synthesishttps://arxiv.org/abs/2204.059992023-04-12
CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesishttps://arxiv.org/abs/2203.134742023-03-25
StarCoder: May the Source be with You!https://arxiv.org/abs/2305.061612023-05-10
Code Llama: Open Foundation Models for Codehttps://arxiv.org/abs/2308.129502023-08-24
RepoBench: Benchmarking Repository-Level Code Auto-Completion Systemshttps://arxiv.org/abs/2306.030912023-06-05
CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completionhttps://arxiv.org/abs/2310.112482023-10-17
StarCoder 2 and The Stack v2: The Next Generationhttps://arxiv.org/abs/2402.191732024-02-29
CodeGemma: Open Code Models Based on Gemmahttps://arxiv.org/abs/2406.114092024-06-17
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligencehttps://arxiv.org/abs/2406.119312024-06-17
Qwen2.5-Coder Technical Reporthttps://arxiv.org/abs/2409.121862024-09-18
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratingshttps://arxiv.org/abs/2501.012572025-01-02
Exploring Code Comprehension in Scientific Programming: Preliminary Insights from Research Scientistshttps://arxiv.org/abs/2501.100372025-01-17
COFFE: A Code Efficiency Benchmark for Code Generationhttps://arxiv.org/abs/2502.028272025-02-05
Evaluating the Generalization Capabilities of Large Language Models on Code Reasoninghttps://arxiv.org/abs/2504.055182025-04-07
rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Datasethttps://arxiv.org/abs/2505.212972025-05-27

Integration of Reasoning Capabilities

Reasoning for Code Generation

Paper TitleURLRelease Date
Chain-of-Thought Prompting Elicits Reasoning in Large Language Modelshttps://arxiv.org/abs/2201.119032022-01-28
Self-planning Code Generation with Large Language Modelshttps://arxiv.org/abs/2303.066892023-03-12
Structured Chain-of-Thought Prompting for Code Generationhttps://arxiv.org/abs/2305.065992023-05-11
CodeCoT: Tackling Code Syntax Errors in CoT Reasoning for Code Generationhttps://arxiv.org/abs/2308.087842023-08-17
CodePlan: Repository-level Coding using LLMs and Planninghttps://arxiv.org/abs/2309.124992023-09-21
Chain-of-Thought in Neural Code Generation: From and For Lightweight Language Modelshttps://arxiv.org/abs/2312.055622023-12-09
Planning In Natural Language Improves LLM Search For Code Generationhttps://arxiv.org/abs/2409.037332024-09-05
Chain of Grounded Objectives: Bridging Process and Goal-oriented Prompting for Code Generationhttps://arxiv.org/abs/2501.139782025-01-23
LLM-Guided Compositional Program Synthesishttps://arxiv.org/abs/2503.155402025-03-12
Modularization is Better: Effective Code Generation with Modular Promptinghttps://arxiv.org/abs/2503.124832025-03-16
Uncertainty-Guided Chain-of-Thought for Code Generation with LLMshttps://arxiv.org/abs/2503.153412025-03-19
MSCoT: Structured Chain-of-Thought Generation for Multiple Programming Languageshttps://arxiv.org/abs/2504.101782025-04-18
Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generationhttps://arxiv.org/abs/2506.069712025-06-08
Reasoning as a Resource: Optimizing Fast and Slow Thinking in Code Generation Modelshttps://arxiv.org/abs/2506.093962025-06-11

Reasoning Over Code

Paper TitleURLRelease Date
CodeQA: A Question Answering Dataset for Source Code Comprehensionhttps://arxiv.org/abs/2109.083652021-09-17
CRUXEval: A Benchmark for Code Reasoning, Understanding and Executionhttps://arxiv.org/abs/2401.030652024-01-05
CodeMind: A Framework to Challenge Large Language Models for Code Reasoninghttps://arxiv.org/abs/2402.096642024-02-15
Reasoning Runtime Behavior of a Program with LLM: How Far Are We?https://arxiv.org/abs/2403.164372024-03-25
NExT: Teaching Large Language Models to Reason about Code Executionhttps://arxiv.org/abs/2404.146622024-04-23
RepoQA: Evaluating Long Context Code Understandinghttps://arxiv.org/abs/2406.060252024-06-10
SelfPiCo: Self-Guided Partial Code Execution with LLMshttps://arxiv.org/abs/2407.169742024-07-24
CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding Capabilitieshttps://arxiv.org/abs/2410.019992024-10-02
What You See Is Not Always What You Get: An Empirical Study of Code Comprehensionhttps://arxiv.org/abs/2412.080982024-12-11
How Accurately Do Large Language Models Understand Code?https://arxiv.org/abs/2504.043722025-04-06

Interactive Programming

Paper TitleURLRelease Date
Interactive Program Synthesishttps://arxiv.org/abs/1703.035392017-03-10
Self-Refine: Iterative Refinement with Self-Feedbackhttps://arxiv.org/abs/2303.176512023-03-30
Teaching Large Language Models to Self-Debughttps://arxiv.org/abs/2304.051282023-04-11
Self-collaboration Code Generation via ChatGPThttps://arxiv.org/abs/2304.075902023-04-15
Self-Edit: Fault-Aware Code Editor for Code Generationhttps://arxiv.org/abs/2305.040872023-05-06
LeTI: Learning to Generate from Textual Interactionshttps://arxiv.org/abs/2305.103142023-05-17
InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedbackhttps://arxiv.org/abs/2306.148982023-06-26
CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-moduleshttps://arxiv.org/abs/2310.089922023-10-13
AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisationhttps://arxiv.org/abs/2312.130102023-12-20
OpenCodeInterpreter: Integrating Code Generation with Execution and Refinementhttps://arxiv.org/abs/2402.146582024-02-22
What Makes Large Language Models Reason in (Multi-Turn) Code Generation?https://arxiv.org/abs/2410.081052024-10-10
Revisit Self-Debugging with Self-Generated Tests for Code Generationhttps://arxiv.org/abs/2501.127932025-01-22
Large Language Model Guided Self-Debugging Code Generationhttps://arxiv.org/abs/2502.029282025-02-05
Interactive Agents to Overcome Ambiguity in Software Engineeringhttps://arxiv.org/abs/2502.130692025-02-18
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environmentshttps://arxiv.org/abs/2502.198522025-02-28
Prompt Alchemy: Automatic Prompt Refinement for Enhancing Code Generationhttps://arxiv.org/abs/2503.110852025-03-14
Humanity's Last Code Exam: Can Advanced LLMs Conquer Human's Hardest Code Competition?https://arxiv.org/abs/2506.127132025-06-12

Code Agents with Complex Reasoning

Paper TitleURLRelease Date
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?https://arxiv.org/abs/2310.067702023-10-10
CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challengeshttps://arxiv.org/abs/2401.073392024-01-14
Executable Code Actions Elicit Better LLM Agentshttps://arxiv.org/abs/2402.010302024-02-01
Cursor AI: The AI Code Editorhttps://www.cursor.com2024-02-17
Devin AI: Autonomous AI Software Engineerhttps://devin.ai2024-03-12
AutoCodeRover: Autonomous Program Improvementhttps://arxiv.org/abs/2404.054272024-04-08
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineeringhttps://arxiv.org/abs/2405.157932024-05-06
Agentless: Demystifying LLM-based Software Engineering Agentshttps://arxiv.org/abs/2407.014892024-07-01
OpenHands: An Open Platform for AI Software Developers as Generalist Agentshttps://arxiv.org/abs/2407.167412024-07-23
SWE-bench Verifiedhttps://openai.com/index/introducing-swe-bench-verified2024-08-13
HyperAgent: Generalist Software Engineering Agents to Solve Coding Tasks at Scalehttps://arxiv.org/abs/2409.162992024-09-09
SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?https://arxiv.org/abs/2410.038592024-10-04
Evaluating Software Development Agents: Patch Patterns, Code Quality, and Issue Complexity in Real-World GitHub Scenarioshttps://arxiv.org/abs/2410.124682024-10-16
Verbal Process Supervision Elicits Better Coding Agentshttps://arxiv.org/abs/2503.184942025-03-24
A Self-Improving Coding Agenthttps://arxiv.org/abs/2504.152282025-04-21
Breakpoint: A Benchmark for Systematic and Scalable Evaluation of Long-Horizon Code Repairhttps://arxiv.org/abs/2506.001722025-05-31
Code Researcher: Deep Research Agent for Large Systems Code and Commit Historyhttps://arxiv.org/abs/2506.110602025-05-27
Coding Agents with Multimodal Browsing are Generalist Problem Solvershttps://arxiv.org/abs/2506.030112025-06-03
UTBoost: Rigorous Evaluation of Coding Agents on SWE-Benchhttps://arxiv.org/abs/2506.092892025-06-10

Contributors

DwanZhang-AI

9 commits

dayuyang1999

3 commits

Leolty

2 commits

iCSawyer

1 commits

dayuyang1999/Awesome-Code-Reasoning

366

15 commits

updated Jun 23, 2025

See the code

README

πŸ§ πŸ’» A Survey on Code Reasoning πŸ€–πŸ”

This is the official repository of our paper: Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs πŸš€

arXiv Maintenance PR's Welcome Awesome

Paper page

Code Reasoning
Taxonomy of interplay between Code and Reasoning

Please do not hesitate to contact us or launch pull requests if you find any related papers that are missing in our paper, and let us know if you discover any mistakes or have suggestions by emailing us: yangdayu1997@gmail.com βœ‰οΈ

News πŸ“°

  • Update on 2025/02/27: Paper is released on arXiv. πŸŽ‰ arXiv
  • Update on 2025/02/11: Updating Reading Lists πŸ“š πŸ“–

Citation πŸ“–

🫢 If you are interested in our work or find this repository helpful, please consider using the following citation format when referencing our paper:

@article{yang2025codereasoning,
  title={Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs},
  author={Yang, Dayu and Liu, Tianyang and Zhang, Daoan and others},
  journal={arXiv preprint arXiv:2502.19411},
  year={2025}
}

Acknowledgements

This is an open collaborative research project among:

Contributors

Following the release of this paper, we have received numerous valuable comments from our readers. We sincerely thank those who have reached out with constructive suggestions and feedback.

This repository is actively maintained, and we welcome your contributions! If you have any questions about this list of resources, please feel free to contact me at yangdayu1997@gmail.com.

Table Of Contents

Code-aided Reasoning

Generating as Code

Paper TitleURLRelease Date
PAL: Program-aided Language Modelshttps://arxiv.org/abs/2211.104352022-11-18
Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Taskshttps://arxiv.org/abs/2211.125882022-11-22
Chain of Code: Reasoning with a Language Model-Augmented Code Emulatorhttps://arxiv.org/abs/2312.044742023-12-07
Program-Aided Reasoners (better) Know What They Knowhttps://arxiv.org/abs/2311.095532023-11-16
When Do Program-of-Thoughts Work for Reasoning?https://arxiv.org/abs/2308.154522023-08-29
MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoninghttps://arxiv.org/abs/2310.037312023-10-05
MathCoder2: Better Math Reasoning from Continued Pretraining on Model-translated Mathematical Codehttps://arxiv.org/abs/2410.081962024-10-10
Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMshttps://arxiv.org/abs/2401.100652024-01-18
Steering Large Language Models between Code Execution and Textual Reasoninghttps://arxiv.org/abs/2410.035242024-10-04
Interactive and Expressive Code-Augmented Planning with Large Language Modelshttps://arxiv.org/abs/2411.138262024-11-21
Gap-Filling Prompting Enhances Code-Assisted Mathematical Reasoninghttps://arxiv.org/abs/2411.054072024-11-08
Can LLMs Reason in the Wild with Programs?https://arxiv.org/abs/2406.137642024-06-19
Planning-Driven Programming: A Large Language Model Programming Workflowhttps://arxiv.org/abs/2411.145032024-11-21
Unlocking Reasoning Potential in Large Language Models by Scaling Code-form Planninghttps://arxiv.org/abs/2409.124522024-09-19
INC-Math: Integrating Natural Language and Code for Enhanced Mathematical Reasoninghttps://arxiv.org/abs/2409.193812024-09-28
Learning to Reason via Program Generation, Emulation, and Searchhttps://arxiv.org/abs/2405.163372024-05-25
NExT: Teaching Large Language Models to Reason about Code Executionhttps://arxiv.org/abs/2404.146622024-04-23
Unlocking Reasoning Potential in Large Language Models by Scaling Code-form Planninghttps://arxiv.org/abs/2409.124522024-09-19
Code Prompting: a Neural Symbolic Method for Complex Reasoning in Large Language Modelshttps://arxiv.org/abs/2305.185072023-05-29
CodeI/O: Condensing Reasoning Patterns via Code Input-Output Predictionhttps://arxiv.org/abs/2502.073162025-02-11
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesishttps://arxiv.org/abs/2503.231452025-03-29
Evaluating Grounded Reasoning by Code-Assisted Large Language Models for Mathematicshttps://arxiv.org/abs/2504.176652025-04-24

Training with Code

Paper TitleURLRelease Date
CodeTrain: Pre-training LLMs with Code-Based Taskshttps://arxiv.org/abs/2401.111112024-01-05
Learning to Reason Through Code Exampleshttps://arxiv.org/abs/2312.222222023-12-15
Code-Augmented Training for Better Reasoninghttps://arxiv.org/abs/2311.333332023-11-20
Language Models of Code are Few-Shot Commonsense Learnershttps://arxiv.org/pdf/2210.071282022-12-06
Logic Distillation: Learning from Code Function by Function for Planning and Decision-makinghttps://arxiv.org/pdf/2407.194052024-07-28
Unlocking Reasoning Potential in Large Langauge Models by Scaling Code-form Planninghttps://arxiv.org/pdf/2409.124522022-10-04
ViStruct: Visual Structural Knowledge Extraction via Curriculum Guided Code-Vision Representationhttps://arxiv.org/pdf/2311.132582023-11-22
Eliciting Better Multilingual Structured Reasoning from LLMs through Codehttps://arxiv.org/pdf/2403.025672024-06-12
LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model Programshttps://arxiv.org/pdf/2312.043722024-04-04
MARIO: MAth Reasoning with code Interpreter Output – A Reproducible Pipelinehttps://arxiv.org/pdf/2401.081902024-02-21
Reasoning Like Program Executorshttps://arxiv.org/pdf/2201.114732022-10-22
SEMCODER: Training Code Language Models with Comprehensive Semantics Reasoninghttps://arxiv.org/pdf/2406.010062024-10-31
CodePMP: Scalable Preference Model Pretraining for Large Language Model Reasoninghttps://arxiv.org/pdf/2410.02229?2024-10-03
Siam: Self-improving code-assisted mathematical reasoning of large language modelshttps://arxiv.org/pdf/2408.15565?2024-08-28
Crystal: Illuminating LLM abilities on language and codehttps://arxiv.org/pdf/2411.041562024-11-06
At which training stage does code data help llms reasoning?https://arxiv.org/pdf/2309.162982023-09-03
Unveiling the Impact of Coding Data Instruction Fine-Tuning on Large Language Models Reasoninghttps://arxiv.org/pdf/2405.205352024-12-12
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learninghttps://arxiv.org/abs/2501.129482025-01-22
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learninghttps://arxiv.org/abs/2505.164002025-05-22
OpenThoughts: A Systematic Investigation of Data Curation for Post-training Reasoning Modelshttps://arxiv.org/abs/2506.041782025-06-05
RuleReasoner: Reinforced Rule-based Reasoning with Dynamic Multi-domain Curriculum Learninghttps://arxiv.org/abs/2506.086722025-06-10
CoRT: Code-integrated Reasoning within Thinkinghttps://arxiv.org/abs/2506.098202025-06-11

Reasoning-enhanced Code Intelligence

Essential Code Intelligence

Paper TitleURLRelease Date
CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generationhttps://arxiv.org/abs/2102.046642021-02-09
Competition-level code generation with AlphaCodehttps://arxiv.org/abs/2108.077322021-08-16
Evaluating Large Language Models Trained on Codehttps://arxiv.org/abs/2107.033742021-07-07
Program Synthesis with Large Language Modelshttps://arxiv.org/abs/2108.077322021-08-16
A Systematic Evaluation of Large Language Models of Codehttps://arxiv.org/abs/2202.131692022-02-26
InCoder: A Generative Model for Code Infilling and Synthesishttps://arxiv.org/abs/2204.059992023-04-12
CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesishttps://arxiv.org/abs/2203.134742023-03-25
StarCoder: May the Source be with You!https://arxiv.org/abs/2305.061612023-05-10
Code Llama: Open Foundation Models for Codehttps://arxiv.org/abs/2308.129502023-08-24
RepoBench: Benchmarking Repository-Level Code Auto-Completion Systemshttps://arxiv.org/abs/2306.030912023-06-05
CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completionhttps://arxiv.org/abs/2310.112482023-10-17
StarCoder 2 and The Stack v2: The Next Generationhttps://arxiv.org/abs/2402.191732024-02-29
CodeGemma: Open Code Models Based on Gemmahttps://arxiv.org/abs/2406.114092024-06-17
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligencehttps://arxiv.org/abs/2406.119312024-06-17
Qwen2.5-Coder Technical Reporthttps://arxiv.org/abs/2409.121862024-09-18
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratingshttps://arxiv.org/abs/2501.012572025-01-02
Exploring Code Comprehension in Scientific Programming: Preliminary Insights from Research Scientistshttps://arxiv.org/abs/2501.100372025-01-17
COFFE: A Code Efficiency Benchmark for Code Generationhttps://arxiv.org/abs/2502.028272025-02-05
Evaluating the Generalization Capabilities of Large Language Models on Code Reasoninghttps://arxiv.org/abs/2504.055182025-04-07
rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Datasethttps://arxiv.org/abs/2505.212972025-05-27

Integration of Reasoning Capabilities

Reasoning for Code Generation

Paper TitleURLRelease Date
Chain-of-Thought Prompting Elicits Reasoning in Large Language Modelshttps://arxiv.org/abs/2201.119032022-01-28
Self-planning Code Generation with Large Language Modelshttps://arxiv.org/abs/2303.066892023-03-12
Structured Chain-of-Thought Prompting for Code Generationhttps://arxiv.org/abs/2305.065992023-05-11
CodeCoT: Tackling Code Syntax Errors in CoT Reasoning for Code Generationhttps://arxiv.org/abs/2308.087842023-08-17
CodePlan: Repository-level Coding using LLMs and Planninghttps://arxiv.org/abs/2309.124992023-09-21
Chain-of-Thought in Neural Code Generation: From and For Lightweight Language Modelshttps://arxiv.org/abs/2312.055622023-12-09
Planning In Natural Language Improves LLM Search For Code Generationhttps://arxiv.org/abs/2409.037332024-09-05
Chain of Grounded Objectives: Bridging Process and Goal-oriented Prompting for Code Generationhttps://arxiv.org/abs/2501.139782025-01-23
LLM-Guided Compositional Program Synthesishttps://arxiv.org/abs/2503.155402025-03-12
Modularization is Better: Effective Code Generation with Modular Promptinghttps://arxiv.org/abs/2503.124832025-03-16
Uncertainty-Guided Chain-of-Thought for Code Generation with LLMshttps://arxiv.org/abs/2503.153412025-03-19
MSCoT: Structured Chain-of-Thought Generation for Multiple Programming Languageshttps://arxiv.org/abs/2504.101782025-04-18
Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generationhttps://arxiv.org/abs/2506.069712025-06-08
Reasoning as a Resource: Optimizing Fast and Slow Thinking in Code Generation Modelshttps://arxiv.org/abs/2506.093962025-06-11

Reasoning Over Code

Paper TitleURLRelease Date
CodeQA: A Question Answering Dataset for Source Code Comprehensionhttps://arxiv.org/abs/2109.083652021-09-17
CRUXEval: A Benchmark for Code Reasoning, Understanding and Executionhttps://arxiv.org/abs/2401.030652024-01-05
CodeMind: A Framework to Challenge Large Language Models for Code Reasoninghttps://arxiv.org/abs/2402.096642024-02-15
Reasoning Runtime Behavior of a Program with LLM: How Far Are We?https://arxiv.org/abs/2403.164372024-03-25
NExT: Teaching Large Language Models to Reason about Code Executionhttps://arxiv.org/abs/2404.146622024-04-23
RepoQA: Evaluating Long Context Code Understandinghttps://arxiv.org/abs/2406.060252024-06-10
SelfPiCo: Self-Guided Partial Code Execution with LLMshttps://arxiv.org/abs/2407.169742024-07-24
CodeMMLU: A Multi-Task Benchmark for Assessing Code Understanding Capabilitieshttps://arxiv.org/abs/2410.019992024-10-02
What You See Is Not Always What You Get: An Empirical Study of Code Comprehensionhttps://arxiv.org/abs/2412.080982024-12-11
How Accurately Do Large Language Models Understand Code?https://arxiv.org/abs/2504.043722025-04-06

Interactive Programming

Paper TitleURLRelease Date
Interactive Program Synthesishttps://arxiv.org/abs/1703.035392017-03-10
Self-Refine: Iterative Refinement with Self-Feedbackhttps://arxiv.org/abs/2303.176512023-03-30
Teaching Large Language Models to Self-Debughttps://arxiv.org/abs/2304.051282023-04-11
Self-collaboration Code Generation via ChatGPThttps://arxiv.org/abs/2304.075902023-04-15
Self-Edit: Fault-Aware Code Editor for Code Generationhttps://arxiv.org/abs/2305.040872023-05-06
LeTI: Learning to Generate from Textual Interactionshttps://arxiv.org/abs/2305.103142023-05-17
InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedbackhttps://arxiv.org/abs/2306.148982023-06-26
CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-moduleshttps://arxiv.org/abs/2310.089922023-10-13
AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisationhttps://arxiv.org/abs/2312.130102023-12-20
OpenCodeInterpreter: Integrating Code Generation with Execution and Refinementhttps://arxiv.org/abs/2402.146582024-02-22
What Makes Large Language Models Reason in (Multi-Turn) Code Generation?https://arxiv.org/abs/2410.081052024-10-10
Revisit Self-Debugging with Self-Generated Tests for Code Generationhttps://arxiv.org/abs/2501.127932025-01-22
Large Language Model Guided Self-Debugging Code Generationhttps://arxiv.org/abs/2502.029282025-02-05
Interactive Agents to Overcome Ambiguity in Software Engineeringhttps://arxiv.org/abs/2502.130692025-02-18
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environmentshttps://arxiv.org/abs/2502.198522025-02-28
Prompt Alchemy: Automatic Prompt Refinement for Enhancing Code Generationhttps://arxiv.org/abs/2503.110852025-03-14
Humanity's Last Code Exam: Can Advanced LLMs Conquer Human's Hardest Code Competition?https://arxiv.org/abs/2506.127132025-06-12

Code Agents with Complex Reasoning

Paper TitleURLRelease Date
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?https://arxiv.org/abs/2310.067702023-10-10
CodeAgent: Enhancing Code Generation with Tool-Integrated Agent Systems for Real-World Repo-level Coding Challengeshttps://arxiv.org/abs/2401.073392024-01-14
Executable Code Actions Elicit Better LLM Agentshttps://arxiv.org/abs/2402.010302024-02-01
Cursor AI: The AI Code Editorhttps://www.cursor.com2024-02-17
Devin AI: Autonomous AI Software Engineerhttps://devin.ai2024-03-12
AutoCodeRover: Autonomous Program Improvementhttps://arxiv.org/abs/2404.054272024-04-08
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineeringhttps://arxiv.org/abs/2405.157932024-05-06
Agentless: Demystifying LLM-based Software Engineering Agentshttps://arxiv.org/abs/2407.014892024-07-01
OpenHands: An Open Platform for AI Software Developers as Generalist Agentshttps://arxiv.org/abs/2407.167412024-07-23
SWE-bench Verifiedhttps://openai.com/index/introducing-swe-bench-verified2024-08-13
HyperAgent: Generalist Software Engineering Agents to Solve Coding Tasks at Scalehttps://arxiv.org/abs/2409.162992024-09-09
SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?https://arxiv.org/abs/2410.038592024-10-04
Evaluating Software Development Agents: Patch Patterns, Code Quality, and Issue Complexity in Real-World GitHub Scenarioshttps://arxiv.org/abs/2410.124682024-10-16
Verbal Process Supervision Elicits Better Coding Agentshttps://arxiv.org/abs/2503.184942025-03-24
A Self-Improving Coding Agenthttps://arxiv.org/abs/2504.152282025-04-21
Breakpoint: A Benchmark for Systematic and Scalable Evaluation of Long-Horizon Code Repairhttps://arxiv.org/abs/2506.001722025-05-31
Code Researcher: Deep Research Agent for Large Systems Code and Commit Historyhttps://arxiv.org/abs/2506.110602025-05-27
Coding Agents with Multimodal Browsing are Generalist Problem Solvershttps://arxiv.org/abs/2506.030112025-06-03
UTBoost: Rigorous Evaluation of Coding Agents on SWE-Benchhttps://arxiv.org/abs/2506.092892025-06-10

Contributors

DwanZhang-AI

9 commits

dayuyang1999

3 commits

Leolty

2 commits

iCSawyer

1 commits