AheadOFpotato/Awesome-LRM-Mechanisms

Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures

35

13 commits

updated Jan 29, 2026

See the code

README

Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures

Awesome Github Survey

πŸš€ News

  • [2026.1.29] Our survey is now available at: https://arxiv.org/abs/2601.19928
  • [2026.1.11] πŸ”₯ We are excited to release our survey on LRM mechanistic studies in this repo.

πŸ“Œ Overview

Our paper provides a comprehensive survey of the mechanistic understanding of LRMs.

Overview of LRM Mechanisms Survey

We organize recent findings into three core dimensions:

  1. Training Dynamics:
    1. Roles of SFT and RL in Post-Training
    2. Understanding RL Training Dynamics
  2. Reasoning Mechanisms:
    1. General Reasoning Structures
    2. Specific Reasoning Behaviors
    3. Internal Mechanisms
  3. Unintended Behaviors:
    1. Hallucination
    2. Unfaithfulness
    3. Overthinking
    4. Unsafety

πŸ”— Citation

If you find our survey helpful, please consider citing our paper:

@misc{hu2026mechanisticunderstandinglargereasoning,
      title={Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures}, 
      author={Yi Hu and Jiaqi Gu and Ruxin Wang and Zijun Yao and Hao Peng and Xiaobao Wu and Jianhui Chen and Muhan Zhang and Liangming Pan},
      year={2026},
      eprint={2601.19928},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2601.19928}, 
}

πŸ—‚οΈ Contents

πŸ“„ Paper List

Training Dynamics

Roles of SFT and RL in Post-Training

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-01Deepseek-R1DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningPaperHugging Face Downloads
2025-04limit of RLVRDoes Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?PaperGitHub Stars
2025-09RL Squeezes, SFT ExpandsRL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMsPaper-
2025-07Invisible LeashThe Invisible Leash: Why RLVR May or May Not Escape Its OriginPaper-
2025-09Representation Geometry of LLMTracing the Representation Geometry of Language Models from Pretraining to Post-trainingPaper-
2025-05Entropy Mechanism of RLThe Entropy Mechanism of Reinforcement Learning for Reasoning Language ModelsPaperGitHub Stars
2025-09Emergent Attention HeadsThinking Sparks!: Emergent Attention Heads in Reasoning Models During Post TrainingPaper-
2025-01SFT Memorizes, RL GeneralizesSFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-trainingPaperGitHub Stars
2025-09CoNetHow LLMs Learn to Reason: A Complex Network PerspectivePaper-
2025-08RL Is Neither a Panacea Nor a MirageRL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMsPaper-
2025-09RL Heals OOD ForgettingRL Fine-Tuning Heals OOD Forgetting in SFTPaperGitHub Stars
2025-05ProRLProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language ModelsPaperHugging Face Downloads
2025-09From f(x) and g(x) to f(g(x))From f(x) and g(x) to f(g(x)): LLMs Learn New Skills in RL by Composing Old OnesPaper-
2025-12From Atomic to CompositeFrom Atomic to Composite: Reinforcement Learning Enables Generalization in Complementary ReasoningPaperGitHub Stars
2025-06OctoThinkerOctoThinker: Mid-training Incentivizes Reinforcement Learning ScalingPaperGitHub Stars
2025-12Interplay LM ReasoningOn the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language ModelsPaperGitHub Stars

Understanding RL Training Dynamics

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-09HICRAEmergent Hierarchical Reasoning in LLMs through Reinforcement LearningPaperGitHub Stars
2025-09CoNetHow LLMs Learn to Reason: A Complex Network PerspectivePaper-
2025-10Two-Stage DynamicThe Debate on RLVR Reasoning Capability Boundary: Shrinkage, Expansion, or Both? A Two-Stage Dynamic ViewPaper-
2025-09RL Enhances ActivationReinforcement Learning Fine-Tuning Enhances Activation Intensity and Diversity in the Internal Circuitry of LLMsPaperGitHub Stars
2025-09Post-Training Structural ChangesUnderstanding Post-Training Structural Changes in Large Language ModelsPaper-
2025-10AlphaRLOn Predictability of Reinforcement Learning Dynamics for Large Language ModelsPaperGitHub Stars
2025-05Entropy Mechanism of RLThe Entropy Mechanism of Reinforcement Learning for Reasoning Language ModelsPaperGitHub Stars
2025-10Reasoning Boundary ParadoxThe Reasoning Boundary Paradox: How Reinforcement Learning Constrains Language ModelsPaperGitHub Stars
2025-09VERLBeyond the Exploration-Exploitation Trade-off: A Hidden State Approach for LLM Reasoning in RLVRPaperGitHub Stars

Reasoning Mechanisms

General Reasoning Structures

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-04DeepSeek-R1 ThoughtologyDeepSeek-R1 Thoughtology: Let's think about LLM ReasoningPaper-
2025-06Functional BlocksBeyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement LearningPaperGitHub Stars
2025-06Thought AnchorsThought Anchors: Which LLM Reasoning Steps Matter?PaperGitHub Stars
2025-09Schoenfeld's Episode TheoryUnderstanding the Thinking Process of Reasoning Models: A Perspective from Schoenfeld's Episode TheoryPaperGitHub Stars
2025-11ReJumpReJump: A Tree-Jump Representation for Analyzing and Improving LLM ReasoningPaperGitHub Stars
2025-05Structural PatternsWhat Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought ReasoningPaper-
2025-06Topology of ReasoningTopology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph PropertiesPaperGitHub Stars
2025-05Mapping the Minds of LLMsMapping the Minds of LLMs: A Graph-Based Analysis of Reasoning LLMPaperGitHub Stars

Specific Reasoning Behaviors

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-06Thought AnchorsThought Anchors: Which LLM Reasoning Steps Matter?PaperGitHub Stars
2025-05Thought AnchorsCognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRsPaperGitHub Stars
2025-04Understanding Aha MomentsUnderstanding Aha Moments: from External Observations to Internal MechanismsPaper-
2025-02May Not Aha MomentThere may not be aha moment in r1-zero-like training - a pilot studyBlogGitHub Stars
2025-10First Try MattersFirst Try Matters: Revisiting the Role of Reflection in Reasoning ModelsPaperGitHub Stars
2025-06The Illusion of ThinkingThe Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem ComplexityPaper-
2025-9Performative ThinkingPerformative Thinking? The Brittle Correlation Between CoT Length and Problem ComplexityPaper-

Internal Mechanisms

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-05Towards Understanding Distilled Reasoning ModelsTowards Understanding Distilled Reasoning Models: A Representational ApproachPaperGitHub Stars
2025-05Sparse AutoencodersI Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse AutoencodersPaperGitHub Stars
2025-04Hood of a Reasoning ModelUnder the Hood of a Reasoning ModelPaper-
2025-06Steering VectorsUnderstanding Reasoning in Thinking Language Models via Steering VectorsPaperGitHub Stars
2025-10Thinking Models Learn When To ReasonBase Models Know How to Reason, Thinking Models Learn WhenPaperGitHub Stars
2025-06Thought AnchorsThought Anchors: Which LLM Reasoning Steps Matter?PaperGitHub Stars
2025-09From Reasoning to AnswerFrom Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 ModelsPaperGitHub Stars
2025-04Understanding Aha MomentsUnderstanding Aha Moments: from External Observations to Internal MechanismsPaper-
2025-12ReflCtrlReflCtrl: Controlling LLM Reflection via Representation EngineeringPaperGitHub Stars
2025-08Latent Directions of ReflectionUnveiling the Latent Directions of Reflection in Large Language ModelsPaperGitHub Stars
2025-07Reasoning-Finetuning Repurposes Latent RepresentationsReasoning-Finetuning Repurposes Latent Representations in Base ModelsPaperGitHub Stars
2025-05Bias-Only AdaptationSteering LLM Reasoning Through Bias-Only AdaptationPaperGitHub Stars
2025-11Rank-1 LoRAsRank-1 LoRAs Encode Interpretable Reasoning SignalsPaperGitHub Stars
2025-06Layer ImportanceLayer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-TrainingPaperGitHub Stars
2025-12Effective DepthWhat Affects the Effective Depth of Large Language Models?PaperGitHub Stars
2025-12Internal PolicyBottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal PoliciesPaperGitHub Stars
2025-06The Emergence of Mutual Information (MI) PeaksDemystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM ReasoningPaperGitHub Stars

Unintended Behaviors

Hallucination

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-05FactEvalAre Reasoning Models More Prone to Hallucination?PaperGitHub Stars
2025-05FSPOReasoning Models Hallucinate More: Factuality-Aware Reinforcement Learning for Large Reasoning ModelsPaperGitHub Stars
2025-05RDHDetection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic PerspectivePaperGitHub Stars
2025-05Hallucination TaxThe Hallucination Tax of Reinforcement FinetuningPaperHugging Face Downloads
2025-09TTS_knowledgeTest-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks YetPaperGitHub Stars
2025-05Meta-CognitiveAuditing Meta-Cognitive Hallucinations in Reasoning Large Language ModelsPaperGitHub Stars
2025-06RACEJoint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning ModelsPaperGitHub Stars

Unfaithfulness

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-05FactEvalAre Reasoning Models More Prone to Hallucination?PaperGitHub Stars
2025-01Cue InsertionAre DeepSeek R1 And Other Reasoning Models More Faithful?Paper-
2025-05Don't Say What They ThinkReasoning Models Don't Always Say What They ThinkPaper-
2025-03Unfaithful CoTChain-of-Thought Reasoning In The Wild Is Not Always FaithfulPaperGitHub Stars
2025-06Strategic DeceptionWhen Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning ModelsPaperGitHub Stars
2025-05Reasoning Intermediate TokenBeyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate TokensPaper-
2025-07Alignment PredictionCan We Predict Alignment Before Models Finish Thinking? Towards Monitoring Misaligned Reasoning ModelsPaperGitHub Stars
2025-10Refusal CliffRefusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?Paper-
2025-03CoT MonitorMonitoring Reasoning Models for Misbehavior and the Risks of Promoting ObfuscationPaper-
2025-06VFTTeaching Models to Verbalize Reward Hacking in Chain-of-Thought ReasoningPaper-

Overthinking

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-05R1 ThoughtologyDeepSeek-R1 Thoughtology: Let’s think about LLM reasoningPaperGitHub Stars
2025-05Between Underthinking and OverthinkingBetween Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMsPaper-
2025-06Does Thinking More Always Help?Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning ModelsPaper-
2025-05Short-m@kDon't Overthink it. Preferring Shorter Thinking Chains for Improved LLM ReasoningPaper-
2025-02Thinking-Optimal ScalingTowards Thinking-Optimal Scaling of Test-Time Compute for LLM ReasoningPaperGitHub Stars
2024-12Do NOT Think That MuchDo NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMsPaperGitHub Stars
2025-07Inverse ScalingInverse Scaling in Test-Time ComputePaperGitHub Stars
2025-01Underthinking Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMsPaper-
2025-02ReasoningAction Dilemma The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic TasksPaper-
2025-10REFRAINStop When Enough: Adaptive Early-Stopping for Chain-of-Thought ReasoningPaper-
2025-05AdaptThinkStop When Enough: Adaptive Early-Stopping for Chain-of-Thought ReasoningPaperGitHub Stars
2025-05Self-Braking TuningLet LRMs Break Free from Overthinking via Self-Braking TuningPaperGitHub Stars
2025-04MiPLet LRMs Break Free from Overthinking via Self-Braking TuningPaperGitHub Stars
2025-05Behavioral DivergenceWhen Can Large Reasoning Models Save Thinking? Mechanistic Analysis of Behavioral Divergence in ReasoningPaper-
2025-04Thought ManipulationThought Manipulation: External Thought Can Be Efficient for Large Reasoning ModelsPaper-
2025-07Disregard Correct SolutionsLarge Reasoning Models are not thinking straight: on the unreliability of thinking trajectoriesPaper-
2025-05Manifold SteeringMitigating Overthinking in Large Reasoning Models via Manifold SteeringPaperGitHub Stars
2025-03RepresentationTowards Understanding Distilled Reasoning Models: A Representational ApproachPaperGitHub Stars
2025-04SEALSEAL: Steerable Reasoning Calibration of Large Language Models for FreePaperGitHub Stars
2025-05Internal BiasThe First Impression Problem: Internal Bias Triggers Overthinking in Reasoning ModelsPaperGitHub Stars
2025-04Self-Verification ProbingReasoning Models Know When They're Right: Probing Hidden States for Self-VerificationPaperGitHub Stars
2025-10Zero-Step ThinkingThe Zero-Step Thinking: An Empirical Study of Mode Selection as Harder Early Exit in Reasoning ModelsPaperGitHub Stars

Unsafety

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-02SafeChainSafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning CapabilitiesPaperGitHub Stars
2025-02Hidden RisksThe Hidden Risks of Large Reasoning Models: A Safety Assessment of R1Paper-
2025-03DeepSeek-R1-SafeSafety Evaluation and Enhancement of DeepSeek Models in Chinese ContextsPaperGitHub Stars
2025-03FreeEvalLMTrade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational CapabilitiesPaperGitHub Stars
2025-02MousetrapA Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative ChaosPaperGitHub Stars
2025-02H-CoTH-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash ThinkingPaperGitHub Stars
2025-08R1-ACTR1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety KnowledgePaperHugging Face Downloads
2025-10COGWhen Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning ModelsPaperGitHub Stars

Training Methods

Combine SFT with RL

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-09Annealed-RLVRHow LLMs Learn to Reason: A Complex Network PerspectivePaper-
2025-06ReLIFTLearning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest QuestionsPaper-
2025-12TRAPOTrust-Region Adaptive Policy OptimizationPaperGitHub Stars
2025-05UFTUFT: Unifying Supervised and Reinforcement Fine-TuningPaperGitHub Stars
2025-09HPTTowards a Unified View of Large Language Model Post-TrainingPaperGitHub Stars
2025-08CHORDOn-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic WeightingPaperGitHub Stars
2025-06SRFTSRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for ReasoningPaperHugging Face Downloads
2025-09BRIDGEBeyond Two-Stage Training: Cooperative SFT and RL for LLM ReasoningPaper-

RL balancing exploration and exploitation

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-03DAPODAPO: An Open-Source LLM Reinforcement Learning System at ScalePaperGitHub Stars
2025-05Entropy Mechanism of RLThe Entropy Mechanism of Reinforcement Learning for Reasoning Language ModelsPaperGitHub Stars
2025-05ProRLProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language ModelsPaperHugging Face Downloads
2025-08CURECURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse PreventionPaperGitHub Stars
2025-10Reasoning Boundary ParadoxThe Reasoning Boundary Paradox: How Reinforcement Learning Constrains Language ModelsPaperGitHub Stars
2025-09outcome-based explorationOutcome-based Exploration for LLM ReasoningPaper-
2025-06entropy-based advantageReasoning with Exploration: An Entropy PerspectivePaper-
2025-10AEPOArbitrary Entropy Policy Optimization Breaks The Exploration Bottleneck of Reinforcement LearningPaperGitHub Stars
2025-09AEntOn Entropy Control in LLM-RL AlgorithmsPaper-
2025-09SIRENRethinking Entropy Regularization in Large Reasoning ModelsPaperGitHub Stars

Contributors

AheadOFpotato

11 commits

ChnQ

1 commits

Trae1ounG

1 commits

AheadOFpotato/Awesome-LRM-Mechanisms

Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures

35

13 commits

updated Jan 29, 2026

See the code

README

Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures

Awesome Github Survey

πŸš€ News

  • [2026.1.29] Our survey is now available at: https://arxiv.org/abs/2601.19928
  • [2026.1.11] πŸ”₯ We are excited to release our survey on LRM mechanistic studies in this repo.

πŸ“Œ Overview

Our paper provides a comprehensive survey of the mechanistic understanding of LRMs.

Overview of LRM Mechanisms Survey

We organize recent findings into three core dimensions:

  1. Training Dynamics:
    1. Roles of SFT and RL in Post-Training
    2. Understanding RL Training Dynamics
  2. Reasoning Mechanisms:
    1. General Reasoning Structures
    2. Specific Reasoning Behaviors
    3. Internal Mechanisms
  3. Unintended Behaviors:
    1. Hallucination
    2. Unfaithfulness
    3. Overthinking
    4. Unsafety

πŸ”— Citation

If you find our survey helpful, please consider citing our paper:

@misc{hu2026mechanisticunderstandinglargereasoning,
      title={Towards a Mechanistic Understanding of Large Reasoning Models: A Survey of Training, Inference, and Failures}, 
      author={Yi Hu and Jiaqi Gu and Ruxin Wang and Zijun Yao and Hao Peng and Xiaobao Wu and Jianhui Chen and Muhan Zhang and Liangming Pan},
      year={2026},
      eprint={2601.19928},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2601.19928}, 
}

πŸ—‚οΈ Contents

πŸ“„ Paper List

Training Dynamics

Roles of SFT and RL in Post-Training

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-01Deepseek-R1DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement LearningPaperHugging Face Downloads
2025-04limit of RLVRDoes Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?PaperGitHub Stars
2025-09RL Squeezes, SFT ExpandsRL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMsPaper-
2025-07Invisible LeashThe Invisible Leash: Why RLVR May or May Not Escape Its OriginPaper-
2025-09Representation Geometry of LLMTracing the Representation Geometry of Language Models from Pretraining to Post-trainingPaper-
2025-05Entropy Mechanism of RLThe Entropy Mechanism of Reinforcement Learning for Reasoning Language ModelsPaperGitHub Stars
2025-09Emergent Attention HeadsThinking Sparks!: Emergent Attention Heads in Reasoning Models During Post TrainingPaper-
2025-01SFT Memorizes, RL GeneralizesSFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-trainingPaperGitHub Stars
2025-09CoNetHow LLMs Learn to Reason: A Complex Network PerspectivePaper-
2025-08RL Is Neither a Panacea Nor a MirageRL Is Neither a Panacea Nor a Mirage: Understanding Supervised vs. Reinforcement Learning Fine-Tuning for LLMsPaper-
2025-09RL Heals OOD ForgettingRL Fine-Tuning Heals OOD Forgetting in SFTPaperGitHub Stars
2025-05ProRLProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language ModelsPaperHugging Face Downloads
2025-09From f(x) and g(x) to f(g(x))From f(x) and g(x) to f(g(x)): LLMs Learn New Skills in RL by Composing Old OnesPaper-
2025-12From Atomic to CompositeFrom Atomic to Composite: Reinforcement Learning Enables Generalization in Complementary ReasoningPaperGitHub Stars
2025-06OctoThinkerOctoThinker: Mid-training Incentivizes Reinforcement Learning ScalingPaperGitHub Stars
2025-12Interplay LM ReasoningOn the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language ModelsPaperGitHub Stars

Understanding RL Training Dynamics

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-09HICRAEmergent Hierarchical Reasoning in LLMs through Reinforcement LearningPaperGitHub Stars
2025-09CoNetHow LLMs Learn to Reason: A Complex Network PerspectivePaper-
2025-10Two-Stage DynamicThe Debate on RLVR Reasoning Capability Boundary: Shrinkage, Expansion, or Both? A Two-Stage Dynamic ViewPaper-
2025-09RL Enhances ActivationReinforcement Learning Fine-Tuning Enhances Activation Intensity and Diversity in the Internal Circuitry of LLMsPaperGitHub Stars
2025-09Post-Training Structural ChangesUnderstanding Post-Training Structural Changes in Large Language ModelsPaper-
2025-10AlphaRLOn Predictability of Reinforcement Learning Dynamics for Large Language ModelsPaperGitHub Stars
2025-05Entropy Mechanism of RLThe Entropy Mechanism of Reinforcement Learning for Reasoning Language ModelsPaperGitHub Stars
2025-10Reasoning Boundary ParadoxThe Reasoning Boundary Paradox: How Reinforcement Learning Constrains Language ModelsPaperGitHub Stars
2025-09VERLBeyond the Exploration-Exploitation Trade-off: A Hidden State Approach for LLM Reasoning in RLVRPaperGitHub Stars

Reasoning Mechanisms

General Reasoning Structures

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-04DeepSeek-R1 ThoughtologyDeepSeek-R1 Thoughtology: Let's think about LLM ReasoningPaper-
2025-06Functional BlocksBeyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement LearningPaperGitHub Stars
2025-06Thought AnchorsThought Anchors: Which LLM Reasoning Steps Matter?PaperGitHub Stars
2025-09Schoenfeld's Episode TheoryUnderstanding the Thinking Process of Reasoning Models: A Perspective from Schoenfeld's Episode TheoryPaperGitHub Stars
2025-11ReJumpReJump: A Tree-Jump Representation for Analyzing and Improving LLM ReasoningPaperGitHub Stars
2025-05Structural PatternsWhat Makes a Good Reasoning Chain? Uncovering Structural Patterns in Long Chain-of-Thought ReasoningPaper-
2025-06Topology of ReasoningTopology of Reasoning: Understanding Large Reasoning Models through Reasoning Graph PropertiesPaperGitHub Stars
2025-05Mapping the Minds of LLMsMapping the Minds of LLMs: A Graph-Based Analysis of Reasoning LLMPaperGitHub Stars

Specific Reasoning Behaviors

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-06Thought AnchorsThought Anchors: Which LLM Reasoning Steps Matter?PaperGitHub Stars
2025-05Thought AnchorsCognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRsPaperGitHub Stars
2025-04Understanding Aha MomentsUnderstanding Aha Moments: from External Observations to Internal MechanismsPaper-
2025-02May Not Aha MomentThere may not be aha moment in r1-zero-like training - a pilot studyBlogGitHub Stars
2025-10First Try MattersFirst Try Matters: Revisiting the Role of Reflection in Reasoning ModelsPaperGitHub Stars
2025-06The Illusion of ThinkingThe Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem ComplexityPaper-
2025-9Performative ThinkingPerformative Thinking? The Brittle Correlation Between CoT Length and Problem ComplexityPaper-

Internal Mechanisms

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-05Towards Understanding Distilled Reasoning ModelsTowards Understanding Distilled Reasoning Models: A Representational ApproachPaperGitHub Stars
2025-05Sparse AutoencodersI Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse AutoencodersPaperGitHub Stars
2025-04Hood of a Reasoning ModelUnder the Hood of a Reasoning ModelPaper-
2025-06Steering VectorsUnderstanding Reasoning in Thinking Language Models via Steering VectorsPaperGitHub Stars
2025-10Thinking Models Learn When To ReasonBase Models Know How to Reason, Thinking Models Learn WhenPaperGitHub Stars
2025-06Thought AnchorsThought Anchors: Which LLM Reasoning Steps Matter?PaperGitHub Stars
2025-09From Reasoning to AnswerFrom Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 ModelsPaperGitHub Stars
2025-04Understanding Aha MomentsUnderstanding Aha Moments: from External Observations to Internal MechanismsPaper-
2025-12ReflCtrlReflCtrl: Controlling LLM Reflection via Representation EngineeringPaperGitHub Stars
2025-08Latent Directions of ReflectionUnveiling the Latent Directions of Reflection in Large Language ModelsPaperGitHub Stars
2025-07Reasoning-Finetuning Repurposes Latent RepresentationsReasoning-Finetuning Repurposes Latent Representations in Base ModelsPaperGitHub Stars
2025-05Bias-Only AdaptationSteering LLM Reasoning Through Bias-Only AdaptationPaperGitHub Stars
2025-11Rank-1 LoRAsRank-1 LoRAs Encode Interpretable Reasoning SignalsPaperGitHub Stars
2025-06Layer ImportanceLayer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-TrainingPaperGitHub Stars
2025-12Effective DepthWhat Affects the Effective Depth of Large Language Models?PaperGitHub Stars
2025-12Internal PolicyBottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal PoliciesPaperGitHub Stars
2025-06The Emergence of Mutual Information (MI) PeaksDemystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM ReasoningPaperGitHub Stars

Unintended Behaviors

Hallucination

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-05FactEvalAre Reasoning Models More Prone to Hallucination?PaperGitHub Stars
2025-05FSPOReasoning Models Hallucinate More: Factuality-Aware Reinforcement Learning for Large Reasoning ModelsPaperGitHub Stars
2025-05RDHDetection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic PerspectivePaperGitHub Stars
2025-05Hallucination TaxThe Hallucination Tax of Reinforcement FinetuningPaperHugging Face Downloads
2025-09TTS_knowledgeTest-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks YetPaperGitHub Stars
2025-05Meta-CognitiveAuditing Meta-Cognitive Hallucinations in Reasoning Large Language ModelsPaperGitHub Stars
2025-06RACEJoint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning ModelsPaperGitHub Stars

Unfaithfulness

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-05FactEvalAre Reasoning Models More Prone to Hallucination?PaperGitHub Stars
2025-01Cue InsertionAre DeepSeek R1 And Other Reasoning Models More Faithful?Paper-
2025-05Don't Say What They ThinkReasoning Models Don't Always Say What They ThinkPaper-
2025-03Unfaithful CoTChain-of-Thought Reasoning In The Wild Is Not Always FaithfulPaperGitHub Stars
2025-06Strategic DeceptionWhen Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning ModelsPaperGitHub Stars
2025-05Reasoning Intermediate TokenBeyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate TokensPaper-
2025-07Alignment PredictionCan We Predict Alignment Before Models Finish Thinking? Towards Monitoring Misaligned Reasoning ModelsPaperGitHub Stars
2025-10Refusal CliffRefusal Falls off a Cliff: How Safety Alignment Fails in Reasoning?Paper-
2025-03CoT MonitorMonitoring Reasoning Models for Misbehavior and the Risks of Promoting ObfuscationPaper-
2025-06VFTTeaching Models to Verbalize Reward Hacking in Chain-of-Thought ReasoningPaper-

Overthinking

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-05R1 ThoughtologyDeepSeek-R1 Thoughtology: Let’s think about LLM reasoningPaperGitHub Stars
2025-05Between Underthinking and OverthinkingBetween Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMsPaper-
2025-06Does Thinking More Always Help?Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning ModelsPaper-
2025-05Short-m@kDon't Overthink it. Preferring Shorter Thinking Chains for Improved LLM ReasoningPaper-
2025-02Thinking-Optimal ScalingTowards Thinking-Optimal Scaling of Test-Time Compute for LLM ReasoningPaperGitHub Stars
2024-12Do NOT Think That MuchDo NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMsPaperGitHub Stars
2025-07Inverse ScalingInverse Scaling in Test-Time ComputePaperGitHub Stars
2025-01Underthinking Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMsPaper-
2025-02ReasoningAction Dilemma The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic TasksPaper-
2025-10REFRAINStop When Enough: Adaptive Early-Stopping for Chain-of-Thought ReasoningPaper-
2025-05AdaptThinkStop When Enough: Adaptive Early-Stopping for Chain-of-Thought ReasoningPaperGitHub Stars
2025-05Self-Braking TuningLet LRMs Break Free from Overthinking via Self-Braking TuningPaperGitHub Stars
2025-04MiPLet LRMs Break Free from Overthinking via Self-Braking TuningPaperGitHub Stars
2025-05Behavioral DivergenceWhen Can Large Reasoning Models Save Thinking? Mechanistic Analysis of Behavioral Divergence in ReasoningPaper-
2025-04Thought ManipulationThought Manipulation: External Thought Can Be Efficient for Large Reasoning ModelsPaper-
2025-07Disregard Correct SolutionsLarge Reasoning Models are not thinking straight: on the unreliability of thinking trajectoriesPaper-
2025-05Manifold SteeringMitigating Overthinking in Large Reasoning Models via Manifold SteeringPaperGitHub Stars
2025-03RepresentationTowards Understanding Distilled Reasoning Models: A Representational ApproachPaperGitHub Stars
2025-04SEALSEAL: Steerable Reasoning Calibration of Large Language Models for FreePaperGitHub Stars
2025-05Internal BiasThe First Impression Problem: Internal Bias Triggers Overthinking in Reasoning ModelsPaperGitHub Stars
2025-04Self-Verification ProbingReasoning Models Know When They're Right: Probing Hidden States for Self-VerificationPaperGitHub Stars
2025-10Zero-Step ThinkingThe Zero-Step Thinking: An Empirical Study of Mode Selection as Harder Early Exit in Reasoning ModelsPaperGitHub Stars

Unsafety

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-02SafeChainSafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning CapabilitiesPaperGitHub Stars
2025-02Hidden RisksThe Hidden Risks of Large Reasoning Models: A Safety Assessment of R1Paper-
2025-03DeepSeek-R1-SafeSafety Evaluation and Enhancement of DeepSeek Models in Chinese ContextsPaperGitHub Stars
2025-03FreeEvalLMTrade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational CapabilitiesPaperGitHub Stars
2025-02MousetrapA Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative ChaosPaperGitHub Stars
2025-02H-CoTH-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash ThinkingPaperGitHub Stars
2025-08R1-ACTR1-ACT: Efficient Reasoning Model Safety Alignment by Activating Safety KnowledgePaperHugging Face Downloads
2025-10COGWhen Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning ModelsPaperGitHub Stars

Training Methods

Combine SFT with RL

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-09Annealed-RLVRHow LLMs Learn to Reason: A Complex Network PerspectivePaper-
2025-06ReLIFTLearning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest QuestionsPaper-
2025-12TRAPOTrust-Region Adaptive Policy OptimizationPaperGitHub Stars
2025-05UFTUFT: Unifying Supervised and Reinforcement Fine-TuningPaperGitHub Stars
2025-09HPTTowards a Unified View of Large Language Model Post-TrainingPaperGitHub Stars
2025-08CHORDOn-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic WeightingPaperGitHub Stars
2025-06SRFTSRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for ReasoningPaperHugging Face Downloads
2025-09BRIDGEBeyond Two-Stage Training: Cooperative SFT and RL for LLM ReasoningPaper-

RL balancing exploration and exploitation

Click to hide/show the paper list
DateNameTitlePaperGithub
2025-03DAPODAPO: An Open-Source LLM Reinforcement Learning System at ScalePaperGitHub Stars
2025-05Entropy Mechanism of RLThe Entropy Mechanism of Reinforcement Learning for Reasoning Language ModelsPaperGitHub Stars
2025-05ProRLProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language ModelsPaperHugging Face Downloads
2025-08CURECURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse PreventionPaperGitHub Stars
2025-10Reasoning Boundary ParadoxThe Reasoning Boundary Paradox: How Reinforcement Learning Constrains Language ModelsPaperGitHub Stars
2025-09outcome-based explorationOutcome-based Exploration for LLM ReasoningPaper-
2025-06entropy-based advantageReasoning with Exploration: An Entropy PerspectivePaper-
2025-10AEPOArbitrary Entropy Policy Optimization Breaks The Exploration Bottleneck of Reinforcement LearningPaperGitHub Stars
2025-09AEntOn Entropy Control in LLM-RL AlgorithmsPaper-
2025-09SIRENRethinking Entropy Regularization in Large Reasoning ModelsPaperGitHub Stars

Contributors

AheadOFpotato

11 commits

ChnQ

1 commits

Trae1ounG

1 commits