A curated list of reinforcement learning with verifiable rewards (continually updated)
328
10 commits
updated Jun 1, 2026
A curated collection of surveys, tutorials, codebases and papers on
Reinforcement Learning with Verifiable Rewards (RLVR)—
a rapidly emerging paradigm that aligns both LLMs and other agents through
objective, externally verifiable signals.
An overview of how Reinforcement Learning with Verifiable Rewards (RLVR) works.
(Figure taken from
“Tülu 3: Pushing Frontiers in
Open Language Model Post-Training”)
RLVR couples reinforcement learning with objective, externally verifiable signals, yielding a training paradigm that is simultaneously powerful and trustworthy:
Through repeated iterations of this loop, the policy learns to maximise the externally verifiable reward while maintaining a clear audit trail for every decision it makes.
Pull requests are welcome 🎉 — see Contributing for guidelines.
[2026-05-06] New! Added 135 papers from ICLR 2026 and ICML 2026 🎉 [2025-07-03] Initial public release of Awesome-RLVR
format:
- [title](paper link) (presentation type)
- main authors or main affiliations
- Key: key problems and insights
- Data Domain: experiment environments
Inference-Time Techniques for LLM Reasoning (Berkeley Lecture 2025)
Learning to Self-Improve & Reason with LLMs (Berkeley Talk 2025)
LLM Reasoning: Key Ideas and Limitations (Tutorial Slides 2024)
Can LLMs Reason & Plan? (ICML Tutorial 2024)
Towards Reasoning in Large Language Models (ACL Tutorial 2023)
From System 1 to System 2: A Survey of Reasoning Large Language Models (arXiv 2025)
Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models (arXiv 2025)
What, How, Where, and How Well? A Survey on Test-Time Scaling in Large Language Models (arXiv 2025)
A Survey of Efficient Reasoning for Large Reasoning Models: Language, Multimodality, and Beyond (arXiv 2025)
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models (arXiv 2025)
A Visual Guide to Reasoning LLMs (Newsletter 2025)
Understanding Reasoning LLMs – Methods and Strategies for Building and Refining Reasoning Models (Blog 2025)
An Illusion of Progress? Assessing the Current State of Web Agents (arXiv 2025)
Agentic Large Language Models, A Survey (arXiv 2025)
A Comprehensive Survey of LLM Alignment Techniques: RLHF, RLAIF, PPO, DPO and More (arXiv 2024)
Self-Improvement of LLM Agents through Reinforcement Learning at Scale (MIT Scale-ML Talk 2024)
Reinforcement Learning from Verifiable Rewards (Book 2026)
Reinforcement Learning from Verifiable Rewards (Blog 2025)
| Project | Stars | Description |
|---|---|---|
| open-r1 | Fully open reproduction of the DeepSeek-R1 pipeline (SFT, distillation, GRPO, evaluation) | |
| OpenRLHF | An Easy-to-use, Scalable and High-performance RLHF Framework based on Ray (PPO & GRPO & REINFORCE++ & vLLM & Ray & Dynamic Sampling & Async Agentic RL)* | |
| verl | a flexible, efficient and production-ready RL training library for large language models | |
| TinyZero | Minimal reproduction of DeepSeek R1-Zero | |
| AReaL | Ant Reasoning Reinforcement Learning for LLMs | |
| Open-Reasoner-Zero | one open source implementation of large-scale reasoning-oriented RL training focusing on scalability, simplicity and accessibility | |
| ROLL | an Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models | |
| slime | an LLM post-training framework for RL scaling with high-performance training and flexible data generation | |
| RAGEN | RAGEN (Reasoning AGENt, pronounced like "region") leverages reinforcement learning (RL) to train LLM reasoning agents in interactive, stochastic environments. | |
| PRIME | PRIME (Process Reinforcement through IMplicit REwards), an open-source solution for online RL with process rewards | |
| rllm | an open-source framework for post-training language agents via reinforcement learning | |
| Nemo-Aligner | Scalable toolkit for efficient model alignment | |
| Trinity-RFT | A unified RFT framework with plug-and-play modules (for algorithms, data pipelines, and synchronization) |
RuCL: Stratified Rubric-Based Curriculum Learning for Multimodal Large Language Model Reasoning
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning
See First, Reason Later: Mutual Information-Guided Reinforcement Learning for Vision-Language Models
G$^2$RPO: Geometric GRPO; Escaping LLM's Reasoning Rut to Break Accuracy--Entropy Trade-off
Rate or Fate? RLV$^{arepsilon}$R: Reinforcement Learning with Verifiable Noisy Rewards
Anchored Policy Optimization: Mitigating Exploration Collapse via Support-Constrained Rectification
Spurious Rewards: Rethinking Training Signals in RLVR
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
The Unlearnability Phenomenon in RLVR for Language Models
Evaluating Parameter Efficient Methods for RLVR
Clipping Bottleneck: Stabilizing RLVR via Stochastic Recovery of Near-Boundary Signals
One-Way Policy Optimization for Self-Evolving LLMs
Experience Augmented Policy Optimization for LLM Reasoning
EchoRL: Reinforcement Learning via Rollout Echoing
Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning
Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
TGPO: Efficient Policy Optimization through Sequence Anchor and Information Gating
Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently
Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs
Beyond Normalization: Rethinking the Partition Function as a Difficulty Scheduler for RLVR
Learning Useful Supervision for Reinforcement Learning in Reasoning Models
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
PLaID++: A Preference Aligned Language Model for Targeted Inorganic Materials Design
Enhancing Multi-Modal LLMs Reasoning via Difficulty-Aware Group Normalization
Escaping the Mode: Multi-Answer Reinforcement Learning in LMs
Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration
Reinforcement Learning via Self-Distillation
A Regret Minimization Framework on Preference Learning in Large Language Models
Geometry-Preserving Orthonormal Initialization for Low-Rank Adaptation in Reinforcement Learning
Scaling the Scaling Logic: Agentic Meta-Synthesis of Logic Reasoning
Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
Outcome-Based Rewards Do Not Guarantee Faithful and Verifiable Reasoning
Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain Reasoning
Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation
Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text
Shrinking the Variance: Shrinkage Baselines for Reinforcement Learning with Verifiable Rewards
Rubric Curriculum RL: Exploiting the Generation-Verification Gap in Creative Writing
On the Learning Dynamics of RLVR at the Edge of Competence
BroRL: Scaling Reinforcement Learning via Broadened Exploration
Probing RLVR Training Instability through the Lens of Objective-Level Hacking
Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
PretrainZero: Reinforcement Active Pretraining
Reward Modeling from Natural Language Human Feedback
RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
Rectifying LLM Thought from Lens of Optimization
EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget
HiPO: Self-Hint Policy Optimization for RLVR
Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs
Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
Generalization of RLVR Using Causal Reasoning as a Testbed
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
References Improve LLM Alignment in Non-Verifiable Domains
Parameter-Efficient Reinforcement Learning using Prefix Optimization
HARDTESTGEN: A High-Quality RL Verifier Generation Pipeline for LLM Algorithmic Coding
Learning to Reason as Action Abstractions with Scalable Mid-Training RL
Diversity-Enhanced Reasoning for Subjective Questions
PROS: Towards Compute-Efficient RLVR via Rollout Prefix Reuse
LongRLVR: Long-Context Reinforcement Learning Requires Verifiable Context Rewards
Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs
Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning
Group Verification-based Policy Optimization for Interactive Coding Agents
Beyond Magnitude: Leveraging Direction of RLVR Updates for LLM Reasoning
Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
ReVeal: Self-Evolving Code Agents via Reliable Self-Verification
Process-Verified Reinforcement Learning for Theorem Proving via Lean
RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment
Breaking Barriers: Do Reinforcement Post Training Gains Transfer To Unseen Domains?
Tina: Tiny Reasoning Models via LoRA
Learning to Reason without External Rewards
$ extbf{Re}^{2}$: Unlocking LLM Reasoning via Reinforcement Learning with Re-solving
Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent
Diversity-Incentivized Exploration for Versatile Reasoning
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
ExGRPO: Learning to Reason from Experience
MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
Learn More with Less: Uncertainty Consistency Guided Query Selection for RLVR
Spotlight on Token Perception for Multimodal Reinforcement Learning
SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
DuPO: Enabling Reliable Self-Verification via Dual Preference Optimization
FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable Reasoning
Conditional Advantage Estimation for Reinforcement Learning in Large Reasoning Models
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
Agentic Reinforced Policy Optimization
Beyond Pass@ 1: Self-Play with Variational Problem Synthesis Sustains RLVR
AnesSuite: A Comprehensive Benchmark and Dataset Suite for Anesthesiology Reasoning in LLMs
Native Reasoning Models: Training Language Models to Reason on Unverifiable Data
Overthinking Reduction with Decoupled Rewards and Curriculum Data Scheduling
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
QuRL: Low-Precision Reinforcement Learning for Efficient Reasoning
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward
Perception-Aware Policy Optimization for Multimodal Reasoning
Sample Lottery: Unsupervised Discovery of Critical Instances for LLM Reasoning
Scheduling Your LLM Reinforcement Learning with Reasoning Trees
Risk-Sensitive Reinforcement Learning for Alleviating Exploration Dilemmas in Large Language Models
QuRL: Rubrics As Judge For Open-Ended Question Answering
Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
How Far Can Unsupervised RLVR Scale LLM Training?
TraPO: A Semi-Supervised Reinforcement Learning Framework for Boosting LLM Reasoning
Search Self-Play: Pushing the Frontier of Agent Capability without Supervision
A Simple "Motivation" Can Enhance Reinforcement Finetuning of Large Reasoning Models
ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Absolute Zero: Reinforced Self-play Reasoning with Zero Data
CURE: Co-Evolving Coders and Unit Testers via Reinforcement Learning
Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles
To Think or Not To Think: A Study of Thinking in Rule-Based Visual Reinforcement Fine-Tuning
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
Beyond Verifiable Rewards: Scaling Reinforcement Learning in Language Models to Unverifiable Data
RLVR-World: Training World Models with Reinforcement Learning
SeRL: Self-play Reinforcement Learning for Large Language Models with Limited Data
Learning to Reason under Off-Policy Guidance
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset
SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Agentic RL Scaling Law: Spontaneous Code Execution for Mathematical Problem Solving
Rethinking Verification for LLM Code Generation: From Generation to Testing
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
Solver-Informed RL: Grounding Large Language Models for Authentic Optimization Modeling
ATLAS: Autoformalizing Theorems through Lifting, Augmentation, and Synthesis of Data
QiMeng-CodeV-R1: Reasoning-Enhanced Verilog Generation
QiMeng-SALV: Signal-Aware Learning for Verilog Code Generation
miniF2F-Lean Revisited: Reviewing Limitations and Charting a Path Forward
SWE-SQL: Illuminating LLM Pathways to Solve User SQL Issues in Real-World Applications
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
VinePPO: Refining Credit Assignment in RL Training of LLMs
Controlling Large Language Model with Latent Action
Emergent Misalignment: Narrow Finetuning Can Produce Broadly Misaligned LLMs
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
Demystifying Long Chain-of-Thought Reasoning
SHIELDAGENT: Shielding Agents via Verifiable Safety Policy Reasoning
TOPLOC: A Locality Sensitive Hashing Scheme for Trustless Verifiable Inference
Brain Bandit: A Biologically Grounded Neural Network for Efficient Control of Exploration
AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
Towards Agentic Self-Learning LLMs in Search Environment
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
Language Models that Think, Chat Better
RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Demystifying Long Chain-of-Thought Reasoning in LLMs
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Kimi K 1.5: Scaling Reinforcement Learning with LLMs
S²R: Teaching LLMs to Self-Verify and Self-Correct via Reinforcement Learning
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search
Process Reward Models That Think
THINKPRUNE: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
GPG: A Simple and Strong Reinforcement Learning Baseline for Model Reasoning
SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
L1: Controlling How Long a Reasoning Model Thinks With Reinforcement Learning
Scaling Test-Time Compute Without Verification or RL is Suboptimal
DAST: Difficulty-Adaptive Slow-Thinking for Large Reasoning Models
Reasoning with Reinforced Functional Token Tuning
Provably Optimal Distributional RL for LLM Post-Training
On the Emergence of Thinking in LLMs I: Searching for the Right Intuition
STP: Self-Play LLM Theorem Provers with Iterative Conjecturing and Proving
A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
Recitation over Reasoning: How Cutting-Edge LMs Fail on Elementary Reasoning Problems
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
(REINFORCE++) A Simple and Efficient Approach for Aligning Large Language Models
ReFT v3: Reasoning with Reinforced Fine-Tuning (ACL 2025)
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
SimPO: Simple Preference Optimization with a Reference-Free Reward
DeepSeek-Prover v1.5: Harnessing Proof Assistant Feedback for RL and MCTS
Tülu 3: Pushing Frontiers in Open Language Model Post-Training
Kimi k1.5: Scaling Reinforcement Learning with LLMs
Model Alignment as Prospect Theoretic Optimization
UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning
GUI-R1: A Generalist R1-Style Vision-Language Action Model for GUI Agents
MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning
Smart-Searcher: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning
Direct Preference Optimization: Your Language Model is Secretly a Reward Model (ICLR 2024)
Math-Shepherd: Verify and Reinforce LLMs Step-by-Step without Human Annotations (NeurIPS 2023)
Let’s Verify Step by Step (ICML 2023)
Solving Olympiad Geometry without Human Demonstrations (Nature 2023)
Training Language Models to Follow Instructions with Human Feedback (NeurIPS 2022)
Awesome-RLVR © 2025 OpenDILab & Contributors Apache 2.0 License
137 followers · starred Sep 2026
A curated list of reinforcement learning with verifiable rewards (continually updated)
328
10 commits
updated Jun 1, 2026
A curated collection of surveys, tutorials, codebases and papers on
Reinforcement Learning with Verifiable Rewards (RLVR)—
a rapidly emerging paradigm that aligns both LLMs and other agents through
objective, externally verifiable signals.
An overview of how Reinforcement Learning with Verifiable Rewards (RLVR) works.
(Figure taken from
“Tülu 3: Pushing Frontiers in
Open Language Model Post-Training”)
RLVR couples reinforcement learning with objective, externally verifiable signals, yielding a training paradigm that is simultaneously powerful and trustworthy:
Through repeated iterations of this loop, the policy learns to maximise the externally verifiable reward while maintaining a clear audit trail for every decision it makes.
Pull requests are welcome 🎉 — see Contributing for guidelines.
[2026-05-06] New! Added 135 papers from ICLR 2026 and ICML 2026 🎉 [2025-07-03] Initial public release of Awesome-RLVR
format:
- [title](paper link) (presentation type)
- main authors or main affiliations
- Key: key problems and insights
- Data Domain: experiment environments
Inference-Time Techniques for LLM Reasoning (Berkeley Lecture 2025)
Learning to Self-Improve & Reason with LLMs (Berkeley Talk 2025)
LLM Reasoning: Key Ideas and Limitations (Tutorial Slides 2024)
Can LLMs Reason & Plan? (ICML Tutorial 2024)
Towards Reasoning in Large Language Models (ACL Tutorial 2023)
From System 1 to System 2: A Survey of Reasoning Large Language Models (arXiv 2025)
Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models (arXiv 2025)
What, How, Where, and How Well? A Survey on Test-Time Scaling in Large Language Models (arXiv 2025)
A Survey of Efficient Reasoning for Large Reasoning Models: Language, Multimodality, and Beyond (arXiv 2025)
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models (arXiv 2025)
A Visual Guide to Reasoning LLMs (Newsletter 2025)
Understanding Reasoning LLMs – Methods and Strategies for Building and Refining Reasoning Models (Blog 2025)
An Illusion of Progress? Assessing the Current State of Web Agents (arXiv 2025)
Agentic Large Language Models, A Survey (arXiv 2025)
A Comprehensive Survey of LLM Alignment Techniques: RLHF, RLAIF, PPO, DPO and More (arXiv 2024)
Self-Improvement of LLM Agents through Reinforcement Learning at Scale (MIT Scale-ML Talk 2024)
Reinforcement Learning from Verifiable Rewards (Book 2026)
Reinforcement Learning from Verifiable Rewards (Blog 2025)
| Project | Stars | Description |
|---|---|---|
| open-r1 | Fully open reproduction of the DeepSeek-R1 pipeline (SFT, distillation, GRPO, evaluation) | |
| OpenRLHF | An Easy-to-use, Scalable and High-performance RLHF Framework based on Ray (PPO & GRPO & REINFORCE++ & vLLM & Ray & Dynamic Sampling & Async Agentic RL)* | |
| verl | a flexible, efficient and production-ready RL training library for large language models | |
| TinyZero | Minimal reproduction of DeepSeek R1-Zero | |
| AReaL | Ant Reasoning Reinforcement Learning for LLMs | |
| Open-Reasoner-Zero | one open source implementation of large-scale reasoning-oriented RL training focusing on scalability, simplicity and accessibility | |
| ROLL | an Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models | |
| slime | an LLM post-training framework for RL scaling with high-performance training and flexible data generation | |
| RAGEN | RAGEN (Reasoning AGENt, pronounced like "region") leverages reinforcement learning (RL) to train LLM reasoning agents in interactive, stochastic environments. | |
| PRIME | PRIME (Process Reinforcement through IMplicit REwards), an open-source solution for online RL with process rewards | |
| rllm | an open-source framework for post-training language agents via reinforcement learning | |
| Nemo-Aligner | Scalable toolkit for efficient model alignment | |
| Trinity-RFT | A unified RFT framework with plug-and-play modules (for algorithms, data pipelines, and synchronization) |
RuCL: Stratified Rubric-Based Curriculum Learning for Multimodal Large Language Model Reasoning
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning
See First, Reason Later: Mutual Information-Guided Reinforcement Learning for Vision-Language Models
G$^2$RPO: Geometric GRPO; Escaping LLM's Reasoning Rut to Break Accuracy--Entropy Trade-off
Rate or Fate? RLV$^{arepsilon}$R: Reinforcement Learning with Verifiable Noisy Rewards
Anchored Policy Optimization: Mitigating Exploration Collapse via Support-Constrained Rectification
Spurious Rewards: Rethinking Training Signals in RLVR
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
The Unlearnability Phenomenon in RLVR for Language Models
Evaluating Parameter Efficient Methods for RLVR
Clipping Bottleneck: Stabilizing RLVR via Stochastic Recovery of Near-Boundary Signals
One-Way Policy Optimization for Self-Evolving LLMs
Experience Augmented Policy Optimization for LLM Reasoning
EchoRL: Reinforcement Learning via Rollout Echoing
Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR
ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning
Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
TGPO: Efficient Policy Optimization through Sequence Anchor and Information Gating
Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently
Do We Need Adam? Surprisingly Strong and Sparse Reinforcement Learning with SGD in LLMs
Beyond Normalization: Rethinking the Partition Function as a Difficulty Scheduler for RLVR
Learning Useful Supervision for Reinforcement Learning in Reasoning Models
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
PLaID++: A Preference Aligned Language Model for Targeted Inorganic Materials Design
Enhancing Multi-Modal LLMs Reasoning via Difficulty-Aware Group Normalization
Escaping the Mode: Multi-Answer Reinforcement Learning in LMs
Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration
Reinforcement Learning via Self-Distillation
A Regret Minimization Framework on Preference Learning in Large Language Models
Geometry-Preserving Orthonormal Initialization for Low-Rank Adaptation in Reinforcement Learning
Scaling the Scaling Logic: Agentic Meta-Synthesis of Logic Reasoning
Single-Rollout Hidden-State Dynamics for Training-Free RLVR Data Selection
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
Outcome-Based Rewards Do Not Guarantee Faithful and Verifiable Reasoning
Reward and Guidance through Rubrics: Promoting Exploration to Improve Multi-Domain Reasoning
Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation
Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text
Shrinking the Variance: Shrinkage Baselines for Reinforcement Learning with Verifiable Rewards
Rubric Curriculum RL: Exploiting the Generation-Verification Gap in Creative Writing
On the Learning Dynamics of RLVR at the Edge of Competence
BroRL: Scaling Reinforcement Learning via Broadened Exploration
Probing RLVR Training Instability through the Lens of Objective-Level Hacking
Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
PretrainZero: Reinforcement Active Pretraining
Reward Modeling from Natural Language Human Feedback
RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
Rectifying LLM Thought from Lens of Optimization
EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget
HiPO: Self-Hint Policy Optimization for RLVR
Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs
Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
Generalization of RLVR Using Causal Reasoning as a Testbed
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
References Improve LLM Alignment in Non-Verifiable Domains
Parameter-Efficient Reinforcement Learning using Prefix Optimization
HARDTESTGEN: A High-Quality RL Verifier Generation Pipeline for LLM Algorithmic Coding
Learning to Reason as Action Abstractions with Scalable Mid-Training RL
Diversity-Enhanced Reasoning for Subjective Questions
PROS: Towards Compute-Efficient RLVR via Rollout Prefix Reuse
LongRLVR: Long-Context Reinforcement Learning Requires Verifiable Context Rewards
Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs
Controllable Exploration in Hybrid-Policy RLVR for Multi-Modal Reasoning
Group Verification-based Policy Optimization for Interactive Coding Agents
Beyond Magnitude: Leveraging Direction of RLVR Updates for LLM Reasoning
Curriculum Reinforcement Learning from Easy to Hard Tasks Improves LLM Reasoning
ReVeal: Self-Evolving Code Agents via Reliable Self-Verification
Process-Verified Reinforcement Learning for Theorem Proving via Lean
RL Squeezes, SFT Expands: A Comparative Study of Reasoning LLMs
Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment
Breaking Barriers: Do Reinforcement Post Training Gains Transfer To Unseen Domains?
Tina: Tiny Reasoning Models via LoRA
Learning to Reason without External Rewards
$ extbf{Re}^{2}$: Unlocking LLM Reasoning via Reinforcement Learning with Re-solving
Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent
Diversity-Incentivized Exploration for Versatile Reasoning
Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards
ExGRPO: Learning to Reason from Experience
MATH-Beyond: A Benchmark for RL to Expand Beyond the Base Model
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
Learn More with Less: Uncertainty Consistency Guided Query Selection for RLVR
Spotlight on Token Perception for Multimodal Reinforcement Learning
SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
DuPO: Enabling Reliable Self-Verification via Dual Preference Optimization
FAPO: Flawed-Aware Policy Optimization for Efficient and Reliable Reasoning
Conditional Advantage Estimation for Reinforcement Learning in Large Reasoning Models
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
Agentic Reinforced Policy Optimization
Beyond Pass@ 1: Self-Play with Variational Problem Synthesis Sustains RLVR
AnesSuite: A Comprehensive Benchmark and Dataset Suite for Anesthesiology Reasoning in LLMs
Native Reasoning Models: Training Language Models to Reason on Unverifiable Data
Overthinking Reduction with Decoupled Rewards and Curriculum Data Scheduling
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
QuRL: Low-Precision Reinforcement Learning for Efficient Reasoning
Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward
Perception-Aware Policy Optimization for Multimodal Reasoning
Sample Lottery: Unsupervised Discovery of Critical Instances for LLM Reasoning
Scheduling Your LLM Reinforcement Learning with Reasoning Trees
Risk-Sensitive Reinforcement Learning for Alleviating Exploration Dilemmas in Large Language Models
QuRL: Rubrics As Judge For Open-Ended Question Answering
Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
How Far Can Unsupervised RLVR Scale LLM Training?
TraPO: A Semi-Supervised Reinforcement Learning Framework for Boosting LLM Reasoning
Search Self-Play: Pushing the Frontier of Agent Capability without Supervision
A Simple "Motivation" Can Enhance Reinforcement Finetuning of Large Reasoning Models
ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Absolute Zero: Reinforced Self-play Reasoning with Zero Data
CURE: Co-Evolving Coders and Unit Testers via Reinforcement Learning
Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles
To Think or Not To Think: A Study of Thinking in Rule-Based Visual Reinforcement Fine-Tuning
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
Beyond Verifiable Rewards: Scaling Reinforcement Learning in Language Models to Unverifiable Data
RLVR-World: Training World Models with Reinforcement Learning
SeRL: Self-play Reinforcement Learning for Large Language Models with Limited Data
Learning to Reason under Off-Policy Guidance
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
rStar-Coder: Scaling Competitive Code Reasoning with a Large-Scale Verified Dataset
SynLogic: Synthesizing Verifiable Reasoning Data at Scale for Learning Logical Reasoning and Beyond
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Agentic RL Scaling Law: Spontaneous Code Execution for Mathematical Problem Solving
Rethinking Verification for LLM Code Generation: From Generation to Testing
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
Scaling Code-Assisted Chain-of-Thoughts and Instructions for Model Reasoning
The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
Solver-Informed RL: Grounding Large Language Models for Authentic Optimization Modeling
ATLAS: Autoformalizing Theorems through Lifting, Augmentation, and Synthesis of Data
QiMeng-CodeV-R1: Reasoning-Enhanced Verilog Generation
QiMeng-SALV: Signal-Aware Learning for Verilog Code Generation
miniF2F-Lean Revisited: Reviewing Limitations and Charting a Path Forward
SWE-SQL: Illuminating LLM Pathways to Solve User SQL Issues in Real-World Applications
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
VinePPO: Refining Credit Assignment in RL Training of LLMs
Controlling Large Language Model with Latent Action
Emergent Misalignment: Narrow Finetuning Can Produce Broadly Misaligned LLMs
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
Demystifying Long Chain-of-Thought Reasoning
SHIELDAGENT: Shielding Agents via Verifiable Safety Policy Reasoning
TOPLOC: A Locality Sensitive Hashing Scheme for Trustless Verifiable Inference
Brain Bandit: A Biologically Grounded Neural Network for Efficient Control of Exploration
AgentPRM: Process Reward Models for LLM Agents via Step-Wise Promise and Progress
Towards Agentic Self-Learning LLMs in Search Environment
Reinforcement Learning with Verifiable yet Noisy Rewards under Imperfect Verifiers
Language Models that Think, Chat Better
RLVE: Scaling Up Reinforcement Learning for Language Models with Adaptive Verifiable Environments
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Demystifying Long Chain-of-Thought Reasoning in LLMs
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Kimi K 1.5: Scaling Reinforcement Learning with LLMs
S²R: Teaching LLMs to Self-Verify and Self-Correct via Reinforcement Learning
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling
QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search
Process Reward Models That Think
THINKPRUNE: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
GPG: A Simple and Strong Reinforcement Learning Baseline for Model Reasoning
SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
L1: Controlling How Long a Reasoning Model Thinks With Reinforcement Learning
Scaling Test-Time Compute Without Verification or RL is Suboptimal
DAST: Difficulty-Adaptive Slow-Thinking for Large Reasoning Models
Reasoning with Reinforced Functional Token Tuning
Provably Optimal Distributional RL for LLM Post-Training
On the Emergence of Thinking in LLMs I: Searching for the Right Intuition
STP: Self-Play LLM Theorem Provers with Iterative Conjecturing and Proving
A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
Recitation over Reasoning: How Cutting-Edge LMs Fail on Elementary Reasoning Problems
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad
(REINFORCE++) A Simple and Efficient Approach for Aligning Large Language Models
ReFT v3: Reasoning with Reinforced Fine-Tuning (ACL 2025)
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
SimPO: Simple Preference Optimization with a Reference-Free Reward
DeepSeek-Prover v1.5: Harnessing Proof Assistant Feedback for RL and MCTS
Tülu 3: Pushing Frontiers in Open Language Model Post-Training
Kimi k1.5: Scaling Reinforcement Learning with LLMs
Model Alignment as Prospect Theoretic Optimization
UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning
GUI-R1: A Generalist R1-Style Vision-Language Action Model for GUI Agents
MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning
Smart-Searcher: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning
Direct Preference Optimization: Your Language Model is Secretly a Reward Model (ICLR 2024)
Math-Shepherd: Verify and Reinforce LLMs Step-by-Step without Human Annotations (NeurIPS 2023)
Let’s Verify Step by Step (ICML 2023)
Solving Olympiad Geometry without Human Demonstrations (Nature 2023)
Training Language Models to Follow Instructions with Human Feedback (NeurIPS 2022)
Awesome-RLVR © 2025 OpenDILab & Contributors Apache 2.0 License
137 followers · starred Sep 2026