
全面综述:RL在现实世界落地的未来方向
Maximum Likelihood Reinforcement Learning
MaxRL建立起RL(pass@1)与最大似然估计MLE(pass@k)之间的桥梁
JustRL: Scaling a 1.5B LLM with a Simple RL Recipe
当我们试着加入一些"应该有用"的优化时,性能反而下降了
Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning
技巧还是陷阱?从bese模型和aligned模型的角度观察


[复旦]PPO-Max:Secrets of RLHF in Large Language Models Part I- PPO [复旦]PPO-Max:github地址
Soft Clip 机制:CISPO-MiniMax-M1- Scaling Test-Time Compute Efficiently with Lightning Attention
FIPO(Future-KL Influenced Policy Optimization)
k3估计器:Approximating KL Divergence
Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
[for LLM]On a few pitfalls in KL divergence gradient estimation for
RL当你的 KL 散度正则化在“裸奔”
A Comedy of Estimators On KL Regularization in RL Training of LLMs
分析了两种主流 KL 估计器(K1 和 K3)在两种放置位置(Reward 和 Loss)下的梯度特性
[熵]1-The Entropy Mechanism of Reinforcement Learning for Reasoning Language Model
关注的是宏观的、全局的“策略熵”
Harnessing Uncertainty Entropy Modulated Policy Gradients for Long-Horizon LLM Agents
TODO Rethinking Entropy Regularization in Large Reasoning model
[推荐]Why Online Reinforcement Learning Forgets Less
核心论点:RL更新的“稀疏性”是个假象
为何在线强化学习能有效缓解灾难性遗忘?
TODO-Reinforcement Learning Finetunes Small Subnetworks in Large Language Models
Reinforcement Learning Finetunes Small Subnetworks in Large Language Models
指出 RL 微调是“局部更新”,而非全局重塑,因此更容易被后续训练干扰
Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs

优势重加权(Advantage Reweighting, AR):通过重新调整不同概率词元的优势(advantage)权重,直接削弱低概率词元的影响力。
低概率词元隔离(Low-Probability Token Isolation, Lopti):将更新过程分解为两个阶段,先更新低概率词元,再更新高概率词元,通过隔离来避免梯度干扰。
既然高概率词元的梯度那么小,我们干脆在更新时忽略它们,只用中低概率的词元不就行了吗?
Mitigating LLM Hallucination via Behaviorally Calibrated Reinforcement Learning
用行为校准 RL 抑制模型幻觉
Why Language Models Hallucinate
为什么大模型出现幻觉?
单奖励
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
多奖励+先归一化后加
Stabilizing Reinforcement Learning with LLMs: Formulation and Practices
一阶近似
【对整个 response 进行裁剪】
Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
强化学习是否真的在Llms中激发了超出基础模型的推理能力?
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
预训练(Pre-Training)、中期训练(Mid-Training)和基于强化学习的后训练(RL Post-Training)
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
爬坡似的接触特定的领域数据:基础领域知识-》专业领域知识
[推荐]Why Online Reinforcement Learning Forgets Less
这段话的核心发现是:人们之前认为 RL 的权重更新是稀疏的(即只有很少的参数被改变),但这其实是一个由训练精度设置引起的错觉。
为何在线强化学习能有效缓解灾难性遗忘?
The Path Not Taken: RLVR Provably Learns Off the Principals
RLVR微调的本质是“非主成分学习”!SFT微调的是“主成分”
The Art of Scaling Reinforcement Learning Compute for LLMs
RL 的 scaling 到底有没有规律可循?
对于视觉理解任务,显式的语言逻辑(Verbalized Logic)可能并不是必须的。


学习者在观察到结果后,会反思发生了什么,形成修正后的内部模型,并在后续的尝试中应用这些修正
使用On-Policy Distillation方案

策略空间非凸是训练不稳定的内在数学根源,而训练不稳定是非凸性在动态优化过程中的外在表现。
待添加

1.节点内部 (8 卡):使用 TP (张量并行)。因为 8 卡之间有 NVLink 高速互联,可以承受 TP 的高频通信,解决单卡存不下大层的问题。
2.节点之间:使用 PP (流水线并行)。将模型层切分到不同的机器组上,减少跨机器的通信频率。
3.整体集群:使用 DP (数据并行:模型复制,数据分片)。将上述的 "TP+PP" 组合视为一个大的“虚拟卡”,然后复制多份这样的组合,处理不同的数据批次,通过 DP 来扩大总吞吐量。

Megatron-LM/
├── megatron/
│ ├── core/ # Megatron Core (kernels, parallelism, building blocks)
│ │ ├── models/ # Transformer models
│ │ ├── transformer/ # Transformer building blocks
│ │ ├── tensor_parallel/ # Tensor parallelism
│ │ ├── pipeline_parallel/ # Pipeline parallelism
│ │ ├── distributed/ # Distributed training (FSDP, DDP)
│ │ ├── optimizer/ # Optimizers
│ │ ├── datasets/ # Dataset loaders
│ │ ├── inference/ # Inference engines and server
│ │ └── export/ # Model export (e.g. TensorRT-LLM)
│ ├── training/ # Training scripts
│ ├── legacy/ # Legacy components
│ ├── post_training/ # Post-training (quantization, distillation, pruning, etc.)
│ └── rl/ # Reinforcement learning (RLHF, etc.)
├── examples/ # Ready-to-use training examples
├── tools/ # Utility tools
├── tests/ # Comprehensive test suite
└── docs/ # Documentation
openrlhf-ds_rank0与vllm_ranks之间的通讯(训练与推理之间的通信)
openrlhf.utils.distributed_util

全面综述:RL在现实世界落地的未来方向
Maximum Likelihood Reinforcement Learning
MaxRL建立起RL(pass@1)与最大似然估计MLE(pass@k)之间的桥梁
JustRL: Scaling a 1.5B LLM with a Simple RL Recipe
当我们试着加入一些"应该有用"的优化时,性能反而下降了
Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning
技巧还是陷阱?从bese模型和aligned模型的角度观察


[复旦]PPO-Max:Secrets of RLHF in Large Language Models Part I- PPO [复旦]PPO-Max:github地址
Soft Clip 机制:CISPO-MiniMax-M1- Scaling Test-Time Compute Efficiently with Lightning Attention
FIPO(Future-KL Influenced Policy Optimization)
k3估计器:Approximating KL Divergence
Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
[for LLM]On a few pitfalls in KL divergence gradient estimation for
RL当你的 KL 散度正则化在“裸奔”
A Comedy of Estimators On KL Regularization in RL Training of LLMs
分析了两种主流 KL 估计器(K1 和 K3)在两种放置位置(Reward 和 Loss)下的梯度特性
[熵]1-The Entropy Mechanism of Reinforcement Learning for Reasoning Language Model
关注的是宏观的、全局的“策略熵”
Harnessing Uncertainty Entropy Modulated Policy Gradients for Long-Horizon LLM Agents
TODO Rethinking Entropy Regularization in Large Reasoning model
[推荐]Why Online Reinforcement Learning Forgets Less
核心论点:RL更新的“稀疏性”是个假象
为何在线强化学习能有效缓解灾难性遗忘?
TODO-Reinforcement Learning Finetunes Small Subnetworks in Large Language Models
Reinforcement Learning Finetunes Small Subnetworks in Large Language Models
指出 RL 微调是“局部更新”,而非全局重塑,因此更容易被后续训练干扰
Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs

优势重加权(Advantage Reweighting, AR):通过重新调整不同概率词元的优势(advantage)权重,直接削弱低概率词元的影响力。
低概率词元隔离(Low-Probability Token Isolation, Lopti):将更新过程分解为两个阶段,先更新低概率词元,再更新高概率词元,通过隔离来避免梯度干扰。
既然高概率词元的梯度那么小,我们干脆在更新时忽略它们,只用中低概率的词元不就行了吗?
Mitigating LLM Hallucination via Behaviorally Calibrated Reinforcement Learning
用行为校准 RL 抑制模型幻觉
Why Language Models Hallucinate
为什么大模型出现幻觉?
单奖励
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
多奖励+先归一化后加
Stabilizing Reinforcement Learning with LLMs: Formulation and Practices
一阶近似
【对整个 response 进行裁剪】
Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
强化学习是否真的在Llms中激发了超出基础模型的推理能力?
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
预训练(Pre-Training)、中期训练(Mid-Training)和基于强化学习的后训练(RL Post-Training)
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
爬坡似的接触特定的领域数据:基础领域知识-》专业领域知识
[推荐]Why Online Reinforcement Learning Forgets Less
这段话的核心发现是:人们之前认为 RL 的权重更新是稀疏的(即只有很少的参数被改变),但这其实是一个由训练精度设置引起的错觉。
为何在线强化学习能有效缓解灾难性遗忘?
The Path Not Taken: RLVR Provably Learns Off the Principals
RLVR微调的本质是“非主成分学习”!SFT微调的是“主成分”
The Art of Scaling Reinforcement Learning Compute for LLMs
RL 的 scaling 到底有没有规律可循?
对于视觉理解任务,显式的语言逻辑(Verbalized Logic)可能并不是必须的。


学习者在观察到结果后,会反思发生了什么,形成修正后的内部模型,并在后续的尝试中应用这些修正
使用On-Policy Distillation方案

策略空间非凸是训练不稳定的内在数学根源,而训练不稳定是非凸性在动态优化过程中的外在表现。
待添加

1.节点内部 (8 卡):使用 TP (张量并行)。因为 8 卡之间有 NVLink 高速互联,可以承受 TP 的高频通信,解决单卡存不下大层的问题。
2.节点之间:使用 PP (流水线并行)。将模型层切分到不同的机器组上,减少跨机器的通信频率。
3.整体集群:使用 DP (数据并行:模型复制,数据分片)。将上述的 "TP+PP" 组合视为一个大的“虚拟卡”,然后复制多份这样的组合,处理不同的数据批次,通过 DP 来扩大总吞吐量。

Megatron-LM/
├── megatron/
│ ├── core/ # Megatron Core (kernels, parallelism, building blocks)
│ │ ├── models/ # Transformer models
│ │ ├── transformer/ # Transformer building blocks
│ │ ├── tensor_parallel/ # Tensor parallelism
│ │ ├── pipeline_parallel/ # Pipeline parallelism
│ │ ├── distributed/ # Distributed training (FSDP, DDP)
│ │ ├── optimizer/ # Optimizers
│ │ ├── datasets/ # Dataset loaders
│ │ ├── inference/ # Inference engines and server
│ │ └── export/ # Model export (e.g. TensorRT-LLM)
│ ├── training/ # Training scripts
│ ├── legacy/ # Legacy components
│ ├── post_training/ # Post-training (quantization, distillation, pruning, etc.)
│ └── rl/ # Reinforcement learning (RLHF, etc.)
├── examples/ # Ready-to-use training examples
├── tools/ # Utility tools
├── tests/ # Comprehensive test suite
└── docs/ # Documentation
openrlhf-ds_rank0与vllm_ranks之间的通讯(训练与推理之间的通信)
openrlhf.utils.distributed_util