A curated list of reinforcement learning with human feedback resources (continually updated)
4,427
84 commits
updated May 20, 2026
This is a collection of research papers for Reinforcement Learning with Human Feedback (RLHF). And the repository will be continuously updated to track the frontier of RLHF.
Welcome to follow and star!
The idea of RLHF is to use methods from reinforcement learning to directly optimize a language model with human feedback. RLHF has enabled language models to begin to align a model trained on a general corpus of text data to that of complex human values.


(The following section was automatically generated by ChatGPT)
RLHF typically refers to "Reinforcement Learning with Human Feedback". Reinforcement Learning (RL) is a type of machine learning that involves training an agent to make decisions based on feedback from its environment. In RLHF, the agent also receives feedback from humans in the form of ratings or evaluations of its actions, which can help it learn more quickly and accurately.
RLHF is an active research area in artificial intelligence, with applications in fields such as robotics, gaming, and personalized recommendation systems. It seeks to address the challenges of RL in scenarios where the agent has limited access to feedback from the environment and requires human input to improve its performance.
Reinforcement Learning with Human Feedback (RLHF) is a rapidly developing area of research in artificial intelligence, and there are several advanced techniques that have been developed to improve the performance of RLHF systems. Here are some examples:
Inverse Reinforcement Learning (IRL): IRL is a technique that allows the agent to learn a reward function from human feedback, rather than relying on pre-defined reward functions. This makes it possible for the agent to learn from more complex feedback signals, such as demonstrations of desired behavior.
Apprenticeship Learning: Apprenticeship learning is a technique that combines IRL with supervised learning to enable the agent to learn from both human feedback and expert demonstrations. This can help the agent learn more quickly and effectively, as it is able to learn from both positive and negative feedback.
Interactive Machine Learning (IML): IML is a technique that involves active interaction between the agent and the human expert, allowing the expert to provide feedback on the agent's actions in real-time. This can help the agent learn more quickly and efficiently, as it can receive feedback on its actions at each step of the learning process.
Human-in-the-Loop Reinforcement Learning (HITLRL): HITLRL is a technique that involves integrating human feedback into the RL process at multiple levels, such as reward shaping, action selection, and policy optimization. This can help to improve the efficiency and effectiveness of the RLHF system by taking advantage of the strengths of both humans and machines.
Here are some examples of Reinforcement Learning with Human Feedback (RLHF):
Game Playing: In game playing, human feedback can help the agent learn strategies and tactics that are effective in different game scenarios. For example, in the popular game of Go, human experts can provide feedback to the agent on its moves, helping it improve its gameplay and decision-making.
Personalized Recommendation Systems: In recommendation systems, human feedback can help the agent learn the preferences of individual users, making it possible to provide personalized recommendations. For example, the agent could use feedback from users on recommended products to learn which features are most important to them.
Robotics: In robotics, human feedback can help the agent learn how to interact with the physical environment in a safe and efficient manner. For example, a robot could learn to navigate a new environment more quickly with feedback from a human operator on the best path to take or which objects to avoid.
Education: In education, human feedback can help the agent learn how to teach students more effectively. For example, an AI-based tutor could use feedback from teachers on which teaching strategies work best with different students, helping to personalize the learning experience.
You can also visit this link to get an AI-enhanced paper reading experience.
format:
- [title](paper link) [links]
- author1, author2, and author3...
- publisher
- keyword
- code
- experiment environments and datasets
Why DPO is a Misspecified Estimator and How to Fix It
What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data
Multiplayer Nash Preference Optimization
Token-Importance Guided Direct Preference Optimization
SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
BaseReward: A Strong Baseline for Multimodal Reward Model
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs
Learning to summarize user information for personalized reinforcement learning from human feedback
Token-Guard: Towards Token-Level Hallucination Control via Self-Checking Decoding
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
Unifying Stable Optimization and Reference Regularization in RLHF
Text2Grad: Reinforcement Learning from Natural Language Feedback
ARMOR: Aligning Secure and Safe Large Language Models via Meticulous Reasoning
Reward Model Routing in Alignment
All Roads Lead to Likelihood: The Value of Reinforcement Learning in Fine-Tuning
General Exploratory Bonus for Optimistic Exploration in RLHF
RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM Alignment
Learning Correlated Reward Models: Statistical Barriers and Opportunities
Verification and Co-Alignment via Heterogeneous Consistency for Preference-Aligned LLM Annotations
Disentangling Length Bias in Preference Learning via Response-Conditioned Modeling
Enforcing Axioms for AI Alignment under Loss-Based Rules
OPPO: Accelerating PPO-based RLHF via Pipeline Overlap
QuRL: Rubrics As Judge For Open-Ended Question Answering
Translate Policy to Language: Flow Matching Generated Rewards for LLM Explanations
Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy
Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game
Swap-guided Preference Learning for Personalized Reinforcement Learning from Human Feedback
Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards
Reward Models Inherit Value Biases from Pretraining
Semantic-aware Wasserstein Policy Regularization for Large Language Model Alignment
RewardBench 2: Advancing Reward Model Evaluation
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
COMAL: A Convergent Meta-Algorithm for Aligning LLMs with General Preferences
Displacement-Resistant Extensions of DPO with Nonconvex $f$-Divergences
Keep the Best, Forget the Rest: Reliable Alignment with Order-Aware Preference Optimization
Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment Dataset
Fair Reinforcement Learning for Just AI
Robust Reward Modeling via Causal Rubrics
Escaping Policy Contraction: Contraction-Aware PPO (CaPPO) for Stable Language Model Fine-Tuning
Beyond Pairwise: Empowering LLM Alignment With (Ranked) Choice Modeling
Learning Ordinal Probabilistic Reward from Preferences
Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment
Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment
Balancing the Experts: Unlocking LoRA-MoE for GRPO via Mechanism-Aware Rewards
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary
Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic Framework
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization
Threshold-Guided Optimization for Visual Generative Models
Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
Controllable and explainable personality sliders for LLMs at inference time
PS-PPO : Prefix-Sampling PPO for Critic-Free RLHF
Unbiased Reward Modeling from Implicit Preference
Real-Time Aligned Reward Model beyond Semantics
Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
B-Spar: Bayesian Sparse-Reward Modeling for RL-based Image Editing
Calibrated Preference Learning: The Case of Label Ranking
Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
DARC: Disagreement-Aware Alignment via Risk-Constrained Decoding
The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
Distributionally Robust Reinforcement Learning with Human Feedback
Automatically Finding Reward Model Biases
Convex Optimization for Alignment and Preference Learning on a Single GPU
MMKU-Bench: A Multimodal Update Benchmark for Diverse Visual Knowledge
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization
Unbiased Alignment for Large Language Models with Noisy Preferences
Unbiased Principles, Robust Rewards
The Secret Engine Behind RLHF: It's Contarstive Learning All Along
When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models
TUR-DPO: Topology- and Uncertainty-Aware Direct Preference Optimization
Asymptotic Universal Alignment: A New Alignment Framework via Test-Time Scaling
Reward Modeling from Natural Language Human Feedback
Efficient Preference Poisoning Attack on Offline RLHF
Position: Agentic Safety is an Epistemic Property, Not a Behavioral One
Position: Large Language Models Should Learn Personalized Rather Than Aggregated Human Preferences
Factored Causal Representation Learning for Robust Reward Modeling in RLHF
Online Compatible Reward Identification from Preference Feedback
$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses
IRPM: Intergroup Relative Preference Modeling for Pointwise Generative Reward Models
Implicit Preference Alignment for Human Image Animation
Multilingual Safety Alignment Via Sparse Weight Editing
Graph-Preference Learning: Debiasing Network-Sampled Human Feedback for Target Welfare Estimation
COLLIE: Guiding Skill Discovery in Semantically Coherent Latent Space
Optimal Transport for Reward Modeling from Noisy Feedback
Reliability-Aware LLM Alignment from Inconsistent Human Feedback
Position: We Need Large Language Models Optimized For Our Well-Being
Implicit Safety Alignment from Crowd Preferences
Contrastive Weak-to-Strong Generalization
Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner
Unifying Adversarial Robustness and Training Across Text Scoring Models
ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning
The Sign Estimator: Preference Modeling for LLM Alignment under Heterogeneity
Leveraging Machine Unlearning for Cost-Efficient Preference Alignment
Regularization in the Axiomatic Approach to Learning from Human Preferences
Conditional Equivalence of DPO and RLHF: Assumptions, Failure Modes, and Provable Alignment
A Regret Minimization Framework on Preference Learning in Large Language Models
Position: Measuring Human Preferences in RLHF is a Social Science Problem
Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling
Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data
Towards Efficient Online Exploration for Reinforcement Learning with Human Feedback
OpenRLHF: A Ray-based Easy-to-use, Scalable and High-performance RLHF Framework
Language Models Learn to Mislead Humans via RLHF
A Simple and Effective Reinforcement Learning Method for Text-to-Image Diffusion Fine-tuning
Differential Information: An Information-Theoretic Perspective on Preference Optimization
Generalist Reward Models: Found Inside Large Language Models
A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback
RLTHF: Targeted Human Feedback for LLM Alignment
Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback
Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model
REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models
DPO Meets PPO: Reinforced Token Optimization for RLHF
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
REvolve: Reward Evolution with Large Language Models using Human Feedback
Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward Inference
Learning Reward and Policy Jointly from Demonstration and Preference Improves Alignment
MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd
Aligning Few-Step Diffusion Models with Dense Reward Difference Learning
HybridFlow: A Flexible and Efficient RLHF Framework
ALaRM: Align Language Models via Hierarchical Rewards Modeling
TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback
Aligning Large Multimodal Models with Factually Augmented RLHF
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHF
Dense Reward for Free in Reinforcement Learning from Human Feedback
A Minimaximalist Approach to Reinforcement Learning from Human Feedback
RLHF Workflow: From Reward Modeling to Online RLHF
MaxMin-RLHF: Towards equitable alignment of large language models with diverse human preferences
Dataset Reset Policy Optimization for RLHF
A Dense Reward View on Aligning Text-to-Image Diffusion with Preference
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
On Diversified Preferences of Large Language Model Alignment
Aligning Crowd Feedback via Distributional Preference Reward Modeling
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!
A Theoretical Analysis of Nash Learning from Human Feedback under General KL-Regularized Preference
Mitigating the Alignment Tax of RLHF
Training Diffusion Models with Reinforcement Learning
AlignDiff: Aligning Diverse Human Preferences via Behavior-Customisable Diffusion Model
Dense Reward for Free in Reinforcement Learning from Human Feedback
Transforming and Combining Rewards for Aligning Large Language Models
Parameter Efficient Reinforcement Learning from Human Feedback
Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble
RIME: Robust Preference-based Reinforcement Learning with Noisy Human Preferences
The Trickle-down Impact of Reward (In-)consistency on RLHF
A General Theoretical Paradigm to Understand Learning from Human Preferences
Fine-Grained Human Feedback Gives Better Rewards for Language Model Training
Preference-grounded Token-level Guidance for Language Model Fine-tuning
Inverse Preference Learning: Preference-based RL without a Reward Function
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
Adversarial Preference Optimization
Sample Efficient Reinforcement Learning from Human Feedback via Active Exploration
Reinforcement Learning from Statistical Feedback: the Journey from AB Testing to ANT Testing
Data-Efficient Alignment of Large Language Models with Human Feedback Through Natural Language
Direct Preference-based Policy Optimization without Reward Modeling
AlignDiff: Aligning Diverse Human Preferences via Behavior-Customisable Diffusion Model
Eureka: Human-Level Reward Design via Coding Large Language Models
Safe RLHF: Safe Reinforcement Learning from Human Feedback
Quality Diversity through Human Feedback
Tuning computer vision models with task rewards
The Wisdom of Hindsight Makes Language Models Better Instruction Followers
Language Instructed Reinforcement Learning for Human-AI Coordination
Aligning Language Models with Offline Reinforcement Learning from Human Feedback
Preference Ranking Optimization for Human Alignment
Bridging the Gap: A Survey on Integrating (Human) Feedback for Natural Language Generation
RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Few-shot Preference Learning for Human-in-the-Loop RL
Better Aligning Text-to-Image Models with Human Preference
ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation
Aligning Text-to-Image Models using Human Feedback
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Pretraining Language Models with Human Preferences (PHF)
Aligning Language Models with Preferences through f-divergence Minimization (f-DPG)
Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons
The Capacity for Moral Self-Correction in Large Language Models
format:
- [title](codebase link) [links]
- author1, author2, and author3...
- keyword
- experiment environments, datasets or tasks
format:
- [title](dataset link) [links]
- author1, author2, and author3...
- keyword
- experiment environments or tasks
Our purpose is to make this repo even better. If you are interested in contributing, please refer to HERE for instructions in contribution.
Awesome RLHF is released under the Apache 2.0 license.
A curated list of reinforcement learning with human feedback resources (continually updated)
4,427
84 commits
updated May 20, 2026
This is a collection of research papers for Reinforcement Learning with Human Feedback (RLHF). And the repository will be continuously updated to track the frontier of RLHF.
Welcome to follow and star!
The idea of RLHF is to use methods from reinforcement learning to directly optimize a language model with human feedback. RLHF has enabled language models to begin to align a model trained on a general corpus of text data to that of complex human values.


(The following section was automatically generated by ChatGPT)
RLHF typically refers to "Reinforcement Learning with Human Feedback". Reinforcement Learning (RL) is a type of machine learning that involves training an agent to make decisions based on feedback from its environment. In RLHF, the agent also receives feedback from humans in the form of ratings or evaluations of its actions, which can help it learn more quickly and accurately.
RLHF is an active research area in artificial intelligence, with applications in fields such as robotics, gaming, and personalized recommendation systems. It seeks to address the challenges of RL in scenarios where the agent has limited access to feedback from the environment and requires human input to improve its performance.
Reinforcement Learning with Human Feedback (RLHF) is a rapidly developing area of research in artificial intelligence, and there are several advanced techniques that have been developed to improve the performance of RLHF systems. Here are some examples:
Inverse Reinforcement Learning (IRL): IRL is a technique that allows the agent to learn a reward function from human feedback, rather than relying on pre-defined reward functions. This makes it possible for the agent to learn from more complex feedback signals, such as demonstrations of desired behavior.
Apprenticeship Learning: Apprenticeship learning is a technique that combines IRL with supervised learning to enable the agent to learn from both human feedback and expert demonstrations. This can help the agent learn more quickly and effectively, as it is able to learn from both positive and negative feedback.
Interactive Machine Learning (IML): IML is a technique that involves active interaction between the agent and the human expert, allowing the expert to provide feedback on the agent's actions in real-time. This can help the agent learn more quickly and efficiently, as it can receive feedback on its actions at each step of the learning process.
Human-in-the-Loop Reinforcement Learning (HITLRL): HITLRL is a technique that involves integrating human feedback into the RL process at multiple levels, such as reward shaping, action selection, and policy optimization. This can help to improve the efficiency and effectiveness of the RLHF system by taking advantage of the strengths of both humans and machines.
Here are some examples of Reinforcement Learning with Human Feedback (RLHF):
Game Playing: In game playing, human feedback can help the agent learn strategies and tactics that are effective in different game scenarios. For example, in the popular game of Go, human experts can provide feedback to the agent on its moves, helping it improve its gameplay and decision-making.
Personalized Recommendation Systems: In recommendation systems, human feedback can help the agent learn the preferences of individual users, making it possible to provide personalized recommendations. For example, the agent could use feedback from users on recommended products to learn which features are most important to them.
Robotics: In robotics, human feedback can help the agent learn how to interact with the physical environment in a safe and efficient manner. For example, a robot could learn to navigate a new environment more quickly with feedback from a human operator on the best path to take or which objects to avoid.
Education: In education, human feedback can help the agent learn how to teach students more effectively. For example, an AI-based tutor could use feedback from teachers on which teaching strategies work best with different students, helping to personalize the learning experience.
You can also visit this link to get an AI-enhanced paper reading experience.
format:
- [title](paper link) [links]
- author1, author2, and author3...
- publisher
- keyword
- code
- experiment environments and datasets
Why DPO is a Misspecified Estimator and How to Fix It
What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data
Multiplayer Nash Preference Optimization
Token-Importance Guided Direct Preference Optimization
SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
BaseReward: A Strong Baseline for Multimodal Reward Model
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs
Learning to summarize user information for personalized reinforcement learning from human feedback
Token-Guard: Towards Token-Level Hallucination Control via Self-Checking Decoding
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
Unifying Stable Optimization and Reference Regularization in RLHF
Text2Grad: Reinforcement Learning from Natural Language Feedback
ARMOR: Aligning Secure and Safe Large Language Models via Meticulous Reasoning
Reward Model Routing in Alignment
All Roads Lead to Likelihood: The Value of Reinforcement Learning in Fine-Tuning
General Exploratory Bonus for Optimistic Exploration in RLHF
RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM Alignment
Learning Correlated Reward Models: Statistical Barriers and Opportunities
Verification and Co-Alignment via Heterogeneous Consistency for Preference-Aligned LLM Annotations
Disentangling Length Bias in Preference Learning via Response-Conditioned Modeling
Enforcing Axioms for AI Alignment under Loss-Based Rules
OPPO: Accelerating PPO-based RLHF via Pipeline Overlap
QuRL: Rubrics As Judge For Open-Ended Question Answering
Translate Policy to Language: Flow Matching Generated Rewards for LLM Explanations
Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy
Stackelberg Learning from Human Feedback: Preference Optimization as a Sequential Game
Swap-guided Preference Learning for Personalized Reinforcement Learning from Human Feedback
Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards
Reward Models Inherit Value Biases from Pretraining
Semantic-aware Wasserstein Policy Regularization for Large Language Model Alignment
RewardBench 2: Advancing Reward Model Evaluation
Causally Robust Reward Learning from Reason-Augmented Preference Feedback
COMAL: A Convergent Meta-Algorithm for Aligning LLMs with General Preferences
Displacement-Resistant Extensions of DPO with Nonconvex $f$-Divergences
Keep the Best, Forget the Rest: Reliable Alignment with Order-Aware Preference Optimization
Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment Dataset
Fair Reinforcement Learning for Just AI
Robust Reward Modeling via Causal Rubrics
Escaping Policy Contraction: Contraction-Aware PPO (CaPPO) for Stable Language Model Fine-Tuning
Beyond Pairwise: Empowering LLM Alignment With (Ranked) Choice Modeling
Learning Ordinal Probabilistic Reward from Preferences
Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment
Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment
Balancing the Experts: Unlocking LoRA-MoE for GRPO via Mechanism-Aware Rewards
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary
Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic Framework
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment
BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
Safety Game: Inference-Time Alignment of Black-Box LLMs via Constrained Optimization
Threshold-Guided Optimization for Visual Generative Models
Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
Controllable and explainable personality sliders for LLMs at inference time
PS-PPO : Prefix-Sampling PPO for Critic-Free RLHF
Unbiased Reward Modeling from Implicit Preference
Real-Time Aligned Reward Model beyond Semantics
Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
B-Spar: Bayesian Sparse-Reward Modeling for RL-based Image Editing
Calibrated Preference Learning: The Case of Label Ranking
Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
DARC: Disagreement-Aware Alignment via Risk-Constrained Decoding
The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs
Distributionally Robust Reinforcement Learning with Human Feedback
Automatically Finding Reward Model Biases
Convex Optimization for Alignment and Preference Learning on a Single GPU
MMKU-Bench: A Multimodal Update Benchmark for Diverse Visual Knowledge
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization
Unbiased Alignment for Large Language Models with Noisy Preferences
Unbiased Principles, Robust Rewards
The Secret Engine Behind RLHF: It's Contarstive Learning All Along
When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models
TUR-DPO: Topology- and Uncertainty-Aware Direct Preference Optimization
Asymptotic Universal Alignment: A New Alignment Framework via Test-Time Scaling
Reward Modeling from Natural Language Human Feedback
Efficient Preference Poisoning Attack on Offline RLHF
Position: Agentic Safety is an Epistemic Property, Not a Behavioral One
Position: Large Language Models Should Learn Personalized Rather Than Aggregated Human Preferences
Factored Causal Representation Learning for Robust Reward Modeling in RLHF
Online Compatible Reward Identification from Preference Feedback
$f$-Divergence Regularized RLHF: Two Tales of Sampling and Unified Analyses
IRPM: Intergroup Relative Preference Modeling for Pointwise Generative Reward Models
Implicit Preference Alignment for Human Image Animation
Multilingual Safety Alignment Via Sparse Weight Editing
Graph-Preference Learning: Debiasing Network-Sampled Human Feedback for Target Welfare Estimation
COLLIE: Guiding Skill Discovery in Semantically Coherent Latent Space
Optimal Transport for Reward Modeling from Noisy Feedback
Reliability-Aware LLM Alignment from Inconsistent Human Feedback
Position: We Need Large Language Models Optimized For Our Well-Being
Implicit Safety Alignment from Crowd Preferences
Contrastive Weak-to-Strong Generalization
Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner
Unifying Adversarial Robustness and Training Across Text Scoring Models
ActiveUltraFeedback: Efficient Preference Data Generation using Active Learning
The Sign Estimator: Preference Modeling for LLM Alignment under Heterogeneity
Leveraging Machine Unlearning for Cost-Efficient Preference Alignment
Regularization in the Axiomatic Approach to Learning from Human Preferences
Conditional Equivalence of DPO and RLHF: Assumptions, Failure Modes, and Provable Alignment
A Regret Minimization Framework on Preference Learning in Large Language Models
Position: Measuring Human Preferences in RLHF is a Social Science Problem
Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling
Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data
Towards Efficient Online Exploration for Reinforcement Learning with Human Feedback
OpenRLHF: A Ray-based Easy-to-use, Scalable and High-performance RLHF Framework
Language Models Learn to Mislead Humans via RLHF
A Simple and Effective Reinforcement Learning Method for Text-to-Image Diffusion Fine-tuning
Differential Information: An Information-Theoretic Perspective on Preference Optimization
Generalist Reward Models: Found Inside Large Language Models
A Unified Pairwise Framework for RLHF: Bridging Generative Reward Modeling and Policy Optimization
Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback
RLTHF: Targeted Human Feedback for LLM Alignment
Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback
Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model
REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models
DPO Meets PPO: Reinforced Token Optimization for RLHF
Reward-Augmented Data Enhances Direct Preference Alignment of LLMs
The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
REvolve: Reward Evolution with Large Language Models using Human Feedback
Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward Inference
Learning Reward and Policy Jointly from Demonstration and Preference Improves Alignment
MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd
Aligning Few-Step Diffusion Models with Dense Reward Difference Learning
HybridFlow: A Flexible and Efficient RLHF Framework
ALaRM: Align Language Models via Hierarchical Rewards Modeling
TLCR: Token-Level Continuous Reward for Fine-grained Reinforcement Learning from Human Feedback
Aligning Large Multimodal Models with Factually Augmented RLHF
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Principled Penalty-based Methods for Bilevel Reinforcement Learning and RLHF
Dense Reward for Free in Reinforcement Learning from Human Feedback
A Minimaximalist Approach to Reinforcement Learning from Human Feedback
RLHF Workflow: From Reward Modeling to Online RLHF
MaxMin-RLHF: Towards equitable alignment of large language models with diverse human preferences
Dataset Reset Policy Optimization for RLHF
A Dense Reward View on Aligning Text-to-Image Diffusion with Preference
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
On Diversified Preferences of Large Language Model Alignment
Aligning Crowd Feedback via Distributional Preference Reward Modeling
Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization
Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!
A Theoretical Analysis of Nash Learning from Human Feedback under General KL-Regularized Preference
Mitigating the Alignment Tax of RLHF
Training Diffusion Models with Reinforcement Learning
AlignDiff: Aligning Diverse Human Preferences via Behavior-Customisable Diffusion Model
Dense Reward for Free in Reinforcement Learning from Human Feedback
Transforming and Combining Rewards for Aligning Large Language Models
Parameter Efficient Reinforcement Learning from Human Feedback
Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble
RIME: Robust Preference-based Reinforcement Learning with Noisy Human Preferences
The Trickle-down Impact of Reward (In-)consistency on RLHF
A General Theoretical Paradigm to Understand Learning from Human Preferences
Fine-Grained Human Feedback Gives Better Rewards for Language Model Training
Preference-grounded Token-level Guidance for Language Model Fine-tuning
Inverse Preference Learning: Preference-based RL without a Reward Function
AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
Adversarial Preference Optimization
Sample Efficient Reinforcement Learning from Human Feedback via Active Exploration
Reinforcement Learning from Statistical Feedback: the Journey from AB Testing to ANT Testing
Data-Efficient Alignment of Large Language Models with Human Feedback Through Natural Language
Direct Preference-based Policy Optimization without Reward Modeling
AlignDiff: Aligning Diverse Human Preferences via Behavior-Customisable Diffusion Model
Eureka: Human-Level Reward Design via Coding Large Language Models
Safe RLHF: Safe Reinforcement Learning from Human Feedback
Quality Diversity through Human Feedback
Tuning computer vision models with task rewards
The Wisdom of Hindsight Makes Language Models Better Instruction Followers
Language Instructed Reinforcement Learning for Human-AI Coordination
Aligning Language Models with Offline Reinforcement Learning from Human Feedback
Preference Ranking Optimization for Human Alignment
Bridging the Gap: A Survey on Integrating (Human) Feedback for Natural Language Generation
RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Few-shot Preference Learning for Human-in-the-Loop RL
Better Aligning Text-to-Image Models with Human Preference
ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation
Aligning Text-to-Image Models using Human Feedback
Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Pretraining Language Models with Human Preferences (PHF)
Aligning Language Models with Preferences through f-divergence Minimization (f-DPG)
Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons
The Capacity for Moral Self-Correction in Large Language Models
format:
- [title](codebase link) [links]
- author1, author2, and author3...
- keyword
- experiment environments, datasets or tasks
format:
- [title](dataset link) [links]
- author1, author2, and author3...
- keyword
- experiment environments or tasks
Our purpose is to make this repo even better. If you are interested in contributing, please refer to HERE for instructions in contribution.
Awesome RLHF is released under the Apache 2.0 license.