A carefully curated collection of papers and resources exploring methods to advance reasoning in Vision-Language Models (VLMs).
| Title | Venue | Date | Code | Star |
|---|---|---|---|---|
| Enhancing LLMs for Physics Problem-Solving using Reinforcement Learning with Human-AI Feedback | arXiv | 2024-12-06 | - | - |
| RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation | arXiv | 2025-02-10 | - | - |
| Title | Venue | Date | Code | Star |
|---|---|---|---|---|
| Continual SFT Matches Multimodal RLHF with Negative Supervision | arXiv | 2024-11-22 | - | - |
| ORPO: Monolithic Preference Optimization without Reference Model | arXiv | 2024-03-14 | - | - |
| Title | Venue | Date | Code | Star |
|---|---|---|---|---|
| VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning | arXiv | 2025-04-13 | GitHub | |
| TimeZero: Temporal Video Grounding with Reasoning-Guided LVLM | arXiv | 2025-03-17 | GitHub | |
| Understanding R1-Zero-Like Training: A Critical Perspective | arXiv | 2025-03-26 | GitHub | |
| DRA-GRPO: Exploring Diversity-Aware Reward Adjustment for R1-Zero-Like Training of Large Language Models | arXiv | 2025-05-14 | GitHub |
| Title | Venue | Date | Code | Star |
|---|---|---|---|---|
| R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning | arXiv | 2025-05-05 | GitHub |
| Title | Venue | Date | Code | Star |
|---|---|---|---|---|
| Learning Future Representation with Synthetic Observations for Sample-efficient Reinforcement Learning | arXiv | 2024-05-20 | - | - |
| The Perfect Blend: Redefining RLHF with Mixture of Judges | arXiv | 2024-09-30 | - | - |
| Fast Best-of-N Decoding via Speculative Rejection | arXiv | 2024-10-31 | GitHub | |
| Distributionally Robust Optimization | arXiv | 2024-11-04 | - | - |
| MM-EUREKA: EXPLORING VISUAL AHA MOMENT WITH RULE-BASED LARGE-SCALE REINFORCEMENTLEARNING | arXiv | 2025-03-10 | GitHub |
| Title | Date | Code | Star |
|---|---|---|---|
| R1-V: Reinforcing Super Generalization Ability in Vision Language Models with Less Than $3 | 2025-02-03 | GitHub |
| Title | Venue | Date | Code | Star |
|---|---|---|---|---|
| A COMPREHENSIVE SURVEY OF LLM ALIGNMENT TECHNIQUES:RLHF, RLAIF, PPO, DPO AND MORE | arXiv | 2024-07-23 | - | - |
| Personalized Generation In Large Model Era: A Survey | arXiv | 2025-03-04 | - | - |
| A Survey on Post-training of Large Language Models | arXiv | 2025-03-08 | - | - |
| Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models | arXiv | 2025-03-13 | GitHub | |
| Multimodal Chain-of-Thought Reasoning:A Comprehensive Survey | arXiv | 2025-03-16 | GitHub | |
| Awesome-Large-Multimodal-Reasoning-Models | arXiv | 2025-05-08 | GitHub |
22 commits
2 commits
A carefully curated collection of papers and resources exploring methods to advance reasoning in Vision-Language Models (VLMs).
| Title | Venue | Date | Code | Star |
|---|---|---|---|---|
| Enhancing LLMs for Physics Problem-Solving using Reinforcement Learning with Human-AI Feedback | arXiv | 2024-12-06 | - | - |
| RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation | arXiv | 2025-02-10 | - | - |
| Title | Venue | Date | Code | Star |
|---|---|---|---|---|
| Continual SFT Matches Multimodal RLHF with Negative Supervision | arXiv | 2024-11-22 | - | - |
| ORPO: Monolithic Preference Optimization without Reference Model | arXiv | 2024-03-14 | - | - |
| Title | Venue | Date | Code | Star |
|---|---|---|---|---|
| VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning | arXiv | 2025-04-13 | GitHub | |
| TimeZero: Temporal Video Grounding with Reasoning-Guided LVLM | arXiv | 2025-03-17 | GitHub | |
| Understanding R1-Zero-Like Training: A Critical Perspective | arXiv | 2025-03-26 | GitHub | |
| DRA-GRPO: Exploring Diversity-Aware Reward Adjustment for R1-Zero-Like Training of Large Language Models | arXiv | 2025-05-14 | GitHub |
| Title | Venue | Date | Code | Star |
|---|---|---|---|---|
| R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning | arXiv | 2025-05-05 | GitHub |
| Title | Venue | Date | Code | Star |
|---|---|---|---|---|
| Learning Future Representation with Synthetic Observations for Sample-efficient Reinforcement Learning | arXiv | 2024-05-20 | - | - |
| The Perfect Blend: Redefining RLHF with Mixture of Judges | arXiv | 2024-09-30 | - | - |
| Fast Best-of-N Decoding via Speculative Rejection | arXiv | 2024-10-31 | GitHub | |
| Distributionally Robust Optimization | arXiv | 2024-11-04 | - | - |
| MM-EUREKA: EXPLORING VISUAL AHA MOMENT WITH RULE-BASED LARGE-SCALE REINFORCEMENTLEARNING | arXiv | 2025-03-10 | GitHub |
| Title | Date | Code | Star |
|---|---|---|---|
| R1-V: Reinforcing Super Generalization Ability in Vision Language Models with Less Than $3 | 2025-02-03 | GitHub |
| Title | Venue | Date | Code | Star |
|---|---|---|---|---|
| A COMPREHENSIVE SURVEY OF LLM ALIGNMENT TECHNIQUES:RLHF, RLAIF, PPO, DPO AND MORE | arXiv | 2024-07-23 | - | - |
| Personalized Generation In Large Model Era: A Survey | arXiv | 2025-03-04 | - | - |
| A Survey on Post-training of Large Language Models | arXiv | 2025-03-08 | - | - |
| Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models | arXiv | 2025-03-13 | GitHub | |
| Multimodal Chain-of-Thought Reasoning:A Comprehensive Survey | arXiv | 2025-03-16 | GitHub | |
| Awesome-Large-Multimodal-Reasoning-Models | arXiv | 2025-05-08 | GitHub |
22 commits
2 commits