Yuan-Li-FNLP/R3-RAG

Python

52

8 commits

updated Oct 5, 2025

See the code

README

R3-RAG: Learning Step-by-Step Reasoning and Retrieval for LLMs via Reinforcement Learning

arXiv Python 3.10+ License

πŸ“– Overview

R3-RAG is a novel framework that uses Reinforcement learning to teach LLMs how to Reason and Retrieve step by step. Unlike traditional RAG methods that rely on human-designed workflows, R3-RAG enables models to autonomously learn optimal reasoning-retrieval strategies through reinforcement learning with both outcome and process rewards.

R3-RAG Training Pipeline

Key Features

  • πŸ”₯ Autonomous Learning: Uses RL to learn reasoning-retrieval strategies instead of relying on fixed human-designed workflows
  • 🎯 Dual Reward System: Combines outcome rewards (answer correctness) with process rewards (document relevance)
  • πŸš€ Strong Performance: Achieves significant improvements over state-of-the-art iterative RAG methods
  • πŸ”„ Transferable: Works across different retrievers (E5, BGE, BM25) with consistent performance
  • πŸ“Š Comprehensive: Evaluated on multiple multi-hop QA datasets (HotpotQA, 2WikiMultiHopQA, MuSiQue)

πŸ“Š Main Results

Our method significantly outperforms existing baselines across three multi-hop QA datasets:

MethodsRetrieverHotpotQA2WikiMultiHopQAMuSiQueAverage
Llama-3.1-8B
CoT-39.228.814.027.3
RAG with CoTE553.332.916.334.2
IRCoTE552.840.616.736.7
R3-RAGE564.461.032.252.6
R3-RAGBGE65.362.133.853.8
Qwen2.5-7B
CoT-34.031.112.725.9
RAG with CoTE552.433.516.934.3
IRCoTE548.435.813.532.6
R3-RAGE565.562.333.653.8
R3-RAGBGE66.463.034.854.8

πŸš€ Quick Start

Environment Setup

We recommend setting up three separate conda environments to avoid dependency conflicts:

  1. FlashRAG Environment (for retrieval tools): Please refer to FlashRAG to set up the environment.

  2. LLaMA-Factory Environment (for cold start training): Please refer to LLaMA-Factory to set up the environment.

  3. OpenRLHF Environment (for RL training): Please refer to OpenRLHF to set up the environment, then install our modified openrlhf code in this repository.

Model Download

Download our pre-trained models from Hugging Face:

# Cold start models
git clone https://huggingface.co/Yuan-Li-FNLP/R3-RAG-CS-Llama
git clone https://huggingface.co/Yuan-Li-FNLP/R3-RAG-CS-Qwen

# Full R3-RAG models
git clone https://huggingface.co/Yuan-Li-FNLP/R3-RAG-Llama
git clone https://huggingface.co/Yuan-Li-FNLP/R3-RAG-Qwen

Quick Demo

Experience R3-RAG with our visualization interface:

# First, start the server
cd startup
bash server.sh

# Then, start the visualization interface
bash startup_visualize.sh

Make sure to configure the model paths and parameters in the startup scripts before running.

πŸ“ Repository Structure

R3-RAG/
β”œβ”€β”€ benchmark/ # Evaluation scripts and benchmarks
β”‚ β”œβ”€β”€ evaluate.py # Main evaluation script
β”‚ β”œβ”€β”€ metrics/ # Evaluation metrics implementation
β”‚ └── datasets/ # Dataset loading and processing
β”œβ”€β”€ data/ # Cold start data construction
β”‚ β”œβ”€β”€ build_coldstart_data.py # Generate high-quality cold start trajectories
β”‚ β”œβ”€β”€ data_processing/ # Data preprocessing utilities
β”‚ └── templates/ # Prompt templates for data generation
β”œβ”€β”€ startup/ # Demo and visualization scripts
β”‚ β”œβ”€β”€ server.sh # Start the server
β”‚ β”œβ”€β”€ startup_visualize.sh # Start visualization interface
β”‚ └── demo_config.py # Configuration for demo
β”œβ”€β”€ tool/ # Retrieval tools and services
β”‚ β”œβ”€β”€ retrieval/ # Retrieval tool implementations
β”‚ β”œβ”€β”€ vllm_service/ # VLLM service code
β”‚ └── utils/ # Utility functions
β”œβ”€β”€ train/ # Training frameworks
β”‚ β”œβ”€β”€ llamafactory/ # SFT training with LLaMA-Factory
β”‚ β”‚ β”œβ”€β”€ sft_training.py # Cold start SFT training script
β”‚ β”‚ └── configs/ # Training configurations
β”‚ └── openrlhf/ # RLHF training with OpenRLHF
β”‚ β”œβ”€β”€ rl_training.py # Reinforcement learning training
β”‚ β”œβ”€β”€ reward_models/ # Reward model implementations
β”‚ └── configs/ # RL training configurations
β”œβ”€β”€ README.md # This file
└── LICENSE # License file

πŸ€— Available Models and Data

Models

Datasets

All models and datasets are available on Hugging Face.

πŸ“„ Citation

If you find our work helpful, please consider citing:

@misc{li2025r3raglearningstepbystepreasoning,
      title={R3-RAG: Learning Step-by-Step Reasoning and Retrieval for LLMs via Reinforcement Learning}, 
      author={Yuan Li and Qi Luo and Xiaonan Li and Bufan Li and Qinyuan Cheng and Bo Wang and Yining Zheng and Yuxin Wang and Zhangyue Yin and Xipeng Qiu},
      year={2025},
      eprint={2505.23794},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2505.23794}, 
}

🀝 Contributing

We welcome contributions! Please feel free to submit issues and pull requests.

πŸ“ License

This project is licensed under the GNU General Public License v3.0 - see the LICENSE file for details.

πŸ™ Acknowledgments

πŸ“ž Contact

For questions or collaborations, please contact:


Made with ❀️ by the Fudan NLP Group

Yuan-Li-FNLP/R3-RAG

Python

52

8 commits

updated Oct 5, 2025

See the code

README

R3-RAG: Learning Step-by-Step Reasoning and Retrieval for LLMs via Reinforcement Learning

arXiv Python 3.10+ License

πŸ“– Overview

R3-RAG is a novel framework that uses Reinforcement learning to teach LLMs how to Reason and Retrieve step by step. Unlike traditional RAG methods that rely on human-designed workflows, R3-RAG enables models to autonomously learn optimal reasoning-retrieval strategies through reinforcement learning with both outcome and process rewards.

R3-RAG Training Pipeline

Key Features

  • πŸ”₯ Autonomous Learning: Uses RL to learn reasoning-retrieval strategies instead of relying on fixed human-designed workflows
  • 🎯 Dual Reward System: Combines outcome rewards (answer correctness) with process rewards (document relevance)
  • πŸš€ Strong Performance: Achieves significant improvements over state-of-the-art iterative RAG methods
  • πŸ”„ Transferable: Works across different retrievers (E5, BGE, BM25) with consistent performance
  • πŸ“Š Comprehensive: Evaluated on multiple multi-hop QA datasets (HotpotQA, 2WikiMultiHopQA, MuSiQue)

πŸ“Š Main Results

Our method significantly outperforms existing baselines across three multi-hop QA datasets:

MethodsRetrieverHotpotQA2WikiMultiHopQAMuSiQueAverage
Llama-3.1-8B
CoT-39.228.814.027.3
RAG with CoTE553.332.916.334.2
IRCoTE552.840.616.736.7
R3-RAGE564.461.032.252.6
R3-RAGBGE65.362.133.853.8
Qwen2.5-7B
CoT-34.031.112.725.9
RAG with CoTE552.433.516.934.3
IRCoTE548.435.813.532.6
R3-RAGE565.562.333.653.8
R3-RAGBGE66.463.034.854.8

πŸš€ Quick Start

Environment Setup

We recommend setting up three separate conda environments to avoid dependency conflicts:

  1. FlashRAG Environment (for retrieval tools): Please refer to FlashRAG to set up the environment.

  2. LLaMA-Factory Environment (for cold start training): Please refer to LLaMA-Factory to set up the environment.

  3. OpenRLHF Environment (for RL training): Please refer to OpenRLHF to set up the environment, then install our modified openrlhf code in this repository.

Model Download

Download our pre-trained models from Hugging Face:

# Cold start models
git clone https://huggingface.co/Yuan-Li-FNLP/R3-RAG-CS-Llama
git clone https://huggingface.co/Yuan-Li-FNLP/R3-RAG-CS-Qwen

# Full R3-RAG models
git clone https://huggingface.co/Yuan-Li-FNLP/R3-RAG-Llama
git clone https://huggingface.co/Yuan-Li-FNLP/R3-RAG-Qwen

Quick Demo

Experience R3-RAG with our visualization interface:

# First, start the server
cd startup
bash server.sh

# Then, start the visualization interface
bash startup_visualize.sh

Make sure to configure the model paths and parameters in the startup scripts before running.

πŸ“ Repository Structure

R3-RAG/
β”œβ”€β”€ benchmark/ # Evaluation scripts and benchmarks
β”‚ β”œβ”€β”€ evaluate.py # Main evaluation script
β”‚ β”œβ”€β”€ metrics/ # Evaluation metrics implementation
β”‚ └── datasets/ # Dataset loading and processing
β”œβ”€β”€ data/ # Cold start data construction
β”‚ β”œβ”€β”€ build_coldstart_data.py # Generate high-quality cold start trajectories
β”‚ β”œβ”€β”€ data_processing/ # Data preprocessing utilities
β”‚ └── templates/ # Prompt templates for data generation
β”œβ”€β”€ startup/ # Demo and visualization scripts
β”‚ β”œβ”€β”€ server.sh # Start the server
β”‚ β”œβ”€β”€ startup_visualize.sh # Start visualization interface
β”‚ └── demo_config.py # Configuration for demo
β”œβ”€β”€ tool/ # Retrieval tools and services
β”‚ β”œβ”€β”€ retrieval/ # Retrieval tool implementations
β”‚ β”œβ”€β”€ vllm_service/ # VLLM service code
β”‚ └── utils/ # Utility functions
β”œβ”€β”€ train/ # Training frameworks
β”‚ β”œβ”€β”€ llamafactory/ # SFT training with LLaMA-Factory
β”‚ β”‚ β”œβ”€β”€ sft_training.py # Cold start SFT training script
β”‚ β”‚ └── configs/ # Training configurations
β”‚ └── openrlhf/ # RLHF training with OpenRLHF
β”‚ β”œβ”€β”€ rl_training.py # Reinforcement learning training
β”‚ β”œβ”€β”€ reward_models/ # Reward model implementations
β”‚ └── configs/ # RL training configurations
β”œβ”€β”€ README.md # This file
└── LICENSE # License file

πŸ€— Available Models and Data

Models

Datasets

All models and datasets are available on Hugging Face.

πŸ“„ Citation

If you find our work helpful, please consider citing:

@misc{li2025r3raglearningstepbystepreasoning,
      title={R3-RAG: Learning Step-by-Step Reasoning and Retrieval for LLMs via Reinforcement Learning}, 
      author={Yuan Li and Qi Luo and Xiaonan Li and Bufan Li and Qinyuan Cheng and Bo Wang and Yining Zheng and Yuxin Wang and Zhangyue Yin and Xipeng Qiu},
      year={2025},
      eprint={2505.23794},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2505.23794}, 
}

🀝 Contributing

We welcome contributions! Please feel free to submit issues and pull requests.

πŸ“ License

This project is licensed under the GNU General Public License v3.0 - see the LICENSE file for details.

πŸ™ Acknowledgments

πŸ“ž Contact

For questions or collaborations, please contact:


Made with ❀️ by the Fudan NLP Group

Languages

Python

98.1%

Shell

1.6%