Post-training framework for large models, from new objectives to new rollout systems.
See the code
Algorithm-first post-training framework for large language models (LLMs) and vision-language models (VLMs).
"What I cannot create, I do not understand." — Richard Feynman
FeynRL (pronounced "FineRL") is an algorithm-first framework for post-training and fine-tuning large models. It works with both text-only large language models (LLMs) and vision-language models (VLMs), and supports supervised fine-tuning (SFT), preference learning (e.g., DPO), and reinforcement learning (e.g., PPO, GRPO, DAPO, CISPO, P3O) across both. It is built for researchers and engineers who want to understand, modify, and develop new methods without fighting the infrastructure.
The main goal of FeynRL is simple: make new algorithms easy to implement, easy to debug, and still possible to train at scale. The codebase is designed so that algorithmic logic stays local and systems logic stays explicit, which makes the framework easier to reason about, easier to extend, and more reliable to debug.
Most post-training frameworks optimize for the largest built-in feature surface, which makes them powerful but hard to modify once you want to try something new. FeynRL makes the opposite trade-off: it optimizes for clarity, locality of change, and algorithm development. Adding a new method usually means writing a single file with its own loss and update logic, not threading changes through the orchestration, rollout, and data layers.
FeynRL is built for anyone who wants to deeply understand training of large models, become an expert, and build new methods and ideas. It's written to be read and learned from, so you can trace exactly how rollouts, losses, and weight syncs fit together, experiment with your own methods, and grow from running recipes to designing them. You can still use it as a ready-made recipe out of the box; the difference is that nothing stays a black box when you're ready to look inside.
This is the first public release, so expect rough edges. We're open-sourcing it as a foundation for post-training research with the community.
For a detailed breakdown of the architecture, see the Architecture Overview.
For RL, Ray orchestrates the full training loop: it schedules DeepSpeed training workers and vLLM rollout workers across nodes, and coordinates weight synchronization between them. In sync mode, each epoch generates all rollouts, trains on them, syncs weights, and repeats — fully on-policy and easy to reason about. In overlap mode (also called async mode), rollout generation and training run concurrently on separate GPU pools so training GPUs don't sit idle waiting for rollouts. Generation is continuous across epoch boundaries and checkpoint saves — the only pauses are brief drains during weight sync, which runs once at the end of every non-final epoch. A configurable staleness budget bounds how off-policy the replay data can drift. Async mode uses NCCL for weight sync; sync mode pushes weights directly through GPU memory and falls back to disk-based reload if that fails. SFT and DPO are simpler because they only require a single model and no rollout workers, so they run directly on DeepSpeed without Ray. All paradigms support full fine-tuning and LoRA, and plug into mixed-dataset sampling, experiment tracking, and standalone evaluation without changing the overall workflow.
The repository is organized so that algorithmic changes usually stay local:
algs/ — Algorithm and optimization logic. Each algorithm (PPO, GRPO, DAPO, CISPO, P3O, DPO, SFT) has its own module with a README documenting the math and pseudocode.rollouts/ — Rollout generation, vLLM engine wrappers, weight sync, and replay buffer.rewards/ — Pluggable reward functions (GSM8K, math verification, and custom).data_feeds/ — Data loading, sampling, and mixed-dataset support.data_prep/ — Dataset preparation scripts.configs/ — YAML configs for RL, SFT, DPO, and evaluation, with full parameter reference.unit_tests/ — Unit and integration tests.Installation & Setup — Configure your environment and dependencies.
Quickstart & How-To — Learn how to launch jobs and run experiments.
Experiments — Reference experiment results and the canonical example configs used to reproduce them.
Configuration Reference — Full parameter guide for RL, SFT, DPO, and evaluation configs.
How the Async Engine Works — An illustrated guide to the overlapped (async) RL engine.
Troubleshooting — Diagnose and fix common issues.
Contributions are welcome! Please see our Contributing Guidelines for details on how to get involved.
Check out the FAQ for common questions and answers.
Special thanks to the Open-Instruct and PipelineRL projects, both of which provided valuable resources for this work and the broader community.
If you use FeynRL in your work, please cite the following:
@misc{FeynRL,
author = {Rasool Fakoor and Murdock Aubry and FeynRL Contributors},
title = {FeynRL: A Modular and Scalable LLM Post-Training Framework},
year = {2026},
howpublished = {\url{https://github.com/FeynRL-project/FeynRL}},
note = {GitHub repository. Corresponding author: Rasool Fakoor}
}
@misc{fakoor2026trustbatchonoffpolicy,
title = {Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training},
author = {Rasool Fakoor and Murdock Aubry and Nicholas Stranges and Alexander J. Smola},
year = {2026},
eprint = {2605.12380},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2605.12380}
}
141 followers · starred Mar 2026
80 followers · starred Mar 2026
169 followers · starred Mar 2026
Post-training framework for large models, from new objectives to new rollout systems.
See the code
Algorithm-first post-training framework for large language models (LLMs) and vision-language models (VLMs).
"What I cannot create, I do not understand." — Richard Feynman
FeynRL (pronounced "FineRL") is an algorithm-first framework for post-training and fine-tuning large models. It works with both text-only large language models (LLMs) and vision-language models (VLMs), and supports supervised fine-tuning (SFT), preference learning (e.g., DPO), and reinforcement learning (e.g., PPO, GRPO, DAPO, CISPO, P3O) across both. It is built for researchers and engineers who want to understand, modify, and develop new methods without fighting the infrastructure.
The main goal of FeynRL is simple: make new algorithms easy to implement, easy to debug, and still possible to train at scale. The codebase is designed so that algorithmic logic stays local and systems logic stays explicit, which makes the framework easier to reason about, easier to extend, and more reliable to debug.
Most post-training frameworks optimize for the largest built-in feature surface, which makes them powerful but hard to modify once you want to try something new. FeynRL makes the opposite trade-off: it optimizes for clarity, locality of change, and algorithm development. Adding a new method usually means writing a single file with its own loss and update logic, not threading changes through the orchestration, rollout, and data layers.
FeynRL is built for anyone who wants to deeply understand training of large models, become an expert, and build new methods and ideas. It's written to be read and learned from, so you can trace exactly how rollouts, losses, and weight syncs fit together, experiment with your own methods, and grow from running recipes to designing them. You can still use it as a ready-made recipe out of the box; the difference is that nothing stays a black box when you're ready to look inside.
This is the first public release, so expect rough edges. We're open-sourcing it as a foundation for post-training research with the community.
For a detailed breakdown of the architecture, see the Architecture Overview.
For RL, Ray orchestrates the full training loop: it schedules DeepSpeed training workers and vLLM rollout workers across nodes, and coordinates weight synchronization between them. In sync mode, each epoch generates all rollouts, trains on them, syncs weights, and repeats — fully on-policy and easy to reason about. In overlap mode (also called async mode), rollout generation and training run concurrently on separate GPU pools so training GPUs don't sit idle waiting for rollouts. Generation is continuous across epoch boundaries and checkpoint saves — the only pauses are brief drains during weight sync, which runs once at the end of every non-final epoch. A configurable staleness budget bounds how off-policy the replay data can drift. Async mode uses NCCL for weight sync; sync mode pushes weights directly through GPU memory and falls back to disk-based reload if that fails. SFT and DPO are simpler because they only require a single model and no rollout workers, so they run directly on DeepSpeed without Ray. All paradigms support full fine-tuning and LoRA, and plug into mixed-dataset sampling, experiment tracking, and standalone evaluation without changing the overall workflow.
The repository is organized so that algorithmic changes usually stay local:
algs/ — Algorithm and optimization logic. Each algorithm (PPO, GRPO, DAPO, CISPO, P3O, DPO, SFT) has its own module with a README documenting the math and pseudocode.rollouts/ — Rollout generation, vLLM engine wrappers, weight sync, and replay buffer.rewards/ — Pluggable reward functions (GSM8K, math verification, and custom).data_feeds/ — Data loading, sampling, and mixed-dataset support.data_prep/ — Dataset preparation scripts.configs/ — YAML configs for RL, SFT, DPO, and evaluation, with full parameter reference.unit_tests/ — Unit and integration tests.Installation & Setup — Configure your environment and dependencies.
Quickstart & How-To — Learn how to launch jobs and run experiments.
Experiments — Reference experiment results and the canonical example configs used to reproduce them.
Configuration Reference — Full parameter guide for RL, SFT, DPO, and evaluation configs.
How the Async Engine Works — An illustrated guide to the overlapped (async) RL engine.
Troubleshooting — Diagnose and fix common issues.
Contributions are welcome! Please see our Contributing Guidelines for details on how to get involved.
Check out the FAQ for common questions and answers.
Special thanks to the Open-Instruct and PipelineRL projects, both of which provided valuable resources for this work and the broader community.
If you use FeynRL in your work, please cite the following:
@misc{FeynRL,
author = {Rasool Fakoor and Murdock Aubry and FeynRL Contributors},
title = {FeynRL: A Modular and Scalable LLM Post-Training Framework},
year = {2026},
howpublished = {\url{https://github.com/FeynRL-project/FeynRL}},
note = {GitHub repository. Corresponding author: Rasool Fakoor}
}
@misc{fakoor2026trustbatchonoffpolicy,
title = {Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training},
author = {Rasool Fakoor and Murdock Aubry and Nicholas Stranges and Alexander J. Smola},
year = {2026},
eprint = {2605.12380},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2605.12380}
}
141 followers · starred Mar 2026
80 followers · starred Mar 2026
169 followers · starred Mar 2026