[ICLR 26] TempFlow-GRPO (Temporal Flow GRPO), a principled GRPO framework that captures and exploits the temporal structure inherent in flow-based generation.
Python
456
75 commits
updated Nov 24, 2025
TempFlow-GRPO (Temporal Flow GRPO), a principled GRPO framework that captures and exploits the temporal structure inherent in flow-based generation.
TempFlow-GRPO (Temporal Flow GRPO), a principled GRPO framework that captures and exploits the temporal structure inherent in flow-based generation. TempFlow-GRPO introduces two key innovations: (i) a trajectory branching mechanism that provides process rewards by concentrating stochasticity at designated branching points, enabling precise credit assignment without requiring specialized intermediate reward models; and (ii) a noise-aware weighting scheme that modulates policy optimization according to the intrinsic exploration potential of each timestep, prioritizing learning during high-impact early stages while ensuring stable refinement in later phases. These innovations endow the model with temporally-aware optimization that respects the underlying generative dynamics, leading to state-of-the-art performance in human preference alignment and standard text-to-image benchmark.
Welcome Ideas and Contributions. Stay tuned!
We have presented an improved Flow-GRPO method, TempFlow-GRPO. We will release our code recently!π₯π₯π₯
To support research and the open-source community, we will release the entire projectβincluding datasets, training pipelines, and model weights. Our code is based on Flow-GRPO!. Thank you for your patience and continued support! π
Note that we use branch=4, per branch exploration=6. You can modify them in our code. We will release a neat code verision in next few days.
Prompt Group: notes the seed group
Batch Group: global_std=True
# Flow-GRPO
bash scripts/multi_node/main.sh
# TempFlow-GRPO
bash scripts/multi_node/train_sd3_pr.sh
# Flow-GRPO
bash scripts/multi_node/train_flux.sh
# TempFlow-GRPO
bash scripts/multi_node/train_flux_pr.sh
# Flow-GRPO
bash scripts/multi_node/train_qwenimage.sh
# TempFlow-GRPO
bash scripts/multi_node/train_qwenimage_pr.sh
Flow-GRPO: The first method integrating online reinforcement learning (RL) into flow matching models.
Work was done at WeChat Vision.
75 commits
Python
99.6%
[ICLR 26] TempFlow-GRPO (Temporal Flow GRPO), a principled GRPO framework that captures and exploits the temporal structure inherent in flow-based generation.
Python
456
75 commits
updated Nov 24, 2025
TempFlow-GRPO (Temporal Flow GRPO), a principled GRPO framework that captures and exploits the temporal structure inherent in flow-based generation.
TempFlow-GRPO (Temporal Flow GRPO), a principled GRPO framework that captures and exploits the temporal structure inherent in flow-based generation. TempFlow-GRPO introduces two key innovations: (i) a trajectory branching mechanism that provides process rewards by concentrating stochasticity at designated branching points, enabling precise credit assignment without requiring specialized intermediate reward models; and (ii) a noise-aware weighting scheme that modulates policy optimization according to the intrinsic exploration potential of each timestep, prioritizing learning during high-impact early stages while ensuring stable refinement in later phases. These innovations endow the model with temporally-aware optimization that respects the underlying generative dynamics, leading to state-of-the-art performance in human preference alignment and standard text-to-image benchmark.
Welcome Ideas and Contributions. Stay tuned!
We have presented an improved Flow-GRPO method, TempFlow-GRPO. We will release our code recently!π₯π₯π₯
To support research and the open-source community, we will release the entire projectβincluding datasets, training pipelines, and model weights. Our code is based on Flow-GRPO!. Thank you for your patience and continued support! π
Note that we use branch=4, per branch exploration=6. You can modify them in our code. We will release a neat code verision in next few days.
Prompt Group: notes the seed group
Batch Group: global_std=True
# Flow-GRPO
bash scripts/multi_node/main.sh
# TempFlow-GRPO
bash scripts/multi_node/train_sd3_pr.sh
# Flow-GRPO
bash scripts/multi_node/train_flux.sh
# TempFlow-GRPO
bash scripts/multi_node/train_flux_pr.sh
# Flow-GRPO
bash scripts/multi_node/train_qwenimage.sh
# TempFlow-GRPO
bash scripts/multi_node/train_qwenimage_pr.sh
Flow-GRPO: The first method integrating online reinforcement learning (RL) into flow matching models.
Work was done at WeChat Vision.
75 commits
Python
99.6%