This is the official repo for AlphaGRPO: Self-Reflective Multimodal Generation via Decompositional Verifiable Reward.
TL;DR: AlphaGRPO enables multimodal generation RL training across text and image generation for AR-Diffusion-native unified multimodal models, such as BAGEL. It supports training on reasoning text-to-image generation and self-reflective refinement.
This codebase flexibly supports different RL methods for image and text generation: FlowGRPO, DiffusionNFT, and AWM for images; GRPO for text.
docs/SPECTRAREWARD.md.AlphaGRPO trains unified multimodal models with a decompositional reward design: complex prompts are broken into verifiable semantic and quality checks, and text/image generation steps can be optimized with separate RL algorithms.
The framework is step-list driven, so each task can explicitly define its own text and image generation sequence. This supports simple reasoning text-to-image generation as well as multi-stage self-reflective refinement without hardcoding a fixed generation pattern.
uv venv .venv
source .venv/bin/activate
uv pip install -r requirements.txt
The bundled alphagrpo20k JSONL dataset files are tracked with Git LFS. Install Git LFS before cloning, or pull LFS files after cloning:
git lfs install
git lfs pull
If the dataset looks unusually small, check that alpha_grpo/dataset/alphagrpo20k/train.jsonl is not a Git LFS pointer file.
Download the base BAGEL model weights before training.
Then update huggingface_models_root in config/bagel.py to point to /path/to/huggingface_models/. The bundled DVReward dataset is under alpha_grpo/dataset/alphagrpo20k/; to use your own data, see docs/DVREWARD.md.
Deploy the reward server in a separate terminal on each reward-server node:
ip=YOUR_IP
port=YOUR_PORT
bash scripts/serve_reward_model.sh $ip $port
Single-node training example:
ip=YOUR_REWARD_SERVER_IP
port=YOUR_REWARD_SERVER_PORT
export PYTHONPATH=$PYTHONPATH:$(pwd)/Bagel/
# IPv6:
export OPENAI_API_URL=http://[${ip}]:${port}/v1
# IPv4:
# export OPENAI_API_URL=http://${ip}:${port}/v1
export OPENAI_MODEL_NAME=Qwen/Qwen3-VL-30B-A3B-Instruct
task=alphagrpo_reflect # alphagrpo_t2iThink is also available
torchrun --nnodes=1 --node_rank=0 --nproc_per_node=8 \
alpha_grpo/train.py --config config/bagel.py:${task}
Multi-node training example:
ip=YOUR_REWARD_SERVER_IP
port=YOUR_REWARD_SERVER_PORT
export PYTHONPATH=$PYTHONPATH:$(pwd)/Bagel/
export OPENAI_API_URL=http://[${ip}]:${port}/v1 # IPv6 reward-server address
export OPENAI_MODEL_NAME=Qwen/Qwen3-VL-30B-A3B-Instruct
# We use 8 nodes × 8 A100: 1 GPU/node for serving reward model, 7 for training.
# Scale gradient_accumulation_steps inversely with GPU count to maintain batch size.
task=alphagrpo_reflect # for self-reflective refinement task
# task=alphagrpo_t2iThink # for reasoning text-to-image task
torchrun --nnodes=$NUM_NODES --node_rank=$RANK --nproc_per_node=7 \
alpha_grpo/train.py --config config/bagel.py:${task}
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation
SpectraReward turns a frozen pretrained MLLM into a training-free reward model for text-to-image RL. It scores how well the prompt can be read back from the generated image, using the mean image-conditioned prompt log-likelihood as the reward. Self-SpectraReward is the unified-model special case, where BAGEL's own understanding branch scores its generation branch, with no external reward model. See docs/SPECTRAREWARD.md for the full method, training setup, backbone support, and results.
External SpectraReward. Run the reward MLLM on a remote server so it does not share GPU memory with BAGEL:
# On the reward-server node(s): serve the reward MLLM (8-way data-parallel here).
bash scripts/serve_spectrareward.sh Qwen/Qwen3-VL-30B-A3B-Instruct 0.0.0.0 18090 8
# On the training side: point to the server and launch as usual.
export PYTHONPATH=$PYTHONPATH:$(pwd)/Bagel/
export SPECTRAREWARD_MODEL_ID=Qwen/Qwen3-VL-30B-A3B-Instruct
export SPECTRAREWARD_URL=http://<reward_server_ip>:18090 # IPv6: http://[${ip}]:18090
torchrun --nnodes=$NUM_NODES --node_rank=$RANK --nproc_per_node=8 \
alpha_grpo/train.py --config config/bagel.py:spectrareward_t2i_awm
Self-SpectraReward. No external reward model is needed; BAGEL scores itself:
export PYTHONPATH=$PYTHONPATH:$(pwd)/Bagel/
torchrun --nnodes=$NUM_NODES --node_rank=$RANK --nproc_per_node=8 \
alpha_grpo/train.py --config config/bagel.py:self_spectrareward_t2i_awm
For the in-process fallback, compatible MLLM backbones, configs, and experimental results, see docs/SPECTRAREWARD.md.
Install eval dependencies:
bash Bagel/eval_install.sh
Run benchmarks:
cd Bagel
# Evaluate on downstream benchmarks (GenEval, TIIF, WISE, DPG, GEdit)
task=geneval # dpg, tiif, wise, gedit
# Evaluate multiple LoRA checkpoints. Configure `jobname` and `checkpoint_list` in the corresponding script before running.
bash scripts/eval/run_${task}_multickpt.sh
# Evaluate one LoRA checkpoint directly.
export BAGEL_LORA_PATH=/path/to/checkpoint/hf_model/
bash scripts/eval/run_${task}.sh
# Conduct self-reflective refinement on downstream tasks.
task=geneval # dpg, tiif
bash scripts/eval/run_${task}_multickpt_reflect.sh
Text-to-image benchmarks. RT2I denotes the AlphaGRPO variant trained on the reasoning text-to-image task. Inf. SRR denotes inference-time self-reflective refinement.
GEdit-Bench-EN. Editing transfer performance: AlphaGRPO improves GEdit scores without training on editing tasks.
To train AlphaGRPO on a custom prompt set with our Decompositional Verifiable Reward (DVReward), see DVREWARD.md for the prompt-decomposition pipeline and how to wire the resulting dataset into training.
Text and image steps support independent RL algorithms (e.g., GRPO for text + AWM or DiffusionNFT for image). New tasks, algorithms, and rewards can be added modularly. See docs/extending.md for details.
This project builds upon BAGEL and Flow-GRPO. We thank the authors for their excellent work.
If you find AlphaGRPO or SpectraReward useful to your research, please consider citing:
@inproceedings{huang2026alphagrpo,
title={AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in Unified Multimodal Models via Decompositional Verifiable Reward},
author={Huang, Runhui and Wu, Jie and Yang, Rui and Liu, Zhe and Zhao, Hengshuang},
booktitle={International Conference on Machine Learning (ICML)},
year={2026}
}
@misc{huang2026readitback,
title={Read It Back: Pretrained {MLLMs} Are Zero-Shot Reward Models for Text-to-Image Generation},
author={Huang, Runhui and Zhang, Qihui and Liu, Zhe and Gao, Yu and Wu, Jie and Zhao, Hengshuang},
year={2026},
eprint={2607.11886},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2607.11886}
}
This project is released under the Apache License 2.0.
8 commits
Python
95.1%
Shell
3.9%
This is the official repo for AlphaGRPO: Self-Reflective Multimodal Generation via Decompositional Verifiable Reward.
TL;DR: AlphaGRPO enables multimodal generation RL training across text and image generation for AR-Diffusion-native unified multimodal models, such as BAGEL. It supports training on reasoning text-to-image generation and self-reflective refinement.
This codebase flexibly supports different RL methods for image and text generation: FlowGRPO, DiffusionNFT, and AWM for images; GRPO for text.
docs/SPECTRAREWARD.md.AlphaGRPO trains unified multimodal models with a decompositional reward design: complex prompts are broken into verifiable semantic and quality checks, and text/image generation steps can be optimized with separate RL algorithms.
The framework is step-list driven, so each task can explicitly define its own text and image generation sequence. This supports simple reasoning text-to-image generation as well as multi-stage self-reflective refinement without hardcoding a fixed generation pattern.
uv venv .venv
source .venv/bin/activate
uv pip install -r requirements.txt
The bundled alphagrpo20k JSONL dataset files are tracked with Git LFS. Install Git LFS before cloning, or pull LFS files after cloning:
git lfs install
git lfs pull
If the dataset looks unusually small, check that alpha_grpo/dataset/alphagrpo20k/train.jsonl is not a Git LFS pointer file.
Download the base BAGEL model weights before training.
Then update huggingface_models_root in config/bagel.py to point to /path/to/huggingface_models/. The bundled DVReward dataset is under alpha_grpo/dataset/alphagrpo20k/; to use your own data, see docs/DVREWARD.md.
Deploy the reward server in a separate terminal on each reward-server node:
ip=YOUR_IP
port=YOUR_PORT
bash scripts/serve_reward_model.sh $ip $port
Single-node training example:
ip=YOUR_REWARD_SERVER_IP
port=YOUR_REWARD_SERVER_PORT
export PYTHONPATH=$PYTHONPATH:$(pwd)/Bagel/
# IPv6:
export OPENAI_API_URL=http://[${ip}]:${port}/v1
# IPv4:
# export OPENAI_API_URL=http://${ip}:${port}/v1
export OPENAI_MODEL_NAME=Qwen/Qwen3-VL-30B-A3B-Instruct
task=alphagrpo_reflect # alphagrpo_t2iThink is also available
torchrun --nnodes=1 --node_rank=0 --nproc_per_node=8 \
alpha_grpo/train.py --config config/bagel.py:${task}
Multi-node training example:
ip=YOUR_REWARD_SERVER_IP
port=YOUR_REWARD_SERVER_PORT
export PYTHONPATH=$PYTHONPATH:$(pwd)/Bagel/
export OPENAI_API_URL=http://[${ip}]:${port}/v1 # IPv6 reward-server address
export OPENAI_MODEL_NAME=Qwen/Qwen3-VL-30B-A3B-Instruct
# We use 8 nodes × 8 A100: 1 GPU/node for serving reward model, 7 for training.
# Scale gradient_accumulation_steps inversely with GPU count to maintain batch size.
task=alphagrpo_reflect # for self-reflective refinement task
# task=alphagrpo_t2iThink # for reasoning text-to-image task
torchrun --nnodes=$NUM_NODES --node_rank=$RANK --nproc_per_node=7 \
alpha_grpo/train.py --config config/bagel.py:${task}
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation
SpectraReward turns a frozen pretrained MLLM into a training-free reward model for text-to-image RL. It scores how well the prompt can be read back from the generated image, using the mean image-conditioned prompt log-likelihood as the reward. Self-SpectraReward is the unified-model special case, where BAGEL's own understanding branch scores its generation branch, with no external reward model. See docs/SPECTRAREWARD.md for the full method, training setup, backbone support, and results.
External SpectraReward. Run the reward MLLM on a remote server so it does not share GPU memory with BAGEL:
# On the reward-server node(s): serve the reward MLLM (8-way data-parallel here).
bash scripts/serve_spectrareward.sh Qwen/Qwen3-VL-30B-A3B-Instruct 0.0.0.0 18090 8
# On the training side: point to the server and launch as usual.
export PYTHONPATH=$PYTHONPATH:$(pwd)/Bagel/
export SPECTRAREWARD_MODEL_ID=Qwen/Qwen3-VL-30B-A3B-Instruct
export SPECTRAREWARD_URL=http://<reward_server_ip>:18090 # IPv6: http://[${ip}]:18090
torchrun --nnodes=$NUM_NODES --node_rank=$RANK --nproc_per_node=8 \
alpha_grpo/train.py --config config/bagel.py:spectrareward_t2i_awm
Self-SpectraReward. No external reward model is needed; BAGEL scores itself:
export PYTHONPATH=$PYTHONPATH:$(pwd)/Bagel/
torchrun --nnodes=$NUM_NODES --node_rank=$RANK --nproc_per_node=8 \
alpha_grpo/train.py --config config/bagel.py:self_spectrareward_t2i_awm
For the in-process fallback, compatible MLLM backbones, configs, and experimental results, see docs/SPECTRAREWARD.md.
Install eval dependencies:
bash Bagel/eval_install.sh
Run benchmarks:
cd Bagel
# Evaluate on downstream benchmarks (GenEval, TIIF, WISE, DPG, GEdit)
task=geneval # dpg, tiif, wise, gedit
# Evaluate multiple LoRA checkpoints. Configure `jobname` and `checkpoint_list` in the corresponding script before running.
bash scripts/eval/run_${task}_multickpt.sh
# Evaluate one LoRA checkpoint directly.
export BAGEL_LORA_PATH=/path/to/checkpoint/hf_model/
bash scripts/eval/run_${task}.sh
# Conduct self-reflective refinement on downstream tasks.
task=geneval # dpg, tiif
bash scripts/eval/run_${task}_multickpt_reflect.sh
Text-to-image benchmarks. RT2I denotes the AlphaGRPO variant trained on the reasoning text-to-image task. Inf. SRR denotes inference-time self-reflective refinement.
GEdit-Bench-EN. Editing transfer performance: AlphaGRPO improves GEdit scores without training on editing tasks.
To train AlphaGRPO on a custom prompt set with our Decompositional Verifiable Reward (DVReward), see DVREWARD.md for the prompt-decomposition pipeline and how to wire the resulting dataset into training.
Text and image steps support independent RL algorithms (e.g., GRPO for text + AWM or DiffusionNFT for image). New tasks, algorithms, and rewards can be added modularly. See docs/extending.md for details.
This project builds upon BAGEL and Flow-GRPO. We thank the authors for their excellent work.
If you find AlphaGRPO or SpectraReward useful to your research, please consider citing:
@inproceedings{huang2026alphagrpo,
title={AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in Unified Multimodal Models via Decompositional Verifiable Reward},
author={Huang, Runhui and Wu, Jie and Yang, Rui and Liu, Zhe and Zhao, Hengshuang},
booktitle={International Conference on Machine Learning (ICML)},
year={2026}
}
@misc{huang2026readitback,
title={Read It Back: Pretrained {MLLMs} Are Zero-Shot Reward Models for Text-to-Image Generation},
author={Huang, Runhui and Zhang, Qihui and Liu, Zhe and Gao, Yu and Wu, Jie and Zhao, Hengshuang},
year={2026},
eprint={2607.11886},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2607.11886}
}
This project is released under the Apache License 2.0.
8 commits
Python
95.1%
Shell
3.9%