deepgenteam/DeepGen-1.0

Model

💡 DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing

180

31 commits

1 linked in READMEs

updated Mar 2, 2026

See the code

README

💡 DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing

DeepGen 1.0 Paper on arXiv Github Github

DeepGen 1.0 is a lightweight unified multimodal model with only 5B parameters (3B VLM + 2B DiT). It integrates five core capabilities—general image generation, general image editing, reasoning image generation, reasoning image editing, and text rendering—within a single model. Across multiple authoritative benchmarks, DeepGen 1.0 is competitive with competitive with or surpassing the state-of-the-art unified multimodal models that are 3× to 16× larger, achieving comprehensive performance, demonstrating that massive scaling is not the sole path to high-performance multimodal generation.

🧠 Method

Our core observation is that a lightweight model, when empowered by synergistic architecture design and data-centric training strategies, can achieve comprehensive capabilities competitive with or even surpassing much larger counterparts. To overcome the limitations of lightweight models in semantic understanding and fine-grained control, we introduce Stacked Channel Bridging (SCB), a deep alignment framework that extracts hierarchical features from multiple VLM layers and fuses them with learnable ``think tokens'' to provide the generative backbone with structured, reasoning-rich guidance. We further design a data-centric training strategy spanning three progressive stages: (1) Alignment Pre-training on large-scale image-text pairs and editing triplets to synchronize VLM and DiT representations, (2) Joint Supervised Fine-tuning on a high-quality mixture of generation, editing, and reasoning tasks to foster omni-capabilities, and (3) Reinforcement Learning with MR-GRPO, which leverages a mixture of reward functions and supervision signals, resulting in substantial gains in generation quality and alignment with human preferences, while maintaining stable training progress and avoiding visual artifacts.

📊 Benchmarks

1. General Image Generation

ModelParamsGeneval ↑DPGBench ↑UniGenBench ↑
OmniGen23B + 4B0.8083.5763.09
BAGEL14B0.8285.1061.53
X-Omni7B + 12B0.8387.65🥉53.77
Lumina-DiMOO8B0.88🥇86.0471.12
Hunyuan-Image-3.080B0.7286.10—
Qwen-Image7B + 20B0.87 🥈88.32 🥇78.81 🥇
LongCat-Image7B + 6B0.87 🥈86.80—
Z-Image-Turbo4B + 6B0.8485.1571.40
GLM-Image9B + 7B—84.78—
DeepGen 1.0 (SFT)3B + 2B0.86 🥉87.0574.18 🥉
DeepGen 1.0 (RL)3B + 2B0.87 🥈87.90 🥈75.74 🥈

2. General Image Editing

ModelParamsGEdit-EN ↑ImgEdit ↑
BAGEL14B6.523.20
Qwen-Image-Edit [2509]7B + 20B7.54 🥈4.35 🥈
LongCat-Image-Edit7B + 6B7.60 🥇4.50 🥇
Mammoth28B + 3B + 2B6.604.06
DeepGen 1.0 (SFT)3B + 2B7.124.09
DeepGen 1.0 (RL)3B + 2B7.17 🥉4.14 🥉

3. Reasoning Image Generation

ModelParamsWISE ↑T2I-CoREBench ↑
OmniGen23B + 4B0.4736.1
BAGEL14B0.70 🥉41.1
Hunyuan-Image-3.080B0.5746.0
Qwen-Image7B + 20B0.6246.3 🥉
LongCat-Image7B + 6B0.6552.2 🥇
Z-Image-Turbo4B + 6B-43.7
DeepGen 1.0 (SFT)3B + 2B0.72 🥈45.7
DeepGen 1.0 (RL)3B + 2B0.73 🥇46.5 🥈

4. Reasoning Image Editing

ModelParamsRISE ↑UniREditBench ↑
OmniGen23B + 4B-43.4
BAGEL14B11.9 🥈51.0
Qwen-Image-Edit [2509]7B + 20B8.956.5 🥉
DeepGen 1.0 (SFT)3B + 2B13.3 🥇77.5 🥇
DeepGen 1.0 (RL)3B + 2B10.8 🥉75.7 🥈

🎨 Quantitative results

🛠️ Usage

Merge ZIP Files

To use the DeepGen checkpoints, please merge the sharded model files first. We release Pre-traning, Supervised Fine-Tuning and Reinforcement Learning checkpoints.

# Merge zip
cat DeepGen_CKPT.zip.part-* > DeepGen_CKPT.zip
# Unzip DeepGen checkpoints 
unzip DeepGen_CKPT.zip
checkpoints/
├── DeepGen_CKPT
    ├──Pretrain├──iter_200000.pth
    ├── SFT├──iter_400000.pth
    ├──RL├──MR-GDPO_final.pt 

if you want only final model state please use model.pt directly , it is same as MR-GDPO_final.pt

the Pretrain├──iter_200000.pth and SFT├──iter_400000.pth can be loaded for continuous training

⭐ Citation

@article{wang2026deepgen,
  title={DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing},
  author={Wang, Dianyi and Li, Ruihang and Han, Feng and Ma, Chaofan and Song, Wei and Wang, Siyuan and Wang, Yibin and Xin, Yi and Liu, Hongjian and Zhang, Zhixiong and others},
  journal={arXiv preprint arXiv:2602.12205},
  year={2026}
}
text-to-image

deepgenteam/DeepGen-1.0

Model

💡 DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing

180

31 commits

1 linked in READMEs

updated Mar 2, 2026

See the code

README

💡 DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing

DeepGen 1.0 Paper on arXiv Github Github

DeepGen 1.0 is a lightweight unified multimodal model with only 5B parameters (3B VLM + 2B DiT). It integrates five core capabilities—general image generation, general image editing, reasoning image generation, reasoning image editing, and text rendering—within a single model. Across multiple authoritative benchmarks, DeepGen 1.0 is competitive with competitive with or surpassing the state-of-the-art unified multimodal models that are 3× to 16× larger, achieving comprehensive performance, demonstrating that massive scaling is not the sole path to high-performance multimodal generation.

🧠 Method

Our core observation is that a lightweight model, when empowered by synergistic architecture design and data-centric training strategies, can achieve comprehensive capabilities competitive with or even surpassing much larger counterparts. To overcome the limitations of lightweight models in semantic understanding and fine-grained control, we introduce Stacked Channel Bridging (SCB), a deep alignment framework that extracts hierarchical features from multiple VLM layers and fuses them with learnable ``think tokens'' to provide the generative backbone with structured, reasoning-rich guidance. We further design a data-centric training strategy spanning three progressive stages: (1) Alignment Pre-training on large-scale image-text pairs and editing triplets to synchronize VLM and DiT representations, (2) Joint Supervised Fine-tuning on a high-quality mixture of generation, editing, and reasoning tasks to foster omni-capabilities, and (3) Reinforcement Learning with MR-GRPO, which leverages a mixture of reward functions and supervision signals, resulting in substantial gains in generation quality and alignment with human preferences, while maintaining stable training progress and avoiding visual artifacts.

📊 Benchmarks

1. General Image Generation

ModelParamsGeneval ↑DPGBench ↑UniGenBench ↑
OmniGen23B + 4B0.8083.5763.09
BAGEL14B0.8285.1061.53
X-Omni7B + 12B0.8387.65🥉53.77
Lumina-DiMOO8B0.88🥇86.0471.12
Hunyuan-Image-3.080B0.7286.10—
Qwen-Image7B + 20B0.87 🥈88.32 🥇78.81 🥇
LongCat-Image7B + 6B0.87 🥈86.80—
Z-Image-Turbo4B + 6B0.8485.1571.40
GLM-Image9B + 7B—84.78—
DeepGen 1.0 (SFT)3B + 2B0.86 🥉87.0574.18 🥉
DeepGen 1.0 (RL)3B + 2B0.87 🥈87.90 🥈75.74 🥈

2. General Image Editing

ModelParamsGEdit-EN ↑ImgEdit ↑
BAGEL14B6.523.20
Qwen-Image-Edit [2509]7B + 20B7.54 🥈4.35 🥈
LongCat-Image-Edit7B + 6B7.60 🥇4.50 🥇
Mammoth28B + 3B + 2B6.604.06
DeepGen 1.0 (SFT)3B + 2B7.124.09
DeepGen 1.0 (RL)3B + 2B7.17 🥉4.14 🥉

3. Reasoning Image Generation

ModelParamsWISE ↑T2I-CoREBench ↑
OmniGen23B + 4B0.4736.1
BAGEL14B0.70 🥉41.1
Hunyuan-Image-3.080B0.5746.0
Qwen-Image7B + 20B0.6246.3 🥉
LongCat-Image7B + 6B0.6552.2 🥇
Z-Image-Turbo4B + 6B-43.7
DeepGen 1.0 (SFT)3B + 2B0.72 🥈45.7
DeepGen 1.0 (RL)3B + 2B0.73 🥇46.5 🥈

4. Reasoning Image Editing

ModelParamsRISE ↑UniREditBench ↑
OmniGen23B + 4B-43.4
BAGEL14B11.9 🥈51.0
Qwen-Image-Edit [2509]7B + 20B8.956.5 🥉
DeepGen 1.0 (SFT)3B + 2B13.3 🥇77.5 🥇
DeepGen 1.0 (RL)3B + 2B10.8 🥉75.7 🥈

🎨 Quantitative results

🛠️ Usage

Merge ZIP Files

To use the DeepGen checkpoints, please merge the sharded model files first. We release Pre-traning, Supervised Fine-Tuning and Reinforcement Learning checkpoints.

# Merge zip
cat DeepGen_CKPT.zip.part-* > DeepGen_CKPT.zip
# Unzip DeepGen checkpoints 
unzip DeepGen_CKPT.zip
checkpoints/
├── DeepGen_CKPT
    ├──Pretrain├──iter_200000.pth
    ├── SFT├──iter_400000.pth
    ├──RL├──MR-GDPO_final.pt 

if you want only final model state please use model.pt directly , it is same as MR-GDPO_final.pt

the Pretrain├──iter_200000.pth and SFT├──iter_400000.pth can be loaded for continuous training

⭐ Citation

@article{wang2026deepgen,
  title={DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing},
  author={Wang, Dianyi and Li, Ruihang and Han, Feng and Ma, Chaofan and Song, Wei and Wang, Siyuan and Wang, Yibin and Xin, Yi and Liu, Hongjian and Zhang, Zhixiong and others},
  journal={arXiv preprint arXiv:2602.12205},
  year={2026}
}
text-to-image