Angusliuuu/Awesome-Controllable-Generative-Models-Papers

A curated list of recent papers (2023–2025) on controllable generative models, covering diffusion-based architectures with fine-grained control, attention interpretation, spectral manipulation, and structure-preserving image editing. Ideal for researchers and developers exploring controllable synthesis.

42

10 commits

updated Apr 22, 2026

See the code

README

🔥 Awesome Controllable Generative Models

A curated and continuously updated collection of recent (2023–2025) research papers on controllable generative models, with a special focus on both UNet-based diffusion models and Transformer-based diffusion architectures.

This list emphasizes core advances in:

  • 🧭 Control mechanisms – including condition injection, adapters, multi-modal control
  • 👁️ Attention interpretation – revealing what diffusion models focus on
  • 🎛️ Frequency-based control – using spectral domain knowledge to guide generation
  • 🔁 Alignment & knowledge transfer – enabling more coherent, faithful, and data-efficient synthesis
  • 🧑‍🎨 Image-to-image (I2I) editing – flexible, structure-preserving transformation across domains

💡 Our goal is not only to track the state-of-the-art in controllable generation, but also to offer a well-organized knowledge map for newcomers and researchers building on top of diffusion models.


🧭 Control Mechanism

PaperVenueLinks
OminiControl: Minimal and Universal Control for Diffusion TransformerICCV 2025 (Highlight)Paper | Code
Rectified Diffusion Guidance for Conditional GenerationCVPR 2025Paper | Code
FlexControl: Computation-Aware Conditional Control with Differentiable RouterICML 2025 (Poster)Paper | Code
CTRL-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion ModelICLR 2025 (Oral)Paper | Code
Ctrl-U: Robust Conditional Image Generation via Uncertainty-aware Reward ModelingICLR 2025Paper | Code
ConceptCtrl: Concept Control of Zero-shot Personalized Image GenerationarXiv 2025Paper | Code
Is Noise Conditioning Necessary for Denoising Generative Models?arXiv 2025Paper | Code
Ctrl‑X: Controlling Structure and Appearance Without GuidanceNeurIPS 2024Paper | Code
ControlNet++: Improving Conditional Controls with Efficient Consistency FeedbackECCV 2024Paper | Code
ControlNet-XS: Rethinking the Control of Text-to-Image Diffusion Models as Feedback-Control SystemsECCV 2024Paper | Code
SmartControl: Enhancing ControlNet for Handling Rough Visual ConditionsECCV 2024 (Poster)Paper | Code
PIXART-δ: Fast and Controllable Image Generation with Latent Consistency ModelsICML 2024 WorkshopPaper | Code
Layout-to-Image Generation with Localized Descriptions using ControlNet with Cross-Attention GuidanceWACV 2024Paper | Code
Controllable Generation with Text-to-Image Diffusion Models: A SurveyT-PAMI 2024Paper | Code
Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion ModelsarXiv 2024Paper | Code
Cocktail : Mixing Multi-Modality Controls for Text-Conditional Image GenerationNeurIPS 2023Paper | Code
T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image DiffusionNeurIPS 2023Paper | Code
Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion ModelsNeurIPS 2023Paper | Code
Adding Conditional Control to Text-to-Image Diffusion ModelsICCV 2023paper | Code
GLIGEN: Open-Set Grounded Text-to-Image GenerationCVPR 2023Paper | Code
MultiDiffusion: Fusing Diffusion Paths for Controlled Image GenerationCVPR 2023Paper | Code
UniControl: A Unified Diffusion Model for Controllable Visual Generation In the WildCVPR 2023Paper | Code
Composer: Creative and Controllable Image Synthesis with Composable ConditionsICML 2023Paper | Code
More Control for Free! Image Synthesis with Semantic Diffusion GuidanceWACV 2023Paper | Code
IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion ModelsarXiv 2023Paper | Code
Sketch-Guided Text-to-Image Diffusion ModelsarXiv 2022Paper | Code
Generative Modeling by Estimating Gradients of the Data DistributionNeurIPS 2019Paper | Code

🧠 These papers push the boundary of how we guide generation, whether through minimal prompts, learned adapters, or uncertainty-aware mechanisms.


👁️ Attention & Interpretability

PaperVenueLinks
ConceptAttention: Diffusion Transformers Learn Highly Interpretable FeaturesICML 2025 (Oral)Paper | Code
ToMA: Token Merge with Attention for Diffusion ModelsICML 2025 (Poster)Paper | Code
Attention in Diffusion Model: A SurveyarXiv 2025Paper | Code
Attention Distillation: A Unified Approach to Visual Characteristics TransferCVPR 2025Paper | Code
What the DAAM: Interpreting Stable Diffusion Using Cross AttentionACL 2024 (Oral)Paper | Code

🔬 Interpretability is not just analysis — it's a step toward transparent and editable generative pipelines.


🎛️ Frequency Domain Control

PaperVenueLinks
DiffFNO: Diffusion Fourier Neural Operator for Arbitrary-Scale Super-ResolutionCVPR 2025 (Oral)Paper | Code
PTDiffusion: Free Lunch for Generating Optical Illusion Hidden Pictures with Phase-Transferred DiffusionCVPR 2025 (Poster)Paper | Code
Diffusion-based Adversarial Purification from the Perspective of the Frequency DomainICML 2025 (Spotlight Poster)Paper | Code
Frequency Autoregressive Image Generation with Continuous TokensarXiv 2025Paper | Code
FreeDiff: Progressive Frequency Truncation for Image Editing with Diffusion ModelsECCV 2024 (Poster)Paper | Code
FreeU: Free Lunch in Diffusion U-NetCVPR 2024Paper | Code
Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image TranslationAAAI 2024Paper | Code
ResDiff: Combining CNN and Diffusion Model for Image Super-ResolutionAAAI 2023Paper | Code

📡 Spectral and signal-level control provides low-level but powerful levers for generative consistency, resolution, and robustness.


🔁 Alignment & Knowledge Transfer

PaperVenueLinks
When Model Knowledge meets Diffusion: Data-free Synthesis with Domain-Class AlignmentICML 2025Paper | Code

🧬 These works align discrete symbolic knowledge with continuous generative priors, aiming for controllability in low-data or zero-shot regimes.


You may also consider including some notable image-to-image (I2I) editing methods in your collection.

Image Editing

PaperVenueLinks
In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion TransformerNeurIPS 2025Paper | Code
AnyEdit: Mastering Unified High-Quality Image Editing for Any IdeaCVPR 2025 (Oral)Paper | Code
Stable Flow: Vital Layers for Training-Free Image EditingCVPR 2025Paper | Code
UniReal: Universal Image Generation and Editing via Learning Real-world DynamicsCVPR 2025Paper | Code
Taming Rectified Flow for Inversion and EditingICML 2025Paper | Code
Semantic Image Inversion and Editing using Rectified Stochastic Differential EquationsICLR 2025Paper | Code
Diffusion Model-Based Image Editing: A SurveyTPAMI 2025Paper | Code
ChronoEdit: Towards Temporal Reasoning for Image Editing and World SimulationarXiv 2025Paper | Code
SAEdit: Token-level control for continuous image editing via Sparse AutoEncoderarXiv 2025Paper | Code
UltraEdit: Instruction-based Fine-Grained Image Editing at ScaleNeurIPS 2024Paper | Code
SmartEdit: Exploring complex instruction-based image editing with multimodal large language modelsCVPR 2024Paper | Code
Direct Inversion: Boosting Diffusion-based Editing with 3 Lines of CodeICLR 2024Paper | Code
MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image EditingNeurIPS 2023Paper | Code
NULL-text inversion for editing real images using guided diffusion modelsCVPR 2023Paper | Code
Imagic: Text-based real image editing with diffusion modelsCVPR 2023Paper | Code
InstructPix2Pix: Learning To Follow Image Editing InstructionsCVPR 2023Paper | Code
Plug-and-Play Diffusion Features for Text-Driven Image-to-Image TranslationCVPR 2023 (Poster)Paper | Code
DiffEdit: Diffusion-based semantic image editing with mask guidanceICLR 2023Paper | Code
Prompt-to-Prompt Image Editing with Cross Attention ControlICLR 2023Paper | Code
Zero-shot Image-to-Image Translation (pix2pix-zero)SIGGRAPH 2023Paper | Code
RePaint: Inpainting using Denoising Diffusion Probabilistic ModelsCVPR 2022Paper | Code

✏️ Editing is arguably where controllability matters most — precision, structure preservation, and user intent must all align.


📬 Contribute

💡 Know a paper we missed? Working on a new controllable generation method?
Feel free to submit a pull request or open an issue — contributions are welcome!

attention-mechanism
controlnet
diffusion-models
generative-model
multimodal

Contributors

Angusliuuu

6 commits

Bili-Sakura

2 commits

cursoragent

2 commits

Angusliuuu/Awesome-Controllable-Generative-Models-Papers

A curated list of recent papers (2023–2025) on controllable generative models, covering diffusion-based architectures with fine-grained control, attention interpretation, spectral manipulation, and structure-preserving image editing. Ideal for researchers and developers exploring controllable synthesis.

42

10 commits

updated Apr 22, 2026

See the code

README

🔥 Awesome Controllable Generative Models

A curated and continuously updated collection of recent (2023–2025) research papers on controllable generative models, with a special focus on both UNet-based diffusion models and Transformer-based diffusion architectures.

This list emphasizes core advances in:

  • 🧭 Control mechanisms – including condition injection, adapters, multi-modal control
  • 👁️ Attention interpretation – revealing what diffusion models focus on
  • 🎛️ Frequency-based control – using spectral domain knowledge to guide generation
  • 🔁 Alignment & knowledge transfer – enabling more coherent, faithful, and data-efficient synthesis
  • 🧑‍🎨 Image-to-image (I2I) editing – flexible, structure-preserving transformation across domains

💡 Our goal is not only to track the state-of-the-art in controllable generation, but also to offer a well-organized knowledge map for newcomers and researchers building on top of diffusion models.


🧭 Control Mechanism

PaperVenueLinks
OminiControl: Minimal and Universal Control for Diffusion TransformerICCV 2025 (Highlight)Paper | Code
Rectified Diffusion Guidance for Conditional GenerationCVPR 2025Paper | Code
FlexControl: Computation-Aware Conditional Control with Differentiable RouterICML 2025 (Poster)Paper | Code
CTRL-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion ModelICLR 2025 (Oral)Paper | Code
Ctrl-U: Robust Conditional Image Generation via Uncertainty-aware Reward ModelingICLR 2025Paper | Code
ConceptCtrl: Concept Control of Zero-shot Personalized Image GenerationarXiv 2025Paper | Code
Is Noise Conditioning Necessary for Denoising Generative Models?arXiv 2025Paper | Code
Ctrl‑X: Controlling Structure and Appearance Without GuidanceNeurIPS 2024Paper | Code
ControlNet++: Improving Conditional Controls with Efficient Consistency FeedbackECCV 2024Paper | Code
ControlNet-XS: Rethinking the Control of Text-to-Image Diffusion Models as Feedback-Control SystemsECCV 2024Paper | Code
SmartControl: Enhancing ControlNet for Handling Rough Visual ConditionsECCV 2024 (Poster)Paper | Code
PIXART-δ: Fast and Controllable Image Generation with Latent Consistency ModelsICML 2024 WorkshopPaper | Code
Layout-to-Image Generation with Localized Descriptions using ControlNet with Cross-Attention GuidanceWACV 2024Paper | Code
Controllable Generation with Text-to-Image Diffusion Models: A SurveyT-PAMI 2024Paper | Code
Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion ModelsarXiv 2024Paper | Code
Cocktail : Mixing Multi-Modality Controls for Text-Conditional Image GenerationNeurIPS 2023Paper | Code
T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image DiffusionNeurIPS 2023Paper | Code
Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion ModelsNeurIPS 2023Paper | Code
Adding Conditional Control to Text-to-Image Diffusion ModelsICCV 2023paper | Code
GLIGEN: Open-Set Grounded Text-to-Image GenerationCVPR 2023Paper | Code
MultiDiffusion: Fusing Diffusion Paths for Controlled Image GenerationCVPR 2023Paper | Code
UniControl: A Unified Diffusion Model for Controllable Visual Generation In the WildCVPR 2023Paper | Code
Composer: Creative and Controllable Image Synthesis with Composable ConditionsICML 2023Paper | Code
More Control for Free! Image Synthesis with Semantic Diffusion GuidanceWACV 2023Paper | Code
IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion ModelsarXiv 2023Paper | Code
Sketch-Guided Text-to-Image Diffusion ModelsarXiv 2022Paper | Code
Generative Modeling by Estimating Gradients of the Data DistributionNeurIPS 2019Paper | Code

🧠 These papers push the boundary of how we guide generation, whether through minimal prompts, learned adapters, or uncertainty-aware mechanisms.


👁️ Attention & Interpretability

PaperVenueLinks
ConceptAttention: Diffusion Transformers Learn Highly Interpretable FeaturesICML 2025 (Oral)Paper | Code
ToMA: Token Merge with Attention for Diffusion ModelsICML 2025 (Poster)Paper | Code
Attention in Diffusion Model: A SurveyarXiv 2025Paper | Code
Attention Distillation: A Unified Approach to Visual Characteristics TransferCVPR 2025Paper | Code
What the DAAM: Interpreting Stable Diffusion Using Cross AttentionACL 2024 (Oral)Paper | Code

🔬 Interpretability is not just analysis — it's a step toward transparent and editable generative pipelines.


🎛️ Frequency Domain Control

PaperVenueLinks
DiffFNO: Diffusion Fourier Neural Operator for Arbitrary-Scale Super-ResolutionCVPR 2025 (Oral)Paper | Code
PTDiffusion: Free Lunch for Generating Optical Illusion Hidden Pictures with Phase-Transferred DiffusionCVPR 2025 (Poster)Paper | Code
Diffusion-based Adversarial Purification from the Perspective of the Frequency DomainICML 2025 (Spotlight Poster)Paper | Code
Frequency Autoregressive Image Generation with Continuous TokensarXiv 2025Paper | Code
FreeDiff: Progressive Frequency Truncation for Image Editing with Diffusion ModelsECCV 2024 (Poster)Paper | Code
FreeU: Free Lunch in Diffusion U-NetCVPR 2024Paper | Code
Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image TranslationAAAI 2024Paper | Code
ResDiff: Combining CNN and Diffusion Model for Image Super-ResolutionAAAI 2023Paper | Code

📡 Spectral and signal-level control provides low-level but powerful levers for generative consistency, resolution, and robustness.


🔁 Alignment & Knowledge Transfer

PaperVenueLinks
When Model Knowledge meets Diffusion: Data-free Synthesis with Domain-Class AlignmentICML 2025Paper | Code

🧬 These works align discrete symbolic knowledge with continuous generative priors, aiming for controllability in low-data or zero-shot regimes.


You may also consider including some notable image-to-image (I2I) editing methods in your collection.

Image Editing

PaperVenueLinks
In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion TransformerNeurIPS 2025Paper | Code
AnyEdit: Mastering Unified High-Quality Image Editing for Any IdeaCVPR 2025 (Oral)Paper | Code
Stable Flow: Vital Layers for Training-Free Image EditingCVPR 2025Paper | Code
UniReal: Universal Image Generation and Editing via Learning Real-world DynamicsCVPR 2025Paper | Code
Taming Rectified Flow for Inversion and EditingICML 2025Paper | Code
Semantic Image Inversion and Editing using Rectified Stochastic Differential EquationsICLR 2025Paper | Code
Diffusion Model-Based Image Editing: A SurveyTPAMI 2025Paper | Code
ChronoEdit: Towards Temporal Reasoning for Image Editing and World SimulationarXiv 2025Paper | Code
SAEdit: Token-level control for continuous image editing via Sparse AutoEncoderarXiv 2025Paper | Code
UltraEdit: Instruction-based Fine-Grained Image Editing at ScaleNeurIPS 2024Paper | Code
SmartEdit: Exploring complex instruction-based image editing with multimodal large language modelsCVPR 2024Paper | Code
Direct Inversion: Boosting Diffusion-based Editing with 3 Lines of CodeICLR 2024Paper | Code
MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image EditingNeurIPS 2023Paper | Code
NULL-text inversion for editing real images using guided diffusion modelsCVPR 2023Paper | Code
Imagic: Text-based real image editing with diffusion modelsCVPR 2023Paper | Code
InstructPix2Pix: Learning To Follow Image Editing InstructionsCVPR 2023Paper | Code
Plug-and-Play Diffusion Features for Text-Driven Image-to-Image TranslationCVPR 2023 (Poster)Paper | Code
DiffEdit: Diffusion-based semantic image editing with mask guidanceICLR 2023Paper | Code
Prompt-to-Prompt Image Editing with Cross Attention ControlICLR 2023Paper | Code
Zero-shot Image-to-Image Translation (pix2pix-zero)SIGGRAPH 2023Paper | Code
RePaint: Inpainting using Denoising Diffusion Probabilistic ModelsCVPR 2022Paper | Code

✏️ Editing is arguably where controllability matters most — precision, structure preservation, and user intent must all align.


📬 Contribute

💡 Know a paper we missed? Working on a new controllable generation method?
Feel free to submit a pull request or open an issue — contributions are welcome!

attention-mechanism
controlnet
diffusion-models
generative-model
multimodal

Contributors

Angusliuuu

6 commits

Bili-Sakura

2 commits

cursoragent

2 commits