A curated list of recent papers (2023–2025) on controllable generative models, covering diffusion-based architectures with fine-grained control, attention interpretation, spectral manipulation, and structure-preserving image editing. Ideal for researchers and developers exploring controllable synthesis.
42
10 commits
updated Apr 22, 2026
A curated and continuously updated collection of recent (2023–2025) research papers on controllable generative models, with a special focus on both UNet-based diffusion models and Transformer-based diffusion architectures.
This list emphasizes core advances in:
💡 Our goal is not only to track the state-of-the-art in controllable generation, but also to offer a well-organized knowledge map for newcomers and researchers building on top of diffusion models.
| Paper | Venue | Links |
|---|---|---|
| OminiControl: Minimal and Universal Control for Diffusion Transformer | ICCV 2025 (Highlight) | Paper | Code |
| Rectified Diffusion Guidance for Conditional Generation | CVPR 2025 | Paper | Code |
| FlexControl: Computation-Aware Conditional Control with Differentiable Router | ICML 2025 (Poster) | Paper | Code |
| CTRL-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model | ICLR 2025 (Oral) | Paper | Code |
| Ctrl-U: Robust Conditional Image Generation via Uncertainty-aware Reward Modeling | ICLR 2025 | Paper | Code |
| ConceptCtrl: Concept Control of Zero-shot Personalized Image Generation | arXiv 2025 | Paper | Code |
| Is Noise Conditioning Necessary for Denoising Generative Models? | arXiv 2025 | Paper | Code |
| Ctrl‑X: Controlling Structure and Appearance Without Guidance | NeurIPS 2024 | Paper | Code |
| ControlNet++: Improving Conditional Controls with Efficient Consistency Feedback | ECCV 2024 | Paper | Code |
| ControlNet-XS: Rethinking the Control of Text-to-Image Diffusion Models as Feedback-Control Systems | ECCV 2024 | Paper | Code |
| SmartControl: Enhancing ControlNet for Handling Rough Visual Conditions | ECCV 2024 (Poster) | Paper | Code |
| PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models | ICML 2024 Workshop | Paper | Code |
| Layout-to-Image Generation with Localized Descriptions using ControlNet with Cross-Attention Guidance | WACV 2024 | Paper | Code |
| Controllable Generation with Text-to-Image Diffusion Models: A Survey | T-PAMI 2024 | Paper | Code |
| Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models | arXiv 2024 | Paper | Code |
| Cocktail : Mixing Multi-Modality Controls for Text-Conditional Image Generation | NeurIPS 2023 | Paper | Code |
| T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion | NeurIPS 2023 | Paper | Code |
| Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion Models | NeurIPS 2023 | Paper | Code |
| Adding Conditional Control to Text-to-Image Diffusion Models | ICCV 2023 | paper | Code |
| GLIGEN: Open-Set Grounded Text-to-Image Generation | CVPR 2023 | Paper | Code |
| MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation | CVPR 2023 | Paper | Code |
| UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild | CVPR 2023 | Paper | Code |
| Composer: Creative and Controllable Image Synthesis with Composable Conditions | ICML 2023 | Paper | Code |
| More Control for Free! Image Synthesis with Semantic Diffusion Guidance | WACV 2023 | Paper | Code |
| IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models | arXiv 2023 | Paper | Code |
| Sketch-Guided Text-to-Image Diffusion Models | arXiv 2022 | Paper | Code |
| Generative Modeling by Estimating Gradients of the Data Distribution | NeurIPS 2019 | Paper | Code |
🧠 These papers push the boundary of how we guide generation, whether through minimal prompts, learned adapters, or uncertainty-aware mechanisms.
| Paper | Venue | Links |
|---|---|---|
| ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features | ICML 2025 (Oral) | Paper | Code |
| ToMA: Token Merge with Attention for Diffusion Models | ICML 2025 (Poster) | Paper | Code |
| Attention in Diffusion Model: A Survey | arXiv 2025 | Paper | Code |
| Attention Distillation: A Unified Approach to Visual Characteristics Transfer | CVPR 2025 | Paper | Code |
| What the DAAM: Interpreting Stable Diffusion Using Cross Attention | ACL 2024 (Oral) | Paper | Code |
🔬 Interpretability is not just analysis — it's a step toward transparent and editable generative pipelines.
| Paper | Venue | Links |
|---|---|---|
| DiffFNO: Diffusion Fourier Neural Operator for Arbitrary-Scale Super-Resolution | CVPR 2025 (Oral) | Paper | Code |
| PTDiffusion: Free Lunch for Generating Optical Illusion Hidden Pictures with Phase-Transferred Diffusion | CVPR 2025 (Poster) | Paper | Code |
| Diffusion-based Adversarial Purification from the Perspective of the Frequency Domain | ICML 2025 (Spotlight Poster) | Paper | Code |
| Frequency Autoregressive Image Generation with Continuous Tokens | arXiv 2025 | Paper | Code |
| FreeDiff: Progressive Frequency Truncation for Image Editing with Diffusion Models | ECCV 2024 (Poster) | Paper | Code |
| FreeU: Free Lunch in Diffusion U-Net | CVPR 2024 | Paper | Code |
| Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image Translation | AAAI 2024 | Paper | Code |
| ResDiff: Combining CNN and Diffusion Model for Image Super-Resolution | AAAI 2023 | Paper | Code |
📡 Spectral and signal-level control provides low-level but powerful levers for generative consistency, resolution, and robustness.
| Paper | Venue | Links |
|---|---|---|
| When Model Knowledge meets Diffusion: Data-free Synthesis with Domain-Class Alignment | ICML 2025 | Paper | Code |
🧬 These works align discrete symbolic knowledge with continuous generative priors, aiming for controllability in low-data or zero-shot regimes.
You may also consider including some notable image-to-image (I2I) editing methods in your collection.
| Paper | Venue | Links |
|---|---|---|
| In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer | NeurIPS 2025 | Paper | Code |
| AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea | CVPR 2025 (Oral) | Paper | Code |
| Stable Flow: Vital Layers for Training-Free Image Editing | CVPR 2025 | Paper | Code |
| UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics | CVPR 2025 | Paper | Code |
| Taming Rectified Flow for Inversion and Editing | ICML 2025 | Paper | Code |
| Semantic Image Inversion and Editing using Rectified Stochastic Differential Equations | ICLR 2025 | Paper | Code |
| Diffusion Model-Based Image Editing: A Survey | TPAMI 2025 | Paper | Code |
| ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation | arXiv 2025 | Paper | Code |
| SAEdit: Token-level control for continuous image editing via Sparse AutoEncoder | arXiv 2025 | Paper | Code |
| UltraEdit: Instruction-based Fine-Grained Image Editing at Scale | NeurIPS 2024 | Paper | Code |
| SmartEdit: Exploring complex instruction-based image editing with multimodal large language models | CVPR 2024 | Paper | Code |
| Direct Inversion: Boosting Diffusion-based Editing with 3 Lines of Code | ICLR 2024 | Paper | Code |
| MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing | NeurIPS 2023 | Paper | Code |
| NULL-text inversion for editing real images using guided diffusion models | CVPR 2023 | Paper | Code |
| Imagic: Text-based real image editing with diffusion models | CVPR 2023 | Paper | Code |
| InstructPix2Pix: Learning To Follow Image Editing Instructions | CVPR 2023 | Paper | Code |
| Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation | CVPR 2023 (Poster) | Paper | Code |
| DiffEdit: Diffusion-based semantic image editing with mask guidance | ICLR 2023 | Paper | Code |
| Prompt-to-Prompt Image Editing with Cross Attention Control | ICLR 2023 | Paper | Code |
| Zero-shot Image-to-Image Translation (pix2pix-zero) | SIGGRAPH 2023 | Paper | Code |
| RePaint: Inpainting using Denoising Diffusion Probabilistic Models | CVPR 2022 | Paper | Code |
✏️ Editing is arguably where controllability matters most — precision, structure preservation, and user intent must all align.
💡 Know a paper we missed? Working on a new controllable generation method?
Feel free to submit a pull request or open an issue — contributions are welcome!
A curated list of recent papers (2023–2025) on controllable generative models, covering diffusion-based architectures with fine-grained control, attention interpretation, spectral manipulation, and structure-preserving image editing. Ideal for researchers and developers exploring controllable synthesis.
42
10 commits
updated Apr 22, 2026
A curated and continuously updated collection of recent (2023–2025) research papers on controllable generative models, with a special focus on both UNet-based diffusion models and Transformer-based diffusion architectures.
This list emphasizes core advances in:
💡 Our goal is not only to track the state-of-the-art in controllable generation, but also to offer a well-organized knowledge map for newcomers and researchers building on top of diffusion models.
| Paper | Venue | Links |
|---|---|---|
| OminiControl: Minimal and Universal Control for Diffusion Transformer | ICCV 2025 (Highlight) | Paper | Code |
| Rectified Diffusion Guidance for Conditional Generation | CVPR 2025 | Paper | Code |
| FlexControl: Computation-Aware Conditional Control with Differentiable Router | ICML 2025 (Poster) | Paper | Code |
| CTRL-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model | ICLR 2025 (Oral) | Paper | Code |
| Ctrl-U: Robust Conditional Image Generation via Uncertainty-aware Reward Modeling | ICLR 2025 | Paper | Code |
| ConceptCtrl: Concept Control of Zero-shot Personalized Image Generation | arXiv 2025 | Paper | Code |
| Is Noise Conditioning Necessary for Denoising Generative Models? | arXiv 2025 | Paper | Code |
| Ctrl‑X: Controlling Structure and Appearance Without Guidance | NeurIPS 2024 | Paper | Code |
| ControlNet++: Improving Conditional Controls with Efficient Consistency Feedback | ECCV 2024 | Paper | Code |
| ControlNet-XS: Rethinking the Control of Text-to-Image Diffusion Models as Feedback-Control Systems | ECCV 2024 | Paper | Code |
| SmartControl: Enhancing ControlNet for Handling Rough Visual Conditions | ECCV 2024 (Poster) | Paper | Code |
| PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models | ICML 2024 Workshop | Paper | Code |
| Layout-to-Image Generation with Localized Descriptions using ControlNet with Cross-Attention Guidance | WACV 2024 | Paper | Code |
| Controllable Generation with Text-to-Image Diffusion Models: A Survey | T-PAMI 2024 | Paper | Code |
| Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models | arXiv 2024 | Paper | Code |
| Cocktail : Mixing Multi-Modality Controls for Text-Conditional Image Generation | NeurIPS 2023 | Paper | Code |
| T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion | NeurIPS 2023 | Paper | Code |
| Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion Models | NeurIPS 2023 | Paper | Code |
| Adding Conditional Control to Text-to-Image Diffusion Models | ICCV 2023 | paper | Code |
| GLIGEN: Open-Set Grounded Text-to-Image Generation | CVPR 2023 | Paper | Code |
| MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation | CVPR 2023 | Paper | Code |
| UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild | CVPR 2023 | Paper | Code |
| Composer: Creative and Controllable Image Synthesis with Composable Conditions | ICML 2023 | Paper | Code |
| More Control for Free! Image Synthesis with Semantic Diffusion Guidance | WACV 2023 | Paper | Code |
| IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models | arXiv 2023 | Paper | Code |
| Sketch-Guided Text-to-Image Diffusion Models | arXiv 2022 | Paper | Code |
| Generative Modeling by Estimating Gradients of the Data Distribution | NeurIPS 2019 | Paper | Code |
🧠 These papers push the boundary of how we guide generation, whether through minimal prompts, learned adapters, or uncertainty-aware mechanisms.
| Paper | Venue | Links |
|---|---|---|
| ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features | ICML 2025 (Oral) | Paper | Code |
| ToMA: Token Merge with Attention for Diffusion Models | ICML 2025 (Poster) | Paper | Code |
| Attention in Diffusion Model: A Survey | arXiv 2025 | Paper | Code |
| Attention Distillation: A Unified Approach to Visual Characteristics Transfer | CVPR 2025 | Paper | Code |
| What the DAAM: Interpreting Stable Diffusion Using Cross Attention | ACL 2024 (Oral) | Paper | Code |
🔬 Interpretability is not just analysis — it's a step toward transparent and editable generative pipelines.
| Paper | Venue | Links |
|---|---|---|
| DiffFNO: Diffusion Fourier Neural Operator for Arbitrary-Scale Super-Resolution | CVPR 2025 (Oral) | Paper | Code |
| PTDiffusion: Free Lunch for Generating Optical Illusion Hidden Pictures with Phase-Transferred Diffusion | CVPR 2025 (Poster) | Paper | Code |
| Diffusion-based Adversarial Purification from the Perspective of the Frequency Domain | ICML 2025 (Spotlight Poster) | Paper | Code |
| Frequency Autoregressive Image Generation with Continuous Tokens | arXiv 2025 | Paper | Code |
| FreeDiff: Progressive Frequency Truncation for Image Editing with Diffusion Models | ECCV 2024 (Poster) | Paper | Code |
| FreeU: Free Lunch in Diffusion U-Net | CVPR 2024 | Paper | Code |
| Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image Translation | AAAI 2024 | Paper | Code |
| ResDiff: Combining CNN and Diffusion Model for Image Super-Resolution | AAAI 2023 | Paper | Code |
📡 Spectral and signal-level control provides low-level but powerful levers for generative consistency, resolution, and robustness.
| Paper | Venue | Links |
|---|---|---|
| When Model Knowledge meets Diffusion: Data-free Synthesis with Domain-Class Alignment | ICML 2025 | Paper | Code |
🧬 These works align discrete symbolic knowledge with continuous generative priors, aiming for controllability in low-data or zero-shot regimes.
You may also consider including some notable image-to-image (I2I) editing methods in your collection.
| Paper | Venue | Links |
|---|---|---|
| In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer | NeurIPS 2025 | Paper | Code |
| AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea | CVPR 2025 (Oral) | Paper | Code |
| Stable Flow: Vital Layers for Training-Free Image Editing | CVPR 2025 | Paper | Code |
| UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics | CVPR 2025 | Paper | Code |
| Taming Rectified Flow for Inversion and Editing | ICML 2025 | Paper | Code |
| Semantic Image Inversion and Editing using Rectified Stochastic Differential Equations | ICLR 2025 | Paper | Code |
| Diffusion Model-Based Image Editing: A Survey | TPAMI 2025 | Paper | Code |
| ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation | arXiv 2025 | Paper | Code |
| SAEdit: Token-level control for continuous image editing via Sparse AutoEncoder | arXiv 2025 | Paper | Code |
| UltraEdit: Instruction-based Fine-Grained Image Editing at Scale | NeurIPS 2024 | Paper | Code |
| SmartEdit: Exploring complex instruction-based image editing with multimodal large language models | CVPR 2024 | Paper | Code |
| Direct Inversion: Boosting Diffusion-based Editing with 3 Lines of Code | ICLR 2024 | Paper | Code |
| MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing | NeurIPS 2023 | Paper | Code |
| NULL-text inversion for editing real images using guided diffusion models | CVPR 2023 | Paper | Code |
| Imagic: Text-based real image editing with diffusion models | CVPR 2023 | Paper | Code |
| InstructPix2Pix: Learning To Follow Image Editing Instructions | CVPR 2023 | Paper | Code |
| Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation | CVPR 2023 (Poster) | Paper | Code |
| DiffEdit: Diffusion-based semantic image editing with mask guidance | ICLR 2023 | Paper | Code |
| Prompt-to-Prompt Image Editing with Cross Attention Control | ICLR 2023 | Paper | Code |
| Zero-shot Image-to-Image Translation (pix2pix-zero) | SIGGRAPH 2023 | Paper | Code |
| RePaint: Inpainting using Denoising Diffusion Probabilistic Models | CVPR 2022 | Paper | Code |
✏️ Editing is arguably where controllability matters most — precision, structure preservation, and user intent must all align.
💡 Know a paper we missed? Working on a new controllable generation method?
Feel free to submit a pull request or open an issue — contributions are welcome!