SiatMMLab/Awesome-Diffusion-Model-Based-Image-Editing-Methods

Diffusion Model-Based Image Editing: A Survey (TPAMI 2025)

714

64 commits

updated Jul 15, 2025

See the code

README

image

Awesome License: MIT Made With Love arXiv visitors

The repository is based on our survey Diffusion Model-Based Image Editing: A Survey (TPAMI 2025).

Yi Huang*, Jiancheng Huang*, Yifan Liu*, Mingfu Yan*, Jiaxi Lv*, Jianzhuang Liu*, Wei Xiong, He Zhang, Liangliang Cao, Shifeng Chen

Shenzhen Institute of Advanced Technology (SIAT), Chinese Academy of Sciences (CAS), Adobe Inc, Apple Inc, Southern University of Science and Technology (SUSTech)

Abstract

Denoising diffusion models have emerged as a powerful tool for various image generation and editing tasks, facilitating the synthesis of visual content in an unconditional or input-conditional manner. The core idea behind them is learning to reverse the process of gradually adding noise to images, allowing them to generate high-quality samples from a complex distribution. In this survey, we provide an exhaustive overview of existing methods using diffusion models for image editing, covering both theoretical and practical aspects in the field. We delve into a thorough analysis and categorization of these works from multiple perspectives, including learning strategies, user-input conditions, and the array of specific editing tasks that can be accomplished. In addition, we pay special attention to image inpainting and outpainting, and explore both earlier traditional context-driven and current multimodal conditional methods, offering a comprehensive analysis of their methodologies. To further evaluate the performance of text-guided image editing algorithms, we propose a systematic benchmark, EditEval, featuring an innovative metric, LMM Score. Finally, we address current limitations and envision some potential directions for future research.

πŸ”– News!!!

πŸ“Œ We are actively tracking the latest research and welcome contributions to our repository and survey paper. If your studies are relevant, please feel free to contact us.

πŸ“° 2025-02-11: πŸ₯³ Congrats, our paper is accepted by TPAMI 2025!!

πŸ“° 2024-10-25: Our benchmark EditEval_v2 is now released.

πŸ“° 2024-03-22: The template of computing LMM Score using GPT-4V, along with a corresponding leaderboard comparing several leading methods, is released.

πŸ“° 2024-03-14: Our benchmark EditEval_v1 is now released.

πŸ“° 2024-03-06: We establish a template for paper submissions. This template is accessible by navigating to the New Issue button within Issues or by clicking here. Once there, please select the Paper Submission Form and complete it following the guidelines provided.

πŸ“° 2024-02-28: Our comprehensive survey paper, summarizing related methods published before February 1, 2024, is now available.

πŸ” BibTeX

If you find this work helpful in your research, welcome to cite the paper and give a ⭐.

@article{huang2025diffusion,
  title={Diffusion Model-Based Image Editing: A Survey},
  author={Huang, Yi and Huang, Jiancheng and Liu, Yifan and Yan, Mingfu and Lv, Jiaxi and Liu, Jianzhuang and Xiong, Wei and Zhang, He and Cao, Liangliang and Chen, Shifeng},
  journal={IEEE Transactions on Pattern Analysis and Machine Intelligence},
  year={2025},
  publisher={IEEE}
}

Table of contents

Papers

Training-Based

Training-Based: Domain-Specific Editing

Training-Based: Reference and Attribute Guided Editing

Training-Based: Instructional Editing

TitlePublicationDate
UIP2P: Unsupervised Instruction-based Image Editing via Edit Reversibility ConstraintICCV 20252024.12
FreeEdit: Mask-free Reference-based Image Editing with Multi-modal InstructionarXiv 20242024.09
EditWorld: Simulating World Dynamics for Instruction-Following Image EditingarXiv 20242024.05
InstructGIE: Towards Generalizable Image EditingarXiv 20242024.03
SmartEdit: Exploring Complex Instruction-based Image Editing with Multimodal Large Language ModelsCVPR 20242023.12
InstructAny2Pix: Flexible Visual Editing via Multimodal Instruction FollowingarXiv 20232023.12
Focus on Your Instruction: Fine-grained and Multi-instruction Image Editing by Attention ModulationCVPR 20242023.12
Emu edit: Precise image editing via recognition and generation tasksarXiv 20232023.11
Guiding instruction-based image editing via multimodal large language modelsICLR 20242023.09
Instructdiffusion: A generalist modeling interface for vision tasksCVPR 20242023.09
MoEController: Instruction-based Arbitrary Image Manipulation with Mixture-of-Expert ControllersarXiv 20232023.09
ImageBrush: Learning Visual In-Context Instructions for Exemplar-Based Image ManipulationNeurIPS 20232023.08
Inst-Inpaint: Instructing to Remove Objects with Diffusion ModelsarXiv 20232023.04
HIVE: Harnessing Human Feedback for Instructional Visual EditingCVPR 20242023.03
DialogPaint: A Dialog-based Image Editing ModelarXiv 20232023.01
Learning to Follow Object-Centric Image Editing Instructions FaithfullyEMNLP 20232023.01
Instructpix2pix: Learning to follow image editing instructionsCVPR 20232022.11

Training-Based: Pseudo-Target Retrieval-Based Editing

Testing-Time Finetuning

Testing-Time Finetuning: Denosing Model Finetuning

Testing-Time Finetuning: Embeddings Finetuning

Testing-Time Finetuning: Guidance with Hypernetworks

Testing-Time Finetuning: Latent Variable Optimization

Testing-Time Finetuning: Hybrid Finetuning

Training and Finetuning Free

Training and Finetuning Free: Input Text Refinement

Training and Finetuning Free: Inversion/Sampling Modification

TitlePublicationDate
FireFlow: Fast Inversion of Rectified Flow for Image Semantic EditingarXiv 20242024.12
Inversion-Free Image Editing with Natural LanguageCVPR 20242023.12
Fixed-point Inversion for Text-to-image diffusion modelsarXiv 20232023.12
Tuning-Free Inversion-Enhanced Control for Consistent Image EditingarXiv 20232023.12
The Blessing of Randomness: SDE Beats ODE in General Diffusion-based Image EditingICLR 20242023.11
LEDITS++: Limitless Image Editing using Text-to-Image ModelsCVPR 20242023.11
A latent space of stochastic diffusion models for zero-shot image editing and guidanceICCV 20232023.10
Effective real image editing with accelerated iterative diffusion inversionICCV 20232023.09
Fec: Three finetuning-free methods to enhance consistency for real image editingarXiv 20232023.09
Iterative multi-granular image editing using diffusion modelsWACV 20242023.09
ProxEdit: Improving Tuning-Free Real Image Editing With Proximal GuidanceWACV 20242023.06
Diffusion self-guidance for controllable image generationNeurIPS 20232023.06
Diffusion Brush: A Latent Diffusion Model-based Editing Tool for AI-generated ImagesarXiv 20232023.06
Null-text guidance in diffusion models is secretly a cartoon-style creatorACM MM 20232023.05
Negative-prompt Inversion: Fast Image Inversion for Editing with Text-guided Diffusion ModelsarXiv 20232023.05
An Edit Friendly DDPM Noise Space: Inversion and ManipulationsCVPR 20242023.04
Training-Free Content Injection Using H-Space in Diffusion ModelsWACV 20242023.03
Edict: Exact diffusion inversion via coupled transformationsCVPR 20232022.11
Direct inversion: Optimization-free text-driven real image editing with diffusion modelsarXiv 20222022.11

Training and Finetuning Free: Attention Modification

Training and Finetuning Free: Mask Guidance

Training and Finetuning Free: Multi-Noise Redirection

Benchmark EditEval_v1

EditEval_v1 is a benchmark tailored for evaluation of general diffusion-model based image editing algorithms. It contains 50 high-quality images selected from Unsplash, each accompanied by a source text prompt, a target editing prompt, and a text editing instruction generated by GPT-4V. This benchmark covers seven most popular specific editing tasks across semantic, stylistic and structural editing defined in our paper: object addition, object replacement, object removal, background change, overall style change, texture change, and action change. Click here to download this dataset!

Benchmark EditEval_v2

EditEval_v2 is an enhanced benchmark designed to evaluate general diffusion-model-based image editing algorithms. This version expands upon its predecessor by including 150 high-quality images selected from Unsplash. Each image is paired with a source text prompt, a target editing prompt, and a text editing instruction generated by GPT-4V. EditEval_v2 continues to cover the seven most popular specific editing tasks across semantic, stylistic, and structural editing as defined in our paper: object addition, object replacement, object removal, background change, overall style change, texture change, and action change. Click here to download this dataset!

Leaderboard

To facilitate a user-friendly application of LMM Score, here we provide a comprehensive template for its implementation in GPT-4V. This template comes with step-by-step instructions and all required materials, making it easy for users to apply. Additionally, we construct a leaderboard comparing various representative methods evaluated using LMM Score on our EditEval_v1 benchmark, which can be found here.

Star History

Star History Chart

Contributors

MingfuYAN

23 commits

Zwette

21 commits

SiatMMLab

7 commits

jiaxilv

5 commits

SiatMMLab/Awesome-Diffusion-Model-Based-Image-Editing-Methods

Diffusion Model-Based Image Editing: A Survey (TPAMI 2025)

714

64 commits

updated Jul 15, 2025

See the code

README

image

Awesome License: MIT Made With Love arXiv visitors

The repository is based on our survey Diffusion Model-Based Image Editing: A Survey (TPAMI 2025).

Yi Huang*, Jiancheng Huang*, Yifan Liu*, Mingfu Yan*, Jiaxi Lv*, Jianzhuang Liu*, Wei Xiong, He Zhang, Liangliang Cao, Shifeng Chen

Shenzhen Institute of Advanced Technology (SIAT), Chinese Academy of Sciences (CAS), Adobe Inc, Apple Inc, Southern University of Science and Technology (SUSTech)

Abstract

Denoising diffusion models have emerged as a powerful tool for various image generation and editing tasks, facilitating the synthesis of visual content in an unconditional or input-conditional manner. The core idea behind them is learning to reverse the process of gradually adding noise to images, allowing them to generate high-quality samples from a complex distribution. In this survey, we provide an exhaustive overview of existing methods using diffusion models for image editing, covering both theoretical and practical aspects in the field. We delve into a thorough analysis and categorization of these works from multiple perspectives, including learning strategies, user-input conditions, and the array of specific editing tasks that can be accomplished. In addition, we pay special attention to image inpainting and outpainting, and explore both earlier traditional context-driven and current multimodal conditional methods, offering a comprehensive analysis of their methodologies. To further evaluate the performance of text-guided image editing algorithms, we propose a systematic benchmark, EditEval, featuring an innovative metric, LMM Score. Finally, we address current limitations and envision some potential directions for future research.

πŸ”– News!!!

πŸ“Œ We are actively tracking the latest research and welcome contributions to our repository and survey paper. If your studies are relevant, please feel free to contact us.

πŸ“° 2025-02-11: πŸ₯³ Congrats, our paper is accepted by TPAMI 2025!!

πŸ“° 2024-10-25: Our benchmark EditEval_v2 is now released.

πŸ“° 2024-03-22: The template of computing LMM Score using GPT-4V, along with a corresponding leaderboard comparing several leading methods, is released.

πŸ“° 2024-03-14: Our benchmark EditEval_v1 is now released.

πŸ“° 2024-03-06: We establish a template for paper submissions. This template is accessible by navigating to the New Issue button within Issues or by clicking here. Once there, please select the Paper Submission Form and complete it following the guidelines provided.

πŸ“° 2024-02-28: Our comprehensive survey paper, summarizing related methods published before February 1, 2024, is now available.

πŸ” BibTeX

If you find this work helpful in your research, welcome to cite the paper and give a ⭐.

@article{huang2025diffusion,
  title={Diffusion Model-Based Image Editing: A Survey},
  author={Huang, Yi and Huang, Jiancheng and Liu, Yifan and Yan, Mingfu and Lv, Jiaxi and Liu, Jianzhuang and Xiong, Wei and Zhang, He and Cao, Liangliang and Chen, Shifeng},
  journal={IEEE Transactions on Pattern Analysis and Machine Intelligence},
  year={2025},
  publisher={IEEE}
}

Table of contents

Papers

Training-Based

Training-Based: Domain-Specific Editing

Training-Based: Reference and Attribute Guided Editing

Training-Based: Instructional Editing

TitlePublicationDate
UIP2P: Unsupervised Instruction-based Image Editing via Edit Reversibility ConstraintICCV 20252024.12
FreeEdit: Mask-free Reference-based Image Editing with Multi-modal InstructionarXiv 20242024.09
EditWorld: Simulating World Dynamics for Instruction-Following Image EditingarXiv 20242024.05
InstructGIE: Towards Generalizable Image EditingarXiv 20242024.03
SmartEdit: Exploring Complex Instruction-based Image Editing with Multimodal Large Language ModelsCVPR 20242023.12
InstructAny2Pix: Flexible Visual Editing via Multimodal Instruction FollowingarXiv 20232023.12
Focus on Your Instruction: Fine-grained and Multi-instruction Image Editing by Attention ModulationCVPR 20242023.12
Emu edit: Precise image editing via recognition and generation tasksarXiv 20232023.11
Guiding instruction-based image editing via multimodal large language modelsICLR 20242023.09
Instructdiffusion: A generalist modeling interface for vision tasksCVPR 20242023.09
MoEController: Instruction-based Arbitrary Image Manipulation with Mixture-of-Expert ControllersarXiv 20232023.09
ImageBrush: Learning Visual In-Context Instructions for Exemplar-Based Image ManipulationNeurIPS 20232023.08
Inst-Inpaint: Instructing to Remove Objects with Diffusion ModelsarXiv 20232023.04
HIVE: Harnessing Human Feedback for Instructional Visual EditingCVPR 20242023.03
DialogPaint: A Dialog-based Image Editing ModelarXiv 20232023.01
Learning to Follow Object-Centric Image Editing Instructions FaithfullyEMNLP 20232023.01
Instructpix2pix: Learning to follow image editing instructionsCVPR 20232022.11

Training-Based: Pseudo-Target Retrieval-Based Editing

Testing-Time Finetuning

Testing-Time Finetuning: Denosing Model Finetuning

Testing-Time Finetuning: Embeddings Finetuning

Testing-Time Finetuning: Guidance with Hypernetworks

Testing-Time Finetuning: Latent Variable Optimization

Testing-Time Finetuning: Hybrid Finetuning

Training and Finetuning Free

Training and Finetuning Free: Input Text Refinement

Training and Finetuning Free: Inversion/Sampling Modification

TitlePublicationDate
FireFlow: Fast Inversion of Rectified Flow for Image Semantic EditingarXiv 20242024.12
Inversion-Free Image Editing with Natural LanguageCVPR 20242023.12
Fixed-point Inversion for Text-to-image diffusion modelsarXiv 20232023.12
Tuning-Free Inversion-Enhanced Control for Consistent Image EditingarXiv 20232023.12
The Blessing of Randomness: SDE Beats ODE in General Diffusion-based Image EditingICLR 20242023.11
LEDITS++: Limitless Image Editing using Text-to-Image ModelsCVPR 20242023.11
A latent space of stochastic diffusion models for zero-shot image editing and guidanceICCV 20232023.10
Effective real image editing with accelerated iterative diffusion inversionICCV 20232023.09
Fec: Three finetuning-free methods to enhance consistency for real image editingarXiv 20232023.09
Iterative multi-granular image editing using diffusion modelsWACV 20242023.09
ProxEdit: Improving Tuning-Free Real Image Editing With Proximal GuidanceWACV 20242023.06
Diffusion self-guidance for controllable image generationNeurIPS 20232023.06
Diffusion Brush: A Latent Diffusion Model-based Editing Tool for AI-generated ImagesarXiv 20232023.06
Null-text guidance in diffusion models is secretly a cartoon-style creatorACM MM 20232023.05
Negative-prompt Inversion: Fast Image Inversion for Editing with Text-guided Diffusion ModelsarXiv 20232023.05
An Edit Friendly DDPM Noise Space: Inversion and ManipulationsCVPR 20242023.04
Training-Free Content Injection Using H-Space in Diffusion ModelsWACV 20242023.03
Edict: Exact diffusion inversion via coupled transformationsCVPR 20232022.11
Direct inversion: Optimization-free text-driven real image editing with diffusion modelsarXiv 20222022.11

Training and Finetuning Free: Attention Modification

Training and Finetuning Free: Mask Guidance

Training and Finetuning Free: Multi-Noise Redirection

Benchmark EditEval_v1

EditEval_v1 is a benchmark tailored for evaluation of general diffusion-model based image editing algorithms. It contains 50 high-quality images selected from Unsplash, each accompanied by a source text prompt, a target editing prompt, and a text editing instruction generated by GPT-4V. This benchmark covers seven most popular specific editing tasks across semantic, stylistic and structural editing defined in our paper: object addition, object replacement, object removal, background change, overall style change, texture change, and action change. Click here to download this dataset!

Benchmark EditEval_v2

EditEval_v2 is an enhanced benchmark designed to evaluate general diffusion-model-based image editing algorithms. This version expands upon its predecessor by including 150 high-quality images selected from Unsplash. Each image is paired with a source text prompt, a target editing prompt, and a text editing instruction generated by GPT-4V. EditEval_v2 continues to cover the seven most popular specific editing tasks across semantic, stylistic, and structural editing as defined in our paper: object addition, object replacement, object removal, background change, overall style change, texture change, and action change. Click here to download this dataset!

Leaderboard

To facilitate a user-friendly application of LMM Score, here we provide a comprehensive template for its implementation in GPT-4V. This template comes with step-by-step instructions and all required materials, making it easy for users to apply. Additionally, we construct a leaderboard comparing various representative methods evaluated using LMM Score on our EditEval_v1 benchmark, which can be found here.

Star History

Star History Chart

Contributors

MingfuYAN

23 commits

Zwette

21 commits

SiatMMLab

7 commits

jiaxilv

5 commits