datar001/Awesome-AD-on-T2IDM

A collection of resources on attacks and defenses targeting text-to-image diffusion models

103

36 commits

updated Dec 20, 2025

See the code

README

Awesome-Attacks and Defenses on T2I Diffusion Models

This repository is a curated collection of research papers focused on $\textbf{Adversarial Attacks and Defenses on Text-to-Image Diffusion Models (AD-on-T2IDM)}$.

We will continuously update this collection to track the latest advancements in the field of AD-on-T2IDM.

Welcome to follow and star! If you have any relevant materials or suggestions, please feel free to contact us (zcy@tju.edu.cn) or submit a pull request.

For more detailed information, please refer to our survey paper: [ARXIV], [Published Version]

News

  • We are pleased to share that our latest work, T2I-RiskyPrompt, has been accepted to AAAI 2026! 🎊As Text-to-Image models evolve, safety becomes paramount. T2I-RiskyPrompt bridges the gap in current safety evaluations with:
    • ✅ Hierarchical Taxonomy: Covers 6 major risky categories & 14 fine-grained subcategories.
    • ✅ Rich Annotations: 6,432 prompts labeled with both hierarchical risky categories and detailed risky reasons
    • ✅ Built-in Detector: An automated pipeline to evaluate the safety of generated images.
    • 🔗 Code & Dataset: [Github], 📄 Paper: [Arxiv]

Citation

@article{zhang2025adversarial,
  title={Adversarial attacks and defenses on text-to-image diffusion models: A survey},
  author={Zhang, Chenyu and Hu, Mingwang and Li, Wenhui and Wang, Lanjun},
  journal={Information Fusion},
  volume={114},
  pages={102701},
  year={2025},
  publisher={Elsevier}
}

Recruiting

Looking for prospective PhD students (enroll in 2026). Ideal candidates are expected to be responsible, organized, and creative. Strong programming skills, solid mathematics knowledge as well as good communication skills are preferred. Please send your CV to wanglanjun AT tju.edu.cn. Note: please attach files directly, not to use 163 Cloud file system.

Content

Abstract

Recently, the text-to-image diffusion model has gained considerable attention from the community due to its exceptional image generation capability. A representative model, Stable Diffusion, amassed more than 10 million users within just two months of its release. This surge in popularity has facilitated studies on the robustness and safety of the model, leading to the proposal of various adversarial attack methods. Simultaneously, there has been a marked increase in research focused on defense methods to improve the robustness and safety of these models. In this survey, we provide a comprehensive review of the literature on adversarial attacks and defenses targeting text-to-image diffusion models. We begin with an overview of popular text-to-image diffusion models, followed by an introduction to a taxonomy of adversarial attacks and an in-depth review of existing attack methods. We then present a detailed analysis of current defense methods that improve model robustness and safety. Finally, we discuss ongoing challenges and explore promising future research directions.

Overview of AD-on-T2IDM

Two key concerns in T2IDM: Robustness and Safety

The robustness ensures that the model can generate images with consistent semantics in response to diverse prompts inputted by users in practice.

The safety prevents the misuse of the model in creating malicious images, such as sexual, violent, and political images, etc.

Adversarial attacks

Based on the intent of the adversary, existing attack methods can be divided into two primary categories: untargeted and targeted attacks.

  • For untargeted attacks, consider a scenario with a prompt input by the user~($\textbf{clean prompt}$) and its corresponding output image~($\textbf{clean image}$). The objective of untargeted attacks is to subtly perturb the clean prompt to craft an $\textbf{adversarial prompt}$, further misleading the victim model to generate an $\textbf{adversarial image}$ with semantics different from the clean image. This type of attack is commonly used to uncover the vulnerability in the robustness of the victim model. Some untargeted attacks are shown as follows:

    untargeted attacks

  • For targeted attacks, assumes that the victim model has built-in $\textbf{safeguards}$ to filter $\textbf{malicious prompts}$ and resultant $\textbf{malicious images}$. These prompts and images often explicitly contain $\textbf{malicious concepts}$, such as 'nudity', 'violence', and other predefined concepts. The objective of targeted attacks is to obtain an $\textbf{adversarial prompt}$, which can bypass these safeguards while inducing the victim model to generate $\textbf{adversarial images}$ containing malicious concepts. This type of attack is typically designed to reveal the vulnerability in the safety of the victim model. Some targeted attacks are shown as follows:

    targeted attacks

Defenses

Based on the defense goal, existing defense methods can be classified into two categories: 1) improving model robustness and 2) improving model safety.

  • The goal of robustness is to ensure that generated images have consistent semantics with diverse input prompts in practical applications. Specifically, according to the adversarial attack, the defense methods are asked to mitigate the robustness vulnerabilities in two types of input prompts: 1) the prompt with multiple objects and attributes, and 2) the grammatically incorrect prompt with the subtle noise.

  • The safety goal is to prevent the generation of malicious images in response to both malicious and adversarial prompts. Specifically, malicious prompts explicitly contain malicious concepts, while adversarial prompts cleverly omit these concepts. Moreover, based on the knowledge of the model, existing safety methods can be classified into two categories: external safeguards and internal safeguards. The external safeguards focus on detecting or correcting the malicious prompt before feeding the prompt into the text-to-image model. In contrast, internal safeguards aim to ensure that the semantics of output images deviate from those of malicious images by modifying internal parameters and features within the model. Some examples of external and internal safeguards are shown as follows:

    external safeguards internal safeguards

Notably, although many methods are proposed to improve the model robustness against the prompt with multiple objects and attributes, this collection omits related papers on this part since there has been related surveys, such as controllable image generation [PDF], the development and advancement of image generation capabilities [PDF-1], [PDF-2], [PDF-3]. Moreover, for grammatically incorrect prompts with subtle noise, mature solutions are still lacking. Therefore, this collection mainly focuses on the defense methods for improving model safety.

:grinning:Paper List

:imp:Adversarial Attacks

:collision:Untargeted Attacks

:pouting_cat:White-Box Attacks

Stable diffusion is unstable

Chengbin Du, Yanxi Li, Zhongwei Qiu, Chang Xu

NeurIPS 2024. [PDF] [CODE]

A pilot study of query-free adversarial attack against stable diffusion

Haomin Zhuang, Yihua Zhang

CVPRW 2023. [PDF] [CODE]

:see_no_evil:Black-Box Attacks

Evaluating the Robustness of Text-to-image Diffusion Models against Real-world Attacks

Hongcheng Gao , Hao Zhang , Yinpeng Dong, Zhijie Deng

arxiv 2023. [PDF]

:anger:Targeted Attacks

:cyclone:White-Box Attacks

Red-Teaming the Stable Diffusion Safety Filter

Javier Rando, Daniel Paleka, David Lindner, Lennart Heim, Florian Tramèr

NeurIPS 2022, WorkShop. [PDF]

Unsafe Diffusion: On the Generation of Unsafe Images and Hateful Memes From Text-To-Image Models

Yiting Qu, Xinyue Shen, Xinlei He, Michael Backes, Savvas Zannettou, Yang Zhang

Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security. [PDF] [CODE]

Ring-A-Bell! How Reliable are Concept Removal Methods for Diffusion Models?

Tsai, Yu-Lin and Hsu, Chia-Yi and Xie, Chulin and Lin, Chih-Hsun and Chen, Jia-You and Li, Bo and Chen, Pin-Yu and Yu, Chia-Mu and Huang, Chun-Ying

ICLR 2024. [PDF]

Riatig: Reliable and imperceptible adversarial text-to-image generation with natural prompts

Han Liu, Yuhao Wu, Shixuan Zhai, Bo Yuan, Ning Zhang

CVPR 2023. [PDF] [CODE]

Mma-diffusion: Multimodal attack on diffusion models

Yang, Yijun and Gao, Ruiyuan and Wang, Xiaosen and Ho, Tsung-Yi and Xu, Nan and Xu, Qiang

CVPR 2024. [PDF] [CODE]

Asymmetric Bias in Text-to-Image Generation with Adversarial Attacks

Haz Sameen Shahgir, Xianghao Kong, Greg Ver Steeg, Yue Dong

arxiv 2023. [PDF] [CODE]

Revealing vulnerabilities in stable diffusion via targeted attacks

Chenyu Zhang, Lanjun Wang, Anan Liu

arxiv 2024. [PDF] [CODE]

To generate or not? safety-driven unlearned diffusion models are still easy to generate unsafe images... for now

Zhang, Yimeng and Jia, Jinghan and Chen, Xin and Chen, Aochuan and Zhang, Yihua and Liu, Jiancheng and Ding, Ke and Liu, Sijia

ECCV 2024. [PDF] [CODE]

Prompting4debugging: Red-teaming text-to-image diffusion models by finding problematic prompts

Chin, Zhi-Yi and Jiang, Chieh-Ming and Huang, Ching-Chun and Chen, Pin-Yu and Chiu, Wei-Chen

ICML 2024. [PDF] [CODE]

ADVI2I: ADVERSARIAL IMAGE ATTACK ON IMAGE-TO-IMAGE DIFFUSION MODELS

Yaopei Zeng, Yuanpu Cao, Bochuan Cao, Yurui Chang, Jinghui Chen, Lu Lin

arxiv 2024. [PDF] [CODE]

Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models

Jiachen Ma, Anda Cao, Zhiqing Xiao, Jie Zhang, Chao Ye, Junbo Zhao

arxiv 2024. [PDF]

Adversarial Attacks on Parts of Speech: An Empirical Study in Text-to-Image Generation

G M Shahariar, Jia Chen, Jiachen Li, Yue Dong

EMNLP 2024. [PDF], [CODE]

:snake:Black-Box Attacks

SneakyPrompt: Evaluating Robustness of Text-to-image Generative Models' Safety Filters

Yuchen Yang, Bo Hui, Haolin Yuan, Neil Gong, Yinzhi Cao

Proceedings of the IEEE Symposium on Security and Privacy 2024. [PDF] [CODE]

ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign Users

Guanlin Li, Kangjie Chen, Shudong Zhang, Jie Zhang, Tianwei Zhang

NeurIPS 2024. [PDF] [CODE]

Coljailbreak: Collaborative generation and editing for jailbreaking text-to-image deep generation

Yizhuo Ma, Shanmin Pang, Qi Guo, Tianyu Wei, Qing Guo

NeurIPS 2024. [PDF]

UPAM: Unified Prompt Attack in Text-to-Image Generation Models Against Both Textual Filters and Visual Checkers

Duo Peng, Qiuhong Ke, Jun Liu

ICML2024, [PDF]

FLIRT: Feedback Loop In-context Red Teaming

Ninareh Mehrabi, Palash Goyal, Christophe Dupuy, Qian Hu, Shalini Ghosh, Richard Zemel, Kai-Wei Chang, Aram Galstyan, Rahul Gupta

EMNLP 2024. [PDF]

Fuzz-testing meets LLM-based agents: An automated and efficient framework for jailbreaking text-to-image generation models

Yingkai Dong, Zheng Li, Xiangtao Meng, Ning Yu, Shanqing Guo

S&P 2025. [PDF], [[CODE]](GitHub - YingkaiD/JailFuzzer)

Automatic Jailbreaking of the Text-to-Image Generative AI Systems

Minseon Kim, Hyomin Lee, Boqing Gong, Huishuai Zhang, Sung Ju Hwang

arxiv 2024. [PDF] [CODE]

Exploiting cultural biases via homoglyphs in text-to-image synthesis

Struppek, Lukas and Hintersdorf, Dom and Friedrich, Felix and Schramowski, Patrick and Kersting, Kristian

Journal of Artificial Intelligence Research 2023. [PDF] [CODE]

Divide-and-Conquer Attack: Harnessing the Power of LLM to Bypass Safety Filters of Text-to-Image Models

Yimo Deng, Huangxun Chen

arxiv 2024. [PDF]

Groot: Adversarial Testing for Generative Text-to-Image Models with Tree-based Semantic Transformation

Yi Liu, Guowei Yang, Gelei Deng, Feiyue Chen, Yuqi Chen, Ling Shi, Tianwei Zhang, Yang Liu

arxiv 2024. [PDF]

BSPA: Exploring Black-box Stealthy Prompt Attacks against Image Generators

Yu Tian, Xiao Yang, Yinpeng Dong, Heming Yang, Hang Su, Jun Zhu

arxiv 2024. [PDF]

Black Box Adversarial Prompting for Foundation Models

Natalie Maus, Patrick Chao, Eric Wong, Jacob Gardner

arxiv 2023. [PDF] [CODE]

Adversarial Attacks on Image Generation With Made-Up Words

Raphaël Millière

arxiv 2022. [PDF]

SurrogatePrompt: Bypassing the Safety Filter of Text-To-Image Models via Substitution

Zhongjie Ba, Jieming Zhong, Jiachen Lei, Peng Cheng, Qinglong Wang, Zhan Qin, Zhibo Wang, Kui Ren

ACM CCS 2024. [PDF]

RT-Attack: Jailbreaking Text-to-Image Models via Random Token

Sensen Gao, Xiaojun Jia, Yihao Huang, Ranjie Duan, Jindong Gu, Yang Liu, Qing Guo

arxiv 2024. [PDF]

Perception-guided Jailbreak against Text-to-Image Models

Yihao Huang, Le Liang, Tianlin Li, Xiaojun Jia, Run Wang, Weikai Miao, Geguang Pu, and Yang Liu

AAAI 2024. [PDF]

DiffZOO: A Purely Query-Based Black-Box Attack for red-teaming Text-to-Image Generative Model via Zeroth Order Optimization

Pucheng Dang, Xing Hu, Dong Li, Rui Zhang, Kaidi Xu, Qi Guo

NAACL 2025. [PDF]

Reason2attack: Jailbreaking text-to-image models via llm reasoning

Chenyu Zhang, Lanjun Wang, Yiwen Ma, Wenhui Li, An-An Liu

AAAI 2026. [PDF]

GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models

Zilong Wang, Xiang Zheng, Xiaosen Wang, Bo Wang, Xingjun Ma, Yu-Gang Jiang

Arxiv 2025. [PDF]

Modifier Unlocked: Jailbreaking Text-to-Image Models Through Prompts

Shuofeng Liu; Mengyao Ma; Minhui Xue; Guangdong Bai

S&P 2025. [PDF]

PLA: Prompt Learning Attack against Text-to-Image Generative Models

Xinqi Lyu, Yihao Liu, Yanjie Li, Bin Xiao

ICCV 2025. [PDF]

Red-Teaming Text-to-Image Systems by Rule-based Preference Modeling

Yichuan Cao, Yibo Miao, Xiao-Shan Gao, Yinpeng Dong

NeurIPS 2025. [PDF]

:pill:Defenses for Improving Safety

:surfer:External Safeguards

:mountain_bicyclist:Prompt Classifier

Latent Guard: a Safety Framework for Text-to-image Generation

Runtao Liu, Ashkan Khakzar, Jindong Gu, Qifeng Chen, Philip Torr, Fabio Pizzati

ECCV 2024. [PDF] [CODE]

AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image Models

Yiming Wang, Jiahao Chen, Qingming Li, Xing Yang, and Shoulin Ji

arxiv 2024. [PDF]

:horse_racing:Prompt Transformation

Universal Prompt Optimizer for Safe Text-to-Image Generation

Zongyu Wu, Hongcheng Gao, Yueze Wang, Xiang Zhang, Suhang Wang

NAACL 2024. [PDF]

GuardT2I: Defending Text-to-Image Models from Adversarial Prompts

Yijun Yang, Ruiyuan Gao, Xiao Yang, Jianyuan Zhong, Qiang Xu

NeurIPS 2024. [PDF]

:hamburger:Internal Safeguards

:fries:Model Editing

Erasing concepts from diffusion models

Gandikota, Rohit and Materzynska, Joanna and Fiotto-Kaufman, Jaden and Bau, David

ICCV 2023. [PDF] [CODE]

Ablating concepts in text-to-image diffusion models

Kumari, Nupur and Zhang, Bingliang and Wang, Sheng-Yu and Shechtman, Eli and Zhang, Richard and Zhu, Jun-Yan

ICCV 2023. [PDF] [CODE]

Unified concept editing in diffusion models

Gandikota, Rohit and Orgad, Hadas and Belinkov, Yonatan and Materzy{'n}ska, Joanna and Bau, David

WACV 2024. [PDF] [CODE]

Editing implicit assumptions in text-to-image diffusion models

Orgad, Hadas and Kawar, Bahjat and Belinkov, Yonatan

ICCV 2023. [PDF] [CODE]

Towards Safe Self-Distillation of Internet-Scale Text-to-Image Diffusion Models

Sanghyun Kim, Seohyeon Jung, Balhae Kim, Moonseok Choi, Jinwoo Shin, Juho Lee

ICML 2023 Workshop on Challenges in Deployable Generative AI. [PDF] [CODE]

Degeneration-tuning: Using scrambled grid shield unwanted concepts from stable diffusion

Ni, Zixuan and Wei, Longhui and Li, Jiacheng and Tang, Siliang and Zhuang, Yueting and Tian, Qi

ACM MM 2023. [PDF]

ReFACT: Updating Text-to-Image Models by Editing the Text Encoder

Dana Arad, Hadas Orgad, Yonatan Belinkov

NAACL 2024. [PDF]

OnMechanistic Knowledge Localization in Text-to-Image Generative Models

Samyadeep Basu, Keivan Rezaei, Priyatham Kattakinda, Vlad I Morariu, Nanxuan Zhao, Ryan A. Rossi, Varun Manjunatha, Soheil Feizi

ICML 2024. [PDF]

Forget-Me-Not: Learning to Forget in Text-to-Image Diffusion Models

Gong Zhang, Kai Wang, Xingqian Xu, Zhangyang Wang, Humphrey Shi

CVPR 2024 [PDF] [CODE]

Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models

Senmao Li, Joost van de Weijer, Taihang Hu, Fahad Shahbaz Khan, Qibin Hou, Yaxing Wang, Jian Yang

ICLR 2024. [PDF] [CODE]

Selective Amnesia: A Continual Learning Approach to Forgetting in Deep Generative Models

Alvin Heng , Harold Soh

NeurIPS 2024, [PDF] [CODE]

Leveraging Catastrophic Forgetting to Develop Safe Diffusion Models against Malicious Finetuning

Jiadong Pan, Hongcheng Gao, Zongyu Wu, Taihang Hu, Li Su, Qingming Huang, Liang Li

NeurIPS 2024. [PDF] [CODE]

Boosting Alignment for Post-Unlearning Text-to-Image Generative Models

Myeongseob Ko, Henry Li, Zhun Wang, Jonathan Patsenker, Jiachen T. Wang, Qinbin Li, Ming Jin, Dawn Song, Ruoxi Jia

NeurIPS 2024. [PDF] [CODE]

Erasing Undesirable Concepts in Diffusion Models with Adversarial Preservation

Anh Bui, Long Vuong, Khanh Doan, Trung Le, Paul Montague, Tamas Abraham, Dinh Phung

NeurIPS 2024. [PDF] [CODE]

Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion Models

Yimeng Zhang, Xin Chen, Jinghan Jia, Yihua Zhang, Chongyu Fan, Jiancheng Liu, Mingyi Hong, Ke Ding, Sijia Liu

NeurIPS 2024. [PDF] [CODE]

All but One: Surgical Concept Erasing with Model Preservation in Text-to-Image Diffusion Models

Hong, Seunghoo and Lee, Juhun and Woo, Simon S

AAAI 2024. [PDF]

SafeGen: Mitigating Unsafe Content Generation in Text-to-Image Models

Xinfeng Li , Yuchen Yang , Jiangyi Deng, Chen Yan , Yanjiao Chen , Xiaoyu Ji , Wenyuan Xu

ACM CCS 2024. [PDF] [CODE]

Direct Unlearning Optimization for Robust and Safe Text-to-Image Models

Yong-Hyun Park, Sangdoo Yun, Jin-Hwa Kim, Junho Kim, Geonhui Jang, Yonghyun Jeong, Junghyo Jo, Gayoung Lee

ICML 2024 Workshop. [PDF]

One-dimensional adapter to rule them all: Concepts diffusion models and erasing applications

Mengyao Lyu, Yuhong Yang, Haiwen Hong, Hui Chen, Xuan Jin, Yuan He, Hui Xue, Jungong Han, Guiguang Ding

CVPR 2024 [PDF] [CODE]

Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models

Poppi, Samuele and Poppi, Tobia and Cocchi, Federico and Cornia, Marcella and Baraldi, Lorenzo and Cucchiara, Rita

ECCV 2024. [PDF] [CODE]

Unlearning Concepts in Diffusion Model via Concept Domain Correction and Concept Preserving Gradient

Yongliang Wu, Shiji Zhou, Mingzhuo Yang, Lianzhe Wang, Wenbo Zhu, Heng Chang, Xiao Zhou, Xu Yang

AAAI 2025. [PDF]

R.A.C.E.: Robust Adversarial Concept Erasure for Secure Text-to-Image Diffusion Model

Changhoon Kim, Kyle Min, Yezhou Yang

ECCV 2024. [PDF] [CODE]

Receler: Reliable Concept Erasing of Text-to-Image Diffusion Models via Lightweight Erasers

Chi-Pin Huang, Kai-Po Chang, Chung-Ting Tsai, Yung-Hsuan Lai, Fu-En Yang, Yu-Chiang Frank Wang

ECCV 2024. [PDF] [CODE]

Reliable and efficient concept erasure of text-to-image diffusion models

Chao Gong, Kai Chen, Zhipeng Wei, Jingjing Chen, and Yu-Gang Jiang

ECCV 2024. [PDF] [CODE]

Safeguard Text-to-Image Diffusion Models with Human Feedback Inversion

Sanghyun Kim, Seohyeon Jung, Balhae Kim, Moonseok Choi, Jinwoo Shin & Juho Lee

ECCV 2024. [PDF] [CODE]

Score Forgetting Distillation: A Swift, Data-Free Method for Machine Unlearning in Diffusion Models

Tianqi Chen, Shujian Zhang, Mingyuan Zhou

ICLR 2025. [PDF] [CODE]

Fantastic Targets for Concept Erasure in Diffusion Models and Where To Find Them

Anh Bui, Trang Vu, Long Vuong, Trung Le, Paul Montague, Tamas Abraham, Junae Kim, Dinh Phung

ICLR 2025. [PDF] [CODE]

Editing Massive Concepts in Text-to-Image Diffusion Models

Tianwei Xiong, Yue Wu, Enze Xie, Yue Wu, Zhenguo Li, Xihui Liu

arxiv 2024. [PDF] [CODE]

ShieldDiff: Suppressing Sexual Content Generation from Diffusion Models through Reinforcement Learning

Dong Han, Salaheldin Mohamed, Yong Li

arxiv 2024. [PDF]

Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned Concepts

Hongcheng Gao, Tianyu Pang, Chao Du, Taihang Hu, Zhijie Deng, Min Lin

ICCV 2025. [PDF] [CODE]

Dark Miner: Defend against undesired generation for text-to-image diffusion models

Zheling Meng, Bo Peng, Xiaochuan Jin, Yue Jiang, Jing Dong, Wei Wang

arxiv 2024. [PDF]

SafetyDPO: Scalable Safety Alignment for Text-to-Image Generation

Runtao Liu, Chen I Chieh, Jindong Gu, Jipeng Zhang, Renjie Pi, Qifeng Chen, Philip Torr, Ashkan Khakzar, Fabio Pizzati

arxiv 2024. [PDF] [CODE]

Continuous Concepts Removal in Text-to-image Diffusion Models

Tingxu Han, Weisong Sun, Yanrong Hu, Chunrong Fang, Yonglong Zhang, Shiqing Ma, Tao Zheng, Zhenyu Chen, Zhenting Wang

arxiv 2025. [PDF]

SPEED: Scalable, Precise, and Efficient Concept Erasure for Diffusion Models

Ouxiang Li, Yuan Wang, Xinting Hu, Houcheng Jiang, Tao Liang, Yanbin Hao, Guojun Ma, Fuli Feng

arxiv 2025. [PDF] [CODE]

Continual Unlearning for Foundational Text-to-Image Models without Generalization Erosion

Kartik Thakral, Tamar Glaser, Tal Hassner, Mayank Vatsa, Richa Singh

arxiv 2025. [PDF]

Sparse Autoencoder as a Zero-Shot Classifier for Concept Erasing in Text-to-Image Diffusion Models

Zhihua Tian, Sirun Nan, Ming Xu, Shengfang Zhai, Wenjie Qu, Jian Liu, Kui Ren, Ruoxi Jia, Jiaheng Zhang

arxiv 2025. [PDF]

SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders

Bartosz Cywiński, Kamil Deja

arxiv 2025. [PDF] [CODE]

CRCE: Coreference-Retention Concept Erasure in Text-to-Image Diffusion Models

Yuyang Xue, Edward Moroshko, Feng Chen, Steven McDonagh, Sotirios A. Tsaftaris

arxiv 2025. [PDF]

TRCE: Towards Reliable Malicious Concept Erasure in Text-to-Image Diffusion Models

Ruidong Chen, Honglin Guo, Lanjun Wang, Chenyu Zhang, Weizhi Nie, An-An Liu

ICCV 2025. [PDF] [CODE]

Concept pinpoint eraser for text-to-image diffusion models via residual attention gate

Byung Hyun Lee, Sungjin Lim, Seunggyu Lee, Dong Un Kang, Se Young Chun

ICLR 2025. [PDF] [CODE]

Concept replacer: Replacing sensitive concepts in diffusion models via precision localization

Lingyun Zhang, Yu Xie, Yanwei Fu, Ping Chen

CVPR 2025. [PDF]

CURE: Concept unlearning via orthogonal representation editing in diffusion models

Tingxu Han, Weisong Sun, Yanrong Hu, Chunrong Fang, Yonglong Zhang, Shiqing Ma, Tao Zheng, Zhenyu Chen, Zhenting Wang

NeurIPS 2025. [PDF]

EraseAnything: Enabling concept erasure in rectified flow transformers

D Gao, S Lu, W Zhou, J Chu, J Zhang, M Jia, B Zhang, Z Fan, W Zhang

ICML 2025. [PDF] [CODE]

Fine-grained erasure in text-to-image diffusion-based foundation models

Kartik Thakral, Tamar Glaser, Tal Hassner, Mayank Vatsa, Richa Singh

CVPR 2025 [PDF]

Precise, fast, and low-cost concept erasure in value space: Orthogonal complement matters

Y Wang, O Li, T Mu, Y Hao, K Liu, X Wang, X He

CVPR 2025. [PDF]

SafeGen: Mitigating sexually explicit content generation in text-to-image models

Xinfeng Li, Yuchen Yang, Jiangyi Deng, Chen Yan, Yanjiao Chen, Xiaoyu Ji, Wenyuan Xu

ACM CCS 2025. [PDF] [CODE]

SafeGuider: Robust and practical content safety control for text-to-image models

P Qi, K Tang, W Zhou, W Zhang, N Yu, T Zhang, Q Guo, J Zhang

ACM CCS 2025. [PDF] [CODE]

:apple:Inference Guidance

Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models

Patrick Schramowski, Manuel Brack, Björn Deiseroth, Kristian Kersting

CVPR 2023. [PDF] [CODE]

Sega: Instructing text-to-image models using semantic guidance

Manuel Brack, Felix Friedrich, Dominik Hintersdorf, Lukas Struppek, Patrick Schramowski, Kristian Kersting

NeurIPS 2023. [PDF] [CODE]

Self-discovering interpretable diffusion latent directions for responsible text-to-image generation

Li, Hang and Shen, Chengzhi and Torr, Philip and Tresp, Volker and Gu, Jindong

CVPR 2024. [PDF] [CODE]

One-dimensional Adapter to Rule Them All: Concepts, Diffusion Models and Erasing Applications

Mengyao Lyu, Yuhong Yang, Haiwen Hong, Hui Chen, Xuan Jin, Yuan He, Hui Xue, Jungong Han, Guiguang Ding

CVPR 2024. [PDF] [CODE]

Safree: Training-free and adaptive guard for safe text-to-image and video generation

Jaehong Yoon, Shoubin Yu, Vaidehi Patil, Huaxiu Yao, Mohit Bansal

ICLR 2025. [PDF] [CODE]

Pruning for Robust Concept Erasing in Diffusion Models

Tianyun Yang, Juan Cao, Chang Xu

arxiv 2024. [PDF]

ConceptPrune: Concept editing in diffusion models via skilled neuron pruning

Ruchika Chavhan, Da Li, Timothy Hospedales

arxiv 2024. [PDF]

TraSCE: Trajectory Steering for Concept Erasure

Anubhav Jain, Yuya Kobayashi, Takashi Shibuya, Yuhta Takida, Nasir Memon, Julian Togelius, Yuki Mitsufuji

arxiv 2025. [PDF] [CODE]

Growth Inhibitors for Suppressing Inappropriate Image Concepts in Diffusion Models

Die Chen, Zhiwen Li, Mingyuan Fan, Cen Chen, Wenmeng Zhou, Yanhao Wang, Yaliang Li

arxiv 2025. [PDF]

Training-Free Safe Denoisers for Safe Use of Diffusion Models

Mingyu Kim, Dongjun Kim, Amman Yusuf, Stefano Ermon, Mi Jung Park

NeurIPS 2025. [PDF]

PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models

Lingzhi Yuan, Xinfeng Li, Chejian Xu, Guanhong Tao, Xiaojun Jia, Yihao Huang, Wei Dong, Yang Liu, XiaoFeng Wang, Bo Li

arxiv 2025. [PDF]

Concept Corrector: Erase concepts on the fly for text-to-image diffusion models

Zheling Meng, Bo Peng, Xiaochuan Jin, Yueming Lyu, Wei Wang, Jing Dong

arxiv 2025. [PDF]

Localized Concept Erasure for Text-to-Image Diffusion Models Using Training-Free Gated Low-Rank Adaptation

Byung Hyun Lee, Sungjin Lim, Se Young Chun

CVPR 2025. [PDF]

Resources

This part provides commonly used datasets and tools in AD-on-T2IDM.

Datasets

Based on the prompt source, existing datasets are categorized into two types: clean and adversarial datasets. The clean dataset consists of clean prompts that are not attacked and typically crafted by human, while the adversarial dataset comprises adversarial prompts generated by attack methods. Moreover, according to the category of prompts involved in the dataset, existing clean datasets are further divided into two types: non-malicious and malicious datasets. The non-malicious dataset contains non-malicious prompts, while the malicious dataset contains explicitly malicious prompts. In this section, we will introduce several non-malicious, malicious, and adversarial datasets, respectively.

Non-Malicious Datasets

  • $\textit{ImageNet}$, which contains images describing 1,000 categories of common objects in the real world, is a significant benchmark in the field of computer vision. As a result, some works craft clean datasets based on the category information in ImageNet. For instance, ATM employs a standardized template: "A photo of {CLASS_NAME}" to generate clean prompts, where "{CLASS_NAME}" denotes the class name in ImageNet.
  • $\textit{MSCOCO}$ [Link]is a cross-modal image-text dataset, a popular benchmark for training and evaluating text-to-image generation models. Specifically, MSCOCO includes 82,783 training images and 40,504 testing images, each with 5 text descriptions.
  • $\textit{LAION-COCO}$ [Link] is a subset of LAION-5B, which is a large-scale image-text dataset in the real world. LAION-COCO includes 600 million images and corresponding text descriptions.
  • $\textit{DiffusionDB}$ [Link] is a large-scale text-to-image prompt dataset, which contains 14 million images generated by Stable Diffusion using prompts from real users.
  • $\textit{VidproM}$[LINK] is a large-scale dataset comprising 1.67 million unique text-to-video prompts from real users, where each prompt is flagged with the probabilities across six aspects of NSFW content, including toxicity, obscenity, identity attack, insult, threat, and sexual explicitness. (NeurIPS 2024 Datasets and Benchmarks Track)

Malicious Datasets

  • $\textit{Unsafe Diffusion}$ [Link] provides 30 manually crafted malicious prompts that describe sexual and bloody content, as well as political figures.
  • $\textit{SneakyPrompt}$ [Link] uses ChatGPT to automatically generate 200 malicious prompts that involve sexual and bloody content.
  • $\textit{I2P}$ [Link] comprises 4,703 malicious prompts, encompassing hate, harassment, violence, self-harm, nudity content, shocking images, and illegal activity. These inappropriate prompts are real-user inputs sourced from an image generation website, Lexica [Link].
  • $\textit{MMA}$ [Link] samples and releases 1,000 malicious prompts from LAION-COCO based on an NSFW~(Not Safe for Work) score. These malicious prompts mainly focus on sexual content.
  • $ART$[Link] follows I2P and collects 15,607 malicious prompts from 7 categories in Lexica [Link].
  • $\textit{ViSU}$ [Link] contains 175k pairs of safe and unsafe data examples. Each example consists of: (1) a safe sentence, (2) a corresponding safe image, (3) an NSFW sentence that is semantically correlated with the safe sentence, and (4) a corresponding NSFW image.
  • $\textit{T2VSafetyBench}$[LINK] contains 4,400 NSFW prompts in 12 classes: Pornography, Borderline Pornography, Violence, Gore, Public Figures, Discrimination, Political Sensitivity, Illegal Activities, Disturbing Content, Misinformation and Falsehoods, Copyright and Trademark Infringement, and Temporal Risk. These prompts are collected from three sources: VidProM[LINK], LLMs, and attack methods. (NeurIPS 2024 Datasets and Benchmarks Track)
  • $\textit{SAFESORA}$[LINK] contains 14,711 unique text prompts for the text-to-video models, with 48.61% potentially inducing harmful videos. Furthermore, these harmful prompts are categorized into 12 classes, including explicit sexual content, animal abuse, child abuse, crime, debated sensitive issue, drug abuse, hateful behavior, violence, racial discrimination, other discrimination, terrorism, and other harmful content. (NeurIPS 2024 Datasets and Benchmarks Track)
  • $\textit{Image Synthesis Style Studies Database}$ [Link] compiles thousands of artists whose styles can be replicated by various text-to-image models, such as Stable Diffusion and Midjourney.
  • $\textit{MACE}$ [Link] provides a dataset comprising 200 celebrities whose portraits, generated using SD v1.4, are recognized with remarkable accuracy (>99%) by the GIPHY Celebrity Detector (GCD) [Link].
  • $\textit{CPDM}$[LINK] is a copyright protection dataset for T2I models, containing 2,100 prompts in 4 classes: style, portrait, artistic creation figure, and licensed illustration.
  • $\textit{T2I-RiskyPrompt}$ [LINK] provides a hierarchical risky taxonomy~(6 risky categories and 14 subcategories) and 6,432 risky prompts labeled with hierarchical labels and detailed risky reasons. Moreover, it also provides a risky image detector used to evaluate the generated images of T2I models using T2I-RiskyPrompt.

Tools

We provide several detectors for detecting malicious prompts and images.

Malicious Prompt Detector

  • NSFW_text_classifier: [Link]

  • distilbert-nsfw-text-classifier: [Link]

  • Detoxify: [Link]

  • Toxic-comment-model: [Link]

  • Meta-Llama-Guard: [Link] (LLM evaluation)

  • Openai-Moderation: [Link] (API)

  • Azure-Moderation: [Link] (API)

Malicious Image Detector

Significant stargazers

Elwood Zonghao Ying

26 followers · starred Apr 2025

datar001/Awesome-AD-on-T2IDM

A collection of resources on attacks and defenses targeting text-to-image diffusion models

103

36 commits

updated Dec 20, 2025

See the code

README

Awesome-Attacks and Defenses on T2I Diffusion Models

This repository is a curated collection of research papers focused on $\textbf{Adversarial Attacks and Defenses on Text-to-Image Diffusion Models (AD-on-T2IDM)}$.

We will continuously update this collection to track the latest advancements in the field of AD-on-T2IDM.

Welcome to follow and star! If you have any relevant materials or suggestions, please feel free to contact us (zcy@tju.edu.cn) or submit a pull request.

For more detailed information, please refer to our survey paper: [ARXIV], [Published Version]

News

  • We are pleased to share that our latest work, T2I-RiskyPrompt, has been accepted to AAAI 2026! 🎊As Text-to-Image models evolve, safety becomes paramount. T2I-RiskyPrompt bridges the gap in current safety evaluations with:
    • ✅ Hierarchical Taxonomy: Covers 6 major risky categories & 14 fine-grained subcategories.
    • ✅ Rich Annotations: 6,432 prompts labeled with both hierarchical risky categories and detailed risky reasons
    • ✅ Built-in Detector: An automated pipeline to evaluate the safety of generated images.
    • 🔗 Code & Dataset: [Github], 📄 Paper: [Arxiv]

Citation

@article{zhang2025adversarial,
  title={Adversarial attacks and defenses on text-to-image diffusion models: A survey},
  author={Zhang, Chenyu and Hu, Mingwang and Li, Wenhui and Wang, Lanjun},
  journal={Information Fusion},
  volume={114},
  pages={102701},
  year={2025},
  publisher={Elsevier}
}

Recruiting

Looking for prospective PhD students (enroll in 2026). Ideal candidates are expected to be responsible, organized, and creative. Strong programming skills, solid mathematics knowledge as well as good communication skills are preferred. Please send your CV to wanglanjun AT tju.edu.cn. Note: please attach files directly, not to use 163 Cloud file system.

Content

Abstract

Recently, the text-to-image diffusion model has gained considerable attention from the community due to its exceptional image generation capability. A representative model, Stable Diffusion, amassed more than 10 million users within just two months of its release. This surge in popularity has facilitated studies on the robustness and safety of the model, leading to the proposal of various adversarial attack methods. Simultaneously, there has been a marked increase in research focused on defense methods to improve the robustness and safety of these models. In this survey, we provide a comprehensive review of the literature on adversarial attacks and defenses targeting text-to-image diffusion models. We begin with an overview of popular text-to-image diffusion models, followed by an introduction to a taxonomy of adversarial attacks and an in-depth review of existing attack methods. We then present a detailed analysis of current defense methods that improve model robustness and safety. Finally, we discuss ongoing challenges and explore promising future research directions.

Overview of AD-on-T2IDM

Two key concerns in T2IDM: Robustness and Safety

The robustness ensures that the model can generate images with consistent semantics in response to diverse prompts inputted by users in practice.

The safety prevents the misuse of the model in creating malicious images, such as sexual, violent, and political images, etc.

Adversarial attacks

Based on the intent of the adversary, existing attack methods can be divided into two primary categories: untargeted and targeted attacks.

  • For untargeted attacks, consider a scenario with a prompt input by the user~($\textbf{clean prompt}$) and its corresponding output image~($\textbf{clean image}$). The objective of untargeted attacks is to subtly perturb the clean prompt to craft an $\textbf{adversarial prompt}$, further misleading the victim model to generate an $\textbf{adversarial image}$ with semantics different from the clean image. This type of attack is commonly used to uncover the vulnerability in the robustness of the victim model. Some untargeted attacks are shown as follows:

    untargeted attacks

  • For targeted attacks, assumes that the victim model has built-in $\textbf{safeguards}$ to filter $\textbf{malicious prompts}$ and resultant $\textbf{malicious images}$. These prompts and images often explicitly contain $\textbf{malicious concepts}$, such as 'nudity', 'violence', and other predefined concepts. The objective of targeted attacks is to obtain an $\textbf{adversarial prompt}$, which can bypass these safeguards while inducing the victim model to generate $\textbf{adversarial images}$ containing malicious concepts. This type of attack is typically designed to reveal the vulnerability in the safety of the victim model. Some targeted attacks are shown as follows:

    targeted attacks

Defenses

Based on the defense goal, existing defense methods can be classified into two categories: 1) improving model robustness and 2) improving model safety.

  • The goal of robustness is to ensure that generated images have consistent semantics with diverse input prompts in practical applications. Specifically, according to the adversarial attack, the defense methods are asked to mitigate the robustness vulnerabilities in two types of input prompts: 1) the prompt with multiple objects and attributes, and 2) the grammatically incorrect prompt with the subtle noise.

  • The safety goal is to prevent the generation of malicious images in response to both malicious and adversarial prompts. Specifically, malicious prompts explicitly contain malicious concepts, while adversarial prompts cleverly omit these concepts. Moreover, based on the knowledge of the model, existing safety methods can be classified into two categories: external safeguards and internal safeguards. The external safeguards focus on detecting or correcting the malicious prompt before feeding the prompt into the text-to-image model. In contrast, internal safeguards aim to ensure that the semantics of output images deviate from those of malicious images by modifying internal parameters and features within the model. Some examples of external and internal safeguards are shown as follows:

    external safeguards internal safeguards

Notably, although many methods are proposed to improve the model robustness against the prompt with multiple objects and attributes, this collection omits related papers on this part since there has been related surveys, such as controllable image generation [PDF], the development and advancement of image generation capabilities [PDF-1], [PDF-2], [PDF-3]. Moreover, for grammatically incorrect prompts with subtle noise, mature solutions are still lacking. Therefore, this collection mainly focuses on the defense methods for improving model safety.

:grinning:Paper List

:imp:Adversarial Attacks

:collision:Untargeted Attacks

:pouting_cat:White-Box Attacks

Stable diffusion is unstable

Chengbin Du, Yanxi Li, Zhongwei Qiu, Chang Xu

NeurIPS 2024. [PDF] [CODE]

A pilot study of query-free adversarial attack against stable diffusion

Haomin Zhuang, Yihua Zhang

CVPRW 2023. [PDF] [CODE]

:see_no_evil:Black-Box Attacks

Evaluating the Robustness of Text-to-image Diffusion Models against Real-world Attacks

Hongcheng Gao , Hao Zhang , Yinpeng Dong, Zhijie Deng

arxiv 2023. [PDF]

:anger:Targeted Attacks

:cyclone:White-Box Attacks

Red-Teaming the Stable Diffusion Safety Filter

Javier Rando, Daniel Paleka, David Lindner, Lennart Heim, Florian Tramèr

NeurIPS 2022, WorkShop. [PDF]

Unsafe Diffusion: On the Generation of Unsafe Images and Hateful Memes From Text-To-Image Models

Yiting Qu, Xinyue Shen, Xinlei He, Michael Backes, Savvas Zannettou, Yang Zhang

Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security. [PDF] [CODE]

Ring-A-Bell! How Reliable are Concept Removal Methods for Diffusion Models?

Tsai, Yu-Lin and Hsu, Chia-Yi and Xie, Chulin and Lin, Chih-Hsun and Chen, Jia-You and Li, Bo and Chen, Pin-Yu and Yu, Chia-Mu and Huang, Chun-Ying

ICLR 2024. [PDF]

Riatig: Reliable and imperceptible adversarial text-to-image generation with natural prompts

Han Liu, Yuhao Wu, Shixuan Zhai, Bo Yuan, Ning Zhang

CVPR 2023. [PDF] [CODE]

Mma-diffusion: Multimodal attack on diffusion models

Yang, Yijun and Gao, Ruiyuan and Wang, Xiaosen and Ho, Tsung-Yi and Xu, Nan and Xu, Qiang

CVPR 2024. [PDF] [CODE]

Asymmetric Bias in Text-to-Image Generation with Adversarial Attacks

Haz Sameen Shahgir, Xianghao Kong, Greg Ver Steeg, Yue Dong

arxiv 2023. [PDF] [CODE]

Revealing vulnerabilities in stable diffusion via targeted attacks

Chenyu Zhang, Lanjun Wang, Anan Liu

arxiv 2024. [PDF] [CODE]

To generate or not? safety-driven unlearned diffusion models are still easy to generate unsafe images... for now

Zhang, Yimeng and Jia, Jinghan and Chen, Xin and Chen, Aochuan and Zhang, Yihua and Liu, Jiancheng and Ding, Ke and Liu, Sijia

ECCV 2024. [PDF] [CODE]

Prompting4debugging: Red-teaming text-to-image diffusion models by finding problematic prompts

Chin, Zhi-Yi and Jiang, Chieh-Ming and Huang, Ching-Chun and Chen, Pin-Yu and Chiu, Wei-Chen

ICML 2024. [PDF] [CODE]

ADVI2I: ADVERSARIAL IMAGE ATTACK ON IMAGE-TO-IMAGE DIFFUSION MODELS

Yaopei Zeng, Yuanpu Cao, Bochuan Cao, Yurui Chang, Jinghui Chen, Lu Lin

arxiv 2024. [PDF] [CODE]

Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models

Jiachen Ma, Anda Cao, Zhiqing Xiao, Jie Zhang, Chao Ye, Junbo Zhao

arxiv 2024. [PDF]

Adversarial Attacks on Parts of Speech: An Empirical Study in Text-to-Image Generation

G M Shahariar, Jia Chen, Jiachen Li, Yue Dong

EMNLP 2024. [PDF], [CODE]

:snake:Black-Box Attacks

SneakyPrompt: Evaluating Robustness of Text-to-image Generative Models' Safety Filters

Yuchen Yang, Bo Hui, Haolin Yuan, Neil Gong, Yinzhi Cao

Proceedings of the IEEE Symposium on Security and Privacy 2024. [PDF] [CODE]

ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign Users

Guanlin Li, Kangjie Chen, Shudong Zhang, Jie Zhang, Tianwei Zhang

NeurIPS 2024. [PDF] [CODE]

Coljailbreak: Collaborative generation and editing for jailbreaking text-to-image deep generation

Yizhuo Ma, Shanmin Pang, Qi Guo, Tianyu Wei, Qing Guo

NeurIPS 2024. [PDF]

UPAM: Unified Prompt Attack in Text-to-Image Generation Models Against Both Textual Filters and Visual Checkers

Duo Peng, Qiuhong Ke, Jun Liu

ICML2024, [PDF]

FLIRT: Feedback Loop In-context Red Teaming

Ninareh Mehrabi, Palash Goyal, Christophe Dupuy, Qian Hu, Shalini Ghosh, Richard Zemel, Kai-Wei Chang, Aram Galstyan, Rahul Gupta

EMNLP 2024. [PDF]

Fuzz-testing meets LLM-based agents: An automated and efficient framework for jailbreaking text-to-image generation models

Yingkai Dong, Zheng Li, Xiangtao Meng, Ning Yu, Shanqing Guo

S&P 2025. [PDF], [[CODE]](GitHub - YingkaiD/JailFuzzer)

Automatic Jailbreaking of the Text-to-Image Generative AI Systems

Minseon Kim, Hyomin Lee, Boqing Gong, Huishuai Zhang, Sung Ju Hwang

arxiv 2024. [PDF] [CODE]

Exploiting cultural biases via homoglyphs in text-to-image synthesis

Struppek, Lukas and Hintersdorf, Dom and Friedrich, Felix and Schramowski, Patrick and Kersting, Kristian

Journal of Artificial Intelligence Research 2023. [PDF] [CODE]

Divide-and-Conquer Attack: Harnessing the Power of LLM to Bypass Safety Filters of Text-to-Image Models

Yimo Deng, Huangxun Chen

arxiv 2024. [PDF]

Groot: Adversarial Testing for Generative Text-to-Image Models with Tree-based Semantic Transformation

Yi Liu, Guowei Yang, Gelei Deng, Feiyue Chen, Yuqi Chen, Ling Shi, Tianwei Zhang, Yang Liu

arxiv 2024. [PDF]

BSPA: Exploring Black-box Stealthy Prompt Attacks against Image Generators

Yu Tian, Xiao Yang, Yinpeng Dong, Heming Yang, Hang Su, Jun Zhu

arxiv 2024. [PDF]

Black Box Adversarial Prompting for Foundation Models

Natalie Maus, Patrick Chao, Eric Wong, Jacob Gardner

arxiv 2023. [PDF] [CODE]

Adversarial Attacks on Image Generation With Made-Up Words

Raphaël Millière

arxiv 2022. [PDF]

SurrogatePrompt: Bypassing the Safety Filter of Text-To-Image Models via Substitution

Zhongjie Ba, Jieming Zhong, Jiachen Lei, Peng Cheng, Qinglong Wang, Zhan Qin, Zhibo Wang, Kui Ren

ACM CCS 2024. [PDF]

RT-Attack: Jailbreaking Text-to-Image Models via Random Token

Sensen Gao, Xiaojun Jia, Yihao Huang, Ranjie Duan, Jindong Gu, Yang Liu, Qing Guo

arxiv 2024. [PDF]

Perception-guided Jailbreak against Text-to-Image Models

Yihao Huang, Le Liang, Tianlin Li, Xiaojun Jia, Run Wang, Weikai Miao, Geguang Pu, and Yang Liu

AAAI 2024. [PDF]

DiffZOO: A Purely Query-Based Black-Box Attack for red-teaming Text-to-Image Generative Model via Zeroth Order Optimization

Pucheng Dang, Xing Hu, Dong Li, Rui Zhang, Kaidi Xu, Qi Guo

NAACL 2025. [PDF]

Reason2attack: Jailbreaking text-to-image models via llm reasoning

Chenyu Zhang, Lanjun Wang, Yiwen Ma, Wenhui Li, An-An Liu

AAAI 2026. [PDF]

GenBreak: Red Teaming Text-to-Image Generators Using Large Language Models

Zilong Wang, Xiang Zheng, Xiaosen Wang, Bo Wang, Xingjun Ma, Yu-Gang Jiang

Arxiv 2025. [PDF]

Modifier Unlocked: Jailbreaking Text-to-Image Models Through Prompts

Shuofeng Liu; Mengyao Ma; Minhui Xue; Guangdong Bai

S&P 2025. [PDF]

PLA: Prompt Learning Attack against Text-to-Image Generative Models

Xinqi Lyu, Yihao Liu, Yanjie Li, Bin Xiao

ICCV 2025. [PDF]

Red-Teaming Text-to-Image Systems by Rule-based Preference Modeling

Yichuan Cao, Yibo Miao, Xiao-Shan Gao, Yinpeng Dong

NeurIPS 2025. [PDF]

:pill:Defenses for Improving Safety

:surfer:External Safeguards

:mountain_bicyclist:Prompt Classifier

Latent Guard: a Safety Framework for Text-to-image Generation

Runtao Liu, Ashkan Khakzar, Jindong Gu, Qifeng Chen, Philip Torr, Fabio Pizzati

ECCV 2024. [PDF] [CODE]

AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image Models

Yiming Wang, Jiahao Chen, Qingming Li, Xing Yang, and Shoulin Ji

arxiv 2024. [PDF]

:horse_racing:Prompt Transformation

Universal Prompt Optimizer for Safe Text-to-Image Generation

Zongyu Wu, Hongcheng Gao, Yueze Wang, Xiang Zhang, Suhang Wang

NAACL 2024. [PDF]

GuardT2I: Defending Text-to-Image Models from Adversarial Prompts

Yijun Yang, Ruiyuan Gao, Xiao Yang, Jianyuan Zhong, Qiang Xu

NeurIPS 2024. [PDF]

:hamburger:Internal Safeguards

:fries:Model Editing

Erasing concepts from diffusion models

Gandikota, Rohit and Materzynska, Joanna and Fiotto-Kaufman, Jaden and Bau, David

ICCV 2023. [PDF] [CODE]

Ablating concepts in text-to-image diffusion models

Kumari, Nupur and Zhang, Bingliang and Wang, Sheng-Yu and Shechtman, Eli and Zhang, Richard and Zhu, Jun-Yan

ICCV 2023. [PDF] [CODE]

Unified concept editing in diffusion models

Gandikota, Rohit and Orgad, Hadas and Belinkov, Yonatan and Materzy{'n}ska, Joanna and Bau, David

WACV 2024. [PDF] [CODE]

Editing implicit assumptions in text-to-image diffusion models

Orgad, Hadas and Kawar, Bahjat and Belinkov, Yonatan

ICCV 2023. [PDF] [CODE]

Towards Safe Self-Distillation of Internet-Scale Text-to-Image Diffusion Models

Sanghyun Kim, Seohyeon Jung, Balhae Kim, Moonseok Choi, Jinwoo Shin, Juho Lee

ICML 2023 Workshop on Challenges in Deployable Generative AI. [PDF] [CODE]

Degeneration-tuning: Using scrambled grid shield unwanted concepts from stable diffusion

Ni, Zixuan and Wei, Longhui and Li, Jiacheng and Tang, Siliang and Zhuang, Yueting and Tian, Qi

ACM MM 2023. [PDF]

ReFACT: Updating Text-to-Image Models by Editing the Text Encoder

Dana Arad, Hadas Orgad, Yonatan Belinkov

NAACL 2024. [PDF]

OnMechanistic Knowledge Localization in Text-to-Image Generative Models

Samyadeep Basu, Keivan Rezaei, Priyatham Kattakinda, Vlad I Morariu, Nanxuan Zhao, Ryan A. Rossi, Varun Manjunatha, Soheil Feizi

ICML 2024. [PDF]

Forget-Me-Not: Learning to Forget in Text-to-Image Diffusion Models

Gong Zhang, Kai Wang, Xingqian Xu, Zhangyang Wang, Humphrey Shi

CVPR 2024 [PDF] [CODE]

Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models

Senmao Li, Joost van de Weijer, Taihang Hu, Fahad Shahbaz Khan, Qibin Hou, Yaxing Wang, Jian Yang

ICLR 2024. [PDF] [CODE]

Selective Amnesia: A Continual Learning Approach to Forgetting in Deep Generative Models

Alvin Heng , Harold Soh

NeurIPS 2024, [PDF] [CODE]

Leveraging Catastrophic Forgetting to Develop Safe Diffusion Models against Malicious Finetuning

Jiadong Pan, Hongcheng Gao, Zongyu Wu, Taihang Hu, Li Su, Qingming Huang, Liang Li

NeurIPS 2024. [PDF] [CODE]

Boosting Alignment for Post-Unlearning Text-to-Image Generative Models

Myeongseob Ko, Henry Li, Zhun Wang, Jonathan Patsenker, Jiachen T. Wang, Qinbin Li, Ming Jin, Dawn Song, Ruoxi Jia

NeurIPS 2024. [PDF] [CODE]

Erasing Undesirable Concepts in Diffusion Models with Adversarial Preservation

Anh Bui, Long Vuong, Khanh Doan, Trung Le, Paul Montague, Tamas Abraham, Dinh Phung

NeurIPS 2024. [PDF] [CODE]

Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion Models

Yimeng Zhang, Xin Chen, Jinghan Jia, Yihua Zhang, Chongyu Fan, Jiancheng Liu, Mingyi Hong, Ke Ding, Sijia Liu

NeurIPS 2024. [PDF] [CODE]

All but One: Surgical Concept Erasing with Model Preservation in Text-to-Image Diffusion Models

Hong, Seunghoo and Lee, Juhun and Woo, Simon S

AAAI 2024. [PDF]

SafeGen: Mitigating Unsafe Content Generation in Text-to-Image Models

Xinfeng Li , Yuchen Yang , Jiangyi Deng, Chen Yan , Yanjiao Chen , Xiaoyu Ji , Wenyuan Xu

ACM CCS 2024. [PDF] [CODE]

Direct Unlearning Optimization for Robust and Safe Text-to-Image Models

Yong-Hyun Park, Sangdoo Yun, Jin-Hwa Kim, Junho Kim, Geonhui Jang, Yonghyun Jeong, Junghyo Jo, Gayoung Lee

ICML 2024 Workshop. [PDF]

One-dimensional adapter to rule them all: Concepts diffusion models and erasing applications

Mengyao Lyu, Yuhong Yang, Haiwen Hong, Hui Chen, Xuan Jin, Yuan He, Hui Xue, Jungong Han, Guiguang Ding

CVPR 2024 [PDF] [CODE]

Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models

Poppi, Samuele and Poppi, Tobia and Cocchi, Federico and Cornia, Marcella and Baraldi, Lorenzo and Cucchiara, Rita

ECCV 2024. [PDF] [CODE]

Unlearning Concepts in Diffusion Model via Concept Domain Correction and Concept Preserving Gradient

Yongliang Wu, Shiji Zhou, Mingzhuo Yang, Lianzhe Wang, Wenbo Zhu, Heng Chang, Xiao Zhou, Xu Yang

AAAI 2025. [PDF]

R.A.C.E.: Robust Adversarial Concept Erasure for Secure Text-to-Image Diffusion Model

Changhoon Kim, Kyle Min, Yezhou Yang

ECCV 2024. [PDF] [CODE]

Receler: Reliable Concept Erasing of Text-to-Image Diffusion Models via Lightweight Erasers

Chi-Pin Huang, Kai-Po Chang, Chung-Ting Tsai, Yung-Hsuan Lai, Fu-En Yang, Yu-Chiang Frank Wang

ECCV 2024. [PDF] [CODE]

Reliable and efficient concept erasure of text-to-image diffusion models

Chao Gong, Kai Chen, Zhipeng Wei, Jingjing Chen, and Yu-Gang Jiang

ECCV 2024. [PDF] [CODE]

Safeguard Text-to-Image Diffusion Models with Human Feedback Inversion

Sanghyun Kim, Seohyeon Jung, Balhae Kim, Moonseok Choi, Jinwoo Shin & Juho Lee

ECCV 2024. [PDF] [CODE]

Score Forgetting Distillation: A Swift, Data-Free Method for Machine Unlearning in Diffusion Models

Tianqi Chen, Shujian Zhang, Mingyuan Zhou

ICLR 2025. [PDF] [CODE]

Fantastic Targets for Concept Erasure in Diffusion Models and Where To Find Them

Anh Bui, Trang Vu, Long Vuong, Trung Le, Paul Montague, Tamas Abraham, Junae Kim, Dinh Phung

ICLR 2025. [PDF] [CODE]

Editing Massive Concepts in Text-to-Image Diffusion Models

Tianwei Xiong, Yue Wu, Enze Xie, Yue Wu, Zhenguo Li, Xihui Liu

arxiv 2024. [PDF] [CODE]

ShieldDiff: Suppressing Sexual Content Generation from Diffusion Models through Reinforcement Learning

Dong Han, Salaheldin Mohamed, Yong Li

arxiv 2024. [PDF]

Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned Concepts

Hongcheng Gao, Tianyu Pang, Chao Du, Taihang Hu, Zhijie Deng, Min Lin

ICCV 2025. [PDF] [CODE]

Dark Miner: Defend against undesired generation for text-to-image diffusion models

Zheling Meng, Bo Peng, Xiaochuan Jin, Yue Jiang, Jing Dong, Wei Wang

arxiv 2024. [PDF]

SafetyDPO: Scalable Safety Alignment for Text-to-Image Generation

Runtao Liu, Chen I Chieh, Jindong Gu, Jipeng Zhang, Renjie Pi, Qifeng Chen, Philip Torr, Ashkan Khakzar, Fabio Pizzati

arxiv 2024. [PDF] [CODE]

Continuous Concepts Removal in Text-to-image Diffusion Models

Tingxu Han, Weisong Sun, Yanrong Hu, Chunrong Fang, Yonglong Zhang, Shiqing Ma, Tao Zheng, Zhenyu Chen, Zhenting Wang

arxiv 2025. [PDF]

SPEED: Scalable, Precise, and Efficient Concept Erasure for Diffusion Models

Ouxiang Li, Yuan Wang, Xinting Hu, Houcheng Jiang, Tao Liang, Yanbin Hao, Guojun Ma, Fuli Feng

arxiv 2025. [PDF] [CODE]

Continual Unlearning for Foundational Text-to-Image Models without Generalization Erosion

Kartik Thakral, Tamar Glaser, Tal Hassner, Mayank Vatsa, Richa Singh

arxiv 2025. [PDF]

Sparse Autoencoder as a Zero-Shot Classifier for Concept Erasing in Text-to-Image Diffusion Models

Zhihua Tian, Sirun Nan, Ming Xu, Shengfang Zhai, Wenjie Qu, Jian Liu, Kui Ren, Ruoxi Jia, Jiaheng Zhang

arxiv 2025. [PDF]

SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders

Bartosz Cywiński, Kamil Deja

arxiv 2025. [PDF] [CODE]

CRCE: Coreference-Retention Concept Erasure in Text-to-Image Diffusion Models

Yuyang Xue, Edward Moroshko, Feng Chen, Steven McDonagh, Sotirios A. Tsaftaris

arxiv 2025. [PDF]

TRCE: Towards Reliable Malicious Concept Erasure in Text-to-Image Diffusion Models

Ruidong Chen, Honglin Guo, Lanjun Wang, Chenyu Zhang, Weizhi Nie, An-An Liu

ICCV 2025. [PDF] [CODE]

Concept pinpoint eraser for text-to-image diffusion models via residual attention gate

Byung Hyun Lee, Sungjin Lim, Seunggyu Lee, Dong Un Kang, Se Young Chun

ICLR 2025. [PDF] [CODE]

Concept replacer: Replacing sensitive concepts in diffusion models via precision localization

Lingyun Zhang, Yu Xie, Yanwei Fu, Ping Chen

CVPR 2025. [PDF]

CURE: Concept unlearning via orthogonal representation editing in diffusion models

Tingxu Han, Weisong Sun, Yanrong Hu, Chunrong Fang, Yonglong Zhang, Shiqing Ma, Tao Zheng, Zhenyu Chen, Zhenting Wang

NeurIPS 2025. [PDF]

EraseAnything: Enabling concept erasure in rectified flow transformers

D Gao, S Lu, W Zhou, J Chu, J Zhang, M Jia, B Zhang, Z Fan, W Zhang

ICML 2025. [PDF] [CODE]

Fine-grained erasure in text-to-image diffusion-based foundation models

Kartik Thakral, Tamar Glaser, Tal Hassner, Mayank Vatsa, Richa Singh

CVPR 2025 [PDF]

Precise, fast, and low-cost concept erasure in value space: Orthogonal complement matters

Y Wang, O Li, T Mu, Y Hao, K Liu, X Wang, X He

CVPR 2025. [PDF]

SafeGen: Mitigating sexually explicit content generation in text-to-image models

Xinfeng Li, Yuchen Yang, Jiangyi Deng, Chen Yan, Yanjiao Chen, Xiaoyu Ji, Wenyuan Xu

ACM CCS 2025. [PDF] [CODE]

SafeGuider: Robust and practical content safety control for text-to-image models

P Qi, K Tang, W Zhou, W Zhang, N Yu, T Zhang, Q Guo, J Zhang

ACM CCS 2025. [PDF] [CODE]

:apple:Inference Guidance

Safe Latent Diffusion: Mitigating Inappropriate Degeneration in Diffusion Models

Patrick Schramowski, Manuel Brack, Björn Deiseroth, Kristian Kersting

CVPR 2023. [PDF] [CODE]

Sega: Instructing text-to-image models using semantic guidance

Manuel Brack, Felix Friedrich, Dominik Hintersdorf, Lukas Struppek, Patrick Schramowski, Kristian Kersting

NeurIPS 2023. [PDF] [CODE]

Self-discovering interpretable diffusion latent directions for responsible text-to-image generation

Li, Hang and Shen, Chengzhi and Torr, Philip and Tresp, Volker and Gu, Jindong

CVPR 2024. [PDF] [CODE]

One-dimensional Adapter to Rule Them All: Concepts, Diffusion Models and Erasing Applications

Mengyao Lyu, Yuhong Yang, Haiwen Hong, Hui Chen, Xuan Jin, Yuan He, Hui Xue, Jungong Han, Guiguang Ding

CVPR 2024. [PDF] [CODE]

Safree: Training-free and adaptive guard for safe text-to-image and video generation

Jaehong Yoon, Shoubin Yu, Vaidehi Patil, Huaxiu Yao, Mohit Bansal

ICLR 2025. [PDF] [CODE]

Pruning for Robust Concept Erasing in Diffusion Models

Tianyun Yang, Juan Cao, Chang Xu

arxiv 2024. [PDF]

ConceptPrune: Concept editing in diffusion models via skilled neuron pruning

Ruchika Chavhan, Da Li, Timothy Hospedales

arxiv 2024. [PDF]

TraSCE: Trajectory Steering for Concept Erasure

Anubhav Jain, Yuya Kobayashi, Takashi Shibuya, Yuhta Takida, Nasir Memon, Julian Togelius, Yuki Mitsufuji

arxiv 2025. [PDF] [CODE]

Growth Inhibitors for Suppressing Inappropriate Image Concepts in Diffusion Models

Die Chen, Zhiwen Li, Mingyuan Fan, Cen Chen, Wenmeng Zhou, Yanhao Wang, Yaliang Li

arxiv 2025. [PDF]

Training-Free Safe Denoisers for Safe Use of Diffusion Models

Mingyu Kim, Dongjun Kim, Amman Yusuf, Stefano Ermon, Mi Jung Park

NeurIPS 2025. [PDF]

PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models

Lingzhi Yuan, Xinfeng Li, Chejian Xu, Guanhong Tao, Xiaojun Jia, Yihao Huang, Wei Dong, Yang Liu, XiaoFeng Wang, Bo Li

arxiv 2025. [PDF]

Concept Corrector: Erase concepts on the fly for text-to-image diffusion models

Zheling Meng, Bo Peng, Xiaochuan Jin, Yueming Lyu, Wei Wang, Jing Dong

arxiv 2025. [PDF]

Localized Concept Erasure for Text-to-Image Diffusion Models Using Training-Free Gated Low-Rank Adaptation

Byung Hyun Lee, Sungjin Lim, Se Young Chun

CVPR 2025. [PDF]

Resources

This part provides commonly used datasets and tools in AD-on-T2IDM.

Datasets

Based on the prompt source, existing datasets are categorized into two types: clean and adversarial datasets. The clean dataset consists of clean prompts that are not attacked and typically crafted by human, while the adversarial dataset comprises adversarial prompts generated by attack methods. Moreover, according to the category of prompts involved in the dataset, existing clean datasets are further divided into two types: non-malicious and malicious datasets. The non-malicious dataset contains non-malicious prompts, while the malicious dataset contains explicitly malicious prompts. In this section, we will introduce several non-malicious, malicious, and adversarial datasets, respectively.

Non-Malicious Datasets

  • $\textit{ImageNet}$, which contains images describing 1,000 categories of common objects in the real world, is a significant benchmark in the field of computer vision. As a result, some works craft clean datasets based on the category information in ImageNet. For instance, ATM employs a standardized template: "A photo of {CLASS_NAME}" to generate clean prompts, where "{CLASS_NAME}" denotes the class name in ImageNet.
  • $\textit{MSCOCO}$ [Link]is a cross-modal image-text dataset, a popular benchmark for training and evaluating text-to-image generation models. Specifically, MSCOCO includes 82,783 training images and 40,504 testing images, each with 5 text descriptions.
  • $\textit{LAION-COCO}$ [Link] is a subset of LAION-5B, which is a large-scale image-text dataset in the real world. LAION-COCO includes 600 million images and corresponding text descriptions.
  • $\textit{DiffusionDB}$ [Link] is a large-scale text-to-image prompt dataset, which contains 14 million images generated by Stable Diffusion using prompts from real users.
  • $\textit{VidproM}$[LINK] is a large-scale dataset comprising 1.67 million unique text-to-video prompts from real users, where each prompt is flagged with the probabilities across six aspects of NSFW content, including toxicity, obscenity, identity attack, insult, threat, and sexual explicitness. (NeurIPS 2024 Datasets and Benchmarks Track)

Malicious Datasets

  • $\textit{Unsafe Diffusion}$ [Link] provides 30 manually crafted malicious prompts that describe sexual and bloody content, as well as political figures.
  • $\textit{SneakyPrompt}$ [Link] uses ChatGPT to automatically generate 200 malicious prompts that involve sexual and bloody content.
  • $\textit{I2P}$ [Link] comprises 4,703 malicious prompts, encompassing hate, harassment, violence, self-harm, nudity content, shocking images, and illegal activity. These inappropriate prompts are real-user inputs sourced from an image generation website, Lexica [Link].
  • $\textit{MMA}$ [Link] samples and releases 1,000 malicious prompts from LAION-COCO based on an NSFW~(Not Safe for Work) score. These malicious prompts mainly focus on sexual content.
  • $ART$[Link] follows I2P and collects 15,607 malicious prompts from 7 categories in Lexica [Link].
  • $\textit{ViSU}$ [Link] contains 175k pairs of safe and unsafe data examples. Each example consists of: (1) a safe sentence, (2) a corresponding safe image, (3) an NSFW sentence that is semantically correlated with the safe sentence, and (4) a corresponding NSFW image.
  • $\textit{T2VSafetyBench}$[LINK] contains 4,400 NSFW prompts in 12 classes: Pornography, Borderline Pornography, Violence, Gore, Public Figures, Discrimination, Political Sensitivity, Illegal Activities, Disturbing Content, Misinformation and Falsehoods, Copyright and Trademark Infringement, and Temporal Risk. These prompts are collected from three sources: VidProM[LINK], LLMs, and attack methods. (NeurIPS 2024 Datasets and Benchmarks Track)
  • $\textit{SAFESORA}$[LINK] contains 14,711 unique text prompts for the text-to-video models, with 48.61% potentially inducing harmful videos. Furthermore, these harmful prompts are categorized into 12 classes, including explicit sexual content, animal abuse, child abuse, crime, debated sensitive issue, drug abuse, hateful behavior, violence, racial discrimination, other discrimination, terrorism, and other harmful content. (NeurIPS 2024 Datasets and Benchmarks Track)
  • $\textit{Image Synthesis Style Studies Database}$ [Link] compiles thousands of artists whose styles can be replicated by various text-to-image models, such as Stable Diffusion and Midjourney.
  • $\textit{MACE}$ [Link] provides a dataset comprising 200 celebrities whose portraits, generated using SD v1.4, are recognized with remarkable accuracy (>99%) by the GIPHY Celebrity Detector (GCD) [Link].
  • $\textit{CPDM}$[LINK] is a copyright protection dataset for T2I models, containing 2,100 prompts in 4 classes: style, portrait, artistic creation figure, and licensed illustration.
  • $\textit{T2I-RiskyPrompt}$ [LINK] provides a hierarchical risky taxonomy~(6 risky categories and 14 subcategories) and 6,432 risky prompts labeled with hierarchical labels and detailed risky reasons. Moreover, it also provides a risky image detector used to evaluate the generated images of T2I models using T2I-RiskyPrompt.

Tools

We provide several detectors for detecting malicious prompts and images.

Malicious Prompt Detector

  • NSFW_text_classifier: [Link]

  • distilbert-nsfw-text-classifier: [Link]

  • Detoxify: [Link]

  • Toxic-comment-model: [Link]

  • Meta-Llama-Guard: [Link] (LLM evaluation)

  • Openai-Moderation: [Link] (API)

  • Azure-Moderation: [Link] (API)

Malicious Image Detector

Significant stargazers

Elwood Zonghao Ying

26 followers · starred Apr 2025