NY1024/Awesome-Trustworthy-GenAI

8

35 commits

updated Jul 5, 2024

See the code

README

Awesome-Trustworthy-GenAI Awesome

GitHub stars GitHub forks


Update[04/07/2024]


Categories

  • Paper
    • LLM
      • Jailbreak
        • Attack
        • Defense
        • Benchmark
        • Survey
        • Others
      • Hallucination
        • Mitigation
        • Detection
        • Benchmark
        • Survey
        • Others
    • MLLM
      • Jailbreak
        • Attack
        • Defense
        • Benchmark
        • Survey
      • Poisoning/Backdoor
        • Attack
        • Defense
      • Adversarial
        • Attack
        • Defense
        • Benchmark
      • Hallucination
      • Bias
      • Others
    • T2I
      • IP
        • Protection
        • Violation
        • Survey
      • Privacy
        • Attack
        • Defense
        • Benchmark
      • Memorization
      • Deepfake
        • Construction
        • Detection
        • Benchmark
      • Bias
      • Backdoor
        • Attack
        • Defense
      • Adversarial
        • Attack
        • Defense
    • Agent
      • Backdoor Attack/Defense
      • Adversarial Attack/Defense
      • Jaikbreak
      • Prompt Injection
      • Hallucination
      • Others
      • Survey
  • Book
  • Tutorial
  • Arena
  • Competition

LLM

Jailbreak

Attack

TitleAuthorPublishDateLinkSource Code
Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow InstructionsFederico BianchiICLR PosterMar-24https://openreview.net/pdf?id=gT5hALch9znull
FuzzLLM: A Novel and Universal Fuzzing Framework for Proactively Discovering Jailbreak Vulnerabilities in Large Language ModelsDongyu YaoICASSPSep-23https://arxiv.org/abs/2309.05274https://arxiv.org/pdf/2309.05274
GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak PromptsJiahao YuarxivSep-23https://arxiv.org/abs/2309.10253https://github.com/sherdencooper/GPTFuzz
Open Sesame! Universal Black Box Jailbreaking of Large Language ModelsRaz LapidarxivSep-23https://arxiv.org/abs/2309.01446null
Red-Teaming Large Language Models using Chain of Utterances for Safety-AlignmentRishabh BhardwajarxivAug-23https://arxiv.org/abs/2308.09662null
Jailbroken: How Does LLM Safety Training Fail?Alexander WeiNeurIPSJul-23https://arxiv.org/abs/2307.02483null
MasterKey: Automated Jailbreak Across Multiple Large Language Model ChatbotsGelei DengNDSSJul-23https://arxiv.org/abs/2307.08715null
Universal and Transferable Adversarial Attacks on Aligned Language ModelsAndy ZouNeurIPSJul 2023https://arxiv.org/abs/2307.15043http://llm-attacks.org/
Adversarial Demonstration Attacks on Large Language ModelsJiongxiao WangarxivMay-23https://arxiv.org/abs/2305.14950null
Catastrophic Jailbreak of Open-source LLMs via Exploiting GenerationYangsibo HuangICLRMar-24https://openreview.net/pdf?id=r42tSSCHPhhttps://github.com/Princeton-SysML/Jailbreak_LLM
GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via CipherYouliang YuanJan-24https://openreview.net/pdf?id=MbfAK4s61Ahttps://github.com/RobustNLP/CipherChat
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng LiuICLRJan-24https://openreview.net/pdf?id=7Jwpw4qKkbhttps://github.com/SheltonLiu-N/AutoDAN
Low-Resource Languages Jailbreak GPT-4Zheng-Xin YongNeurIPS WorkshopOct-23https://arxiv.org/abs/2310.02446null
Multi-step Jailbreaking Privacy Attacks on ChatGPTHaoran LiEMNLPApr-23https://arxiv.org/abs/2304.05197null
Attack Prompt Generation for Red Teaming and Defending Large Language ModelsBoyi DengEMNLPOct-23https://arxiv.org/abs/2310.12505https://github.com/Aatrox103/SAP
Multilingual Jailbreak Challenges in Large Language ModelsYue DengICLROct-23https://arxiv.org/abs/2310.06474https://github.com/DAMO-NLP-SG/multilingual-safety-for-LLMs%7D
Weak-to-Strong Jailbreaking on Large Language ModelsXuandong ZhaoarxivJan-24https://arxiv.org/abs/2401.17256https://github.com/XuandongZhao/weak-to-strong
How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMsYi ZengACLJan-24https://arxiv.org/abs/2401.06373https://chats-lab.github.io/persuasive_jailbreaker/
Jailbreaking GPT-4V via Self-Adversarial Attacks with System PromptsYuanwei WuarxivNov-23https://arxiv.org/abs/2311.09127null
Safety Alignment in NLP Tasks: Weakly Aligned Summarization as an In-Context AttackYu FuACLDec-23https://arxiv.org/abs/2312.06924null
Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay MehrotraarxivDec-23https://arxiv.org/abs/2312.02119https://github.com/RICommunity/TAP
A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models EasilyPeng DingNAACLNov-23https://arxiv.org/abs/2311.08268https://github.com/NJUNLP/ReNeLLM
Goal-Oriented Prompt Attack and Safety Evaluation for LLMsChengyuan LiuarxivSep-23https://arxiv.org/abs/2309.11830https://github.com/liuchengyuan123/CPAD
Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit CluesZhiyuan ChangarxivFeb-24https://arxiv.org/abs/2402.09091null
A Cross-Language Investigation into Jailbreak Attacks in Large Language ModelsJie LiarxivJan-24https://arxiv.org/abs/2401.16765null
Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven JailbreakYanrui DuarxivDec-23https://arxiv.org/abs/2312.04127null
All in How You Ask for It: Simple Black-Box Method for Jailbreak AttacksKazuhiro TakemotoarxivJan 2024https://arxiv.org/abs/2401.09798null
DeepInception: Hypnotize Large Language Model to Be JailbreakerXuan LiarxivNov-23https://arxiv.org/abs/2311.03191https://github.com/tmlr-group/DeepInception
Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona ModulationRusheb ShaharxivNov-23https://arxiv.org/abs/2311.03348null
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-RailsNeal MangaokararxivFeb-24https://arxiv.org/abs/2402.15911null
Semantic Mirror Jailbreak: Genetic Algorithm Based Jailbreak Prompts Against Open-source LLMsXiaoxia LiarxivFeb-24https://arxiv.org/abs/2402.14872null
ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix EmbeddingsHao WangarxivFeb-24https://arxiv.org/abs/2402.16006null
A StrongREJECT for Empty JailbreaksAlexandra SoulyarxivFeb-24https://arxiv.org/abs/2402.10260https://github.com/alexandrasouly/strongreject
Jailbreaking Black Box Large Language Models in Twenty QueriesPatrick ChaoarxivOct-23https://arxiv.org/abs/2310.08419https://github.com/patrickrchao/JailbreakingLLMs
Jailbreak and Guard Aligned Language Models with Only Few In-Context DemonstrationsZeming WeiarxivOct-23https://arxiv.org/abs/2310.06387null
AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language ModelsSicheng ZhuarxivOct-23https://arxiv.org/abs/2310.15140null
Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security AttacksDaniel KangarxivFeb 2023https://arxiv.org/abs/2302.05733null
Pandora: Jailbreak GPTs by Retrieval Augmented Generation PoisoningGelei DengarxivFeb-24https://arxiv.org/abs/2402.08416null
Jailbreaking Proprietary Large Language Models using Word Substitution CipherDivij HandaarxivFeb-24https://arxiv.org/abs/2402.10601null
PAL: Proxy-Guided Black-Box Attack on Large Language ModelsChawin SitawarinarxivFeb-24https://arxiv.org/abs/2402.09674https://github.com/chawins/pal
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMsFengqing JiangarxivFeb-24https://arxiv.org/abs/2402.11753https://github.com/uw-nsl/ArtPrompt
Query-Based Adversarial Prompt GenerationJonathan HayasearxivFeb-24https://arxiv.org/abs/2402.12329null
Coercing LLMs to do and reveal (almost) anythingJonas GeipingarxivFeb-24https://arxiv.org/abs/2402.14020https://github.com/JonasGeiping/carving
COLD-Attack: Jailbreaking LLMs with Stealthiness and ControllabilityXingang GuoICMLFeb-24https://arxiv.org/abs/2402.08679https://github.com/Yu-Fangxu/COLD-Attack
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive AttacksMaksym AndriushchenkoarxivApr-24https://arxiv.org/abs/2404.02151https://github.com/tml-epfl/llm-adaptive-attacks
Sandwich attack: Multi-language Mixture Adaptive Attack on LLMsBibek UpadhayayarxivApr 2024https://arxiv.org/abs/2404.07242null
Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language ModelsZhiyuan YuUSENIX SecurityMar-24https://arxiv.org/abs/2403.17336null
CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code CompletionQibing RenACLMar-24https://arxiv.org/abs/2403.07865https://github.com/renqibing/CodeAttack
Tastle: Distract Large Language Models for Automatic Jailbreak AttackZeguan XiaoarxivMar-24https://arxiv.org/abs/2403.08424null
AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMsZeyi LiaoarxivApr-24https://arxiv.org/abs/2404.07921https://github.com/OSU-NLP-Group/AmpleGCG
Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLMXikang YangarxivMay-24https://arxiv.org/abs/2405.05610null
GUARD: Role-playing to Generate Natural-language Jailbreakings to Test Guideline Adherence of Large Language ModelsHaibo JinarxivFeb-24https://arxiv.org/abs/2402.03299null
Don't Say No: Jailbreaking LLM by Suppressing RefusalYukai ZhouarxivApr-24https://arxiv.org/abs/2404.16369null
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMsAnselm PaulusarxivApr-24https://arxiv.org/abs/2404.16873https://github.com/facebookresearch/advprompter
Lockpicking LLMs: A Logit-Based Jailbreak Using Token-level ManipulationYuxi LiarxivMay-24https://arxiv.org/abs/2405.13068null
Can LLMs Deeply Detect Complex Malicious Queries? A Framework for Jailbreaking via Obfuscating IntentShang ShangarxivMay-24https://arxiv.org/abs/2405.03654null
GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-ExplanationGovind RamesharxivMay-24https://arxiv.org/abs/2405.13077null
Poisoned LangChain: Jailbreak LLMs by LangChainZiqiu WangarxivJun-24https://arxiv.org/abs/2406.18122https://github.com/CAM-FSS/jailbreak-langchain
Covert Malicious Finetuning: Challenges in Safeguarding LLM AdaptationDanny HalawiICMLJun-24https://arxiv.org/abs/2406.20053null
Voice Jailbreak Attacks Against GPT-4oXinyue ShenarxivMay-24https://arxiv.org/abs/2405.19103https://github.com/TrustAIRLab/VoiceJailbreakAttack
StructuralSleight: Automated Jailbreak Attacks on Large Language Models Utilizing Uncommon Text-Encoded StructureBangxin LiarxivJun-24https://arxiv.org/abs/2406.08754null
Knowledge-to-Jailbreak: One Knowledge Point Worth One AttackShangqing TuarxivJun-24https://arxiv.org/abs/2406.11682https://github.com/THU-KEG/Knowledge-to-Jailbreak/
CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language ModelsHuijie LvarxivFeb-24https://arxiv.org/abs/2402.16717https://github.com/huizhang-L/CodeChameleon
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM JailbreakersXirui LiarxivFeb-24https://arxiv.org/abs/2402.16914https://github.com/xirui-li/DrAttack
Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and ReconstructionTong LiuUSENIX SecurityFeb-24https://arxiv.org/abs/2402.18104https://github.com/LLM-DRA/DRA
Leveraging the Context through Multi-Round Interactions for Jailbreaking AttacksYixin ChengarxivFeb-24https://arxiv.org/abs/2402.09177null
Jailbreaking Large Language Models Against Moderation Guardrails via Cipher CharactersHaibo JinarxivMay-24https://arxiv.org/abs/2405.20413null
RL-JACK: Reinforcement Learning-powered Black-box Jailbreaking Attack against LLMsXuan ChenarxivJun-24https://arxiv.org/abs/2406.08725null
Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their DefensesXiaosen ZhengarxivJun-24https://arxiv.org/abs/2406.01288https://github.com/sail-sg/I-FSJ
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency LensLin LuarxivJun-24https://arxiv.org/abs/2406.03805null
Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak AttackMark RussinovicharxivApr-24https://arxiv.org/abs/2404.01833null
MART: Improving LLM Safety with Multi-round Automatic Red-TeamingSuyu GearxivNov-23https://arxiv.org/abs/2311.07689null
Virtual Context: Enhancing Jailbreak Attacks with Special Token InjectionYuqi ZhouarxivJun 2024https://arxiv.org/abs/2406.19845null
When LLM Meets DRL: Advancing Jailbreaking Efficiency via DRL-guided SearchXuan ChenarxivJun-24https://arxiv.org/abs/2406.08705null
Improved Techniques for Optimization-Based Jailbreaking on Large Language ModelsXiaojun JiaarxivMay-24https://arxiv.org/abs/2405.21018https://github.com/jiaxiaojunQAQ/I-GCG

Defense

TitleAuthorPublishDateLinkSource Code
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHFAnand SiththaranjanICLRApr-24https://openreview.net/pdf?id=0tWTxYYPnWnull
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLMBochuan CaoACL2024https://arxiv.org/abs/2309.14348null
Detecting Language Model Attacks with PerplexityGabriel AlonarxivAug-23https://arxiv.org/abs/2308.14132null
Certifying LLM Safety against Adversarial PromptingAounon KumararxivSep-23https://arxiv.org/abs/2309.02705https://github.com/aounon/certified-llm-safety
Baseline Defenses for Adversarial Attacks Against Aligned Language ModelsNeel JainarxivSep-23https://arxiv.org/abs/2309.00614null
SmoothLLM: Defending Large Language Models Against Jailbreaking AttacksAlexander RobeyarxivOct-23https://arxiv.org/abs/2310.03684https://github.com/arobey1/smooth-llm
Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting JailbreaksAbhinav RaoLREC-COLINGMay-23https://arxiv.org/abs/2305.14965null
Intention Analysis Makes LLMs A Good Jailbreak DefenderYuqi ZhangarxivJan-24https://arxiv.org/abs/2401.06561https://github.com/alphadl/SafeLLM_with_IntentionAnalysis
On Prompt-Driven Safeguarding for Large Language ModelsChujie ZhengICMLJan-24https://arxiv.org/abs/2401.18018https://github.com/chujiezheng/LLM-Safeguard
Robust Prompt Optimization for Defending Language Models Against Jailbreaking AttacksAndy ZhouarxivJan-24https://arxiv.org/abs/2401.17263https://github.com/lapisrocks/rpo
SPML: A DSL for Defending Language Models Against Prompt AttacksReshabh K SharmaarxivFeb-24https://arxiv.org/abs/2402.11755https://prompt-compiler.github.io/SPML/
LLMs Can Defend Themselves Against Jailbreaking in a Practical Manner: A Vision PaperDaoyuan WuarxivFeb-24https://arxiv.org/abs/2402.15727null
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware DecodingZhangchen XuACLFeb-24https://arxiv.org/abs/2402.08983https://github.com/uw-nsl/SafeDecoding
GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient AnalysisYueqi XieACLFeb-24https://arxiv.org/abs/2402.13494https://github.com/xyq7/GradSafe
Defending Large Language Models against Jailbreak Attacks via Semantic SmoothingJiabao JiarxivFeb-24https://arxiv.org/abs/2402.16192https://github.com/UCSB-NLP-Chang/SemanticSmooth
Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-TuningAdib HasanarxivJan-24https://arxiv.org/abs/2401.10862null
Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-RefinementHeegyu KimarxivFeb-24https://arxiv.org/abs/2402.15180null
Protecting Your LLMs with Information BottleneckZichuan LiuarxivMay-24https://arxiv.org/abs/2404.13968https://github.com/zichuan-liu/IB4LLMs
Eraser: Jailbreaking Defense in Large Language Models via Unlearning Harmful KnowledgeWeikai LuarxivApr-24https://arxiv.org/abs/2404.05880https://github.com/ZeroNLP/Eraser
RigorLLM: Resilient Guardrails for Large Language Models against Undesired ContentZhuowen YuanarxivMar-24https://arxiv.org/abs/2403.13031null
Detoxifying Large Language Models via Knowledge EditingMengru WangACLMar-24https://arxiv.org/abs/2403.14472https://github.com/zjunlp/EasyEdit
AutoDefense: Multi-Agent LLM Defense against Jailbreak AttacksYifan ZengarxivMar-24https://arxiv.org/abs/2403.04783https://github.com/XHMY/AutoDefense
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss LandscapesXiaomeng HuarxivMar-24https://arxiv.org/abs/2403.00867https://huggingface.co/spaces/TrustSafeAI/GradientCuff-Jailbreak-Defense
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMsFan LiuarxivJun-24https://arxiv.org/abs/2406.06622null
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language ModelsLiwei JiangarxivJun-24https://arxiv.org/abs/2406.18510null
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMsSeungju HanarxivJun-24https://arxiv.org/abs/2406.18495null
SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical MannerXunguang WangarxivJun-24https://arxiv.org/abs/2406.05498null
Merging Improves Self-Critique Against Jailbreak AttacksVictor GallegoarxivJun-24https://arxiv.org/abs/2406.07188https://github.com/vicgalle/merging-self-critique-jailbreaks
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak AttacksChen XiongarxivMay-24https://arxiv.org/abs/2405.20099null
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety AlignmentJiongxiao WangarxivFeb-24https://arxiv.org/abs/2402.14968https://jayfeather1024.github.io/Finetuning-Jailbreak-Defense/
Improving Alignment and Robustness with Circuit BreakersAndy ZouarxivJun-24https://arxiv.org/abs/2406.04313https://github.com/blackswan-ai/circuit-breakers
Robustifying Safety-Aligned Large Language Models through Clean Data CurationXiaoqun LiuarxivMay-24https://arxiv.org/abs/2405.19358null
Efficient Adversarial Training in LLMs with Continuous AttacksSophie XhonneuxarxivMay-24https://arxiv.org/abs/2405.15589https://github.com/sophie-xhonneux/Continuous-AdvTrain
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity GuidanceCaishuang HuangarxivJun-24https://arxiv.org/abs/2406.18118null
Defending Large Language Models Against Jailbreak Attacks via Layer-specific EditingWei ZhaoarxivMay-24https://arxiv.org/abs/2405.18166https://github.com/ledllm/ledllm
Cross-Task Defense: Instruction-Tuning LLMs for Content SafetyYu FuNAACLMay-24https://arxiv.org/abs/2405.15202https://github.com/FYYFU/safety-defense

Benchmark

TitleAuthorPublishYearLinkSource Code
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMsZhao Xuarxiv2024https://arxiv.org/abs/2406.09324https://github.com/usail-hkust/Bag_of_Tricks_for_LLM_Jailbreaking
BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM SafeguardsDiego Dornarxiv2024https://arxiv.org/abs/2406.01364null
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language ModelsPatrick Chaoarxiv2024https://arxiv.org/abs/2404.01318https://github.com/JailbreakBench/jailbreakbench
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeikaarxiv2024https://arxiv.org/abs/2402.04249https://github.com/centerforaisafety/HarmBench
SC-Safety: A Multi-round Open-ended Question Adversarial Safety Benchmark for Large Language Models in ChineseLiang Xuarxiv2024https://arxiv.org/abs/2310.05818https://www.cluebenchmarks.com/
Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language ModelsHuachuan Qiuarxiv2024https://arxiv.org/abs/2307.08487https://github.com/qiuhuachuan/latent-jailbreak
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language ModelsPaul RöttgerNAACLAug-23https://arxiv.org/abs/2308.01263null
Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language ModelsHuachuan QiuarxivJul-23https://arxiv.org/abs/2307.08487https://github.com/qiuhuachuan/latent-jailbreak
SC-Safety: A Multi-round Open-ended Question Adversarial Safety Benchmark for Large Language Models in ChineseLiang XuarxivOct-23https://arxiv.org/abs/2310.05818https://www.cluebenchmarks.com/
AttackEval: How to Evaluate the Effectiveness of Jailbreak Attacking on Large Language ModelsDong shuarxivJan-24https://arxiv.org/abs/2401.09002null
How (un)ethical are instruction-centric responses of LLMs? Unveiling the vulnerabilities of safety guardrails to harmful queriesSomnath BanerjeearxivFeb-24https://arxiv.org/abs/2402.15302https://huggingface.co/datasets/SoftMINER-Group/TechHazardQA
AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM ExpertsShaona GhosharxivApr-24https://arxiv.org/abs/2404.05993
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMsZhao XuarxivJun-24https://arxiv.org/abs/2406.09324null
BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM SafeguardsDiego DornarxivJun-24https://arxiv.org/abs/2406.01364null
Improved Generation of Adversarial Examples Against Safety-aligned LLMsQizhang LiarxivMay-24https://arxiv.org/abs/2405.20778null
JailBreakV-28K: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak AttacksWeidi LuoarxivApr-24https://arxiv.org/abs/2404.03027https://github.com/EddyLuo1232/JailBreakV_28K
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language ModelsPatrick ChaoarxivMar-24https://arxiv.org/abs/2404.01318https://github.com/JailbreakBench/jailbreakbench
JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language ModelsDelong RanarxivJun-24https://arxiv.org/abs/2406.09321https://github.com/ThuCCSLab/JailbreakEval

Survey

TitleAuthorPublishYearLinkSource Code
Unique Security and Privacy Threats of Large Language Model: A Comprehensive SurveyShang Wangarxiv2024https://arxiv.org/abs/2406.07973null
Exploring Vulnerabilities and Protections in Large Language Models: A SurveyFrank Weizhen Liuarxiv2024https://arxiv.org/abs/2406.00240null
Safeguarding Large Language Models: A SurveyYi Dongarxiv2024https://arxiv.org/abs/2406.02622null
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under CompressionJunyuan Hongarxiv2024https://arxiv.org/abs/2403.15447null
Securing Large Language Models: Threats, Vulnerabilities and Responsible PracticesSara Abdaliarxiv2024https://arxiv.org/abs/2403.12503null
Breaking Down the Defenses: A Comparative Survey of Attacks on Large Language ModelsArijit Ghosh Chowdhuryarxiv2024https://arxiv.org/abs/2403.04786null
Attacks, Defenses and Evaluations for LLM Conversation Safety: A SurveyZhichen Dongarxiv2024https://arxiv.org/abs/2402.09283null
Security and Privacy Challenges of Large Language Models: A SurveyBadhan Chandra Dasarxiv2024https://arxiv.org/abs/2402.00888null
Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model SystemsTianyu Cuiarxiv2024https://arxiv.org/abs/2401.05778null
TrustLLM: Trustworthiness in Large Language ModelsLichao Sunarxiv2024https://arxiv.org/abs/2401.05561null
A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the UglyYifan Yaoarxiv2024https://arxiv.org/abs/2312.02003null
A Comprehensive Overview of Large Language ModelsHumza Naveedarxiv2024https://arxiv.org/abs/2307.06435null
Safety Assessment of Chinese Large Language ModelsHao Sunarxiv2024https://arxiv.org/abs/2304.10436null
Holistic Evaluation of Language ModelsPercy Liangarxiv2024https://arxiv.org/abs/2211.09110null
Survey of Vulnerabilities in Large Language Models Revealed by Adversarial AttacksErfan ShayeganiACLOct-23https://arxiv.org/abs/2310.10844https://llm-vulnerability.github.io/
A Comprehensive Survey of Attack Techniques, Implementation, and Mitigation Strategies in Large Language ModelsAysan EsmradiUbiSecDec-23https://arxiv.org/abs/2312.10982null
Against The Achilles' Heel: A Survey on Red Teaming for Generative ModelsLizhi LinarxivMar-24https://arxiv.org/abs/2404.00629null

Others

TitleAuthorPublishYearLinkSource Code
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language ModelsXinyue ShenCCSAug-23https://arxiv.org/abs/2308.03825https://github.com/verazuo/jailbreak_llms
Jailbreaking ChatGPT via Prompt Engineering: An Empirical StudyYi LiuarxivMay-23https://arxiv.org/abs/2305.13860null
Comprehensive Assessment of Jailbreak Attacks Against LLMsJunjie ChuarxivFeb-24https://arxiv.org/abs/2402.05668null
Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming in the WildNanna IniearxivNov-23https://arxiv.org/abs/2311.06237null
A Comprehensive Study of Jailbreak Attack versus Defense for Large Language ModelsZihao XuACLFeb-24https://arxiv.org/abs/2402.13457null
Sowing the Wind, Reaping the Whirlwind: The Impact of Editing Language ModelsRima HazraACLJan-24https://arxiv.org/abs/2401.10647null
Is the System Message Really Important to Jailbreaks in Large Language Models?Xiaotian ZouarxivFeb-24https://arxiv.org/abs/2402.14857null
Testing the Limits of Jailbreaking Defenses with the Purple ProblemTaeyoun KimarxivMar-24https://arxiv.org/abs/2403.14725null
JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language ModelsYingchaojie FengarxivApr-24https://arxiv.org/abs/2404.08793
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMsJavier RandoarxivApr-24https://arxiv.org/abs/2404.14461null
Universal Adversarial Triggers Are Not UniversalNicholas MeadearxivApr-24https://arxiv.org/abs/2404.16020null
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden StatesZhenhong ZhouarxivJun-24https://arxiv.org/abs/2406.05644https://github.com/ydyjya/LLM-IHS-Explanation
Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language ModelsSarah BallarxivJun-24https://arxiv.org/abs/2406.09289null
Badllama 3: removing safety finetuning from Llama 3 in minutesDmitrii VolkovarxivJul-24https://arxiv.org/abs/2407.01376null
"Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' JailbreakLingrui MeiarxivJun-24https://arxiv.org/abs/2406.11668null
Rethinking How to Evaluate Language Model JailbreakHongyu CaiarxivApr-24https://arxiv.org/abs/2404.06407
Jailbreak Paradox: The Achilles' Heel of LLMsAbhinav RaoarxivJun-24https://arxiv.org/abs/2406.12702null
Hacc-Man: An Arcade Game for Jailbreaking LLMsMatheus ValentimarxivMay-24https://arxiv.org/abs/2405.15902null

Hallucination

Mitigation

TitleAuthorPublishDateLinkSource Code
BTR: Binary Token Representations for Efficient Retrieval Augmented Language ModelsQingqing CaoICLRJan-24https://openreview.net/pdf?id=3TO3TtnOFlhttps://github.com/csarron/BTR
Reasoning on Graphs: Faithful and Interpretable Large Language Model ReasoningLINHAO LUOICLRJan-24https://openreview.net/pdf?id=ZGNWW7xZ6Qhttps://github.com/RManLuo/reasoning-on-graphs
Knowledge of Knowledge: Exploring Known-Unknowns Uncertainty with Large Language ModelsAlfonso AmayuelasarxivMay-23https://arxiv.org/abs/2305.13712null
Conformal Language ModelingVictor QuachICLRJun 2023https://arxiv.org/pdf/2306.10193null
Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth LiNeurIPSJun 2023https://arxiv.org/abs/2306.03341https://github.com/likenneth/honest_llama
Supervised Knowledge Makes Large Language Models Better In-context LearnersLinyi YangICLRDec-23https://arxiv.org/abs/2312.15918https://github.com/YangLinyi/Supervised-Knowledge-Makes-Large-Language-Models-Better-In-context-Learners
Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction TuningFuxiao LiuICLRJun 2023https://arxiv.org/abs/2306.14565https://github.com/FuxiaoLiu/LRV-Instruction
Chain-of-Knowledge: Grounding Large Language Models via Dynamic Knowledge Adapting over Heterogeneous SourcesXingxuan LiICLRMay-23https://arxiv.org/abs/2305.13269https://github.com/DAMO-NLP-SG/chain-of-knowledge
Fine-Tuning Language Models for FactualityKatherine TianICLRNov-23https://arxiv.org/abs/2311.08401https://github.com/kttian/llm_factuality_tuning
Chain-of-Table: Evolving Tables in the Reasoning Chain for Table UnderstandingZilong WangICLRJan-24https://arxiv.org/abs/2401.04398null
MetaGPT: Meta Programming for A Multi-Agent Collaborative FrameworkSirui HongICLRAug-23https://arxiv.org/abs/2308.00352https://github.com/geekan/MetaGPT
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingZhibin GouICLRMay-23https://arxiv.org/pdf/2305.11738https://github.com/microsoft/ProphetNet/tree/master/CRITIC
Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun DuarxivMay-23https://arxiv.org/abs/2305.14325https://composable-models.github.io/llm_debate/
Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language ModelsJinhao DuanarxivJul-23https://arxiv.org/abs/2307.01379https://github.com/jinhaoduan/SAR
RAPPER: Reinforced Rationale-Prompted Paradigm for Natural Language Explanation in Visual Question AnsweringKai-Po ChangICLRJan 2024https://openreview.net/pdf?id=bshfchPM9Hnull
Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated FeedbackBaolin PengarxivFeb 2023https://arxiv.org/abs/2302.12813https://github.com/pengbaolin/LLM-Augmenter
The Knowledge Alignment Problem: Bridging Human and External Knowledge for Large Language ModelsShuo ZhangarxivMay-23https://arxiv.org/abs/2305.13669https://github.com/ShuoZhangXJTU/MixAlign
Trusting Your Evidence: Hallucinate Less with Context-aware DecodingWeijia ShiarxivMay-23https://arxiv.org/abs/2305.14739null
Reasoning on Graphs: Faithful and Interpretable Large Language Model ReasoningLINHAO LUOICLROct-23https://arxiv.org/abs/2310.01061https://github.com/RManLuo/reasoning-on-graphs
Explainable Claim Verification via Knowledge-Grounded Reasoning with Large Language ModelsHaoran WangEMNLPOct-23https://arxiv.org/abs/2310.05253https://github.com/wang2226/FOLK
Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learningMustafa ShukorICLROct-23https://arxiv.org/abs/2310.00647https://github.com/mshukor/EvALign-ICL
Mitigating Large Language Model Hallucinations via Autonomous Knowledge Graph-based RetrofittingXinyan GuanarxivNov-23https://arxiv.org/abs/2311.13314null
Knowledge Verification to Nip Hallucination in the BudFanqi WanarxivJan 2024https://arxiv.org/abs/2401.10768https://github.com/fanqiwan/KCA
Model Editing Harms General Abilities of Large Language Models: Regularization to the RescueJia-Chen GuarxivJan 2024https://arxiv.org/abs/2401.04700null
Reducing Hallucinations in Entity Abstract Summarization with Facts-Template DecompositionFangwei ZhuarxivFeb-24https://arxiv.org/abs/2402.18873null
Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language ModelsWenhao YuarxivNov-23https://arxiv.org/abs/2311.09210null
Reducing hallucination in structured outputs via Retrieval-Augmented GenerationPatrice BéchardNAACLApr 2024https://arxiv.org/abs/2404.08189null
Improving Factual Error Correction by Learning to Inject Factual ErrorsXingwei HeAAAIDec-23https://arxiv.org/abs/2312.07049null
Strong hallucinations from negation and how to fix themNicholas AsherarxivFeb-24https://arxiv.org/abs/2402.10543null
Retrieve Only When It Needs: Adaptive Retrieval Augmentation for Hallucination Mitigation in Large Language ModelsHanxing DingarxivFeb-24https://arxiv.org/abs/2402.10612null
Truth-Aware Context Selection: Mitigating Hallucinations of Large Language Models Being Misled by Untruthful ContextsTian YuACLMar-24https://arxiv.org/abs/2403.07556https://github.com/ictnlp/TACS
Enhancing LLM Factual Accuracy with RAG to Counter Hallucinations: A Case Study on Domain-Specific Queries in Private Knowledge-BasesJiarui LiarxivMar-24https://arxiv.org/abs/2403.10446https://github.com/anlp-team/LTI_Neural_Navigator
Uncertainty-Based Abstention in LLMs Improves Safety and Reduces HallucinationsChristian TomaniarxivApr-24https://arxiv.org/abs/2404.10960null
TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful SpaceShaolei ZhangACLFeb-24https://arxiv.org/abs/2402.17811https://ictnlp.github.io/TruthX-site/
Measuring and Reducing LLM Hallucination without Gold-Standard AnswersJiaheng WeiarxivFeb-24https://arxiv.org/abs/2402.10412null
Mitigating LLM Hallucinations via Conformal AbstentionYasin Abbasi YadkoriarxivApr-24https://arxiv.org/abs/2405.01563null
Understanding the Effects of Iterative Prompting on TruthfulnessSatyapriya KrishnaarxivFeb-24https://arxiv.org/abs/2402.06625null
Mitigating Hallucinations in Large Language Models via Self-Refinement-Enhanced Knowledge RetrievalMengjia NiuarxivMay-24https://arxiv.org/abs/2405.06545null
Can LLMs Produce Faithful Explanations For Fact-checking? Towards Faithful Explainable Fact-Checking via Multi-Agent DebateKyungha KimarxivFeb-24https://arxiv.org/abs/2402.07401null

Detection

TitleAuthorVenueDateLinkSource Code
SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee ManakulEMNLPMar-23https://arxiv.org/abs/2303.08896https://github.com/potsawee/selfcheckgpt
A Stitch in Time Saves Nine: Detecting and Mitigating Hallucinations of LLMs by Validating Low-Confidence GenerationNeeraj VarshneyarxivJul-23https://arxiv.org/abs/2307.03987
Self-contradictory Hallucinations of Large https://arxiv.org/abs/2305.15852 Models: Evaluation, Detection and MitigationNiels MündlerICLRMay-23https://arxiv.org/abs/2305.15852https://chatprotect.ai/
Fact-Checking Complex Claims with Program-Guided ReasoningLiangming PanACLMay-23https://arxiv.org/abs/2305.12744https://github.com/mbzuai-nlp/ProgramFC
INSIDE: LLMs' Internal States Retain the Power of Hallucination DetectionChao ChenICLRFeb-24https://arxiv.org/abs/2402.03744null
Detecting Hallucinations in Large Language Model Generation: A Token Probability ApproachErnesto QuevedoICAIMay-24https://arxiv.org/abs/2405.19648https://github.com/Baylor-AI/HalluDetect
Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language ModelsWeihang SuMar-24https://arxiv.org/abs/2403.06448null
Enhancing Uncertainty-Based Hallucination Detection with Stronger FocusTianhang ZhangEMNLPNov-23https://arxiv.org/abs/2311.13230null
In Search of Truth: An Interrogation Approach to Hallucination DetectionYakir YehudaarxivMar-24https://arxiv.org/abs/2403.02889null
PoLLMgraph: Unraveling Hallucinations in Large Language Models via State Transition DynamicsDerui ZhuarxivApr-24https://arxiv.org/abs/2404.04722null
Fact-Checking the Output of Large Language Models via Token-Level Uncertainty QuantificationEkaterina FadeevaACLMar-24https://arxiv.org/pdf/2403.04696null
Can We Verify Step by Step for Incorrect Answer Detection?Xin XuarxivFeb-24https://arxiv.org/abs/2402.10528https://github.com/XinXU-USTC/R2PE
Comparing Hallucination Detection Metrics for Multilingual GenerationHaoqiang KangarxivFeb-24https://arxiv.org/abs/2402.10496null

Benchmark

TitleAuthorPublishDateLinkSource Code
Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language ModelsYue ZhangarxivSep-23https://arxiv.org/abs/2309.01219null
LitCab: Lightweight Language Model Calibration over Short- and Long-form ResponsesXin LiuICLROct-23https://arxiv.org/abs/2310.19208https://github.com/launchnlp/LitCab
Towards Understanding Factual Knowledge of Large Language ModelsXuming HuICLRJan-24https://openreview.net/pdf?id=9OevMUdodshttps://github.com/THU-BPM/Pinocchio
HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsJunyi LiEMNLPMay-23https://arxiv.org/abs/2305.11747https://github.com/RUCAIBox/HaluEval
UHGEval: Benchmarking the Hallucination of Chinese Large Language Models via Unconstrained GenerationXun LiangACLNov-23https://arxiv.org/abs/2311.15296https://iaar-shanghai.github.io/UHGEval/
HaluEval-Wild: Evaluating Hallucinations of Language Models in the WildZhiying ZhuarxivMar-24https://arxiv.org/abs/2403.04307https://github.com/Dianezzy/HaluEval-Wild
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMsAdi SimhiarxivApr-24https://arxiv.org/abs/2404.09971https://github.com/technion-cs-nlp/hallucination-mitigation
DelucionQA: Detecting Hallucinations in Domain-specific Question AnsweringMobashir SadatEMNLPDec-23https://arxiv.org/abs/2312.05200https://github.com/boschresearch/DelucionQA
DiaHalu: A Dialogue-level Hallucination Evaluation Benchmark for Large Language ModelsKedi ChenarxivMar-24https://arxiv.org/abs/2403.00896https://github.com/141forever/DiaHalu
ERBench: An Entity-Relationship based Automatically Verifiable Hallucination Benchmark for Large Language ModelsJio OharxivMar-24https://arxiv.org/abs/2403.05266https://github.com/DILAB-KAIST/ERBench

Survey

TitleAuthorPublishDateLinkSource Code
Survey of Hallucination in Natural Language GenerationZiwei JiarxivFeb 2022https://arxiv.org/abs/2202.03629null
Survey on Factuality in Large Language Models: Knowledge, Retrieval and Domain-SpecificityCunxiang WangarxivOct-23https://arxiv.org/abs/2310.07521null
The Human Factor in Detecting Errors of Large Language Models: A Systematic Literature Review and Future Research DirectionsChristian A. SchillerarxivMar-24https://arxiv.org/abs/2403.09743null
A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open QuestionsLei Huang,arxivNov-23https://arxiv.org/abs/2311.05232null

Other

TitleAuthorPublishDateLinkSource Code
Sources of Hallucination by Large Language Models on Inference TasksNick McKennaEMNLPMay-23https://arxiv.org/abs/2305.14552null
Teaching Language Models to Hallucinate Less with Synthetic TasksErik JonesICLROct-23https://arxiv.org/abs/2310.06827null
In ChatGPT We Trust? Measuring and Characterizing the Reliability of ChatGPTXinyue ShenarxivApr-23https://arxiv.org/abs/2304.08979null
When Large Language Models contradict humans? Large Language Models' Sycophantic BehaviourLeonardo RanaldiarxivNov-23https://arxiv.org/abs/2311.09410null
Calibrated Language Models Must HallucinateAdam Tauman KalaiSTOCNov-23https://arxiv.org/abs/2311.14648null
C-RAG: Certified Generation Risks for Retrieval-Augmented Language ModelsMintong KangICMLFeb-24https://arxiv.org/abs/2402.03181null
Large Language Models are Null-Shot LearnersPittawat TaveekitworachaiarxivJan 2024https://arxiv.org/abs/2401.08273null
Hallucination is Inevitable: An Innate Limitation of Large Language ModelsZiwei XuarxivJan-24https://arxiv.org/abs/2401.11817null
Relying on the Unreliable: The Impact of Language Models' Reluctance to Express UncertaintyKaitlyn ZhouarxivJan 2024https://arxiv.org/abs/2401.06730null
Deficiency of Large Language Models in Finance: An Empirical Examination of HallucinationHaoqiang KangarxivNov-23https://arxiv.org/abs/2311.15548null
Banishing LLM Hallucinations Requires Rethinking GeneralizationJohnny LiarxivJun-24https://arxiv.org/abs/2406.17642null
Seven Failure Points When Engineering a Retrieval Augmented Generation SystemScott BarnettarxivJan-24https://arxiv.org/abs/2401.05856null
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?Zorik GekhmanarxivMay-24https://arxiv.org/abs/2405.05904null
The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive ConversationRongwu XuACLDec-23https://arxiv.org/abs/2312.09085https://llms-believe-the-earth-is-flat.github.io/
Tell me the truth: A system to measure the trustworthiness of Large Language ModelCarlo LipizziarxivMar-24https://arxiv.org/abs/2403.04964null

MLLM

Jailbreak

Attack

TitleAuthorPublishYearLinkSource Code
Visual Adversarial Examples Jailbreak Large Language ModelsXiangyu QiAAAI2024AAAIgithub
Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak AttacksZonghao Yingarxiv2024https://arxiv.org/abs/2406.06302https://github.com/ny1024/jailbreak_gpt4o
Jailbreak Vision Language Models via Bi-Modal Adversarial PromptZonghao Yingarxiv2024https://arxiv.org/abs/2406.04031https://github.com/NY1024/BAP-Jailbreak-Vision-Language-Models-via-Bi-Modal-Adversarial-Prompt
Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language ModelsErfan ShayeganiICLR2024https://openreview.net/forum?id=plmBsXHxgRnull
FigStep: Jailbreaking Large Vision-language Models via Typographic Visual PromptsYichen Gongarxiv2023https://arxiv.org/abs/2311.05608https://github.com/ThuCCSLab/FigStep
Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language ModelsYifan Liarxiv2024https://arxiv.org/abs/2403.09792https://github.com/AoiDragon/HADES
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?Shuo Chenarxiv2024https://arxiv.org/abs/2404.03411null
Cross-Modality Jailbreak and Mismatched Attacks on Medical Multimodal Large Language ModelsXijie Huangarxiv2024https://arxiv.org/abs/2405.20775https://github.com/dirtycomputer/O2M_attack
Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image CharacterSiyuan Maarxiv2024https://arxiv.org/abs/2405.20773null
Jailbreaking Attack against Multimodal Large Language ModelZhenxing Niuarxiv2024https://arxiv.org/abs/2402.02309https://github.com/abc03570128/Jailbreaking-Attack-against-Multimodal-Large-Language-Model

Defense

TitleAuthorPublishYearLinkSource Code
MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language ModelsTianle Guarxiv2024https://arxiv.org/abs/2406.07594https://github.com/Carol-gutianle/MLLMGuard
Cross-Modal Safety Alignment: Is textual unlearning all you need?Trishna Chakrabortyarxiv2024https://arxiv.org/abs/2406.02575null
Unbridled Icarus: A Survey of the Potential Perils of Image Inputs in Multimodal Large Language Model SecurityYihe Fanarxiv2024https://arxiv.org/abs/2404.05264null
AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield PromptingYu Wangarxiv2024https://arxiv.org/abs/2403.09513https://github.com/rain305f/AdaShield
Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language ModelsYongshuo Zongarxiv2024https://arxiv.org/abs/2402.02207https://github.com/ys-zong/VLGuard
JailGuard: A Universal Detection Framework for LLM Prompt-based AttacksXiaoyu Zhangarxiv2024https://arxiv.org/abs/2312.10766null

Benchmark

TitleAuthorPublishYearLinkSource Code
JailBreakV-28K: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak AttacksWeidi Luoarxiv2024https://arxiv.org/abs/2404.03027https://github.com/EddyLuo1232/JailBreakV_28K
MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language ModelsXin Liuarxiv2023https://arxiv.org/abs/2311.17600https://github.com/isXinLiu/MM-SafetyBench
Red Teaming Visual Language ModelsMukai Liarxiv2024https://arxiv.org/abs/2401.12915null
How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMsHaoqin Tuarxiv2023https://arxiv.org/abs/2311.16101https://github.com/UCSC-VLAA/vllm-safety-benchmark
Benchmarking Trustworthiness of Multimodal Large Language Models: A Comprehensive StudyYichi Zhangarxiv2024https://arxiv.org/abs/2406.07057https://multi-trust.github.io/
SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language ModelYongting Zhangarxiv2024https://arxiv.org/abs/2406.12030https://github.com/EchoseChen/SPA-VL-RLHF
MOSSBench: Is Your Multimodal Language Model Oversensitive to Safe Queries?Xirui Liarxiv2024https://arxiv.org/abs/2406.17806https://turningpoint-ai.github.io/MOSSBench/

Survey

TitleAuthorPublishYearLinkSource Code
Safety of Multimodal Large Language Models on Images and TextsXin Liuarxiv2024https://arxiv.org//abs/2402.00357null

Poisoning/Backdoor

Attack

TitleFirst AuthorPublishYearLinkSource Code
Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language ModelsYuancheng Xuarxiv2024https://arxiv.org/abs/2402.06659https://github.com/umd-huang-lab/VLM-Poisoning
Test-Time Backdoor Attacks on Multimodal Large Language ModelsDong Luarxiv2024https://arxiv.org/abs/2402.08577https://sail-sg.github.io/AnyDoor/
VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language ModelsJiawei Liangarxiv2024https://arxiv.org/abs/2402.13851null
Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language ModelsZhenyang Niarxiv2024https://arxiv.org/abs/2404.12916null

Adversarial

Attack

TitleFirst AuthorPublishYearLinkSource Code
Understanding Zero-Shot Adversarial Robustness for Large-Scale ModelsChengzhi Maoarxiv2022https://arxiv.org/abs/2212.07016null
On Evaluating Adversarial Robustness of Large Vision-Language ModelsYunqing Zhaoarxiv2023https://arxiv.org/abs/2305.16934https://github.com/yunqing-me/AttackVLM
On the Adversarial Robustness of Multi-Modal Foundation ModelsChristian SchlarmannICCV AROW2023https://arxiv.org/abs/2308.10741null
Adversarial Illusions in Multi-Modal EmbeddingsTingwei ZhangUSENIX Security2023https://arxiv.org/abs/2308.11804null
Image Hijacks: Adversarial Images can Control Generative Models at RuntimeLuke Baileyarxiv2023https://arxiv.org/abs/2309.00236null
How Robust is Google's Bard to Adversarial Image Attacks?Yinpeng Dongarxiv2023https://arxiv.org/abs/2309.11751https://github.com/thu-ml/Attack-Bard
An Image Is Worth 1000 Lies: Transferability of Adversarial Images across Prompts on Vision-Language ModelsHaochen LuoICLR2024https://openreview.net/pdf?id=nc5GgFAvtkhttps://github.com/Haochen-Luo/CroPA
Inducing High Energy-Latency of Large Vision-Language Models with Verbose ImagesKuofeng GaoICLR2024https://openreview.net/pdf?id=BteuUysuXXhttps://github.com/KuofengGao/Verbose_Images
Hijacking Context in Large Multi-modal ModelsJoonhyun Jeongarxiv2023https://arxiv.org/abs/2312.07553null
On the Robustness of Large Multimodal Models Against Image Adversarial AttacksXuanming Cuiarxiv2023https://arxiv.org/abs/2312.03777null
InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language ModelsXunguang Wangarxiv2023https://arxiv.org/abs/2312.01886https://github.com/xunguangwang/InstructTA
Stop Reasoning! When Multimodal LLMs with Chain-of-Thought Reasoning Meets Adversarial Imagesarxiv2024https://arxiv.org/abs/2402.14899null

Benchmark

TitleFirst AuthorPublishYearLinkSource Code
AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Adversarial Visual-InstructionsHao Zhangarxiv2024https://arxiv.org/abs/2403.09346null
Transferable Multimodal Attack on Vision-Language Pre-training ModelsHaodi WangS&P2024https://www.computer.org/csdl/proceedings-article/sp/2024/313000a102/1Ub239H4xygnull

Hallucination

TitleFirst AuthorPublishYearLinkSource Code
Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction TuningFuxiao LiuICLR2024https://openreview.net/pdf?id=J44HfH4JCghttps://github.com/FuxiaoLiu/LRV-Instruction
Analyzing and Mitigating Object Hallucination in Large Vision-Language ModelsYiyang ZhouICLR2024https://openreview.net/pdf?id=oZDJKTlOUehttps://github.com/YiyangZhou/LURE
A Survey on Hallucination in Large Vision-Language ModelsHanchao Liuarxiv2024https://arxiv.org/abs/2402.00253null
Mitigating Object Hallucination in Large Vision-Language Models via Classifier-Free GuidanceLinxi Zhaoarxiv2024https://arxiv.org/abs/2402.08680null
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided DecodingAilin Dengarxiv2024https://arxiv.org/abs/2402.15300https://github.com/d-ailin/CLIP-Guided-Decoding
Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI FeedbackWenyi Xiaoarxiv2024https://arxiv.org/abs/2404.14233null
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive DecodingXintong Wangarxiv2024https://arxiv.org/abs/2403.18715null
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language ModelsTianrui GuanCVPR2024https://arxiv.org/abs/2310.14566https://github.com/tianyi-lab/HallusionBench

Bias

TitleFirst AuthorPublishYearLinkSource Code
Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference ChallengesChenhang Cuiarxiv2023https://arxiv.org/abs/2311.03287https://github.com/gzcch/Bingo

Others

TitleFirst AuthorPublishYearLinkSource Code
Zero shot VLMs for hate meme detection: Are we there yet?Naquee Rizwanarxiv2024https://arxiv.org/abs/2402.12198null
MemeCraft: Contextual and Stance-Driven Multimodal Meme GenerationHan WangACM MM2024https://arxiv.org/abs/2403.14652null
Moderating Illicit Online Image Promotion for Unsafe User-Generated Content Games Using Large Vision-Language ModelsKeyan GuoUSENIX Security2024https://arxiv.org/abs/2403.18957null

T2I

IP

Protection

TitleAuthorPublishYearLinkSource Code
AIGC-Chain: A Blockchain-Enabled Full Lifecycle Recording System for AIGC Product Copyright ManagementJiajia Jiangarixv2024https://arxiv.org/abs/2406.14966null
PID: Prompt-Independent Data Protection Against Latent Diffusion ModelsAng Liarxiv2024https://arxiv.org/abs/2406.15305null
Glaze: Protecting Artists from Style Mimicry by Text-to-Image ModelsShawn ShanUSENIX Security2023https://arxiv.org/abs/2302.04222null
Generative Watermarking Against Unauthorized Subject-Driven Image SynthesisYihan Maarxiv2023https://arxiv.org/abs/2306.07754null
Toward effective protection against diffusion based mimicry through score distillationHaotian Xuearxiv2023https://arxiv.org/abs/2311.12832https://github.com/xavihart/Diff-Protect
A Watermark-Conditioned Diffusion Model for IP ProtectionRui Minarxiv2024https://arxiv.org/abs/2403.10893null
RAW: A Robust and Agile Plug-and-Play Watermark Framework for AI-Generated Images with Provable GuaranteesXun Xianarxiv2024https://arxiv.org/abs/2403.18774null
A Training-Free Plug-and-Play Watermark Framework for Stable DiffusionGuokai Zhangarxiv2024https://arxiv.org/abs/2404.05607null
Gaussian Shading: Provable Performance-Lossless Image Watermarking for Diffusion ModelsZijin Yangarxiv2024https://arxiv.org/abs/2404.04956null
Lazy Layers to Make Fine-Tuned Diffusion Models More TraceableHaozhe Liuarxiv2024https://arxiv.org/abs/2405.00466null
DiffuseTrace: A Transparent and Flexible Watermarking Scheme for Latent Diffusion ModelLiangqi Leiarxiv2024https://arxiv.org/abs/2405.02696null
AquaLoRA: Toward White-box Protection for Customized Stable Diffusion Models via Watermark LoRAWeitao Fengarxiv2024https://arxiv.org/abs/2405.11135null
A Recipe for Watermarking Diffusion ModelsYunqing Zhaoarixv2023https://arxiv.org/abs/2303.10137https://github.com/yunqing-me/WatermarkDM
Watermarking Diffusion ModelYugeng Liuarxiv2023https://arxiv.org/abs/2305.12502null
Tree-Ring Watermarks: Fingerprints for Diffusion Images that are Invisible and RobustYuxin Wenarxiv2023https://arxiv.org/abs/2305.20030https://github.com/YuxinWenRick/tree-ring-watermark
Generative Models are Self-Watermarked: Declaring Model Authentication through Re-GenerationAditya Desuarxiv2024https://arxiv.org/abs/2402.16889null
Erasing Concepts from Diffusion ModelsRohit Gandikotaarxiv2023https://arxiv.org/abs/2303.07345https://erasing.baulab.info/
Adversarial Example Does Good: Preventing Painting Imitation from Diffusion Models via Adversarial ExamplesChumeng LiangICML2023https://arxiv.org/abs/2302.04578https://github.com/mist-project/mist.git

Violation

TitleAuthorPublishYearLinkSource Code
A Transfer Attack to Image WatermarksYuepeng Huarxiv2024https://arxiv.org/abs/2403.15365null
Disguised Copyright Infringement of Latent Diffusion ModelsYiwei Luarxiv2024https://arxiv.org/abs/2404.06737https://github.com/watml/disguised_copyright_infringement
Stable Signature is Unstable: Removing Image Watermark from Diffusion ModelsYuepeng Huarxiv2024https://arxiv.org/abs/2405.07145null
UnMarker: A Universal Attack on Defensive WatermarkingAndre Kassisarxiv2024https://arxiv.org/abs/2405.08363null
FreezeAsGuard: Mitigating Illegal Adaptation of Diffusion Models via Selective Tensor FreezingKai Huangarxiv2024https://arxiv.org/abs/2405.17472null
Adversarial Perturbations Cannot Reliably Protect Artists From Generative AIRobert Hönigarxiv2024https://arxiv.org/abs/2406.12027null
EnTruth: Enhancing the Traceability of Unauthorized Dataset Usage in Text-to-image Diffusion Models with Minimal and Robust AlterationsJie Renarxiv2024https://arxiv.org/abs/2406.13933null
Leveraging Optimization for Adaptive Attacks on Image WatermarksNils LukasICLR2024https://openreview.net/pdf?id=O9PArxKLe1null

Survey

TitleAuthorPublishYearLinkSource Code
Copyright Protection in Generative AI: A Technical PerspectiveJie Renarxiv2024https://arxiv.org/abs/2402.02333null

Privacy

Attack

TitleAuthorPublishYearLinkSource Code
Recovering the Pre-Fine-Tuning Weights of Generative ModelsEliahu Horwitzarxiv2024https://arxiv.org/abs/2402.10208null
Membership Inference Attacks Against Text-to-image Generation ModelsYixin Wuarxiv2022https://arxiv.org/abs/2210.00968null
Class Attribute Inference Attacks: Inferring Sensitive Class Information by Diffusion-Based Attribute ManipulationsLukas Struppekarxiv2023https://arxiv.org/abs/2303.09289null
White-box Membership Inference Attacks against Diffusion ModelsYan Pangarxiv2023https://arxiv.org/abs/2308.06405null
An Efficient Membership Inference Attack for the Diffusion Model by Proximal InitializationFei KongICLR2024https://openreview.net/pdf?id=rpH9FcCEV6https://github.com/kong13661/PIA
Black-box Membership Inference Attacks against Fine-tuned Diffusion ModelsYan Pangarxiv2023https://arxiv.org/abs/2312.08207null
Prompt Stealing Attacks Against Text-to-Image Generation ModelsXinyue Shenarxiv2023https://arxiv.org/abs/2302.09923null
Shake to Leak: Fine-tuning Diffusion Models Can Amplify the Generative Privacy RiskZhangheng Liarxiv2024https://arxiv.org/html/2403.09450v1https://github.com/VITA-Group/Shake-to-Leak
Is Diffusion Model Safe? Severe Data Leakage via Gradient-Guided Diffusion ModelJiayang Mengarxiv2024https://arxiv.org/abs/2406.09484null
Extracting Training Data from Unconditional Diffusion ModelsYunhao Chenarxiv2024https://arxiv.org/abs/2406.12752null
Extracting Training Data from Diffusion ModelsNicholas Carliniarxiv2023https://arxiv.org/abs/2301.13188null
Towards Black-Box Membership Inference Attack for Diffusion ModelsJingwei Liarxiv2024https://arxiv.org/abs/2405.20771null
Visual Privacy Auditing with Diffusion ModelsKristian Schwethelmarxiv2024https://arxiv.org/abs/2403.07588null
Membership Inference on Text-to-Image Diffusion Models via Conditional Likelihood DiscrepancyShengfang Zhaiarxiv2024https://arxiv.org/abs/2405.14800null
Extracting Prompts by Inverting LLM OutputsCollin Zhangarxiv2024https://arxiv.org/abs/2405.15012null

Defense

TitleAuthorPublishYearLinkSource Code
Differentially Private Fine-Tuning of Diffusion ModelsYu-Lin Tsaiarxiv2024https://arxiv.org/abs/2406.01355https://anonymous.4open.science/r/DP-LORA-F02F
Privacy-Preserving Diffusion Model Using Homomorphic EncryptionYaojian Chenarxiv2024https://arxiv.org/abs/2403.05794null
Efficient Differentially Private Fine-Tuning of Diffusion ModelsJing Liuarxiv2024https://arxiv.org/abs/2406.05257null
Differentially Private Synthetic Data via Foundation Model APIs 1: ImagesZinan LinICLR2024https://openreview.net/pdf?id=YEhQs8POIohttps://github.com/microsoft/DPSDA
Anti-DreamBooth: Protecting users from personalized text-to-image synthesisThanh Van LeICCV2023https://arxiv.org/abs/2303.15433https://github.com/VinAIResearch/Anti-DreamBooth.git
Unlearnable Examples for Diffusion Models: Protect Data from Unauthorized ExploitationZhengyue Zhaoarxiv2023https://arxiv.org/abs/2306.01902null
Can Protective Perturbation Safeguard Personal Data from Being Exploited by Stable Diffusion?Zhengyue Zhaoarxiv2023https://arxiv.org/abs/2312.00084null
MetaCloak: Preventing Unauthorized Subject-driven Text-to-image Diffusion-based Synthesis via Meta-learningYixin Liuarxiv2023https://arxiv.org/abs/2311.13127https://github.com/liuyixin-louis/MetaCloak

Benchmark

TitleAuthorPublishYearLinkSource Code
Raccoon: Prompt Extraction Benchmark of LLM-Integrated ApplicationsJunlin Wangarxiv2024https://arxiv.org/abs/2406.06737https://github.com/M0gician/RaccoonBench

Memorization

TitleAuthorPublishYearLinkSource Code
SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and GenerationChongyu FanICLR2024https://openreview.net/pdf?id=gn0mIhQGNMhttps://github.com/OPTML-Group/Unlearn-Saliency
Ring-A-Bell! How Reliable are Concept Removal Methods For Diffusion Models?Yu-Lin TsaiICLR2024https://openreview.net/pdf?id=lm7MRcsFiShttps://github.com/chiayi-hsu/Ring-A-Bell
Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion ModelsYimeng Zhangarxiv2024https://arxiv.org/abs/2405.15234https://github.com/OPTML-Group/AdvUnlearn
Espresso: Robust Concept Filtering in Text-to-Image ModelsAnudeep Dasarxiv2024https://arxiv.org/abs/2404.19227null
Machine Unlearning for Image-to-Image Generative ModelsGuihong Liarxiv2024https://arxiv.org/abs/2402.00351https://github.com/jpmorganchase/l2l-generator-unlearning
MACE: Mass Concept Erasure in Diffusion ModelsShilin Luarxiv2024https://arxiv.org/abs/2403.06135https://github.com/Shilin-LU/MACE
Unveiling and Mitigating Memorization in Text-to-image Diffusion Models through Cross AttentionJie Renarxiv2024https://arxiv.org/abs/2403.11052https://github.com/renjie3/MemAttn
Could It Be Generated? Towards Practical Analysis of Memorization in Text-To-Image Diffusion ModelsZhe Maarxiv2024https://arxiv.org/abs/2405.05846null

Deepfake

Construction

TitleAuthorPublishYearLinkSource Code
An Analysis of Recent Advances in Deepfake Image Detection in an Evolving Threat LandscapeSifat Muhammad AbdullahS&P2024https://arxiv.org/abs/2404.16212null
Robustness of AI-Image Detectors: Fundamental Limits and Practical AttacksMehrdad SaberiICLR2024https://openreview.net/pdf?id=dLoAdIKENchttps://github.com/mehrdadsaberi/watermark_robustness

Detection

TitleAuthorPublishYearLinkSource Code
DeepFake-O-Meter v2.0: An Open Platform for DeepFake DetectionYan Juarxiv2024https://arxiv.org/abs/2404.13146null
DE-FAKE: Detection and Attribution of Fake Images Generated by Text-to-Image Generation ModelsZeyang Shaarxiv2022https://arxiv.org/abs/2210.06998null
Organic or Diffused: Can We Distinguish Human Art from AI-generated Images?Anna Yoo Jeong Haarxiv2024https://arxiv.org/abs/2402.03214null
Watermark-based Detection and Attribution of AI-Generated ContentZhengyuan Jiangarxiv2024https://arxiv.org/abs/2404.04254null
An Analysis of Recent Advances in Deepfake Image Detection in an Evolving Threat LandscapeSifat Muhammad AbdullahS&P2024https://arxiv.org/abs/2404.16212null

Benchmark

TitleAuthorPublishYearLinkSource Code
The Adversarial AI-Art: Understanding, Generation, Detection, and BenchmarkingYuying Liarxiv2024https://arxiv.org/abs/2404.14581null

Bias

TitleAuthorPublishYearLinkSource Code
Finetuning Text-to-Image Diffusion Models for FairnessXudong ShenICLR2024https://openreview.net/pdf?id=hnrB5YHoYuhttps://sail-sg.github.io/finetune-fair-diffusion/
ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image GenerationAkshita Jhaarxiv2024https://arxiv.org/abs/2401.06310null

Backdoor

Attack

TitleAuthorPublishYearLinkSource Code
Rickrolling the Artist: Injecting Backdoors into Text Encoders for Text-to-Image SynthesisLukas StruppekICCV2023https://arxiv.org/abs/2211.02408null
Generating Potent Poisons and Backdoors from Scratch with Guided DiffusionHossein Souriarxiv2024https://arxiv.org/abs/2403.16365null

Defense

TitleAuthorPublishYearLinkSource Code
Diffusion Denoising as a Certified Defense against Clean-label PoisoningSanghyun Hongarxiv2024https://arxiv.org/abs/2403.11981null
Leveraging Diffusion-Based Image Variations for Robust Training on Poisoned DataNeurIPS Workshop2023https://arxiv.org/abs/2310.06372null
UFID: A Unified Framework for Input-level Backdoor Detection on Diffusion ModelsZihan Guanarxiv2024https://arxiv.org/abs/2404.01101https://github.com/GuanZihan/official_UFID
Invisible Backdoor Attacks on Diffusion Modelsarxiv2024https://arxiv.org/abs/2406.00816https://github.com/invisibleTriggerDiffusion/invisible_triggers_for_diffusion
Watch the Watcher! Backdoor Attacks on Security-Enhancing Diffusion ModelsChangjiang Liarxiv2024https://arxiv.org/abs/2406.09669null
Injecting Bias in Text-To-Image Models via Composite-Trigger BackdoorsAli Naseharxiv2024https://arxiv.org/abs/2406.15213null

Adversarial

Attack

TitleAuthorPublishYearLinkSource Code
Diffusion-Based Adversarial Sample Generation for Improved Stealthiness and ControllabilityHaotian XueNeurIPS2023https://arxiv.org/abs/2305.16494https://github.com/xavihart/Diff-PGD
Stable Diffusion is UnstableChengbin Duarxiv2023https://arxiv.org/abs/2306.02583null
DiffAttack: Evasion Attacks Against Diffusion-Based Adversarial PurificationMintong KangNeurIPS2023https://arxiv.org/abs/2311.16124null
Exploring Adversarial Attacks against Latent Diffusion Model from the Perspective of Adversarial TransferabilityJunxi Chenarxiv2024https://arxiv.org/abs/2401.07087null
Revealing Vulnerabilities in Stable Diffusion via Targeted AttacksChenyu Zhangarxiv2024https://arxiv.org/abs/2401.08725https://github.com/datar001/Revealing-Vulnerabilities-in-Stable-Diffusion-via-Targeted-Attacks
Cheating Suffix: Targeted Attack to Text-To-Image Diffusion Models with Multi-Modal PriorsDingcheng Yangarxiv2024https://arxiv.org/abs/2402.01369https://github.com/ydc123/MMP-Attack
Groot: Adversarial Testing for Generative Text-to-Image Models with Tree-based Semantic TransformationYi Liuarxiv2024https://arxiv.org/abs/2402.12100null
BSPA: Exploring Black-box Stealthy Prompt Attacks against Image GeneratorsYu Tianarxiv2024https://arxiv.org/abs/2402.15218null
Perturbing Attention Gives You More Bang for the Buck: Subtle Imaging Perturbations That Efficiently Fool Customized Diffusion ModelsJingyao Xuarxiv2024https://arxiv.org/abs/2404.15081null
Investigating and Defending Shortcut Learning in Personalized Diffusion ModelsYixin Liuarxiv2024https://arxiv.org/abs/2406.18944null
Unsafe Diffusion: On the Generation of Unsafe Images and Hateful Memes From Text-To-Image ModelsYiting Quarxiv2024https://arxiv.org/abs/2305.13873null
ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign Usersarxiv2024https://arxiv.org/abs/2405.19360https://github.com/GuanlinLee/ART

Defense

TitleAuthorPublishYearLinkSource Code
Raising the Cost of Malicious AI-Powered Image EditingHadi Salmanarxiv2023https://arxiv.org/abs/2302.06588null
Adversarial Examples are Misaligned in Diffusion Model ManifoldsPeter Lorenzarxiv2024https://arxiv.org/abs/2401.06637null
Universal Prompt Optimizer for Safe Text-to-Image GenerationZongyu WuNAACL2024https://arxiv.org/abs/2402.10882https://github.com/wzongyu/POSI
Adversarial Nibbler: An Open Red-Teaming Method for Identifying Diverse Harms in Text-to-Image GenerationJessica Quayearxiv2024https://arxiv.org/abs/2403.12075null
SafeGen: Mitigating Unsafe Content Generation in Text-to-Image ModelsXinfeng Liarxiv2024https://arxiv.org/abs/2404.06666null

Agent

Backdoor Attack/Defense

TitleAuthorVenueYearLinkSource Code
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety TrainingEvan Hubingerarxiv2024https://arxiv.org/abs/2401.05566null
BadAgent: Inserting and Activating Backdoor Attacks in LLM AgentsYifei WangACL2024https://arxiv.org/abs/2406.03007https://github.com/DPamK/BadAgent

Adversarial Attack/Defense

TitleAuthorVenueYearLinkSource Code
Large Language Model Sentinel: Advancing Adversarial Robustness by LLM AgentGuang Linarxiv2024https://arxiv.org/abs/2405.20770null
Adversarial Attacks on Multimodal AgentsChen Henry Wuarxiv2024https://arxiv.org/abs/2406.12814https://github.com/ChenWu98/agent-attack
AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM AgentsEdoardo Debenedettiarxiv2024https://arxiv.org/abs/2406.13352https://github.com/ethz-spylab/agentdojo
GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled ReasoningZhen Xiangarxiv2024https://arxiv.org/abs/2406.09187null

Jailbreak

TitleAuthorVenueYearLinkSource Code
Evil Geniuses: Delving into the Safety of LLM-based AgentsYu Tianarxiv2023https://arxiv.org/abs/2311.11855https://github.com/T1aNS1R/Evil-Geniuses
Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially FastXiangming GuICML2024https://arxiv.org/abs/2402.08567https://sail-sg.github.io/Agent-Smith/
AutoDefense: Multi-Agent LLM Defense against Jailbreak AttacksYifan Zengarxiv2024https://arxiv.org/abs/2403.04783https://github.com/XHMY/AutoDefense

Prompt Injection

TitleAuthorVenueYearLinkSource Code
WIPI: A New Web Threat for LLM-Driven Web AgentsFangzhou Wuarxiv2024https://arxiv.org/abs/2402.16965null
InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model AgentsQiusi Zhanarxiv2024https://arxiv.org/abs/2403.02691https://github.com/uiuc-kang-lab/InjecAgent

Hallucination

TitleAuthorVenueYearLinkSource Code
Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Duarxiv2023https://arxiv.org/abs/2305.14325null
MetaGPT: Meta Programming for A Multi-Agent Collaborative FrameworkICLR2024https://openreview.net/pdf?id=VtmBAGCN7onull
Can LLMs Produce Faithful Explanations For Fact-checking? Towards Faithful Explainable Fact-Checking via Multi-Agent DebateKyungha Kimarxiv2024https://arxiv.org/abs/2402.07401null

Others

TitleAuthorVenueYearLinkSource Code
Achieving Fairness in Multi-Agent MDP Using Reinforcement LearningPeizhong JuICLR2024https://openreview.net/pdf?id=yoVq2BGQdPnull
Agent Alignment in Evolving Social NormsShimin Liarxiv2024https://arxiv.org/abs/2401.04620null
Air Gap: Protecting Privacy-Conscious Conversational AgentsEugene Bagdasaryanarxiv2024https://arxiv.org/abs/2405.05175null
Secret Collusion Among Generative AI AgentsSumeet Ramesh Motwaniarxiv2024https://arxiv.org/abs/2402.07510v1null

Survey

TitleAuthorVenueYearLinkSource Code
Security of AI AgentsYifeng Hearxiv2024https://arxiv.org/abs/2406.08689null

Competition

TitleOrganizerYearLinkCategory
Machine Learning Model Attribution ChallengeMITRE2022https://mlmac.io/Privacy
Training Data Extraction ChallengeGoogle2022https://github.com/google-research/lm-extraction-benchmarkPrivacy
Find the Trojan: Universal Backdoor Detection in Aligned LLMsETHZ2024https://github.com/ethz-spylab/rlhf_trojan_competitionSecurity
LLM - Detect AI Generated TextThe Learning Agency Lab2023https://www.kaggle.com/competitions/llm-detect-ai-generated-text/overviewDeepfake
Deepfake Detection ChallengeKaggle2019https://www.kaggle.com/c/deepfake-detection-challengeDeepfake
Large Language Model Capture-the-Flag (LLM CTF) CompetitionKaggle2024https://ctf.spylab.ai/Safety

Leaderboard

TitleFirst AuthorPublishYearLink
LLM Safety LeaderboardBoxin WangNeurIPS2023https://huggingface.co/spaces/AI-Secure/llm-trustworthy-leaderboard
Hallucinations LeaderboardPasquale Minervininull2023https://huggingface.co/spaces/hallucinations-leaderboard/leaderboard
JailbreakBenchPatrick Chaoarxiv2024https://jailbreakbench.github.io/
A Comprehensive Study of Trustworthiness in Large Language Models.Lichao Sunarxiv2024https://trustllmbenchmark.github.io/TrustLLM-Website/leaderboard.html
PromptBench: Towards Evaluating the Robustness of Large Language Models on Adversarial PromptsKaijie Zhuarxiv2023https://llm-eval.github.io/pages/leaderboard/advprompt.html
Leaderboard Comparing LLM Performance at Producing Hallucinations when Summarizing Short DocumentsVectaragithub2023https://github.com/vectara/hallucination-leaderboard

Arena

TitleInstitutionYearLink
LMSYS Chatbot Arena (Multimodal): Benchmarking LLMs and VLMs in the WildLMSYS Org2024https://arena.lmsys.org/
中文大模型竞技场ModelScope2024https://modelscope.cn/studios/LLMZOO/Chinese-Arena/summary
司南 OpenCompass 大模型竞技场Shanghai AI Lab2024https://opencompass.org.cn/arena
模型广场Coze2024https://www.coze.cn/model/arena

Book

TitleAuthorPublishYearLink
Adversarial Machine Learning A Taxonomy and Terminology of Attacks and MitigationsApostol Vassilevonline2024https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-2e2023.pdf
人工智能安全方滨兴电子工业出版社2022null
人工智能安全曾剑平清华大学出版社2022null
人工智能安全陈左宁电子工业出版社2024null
AI安全:技术与实战腾讯安全朱雀实验室电子工业出版社2022null
Trustworthy Machine LearningKush R. Varshneynull2022null
人工智能:数据与模型安全姜育刚机械工业出版社2024null


Star History Chart


Contributors

NY1024

35 commits

NY1024/Awesome-Trustworthy-GenAI

8

35 commits

updated Jul 5, 2024

See the code

README

Awesome-Trustworthy-GenAI Awesome

GitHub stars GitHub forks


Update[04/07/2024]


Categories

  • Paper
    • LLM
      • Jailbreak
        • Attack
        • Defense
        • Benchmark
        • Survey
        • Others
      • Hallucination
        • Mitigation
        • Detection
        • Benchmark
        • Survey
        • Others
    • MLLM
      • Jailbreak
        • Attack
        • Defense
        • Benchmark
        • Survey
      • Poisoning/Backdoor
        • Attack
        • Defense
      • Adversarial
        • Attack
        • Defense
        • Benchmark
      • Hallucination
      • Bias
      • Others
    • T2I
      • IP
        • Protection
        • Violation
        • Survey
      • Privacy
        • Attack
        • Defense
        • Benchmark
      • Memorization
      • Deepfake
        • Construction
        • Detection
        • Benchmark
      • Bias
      • Backdoor
        • Attack
        • Defense
      • Adversarial
        • Attack
        • Defense
    • Agent
      • Backdoor Attack/Defense
      • Adversarial Attack/Defense
      • Jaikbreak
      • Prompt Injection
      • Hallucination
      • Others
      • Survey
  • Book
  • Tutorial
  • Arena
  • Competition

LLM

Jailbreak

Attack

TitleAuthorPublishDateLinkSource Code
Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow InstructionsFederico BianchiICLR PosterMar-24https://openreview.net/pdf?id=gT5hALch9znull
FuzzLLM: A Novel and Universal Fuzzing Framework for Proactively Discovering Jailbreak Vulnerabilities in Large Language ModelsDongyu YaoICASSPSep-23https://arxiv.org/abs/2309.05274https://arxiv.org/pdf/2309.05274
GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak PromptsJiahao YuarxivSep-23https://arxiv.org/abs/2309.10253https://github.com/sherdencooper/GPTFuzz
Open Sesame! Universal Black Box Jailbreaking of Large Language ModelsRaz LapidarxivSep-23https://arxiv.org/abs/2309.01446null
Red-Teaming Large Language Models using Chain of Utterances for Safety-AlignmentRishabh BhardwajarxivAug-23https://arxiv.org/abs/2308.09662null
Jailbroken: How Does LLM Safety Training Fail?Alexander WeiNeurIPSJul-23https://arxiv.org/abs/2307.02483null
MasterKey: Automated Jailbreak Across Multiple Large Language Model ChatbotsGelei DengNDSSJul-23https://arxiv.org/abs/2307.08715null
Universal and Transferable Adversarial Attacks on Aligned Language ModelsAndy ZouNeurIPSJul 2023https://arxiv.org/abs/2307.15043http://llm-attacks.org/
Adversarial Demonstration Attacks on Large Language ModelsJiongxiao WangarxivMay-23https://arxiv.org/abs/2305.14950null
Catastrophic Jailbreak of Open-source LLMs via Exploiting GenerationYangsibo HuangICLRMar-24https://openreview.net/pdf?id=r42tSSCHPhhttps://github.com/Princeton-SysML/Jailbreak_LLM
GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via CipherYouliang YuanJan-24https://openreview.net/pdf?id=MbfAK4s61Ahttps://github.com/RobustNLP/CipherChat
AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng LiuICLRJan-24https://openreview.net/pdf?id=7Jwpw4qKkbhttps://github.com/SheltonLiu-N/AutoDAN
Low-Resource Languages Jailbreak GPT-4Zheng-Xin YongNeurIPS WorkshopOct-23https://arxiv.org/abs/2310.02446null
Multi-step Jailbreaking Privacy Attacks on ChatGPTHaoran LiEMNLPApr-23https://arxiv.org/abs/2304.05197null
Attack Prompt Generation for Red Teaming and Defending Large Language ModelsBoyi DengEMNLPOct-23https://arxiv.org/abs/2310.12505https://github.com/Aatrox103/SAP
Multilingual Jailbreak Challenges in Large Language ModelsYue DengICLROct-23https://arxiv.org/abs/2310.06474https://github.com/DAMO-NLP-SG/multilingual-safety-for-LLMs%7D
Weak-to-Strong Jailbreaking on Large Language ModelsXuandong ZhaoarxivJan-24https://arxiv.org/abs/2401.17256https://github.com/XuandongZhao/weak-to-strong
How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMsYi ZengACLJan-24https://arxiv.org/abs/2401.06373https://chats-lab.github.io/persuasive_jailbreaker/
Jailbreaking GPT-4V via Self-Adversarial Attacks with System PromptsYuanwei WuarxivNov-23https://arxiv.org/abs/2311.09127null
Safety Alignment in NLP Tasks: Weakly Aligned Summarization as an In-Context AttackYu FuACLDec-23https://arxiv.org/abs/2312.06924null
Tree of Attacks: Jailbreaking Black-Box LLMs AutomaticallyAnay MehrotraarxivDec-23https://arxiv.org/abs/2312.02119https://github.com/RICommunity/TAP
A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models EasilyPeng DingNAACLNov-23https://arxiv.org/abs/2311.08268https://github.com/NJUNLP/ReNeLLM
Goal-Oriented Prompt Attack and Safety Evaluation for LLMsChengyuan LiuarxivSep-23https://arxiv.org/abs/2309.11830https://github.com/liuchengyuan123/CPAD
Play Guessing Game with LLM: Indirect Jailbreak Attack with Implicit CluesZhiyuan ChangarxivFeb-24https://arxiv.org/abs/2402.09091null
A Cross-Language Investigation into Jailbreak Attacks in Large Language ModelsJie LiarxivJan-24https://arxiv.org/abs/2401.16765null
Analyzing the Inherent Response Tendency of LLMs: Real-World Instructions-Driven JailbreakYanrui DuarxivDec-23https://arxiv.org/abs/2312.04127null
All in How You Ask for It: Simple Black-Box Method for Jailbreak AttacksKazuhiro TakemotoarxivJan 2024https://arxiv.org/abs/2401.09798null
DeepInception: Hypnotize Large Language Model to Be JailbreakerXuan LiarxivNov-23https://arxiv.org/abs/2311.03191https://github.com/tmlr-group/DeepInception
Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona ModulationRusheb ShaharxivNov-23https://arxiv.org/abs/2311.03348null
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-RailsNeal MangaokararxivFeb-24https://arxiv.org/abs/2402.15911null
Semantic Mirror Jailbreak: Genetic Algorithm Based Jailbreak Prompts Against Open-source LLMsXiaoxia LiarxivFeb-24https://arxiv.org/abs/2402.14872null
ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix EmbeddingsHao WangarxivFeb-24https://arxiv.org/abs/2402.16006null
A StrongREJECT for Empty JailbreaksAlexandra SoulyarxivFeb-24https://arxiv.org/abs/2402.10260https://github.com/alexandrasouly/strongreject
Jailbreaking Black Box Large Language Models in Twenty QueriesPatrick ChaoarxivOct-23https://arxiv.org/abs/2310.08419https://github.com/patrickrchao/JailbreakingLLMs
Jailbreak and Guard Aligned Language Models with Only Few In-Context DemonstrationsZeming WeiarxivOct-23https://arxiv.org/abs/2310.06387null
AutoDAN: Interpretable Gradient-Based Adversarial Attacks on Large Language ModelsSicheng ZhuarxivOct-23https://arxiv.org/abs/2310.15140null
Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security AttacksDaniel KangarxivFeb 2023https://arxiv.org/abs/2302.05733null
Pandora: Jailbreak GPTs by Retrieval Augmented Generation PoisoningGelei DengarxivFeb-24https://arxiv.org/abs/2402.08416null
Jailbreaking Proprietary Large Language Models using Word Substitution CipherDivij HandaarxivFeb-24https://arxiv.org/abs/2402.10601null
PAL: Proxy-Guided Black-Box Attack on Large Language ModelsChawin SitawarinarxivFeb-24https://arxiv.org/abs/2402.09674https://github.com/chawins/pal
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMsFengqing JiangarxivFeb-24https://arxiv.org/abs/2402.11753https://github.com/uw-nsl/ArtPrompt
Query-Based Adversarial Prompt GenerationJonathan HayasearxivFeb-24https://arxiv.org/abs/2402.12329null
Coercing LLMs to do and reveal (almost) anythingJonas GeipingarxivFeb-24https://arxiv.org/abs/2402.14020https://github.com/JonasGeiping/carving
COLD-Attack: Jailbreaking LLMs with Stealthiness and ControllabilityXingang GuoICMLFeb-24https://arxiv.org/abs/2402.08679https://github.com/Yu-Fangxu/COLD-Attack
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive AttacksMaksym AndriushchenkoarxivApr-24https://arxiv.org/abs/2404.02151https://github.com/tml-epfl/llm-adaptive-attacks
Sandwich attack: Multi-language Mixture Adaptive Attack on LLMsBibek UpadhayayarxivApr 2024https://arxiv.org/abs/2404.07242null
Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language ModelsZhiyuan YuUSENIX SecurityMar-24https://arxiv.org/abs/2403.17336null
CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code CompletionQibing RenACLMar-24https://arxiv.org/abs/2403.07865https://github.com/renqibing/CodeAttack
Tastle: Distract Large Language Models for Automatic Jailbreak AttackZeguan XiaoarxivMar-24https://arxiv.org/abs/2403.08424null
AmpleGCG: Learning a Universal and Transferable Generative Model of Adversarial Suffixes for Jailbreaking Both Open and Closed LLMsZeyi LiaoarxivApr-24https://arxiv.org/abs/2404.07921https://github.com/OSU-NLP-Group/AmpleGCG
Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLMXikang YangarxivMay-24https://arxiv.org/abs/2405.05610null
GUARD: Role-playing to Generate Natural-language Jailbreakings to Test Guideline Adherence of Large Language ModelsHaibo JinarxivFeb-24https://arxiv.org/abs/2402.03299null
Don't Say No: Jailbreaking LLM by Suppressing RefusalYukai ZhouarxivApr-24https://arxiv.org/abs/2404.16369null
AdvPrompter: Fast Adaptive Adversarial Prompting for LLMsAnselm PaulusarxivApr-24https://arxiv.org/abs/2404.16873https://github.com/facebookresearch/advprompter
Lockpicking LLMs: A Logit-Based Jailbreak Using Token-level ManipulationYuxi LiarxivMay-24https://arxiv.org/abs/2405.13068null
Can LLMs Deeply Detect Complex Malicious Queries? A Framework for Jailbreaking via Obfuscating IntentShang ShangarxivMay-24https://arxiv.org/abs/2405.03654null
GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-ExplanationGovind RamesharxivMay-24https://arxiv.org/abs/2405.13077null
Poisoned LangChain: Jailbreak LLMs by LangChainZiqiu WangarxivJun-24https://arxiv.org/abs/2406.18122https://github.com/CAM-FSS/jailbreak-langchain
Covert Malicious Finetuning: Challenges in Safeguarding LLM AdaptationDanny HalawiICMLJun-24https://arxiv.org/abs/2406.20053null
Voice Jailbreak Attacks Against GPT-4oXinyue ShenarxivMay-24https://arxiv.org/abs/2405.19103https://github.com/TrustAIRLab/VoiceJailbreakAttack
StructuralSleight: Automated Jailbreak Attacks on Large Language Models Utilizing Uncommon Text-Encoded StructureBangxin LiarxivJun-24https://arxiv.org/abs/2406.08754null
Knowledge-to-Jailbreak: One Knowledge Point Worth One AttackShangqing TuarxivJun-24https://arxiv.org/abs/2406.11682https://github.com/THU-KEG/Knowledge-to-Jailbreak/
CodeChameleon: Personalized Encryption Framework for Jailbreaking Large Language ModelsHuijie LvarxivFeb-24https://arxiv.org/abs/2402.16717https://github.com/huizhang-L/CodeChameleon
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM JailbreakersXirui LiarxivFeb-24https://arxiv.org/abs/2402.16914https://github.com/xirui-li/DrAttack
Making Them Ask and Answer: Jailbreaking Large Language Models in Few Queries via Disguise and ReconstructionTong LiuUSENIX SecurityFeb-24https://arxiv.org/abs/2402.18104https://github.com/LLM-DRA/DRA
Leveraging the Context through Multi-Round Interactions for Jailbreaking AttacksYixin ChengarxivFeb-24https://arxiv.org/abs/2402.09177null
Jailbreaking Large Language Models Against Moderation Guardrails via Cipher CharactersHaibo JinarxivMay-24https://arxiv.org/abs/2405.20413null
RL-JACK: Reinforcement Learning-powered Black-box Jailbreaking Attack against LLMsXuan ChenarxivJun-24https://arxiv.org/abs/2406.08725null
Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their DefensesXiaosen ZhengarxivJun-24https://arxiv.org/abs/2406.01288https://github.com/sail-sg/I-FSJ
AutoJailbreak: Exploring Jailbreak Attacks and Defenses through a Dependency LensLin LuarxivJun-24https://arxiv.org/abs/2406.03805null
Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak AttackMark RussinovicharxivApr-24https://arxiv.org/abs/2404.01833null
MART: Improving LLM Safety with Multi-round Automatic Red-TeamingSuyu GearxivNov-23https://arxiv.org/abs/2311.07689null
Virtual Context: Enhancing Jailbreak Attacks with Special Token InjectionYuqi ZhouarxivJun 2024https://arxiv.org/abs/2406.19845null
When LLM Meets DRL: Advancing Jailbreaking Efficiency via DRL-guided SearchXuan ChenarxivJun-24https://arxiv.org/abs/2406.08705null
Improved Techniques for Optimization-Based Jailbreaking on Large Language ModelsXiaojun JiaarxivMay-24https://arxiv.org/abs/2405.21018https://github.com/jiaxiaojunQAQ/I-GCG

Defense

TitleAuthorPublishDateLinkSource Code
Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHFAnand SiththaranjanICLRApr-24https://openreview.net/pdf?id=0tWTxYYPnWnull
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLMBochuan CaoACL2024https://arxiv.org/abs/2309.14348null
Detecting Language Model Attacks with PerplexityGabriel AlonarxivAug-23https://arxiv.org/abs/2308.14132null
Certifying LLM Safety against Adversarial PromptingAounon KumararxivSep-23https://arxiv.org/abs/2309.02705https://github.com/aounon/certified-llm-safety
Baseline Defenses for Adversarial Attacks Against Aligned Language ModelsNeel JainarxivSep-23https://arxiv.org/abs/2309.00614null
SmoothLLM: Defending Large Language Models Against Jailbreaking AttacksAlexander RobeyarxivOct-23https://arxiv.org/abs/2310.03684https://github.com/arobey1/smooth-llm
Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting JailbreaksAbhinav RaoLREC-COLINGMay-23https://arxiv.org/abs/2305.14965null
Intention Analysis Makes LLMs A Good Jailbreak DefenderYuqi ZhangarxivJan-24https://arxiv.org/abs/2401.06561https://github.com/alphadl/SafeLLM_with_IntentionAnalysis
On Prompt-Driven Safeguarding for Large Language ModelsChujie ZhengICMLJan-24https://arxiv.org/abs/2401.18018https://github.com/chujiezheng/LLM-Safeguard
Robust Prompt Optimization for Defending Language Models Against Jailbreaking AttacksAndy ZhouarxivJan-24https://arxiv.org/abs/2401.17263https://github.com/lapisrocks/rpo
SPML: A DSL for Defending Language Models Against Prompt AttacksReshabh K SharmaarxivFeb-24https://arxiv.org/abs/2402.11755https://prompt-compiler.github.io/SPML/
LLMs Can Defend Themselves Against Jailbreaking in a Practical Manner: A Vision PaperDaoyuan WuarxivFeb-24https://arxiv.org/abs/2402.15727null
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware DecodingZhangchen XuACLFeb-24https://arxiv.org/abs/2402.08983https://github.com/uw-nsl/SafeDecoding
GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient AnalysisYueqi XieACLFeb-24https://arxiv.org/abs/2402.13494https://github.com/xyq7/GradSafe
Defending Large Language Models against Jailbreak Attacks via Semantic SmoothingJiabao JiarxivFeb-24https://arxiv.org/abs/2402.16192https://github.com/UCSB-NLP-Chang/SemanticSmooth
Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-TuningAdib HasanarxivJan-24https://arxiv.org/abs/2401.10862null
Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-RefinementHeegyu KimarxivFeb-24https://arxiv.org/abs/2402.15180null
Protecting Your LLMs with Information BottleneckZichuan LiuarxivMay-24https://arxiv.org/abs/2404.13968https://github.com/zichuan-liu/IB4LLMs
Eraser: Jailbreaking Defense in Large Language Models via Unlearning Harmful KnowledgeWeikai LuarxivApr-24https://arxiv.org/abs/2404.05880https://github.com/ZeroNLP/Eraser
RigorLLM: Resilient Guardrails for Large Language Models against Undesired ContentZhuowen YuanarxivMar-24https://arxiv.org/abs/2403.13031null
Detoxifying Large Language Models via Knowledge EditingMengru WangACLMar-24https://arxiv.org/abs/2403.14472https://github.com/zjunlp/EasyEdit
AutoDefense: Multi-Agent LLM Defense against Jailbreak AttacksYifan ZengarxivMar-24https://arxiv.org/abs/2403.04783https://github.com/XHMY/AutoDefense
Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss LandscapesXiaomeng HuarxivMar-24https://arxiv.org/abs/2403.00867https://huggingface.co/spaces/TrustSafeAI/GradientCuff-Jailbreak-Defense
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMsFan LiuarxivJun-24https://arxiv.org/abs/2406.06622null
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language ModelsLiwei JiangarxivJun-24https://arxiv.org/abs/2406.18510null
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMsSeungju HanarxivJun-24https://arxiv.org/abs/2406.18495null
SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical MannerXunguang WangarxivJun-24https://arxiv.org/abs/2406.05498null
Merging Improves Self-Critique Against Jailbreak AttacksVictor GallegoarxivJun-24https://arxiv.org/abs/2406.07188https://github.com/vicgalle/merging-self-critique-jailbreaks
Defensive Prompt Patch: A Robust and Interpretable Defense of LLMs against Jailbreak AttacksChen XiongarxivMay-24https://arxiv.org/abs/2405.20099null
Mitigating Fine-tuning based Jailbreak Attack with Backdoor Enhanced Safety AlignmentJiongxiao WangarxivFeb-24https://arxiv.org/abs/2402.14968https://jayfeather1024.github.io/Finetuning-Jailbreak-Defense/
Improving Alignment and Robustness with Circuit BreakersAndy ZouarxivJun-24https://arxiv.org/abs/2406.04313https://github.com/blackswan-ai/circuit-breakers
Robustifying Safety-Aligned Large Language Models through Clean Data CurationXiaoqun LiuarxivMay-24https://arxiv.org/abs/2405.19358null
Efficient Adversarial Training in LLMs with Continuous AttacksSophie XhonneuxarxivMay-24https://arxiv.org/abs/2405.15589https://github.com/sophie-xhonneux/Continuous-AdvTrain
SafeAligner: Safety Alignment against Jailbreak Attacks via Response Disparity GuidanceCaishuang HuangarxivJun-24https://arxiv.org/abs/2406.18118null
Defending Large Language Models Against Jailbreak Attacks via Layer-specific EditingWei ZhaoarxivMay-24https://arxiv.org/abs/2405.18166https://github.com/ledllm/ledllm
Cross-Task Defense: Instruction-Tuning LLMs for Content SafetyYu FuNAACLMay-24https://arxiv.org/abs/2405.15202https://github.com/FYYFU/safety-defense

Benchmark

TitleAuthorPublishYearLinkSource Code
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMsZhao Xuarxiv2024https://arxiv.org/abs/2406.09324https://github.com/usail-hkust/Bag_of_Tricks_for_LLM_Jailbreaking
BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM SafeguardsDiego Dornarxiv2024https://arxiv.org/abs/2406.01364null
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language ModelsPatrick Chaoarxiv2024https://arxiv.org/abs/2404.01318https://github.com/JailbreakBench/jailbreakbench
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeikaarxiv2024https://arxiv.org/abs/2402.04249https://github.com/centerforaisafety/HarmBench
SC-Safety: A Multi-round Open-ended Question Adversarial Safety Benchmark for Large Language Models in ChineseLiang Xuarxiv2024https://arxiv.org/abs/2310.05818https://www.cluebenchmarks.com/
Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language ModelsHuachuan Qiuarxiv2024https://arxiv.org/abs/2307.08487https://github.com/qiuhuachuan/latent-jailbreak
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language ModelsPaul RöttgerNAACLAug-23https://arxiv.org/abs/2308.01263null
Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language ModelsHuachuan QiuarxivJul-23https://arxiv.org/abs/2307.08487https://github.com/qiuhuachuan/latent-jailbreak
SC-Safety: A Multi-round Open-ended Question Adversarial Safety Benchmark for Large Language Models in ChineseLiang XuarxivOct-23https://arxiv.org/abs/2310.05818https://www.cluebenchmarks.com/
AttackEval: How to Evaluate the Effectiveness of Jailbreak Attacking on Large Language ModelsDong shuarxivJan-24https://arxiv.org/abs/2401.09002null
How (un)ethical are instruction-centric responses of LLMs? Unveiling the vulnerabilities of safety guardrails to harmful queriesSomnath BanerjeearxivFeb-24https://arxiv.org/abs/2402.15302https://huggingface.co/datasets/SoftMINER-Group/TechHazardQA
AEGIS: Online Adaptive AI Content Safety Moderation with Ensemble of LLM ExpertsShaona GhosharxivApr-24https://arxiv.org/abs/2404.05993
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMsZhao XuarxivJun-24https://arxiv.org/abs/2406.09324null
BELLS: A Framework Towards Future Proof Benchmarks for the Evaluation of LLM SafeguardsDiego DornarxivJun-24https://arxiv.org/abs/2406.01364null
Improved Generation of Adversarial Examples Against Safety-aligned LLMsQizhang LiarxivMay-24https://arxiv.org/abs/2405.20778null
JailBreakV-28K: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak AttacksWeidi LuoarxivApr-24https://arxiv.org/abs/2404.03027https://github.com/EddyLuo1232/JailBreakV_28K
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language ModelsPatrick ChaoarxivMar-24https://arxiv.org/abs/2404.01318https://github.com/JailbreakBench/jailbreakbench
JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language ModelsDelong RanarxivJun-24https://arxiv.org/abs/2406.09321https://github.com/ThuCCSLab/JailbreakEval

Survey

TitleAuthorPublishYearLinkSource Code
Unique Security and Privacy Threats of Large Language Model: A Comprehensive SurveyShang Wangarxiv2024https://arxiv.org/abs/2406.07973null
Exploring Vulnerabilities and Protections in Large Language Models: A SurveyFrank Weizhen Liuarxiv2024https://arxiv.org/abs/2406.00240null
Safeguarding Large Language Models: A SurveyYi Dongarxiv2024https://arxiv.org/abs/2406.02622null
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under CompressionJunyuan Hongarxiv2024https://arxiv.org/abs/2403.15447null
Securing Large Language Models: Threats, Vulnerabilities and Responsible PracticesSara Abdaliarxiv2024https://arxiv.org/abs/2403.12503null
Breaking Down the Defenses: A Comparative Survey of Attacks on Large Language ModelsArijit Ghosh Chowdhuryarxiv2024https://arxiv.org/abs/2403.04786null
Attacks, Defenses and Evaluations for LLM Conversation Safety: A SurveyZhichen Dongarxiv2024https://arxiv.org/abs/2402.09283null
Security and Privacy Challenges of Large Language Models: A SurveyBadhan Chandra Dasarxiv2024https://arxiv.org/abs/2402.00888null
Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model SystemsTianyu Cuiarxiv2024https://arxiv.org/abs/2401.05778null
TrustLLM: Trustworthiness in Large Language ModelsLichao Sunarxiv2024https://arxiv.org/abs/2401.05561null
A Survey on Large Language Model (LLM) Security and Privacy: The Good, the Bad, and the UglyYifan Yaoarxiv2024https://arxiv.org/abs/2312.02003null
A Comprehensive Overview of Large Language ModelsHumza Naveedarxiv2024https://arxiv.org/abs/2307.06435null
Safety Assessment of Chinese Large Language ModelsHao Sunarxiv2024https://arxiv.org/abs/2304.10436null
Holistic Evaluation of Language ModelsPercy Liangarxiv2024https://arxiv.org/abs/2211.09110null
Survey of Vulnerabilities in Large Language Models Revealed by Adversarial AttacksErfan ShayeganiACLOct-23https://arxiv.org/abs/2310.10844https://llm-vulnerability.github.io/
A Comprehensive Survey of Attack Techniques, Implementation, and Mitigation Strategies in Large Language ModelsAysan EsmradiUbiSecDec-23https://arxiv.org/abs/2312.10982null
Against The Achilles' Heel: A Survey on Red Teaming for Generative ModelsLizhi LinarxivMar-24https://arxiv.org/abs/2404.00629null

Others

TitleAuthorPublishYearLinkSource Code
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language ModelsXinyue ShenCCSAug-23https://arxiv.org/abs/2308.03825https://github.com/verazuo/jailbreak_llms
Jailbreaking ChatGPT via Prompt Engineering: An Empirical StudyYi LiuarxivMay-23https://arxiv.org/abs/2305.13860null
Comprehensive Assessment of Jailbreak Attacks Against LLMsJunjie ChuarxivFeb-24https://arxiv.org/abs/2402.05668null
Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming in the WildNanna IniearxivNov-23https://arxiv.org/abs/2311.06237null
A Comprehensive Study of Jailbreak Attack versus Defense for Large Language ModelsZihao XuACLFeb-24https://arxiv.org/abs/2402.13457null
Sowing the Wind, Reaping the Whirlwind: The Impact of Editing Language ModelsRima HazraACLJan-24https://arxiv.org/abs/2401.10647null
Is the System Message Really Important to Jailbreaks in Large Language Models?Xiaotian ZouarxivFeb-24https://arxiv.org/abs/2402.14857null
Testing the Limits of Jailbreaking Defenses with the Purple ProblemTaeyoun KimarxivMar-24https://arxiv.org/abs/2403.14725null
JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language ModelsYingchaojie FengarxivApr-24https://arxiv.org/abs/2404.08793
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMsJavier RandoarxivApr-24https://arxiv.org/abs/2404.14461null
Universal Adversarial Triggers Are Not UniversalNicholas MeadearxivApr-24https://arxiv.org/abs/2404.16020null
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden StatesZhenhong ZhouarxivJun-24https://arxiv.org/abs/2406.05644https://github.com/ydyjya/LLM-IHS-Explanation
Understanding Jailbreak Success: A Study of Latent Space Dynamics in Large Language ModelsSarah BallarxivJun-24https://arxiv.org/abs/2406.09289null
Badllama 3: removing safety finetuning from Llama 3 in minutesDmitrii VolkovarxivJul-24https://arxiv.org/abs/2407.01376null
"Not Aligned" is Not "Malicious": Being Careful about Hallucinations of Large Language Models' JailbreakLingrui MeiarxivJun-24https://arxiv.org/abs/2406.11668null
Rethinking How to Evaluate Language Model JailbreakHongyu CaiarxivApr-24https://arxiv.org/abs/2404.06407
Jailbreak Paradox: The Achilles' Heel of LLMsAbhinav RaoarxivJun-24https://arxiv.org/abs/2406.12702null
Hacc-Man: An Arcade Game for Jailbreaking LLMsMatheus ValentimarxivMay-24https://arxiv.org/abs/2405.15902null

Hallucination

Mitigation

TitleAuthorPublishDateLinkSource Code
BTR: Binary Token Representations for Efficient Retrieval Augmented Language ModelsQingqing CaoICLRJan-24https://openreview.net/pdf?id=3TO3TtnOFlhttps://github.com/csarron/BTR
Reasoning on Graphs: Faithful and Interpretable Large Language Model ReasoningLINHAO LUOICLRJan-24https://openreview.net/pdf?id=ZGNWW7xZ6Qhttps://github.com/RManLuo/reasoning-on-graphs
Knowledge of Knowledge: Exploring Known-Unknowns Uncertainty with Large Language ModelsAlfonso AmayuelasarxivMay-23https://arxiv.org/abs/2305.13712null
Conformal Language ModelingVictor QuachICLRJun 2023https://arxiv.org/pdf/2306.10193null
Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth LiNeurIPSJun 2023https://arxiv.org/abs/2306.03341https://github.com/likenneth/honest_llama
Supervised Knowledge Makes Large Language Models Better In-context LearnersLinyi YangICLRDec-23https://arxiv.org/abs/2312.15918https://github.com/YangLinyi/Supervised-Knowledge-Makes-Large-Language-Models-Better-In-context-Learners
Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction TuningFuxiao LiuICLRJun 2023https://arxiv.org/abs/2306.14565https://github.com/FuxiaoLiu/LRV-Instruction
Chain-of-Knowledge: Grounding Large Language Models via Dynamic Knowledge Adapting over Heterogeneous SourcesXingxuan LiICLRMay-23https://arxiv.org/abs/2305.13269https://github.com/DAMO-NLP-SG/chain-of-knowledge
Fine-Tuning Language Models for FactualityKatherine TianICLRNov-23https://arxiv.org/abs/2311.08401https://github.com/kttian/llm_factuality_tuning
Chain-of-Table: Evolving Tables in the Reasoning Chain for Table UnderstandingZilong WangICLRJan-24https://arxiv.org/abs/2401.04398null
MetaGPT: Meta Programming for A Multi-Agent Collaborative FrameworkSirui HongICLRAug-23https://arxiv.org/abs/2308.00352https://github.com/geekan/MetaGPT
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive CritiquingZhibin GouICLRMay-23https://arxiv.org/pdf/2305.11738https://github.com/microsoft/ProphetNet/tree/master/CRITIC
Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun DuarxivMay-23https://arxiv.org/abs/2305.14325https://composable-models.github.io/llm_debate/
Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language ModelsJinhao DuanarxivJul-23https://arxiv.org/abs/2307.01379https://github.com/jinhaoduan/SAR
RAPPER: Reinforced Rationale-Prompted Paradigm for Natural Language Explanation in Visual Question AnsweringKai-Po ChangICLRJan 2024https://openreview.net/pdf?id=bshfchPM9Hnull
Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated FeedbackBaolin PengarxivFeb 2023https://arxiv.org/abs/2302.12813https://github.com/pengbaolin/LLM-Augmenter
The Knowledge Alignment Problem: Bridging Human and External Knowledge for Large Language ModelsShuo ZhangarxivMay-23https://arxiv.org/abs/2305.13669https://github.com/ShuoZhangXJTU/MixAlign
Trusting Your Evidence: Hallucinate Less with Context-aware DecodingWeijia ShiarxivMay-23https://arxiv.org/abs/2305.14739null
Reasoning on Graphs: Faithful and Interpretable Large Language Model ReasoningLINHAO LUOICLROct-23https://arxiv.org/abs/2310.01061https://github.com/RManLuo/reasoning-on-graphs
Explainable Claim Verification via Knowledge-Grounded Reasoning with Large Language ModelsHaoran WangEMNLPOct-23https://arxiv.org/abs/2310.05253https://github.com/wang2226/FOLK
Beyond task performance: evaluating and reducing the flaws of large multimodal models with in-context-learningMustafa ShukorICLROct-23https://arxiv.org/abs/2310.00647https://github.com/mshukor/EvALign-ICL
Mitigating Large Language Model Hallucinations via Autonomous Knowledge Graph-based RetrofittingXinyan GuanarxivNov-23https://arxiv.org/abs/2311.13314null
Knowledge Verification to Nip Hallucination in the BudFanqi WanarxivJan 2024https://arxiv.org/abs/2401.10768https://github.com/fanqiwan/KCA
Model Editing Harms General Abilities of Large Language Models: Regularization to the RescueJia-Chen GuarxivJan 2024https://arxiv.org/abs/2401.04700null
Reducing Hallucinations in Entity Abstract Summarization with Facts-Template DecompositionFangwei ZhuarxivFeb-24https://arxiv.org/abs/2402.18873null
Chain-of-Note: Enhancing Robustness in Retrieval-Augmented Language ModelsWenhao YuarxivNov-23https://arxiv.org/abs/2311.09210null
Reducing hallucination in structured outputs via Retrieval-Augmented GenerationPatrice BéchardNAACLApr 2024https://arxiv.org/abs/2404.08189null
Improving Factual Error Correction by Learning to Inject Factual ErrorsXingwei HeAAAIDec-23https://arxiv.org/abs/2312.07049null
Strong hallucinations from negation and how to fix themNicholas AsherarxivFeb-24https://arxiv.org/abs/2402.10543null
Retrieve Only When It Needs: Adaptive Retrieval Augmentation for Hallucination Mitigation in Large Language ModelsHanxing DingarxivFeb-24https://arxiv.org/abs/2402.10612null
Truth-Aware Context Selection: Mitigating Hallucinations of Large Language Models Being Misled by Untruthful ContextsTian YuACLMar-24https://arxiv.org/abs/2403.07556https://github.com/ictnlp/TACS
Enhancing LLM Factual Accuracy with RAG to Counter Hallucinations: A Case Study on Domain-Specific Queries in Private Knowledge-BasesJiarui LiarxivMar-24https://arxiv.org/abs/2403.10446https://github.com/anlp-team/LTI_Neural_Navigator
Uncertainty-Based Abstention in LLMs Improves Safety and Reduces HallucinationsChristian TomaniarxivApr-24https://arxiv.org/abs/2404.10960null
TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful SpaceShaolei ZhangACLFeb-24https://arxiv.org/abs/2402.17811https://ictnlp.github.io/TruthX-site/
Measuring and Reducing LLM Hallucination without Gold-Standard AnswersJiaheng WeiarxivFeb-24https://arxiv.org/abs/2402.10412null
Mitigating LLM Hallucinations via Conformal AbstentionYasin Abbasi YadkoriarxivApr-24https://arxiv.org/abs/2405.01563null
Understanding the Effects of Iterative Prompting on TruthfulnessSatyapriya KrishnaarxivFeb-24https://arxiv.org/abs/2402.06625null
Mitigating Hallucinations in Large Language Models via Self-Refinement-Enhanced Knowledge RetrievalMengjia NiuarxivMay-24https://arxiv.org/abs/2405.06545null
Can LLMs Produce Faithful Explanations For Fact-checking? Towards Faithful Explainable Fact-Checking via Multi-Agent DebateKyungha KimarxivFeb-24https://arxiv.org/abs/2402.07401null

Detection

TitleAuthorVenueDateLinkSource Code
SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language ModelsPotsawee ManakulEMNLPMar-23https://arxiv.org/abs/2303.08896https://github.com/potsawee/selfcheckgpt
A Stitch in Time Saves Nine: Detecting and Mitigating Hallucinations of LLMs by Validating Low-Confidence GenerationNeeraj VarshneyarxivJul-23https://arxiv.org/abs/2307.03987
Self-contradictory Hallucinations of Large https://arxiv.org/abs/2305.15852 Models: Evaluation, Detection and MitigationNiels MündlerICLRMay-23https://arxiv.org/abs/2305.15852https://chatprotect.ai/
Fact-Checking Complex Claims with Program-Guided ReasoningLiangming PanACLMay-23https://arxiv.org/abs/2305.12744https://github.com/mbzuai-nlp/ProgramFC
INSIDE: LLMs' Internal States Retain the Power of Hallucination DetectionChao ChenICLRFeb-24https://arxiv.org/abs/2402.03744null
Detecting Hallucinations in Large Language Model Generation: A Token Probability ApproachErnesto QuevedoICAIMay-24https://arxiv.org/abs/2405.19648https://github.com/Baylor-AI/HalluDetect
Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language ModelsWeihang SuMar-24https://arxiv.org/abs/2403.06448null
Enhancing Uncertainty-Based Hallucination Detection with Stronger FocusTianhang ZhangEMNLPNov-23https://arxiv.org/abs/2311.13230null
In Search of Truth: An Interrogation Approach to Hallucination DetectionYakir YehudaarxivMar-24https://arxiv.org/abs/2403.02889null
PoLLMgraph: Unraveling Hallucinations in Large Language Models via State Transition DynamicsDerui ZhuarxivApr-24https://arxiv.org/abs/2404.04722null
Fact-Checking the Output of Large Language Models via Token-Level Uncertainty QuantificationEkaterina FadeevaACLMar-24https://arxiv.org/pdf/2403.04696null
Can We Verify Step by Step for Incorrect Answer Detection?Xin XuarxivFeb-24https://arxiv.org/abs/2402.10528https://github.com/XinXU-USTC/R2PE
Comparing Hallucination Detection Metrics for Multilingual GenerationHaoqiang KangarxivFeb-24https://arxiv.org/abs/2402.10496null

Benchmark

TitleAuthorPublishDateLinkSource Code
Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language ModelsYue ZhangarxivSep-23https://arxiv.org/abs/2309.01219null
LitCab: Lightweight Language Model Calibration over Short- and Long-form ResponsesXin LiuICLROct-23https://arxiv.org/abs/2310.19208https://github.com/launchnlp/LitCab
Towards Understanding Factual Knowledge of Large Language ModelsXuming HuICLRJan-24https://openreview.net/pdf?id=9OevMUdodshttps://github.com/THU-BPM/Pinocchio
HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language ModelsJunyi LiEMNLPMay-23https://arxiv.org/abs/2305.11747https://github.com/RUCAIBox/HaluEval
UHGEval: Benchmarking the Hallucination of Chinese Large Language Models via Unconstrained GenerationXun LiangACLNov-23https://arxiv.org/abs/2311.15296https://iaar-shanghai.github.io/UHGEval/
HaluEval-Wild: Evaluating Hallucinations of Language Models in the WildZhiying ZhuarxivMar-24https://arxiv.org/abs/2403.04307https://github.com/Dianezzy/HaluEval-Wild
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMsAdi SimhiarxivApr-24https://arxiv.org/abs/2404.09971https://github.com/technion-cs-nlp/hallucination-mitigation
DelucionQA: Detecting Hallucinations in Domain-specific Question AnsweringMobashir SadatEMNLPDec-23https://arxiv.org/abs/2312.05200https://github.com/boschresearch/DelucionQA
DiaHalu: A Dialogue-level Hallucination Evaluation Benchmark for Large Language ModelsKedi ChenarxivMar-24https://arxiv.org/abs/2403.00896https://github.com/141forever/DiaHalu
ERBench: An Entity-Relationship based Automatically Verifiable Hallucination Benchmark for Large Language ModelsJio OharxivMar-24https://arxiv.org/abs/2403.05266https://github.com/DILAB-KAIST/ERBench

Survey

TitleAuthorPublishDateLinkSource Code
Survey of Hallucination in Natural Language GenerationZiwei JiarxivFeb 2022https://arxiv.org/abs/2202.03629null
Survey on Factuality in Large Language Models: Knowledge, Retrieval and Domain-SpecificityCunxiang WangarxivOct-23https://arxiv.org/abs/2310.07521null
The Human Factor in Detecting Errors of Large Language Models: A Systematic Literature Review and Future Research DirectionsChristian A. SchillerarxivMar-24https://arxiv.org/abs/2403.09743null
A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open QuestionsLei Huang,arxivNov-23https://arxiv.org/abs/2311.05232null

Other

TitleAuthorPublishDateLinkSource Code
Sources of Hallucination by Large Language Models on Inference TasksNick McKennaEMNLPMay-23https://arxiv.org/abs/2305.14552null
Teaching Language Models to Hallucinate Less with Synthetic TasksErik JonesICLROct-23https://arxiv.org/abs/2310.06827null
In ChatGPT We Trust? Measuring and Characterizing the Reliability of ChatGPTXinyue ShenarxivApr-23https://arxiv.org/abs/2304.08979null
When Large Language Models contradict humans? Large Language Models' Sycophantic BehaviourLeonardo RanaldiarxivNov-23https://arxiv.org/abs/2311.09410null
Calibrated Language Models Must HallucinateAdam Tauman KalaiSTOCNov-23https://arxiv.org/abs/2311.14648null
C-RAG: Certified Generation Risks for Retrieval-Augmented Language ModelsMintong KangICMLFeb-24https://arxiv.org/abs/2402.03181null
Large Language Models are Null-Shot LearnersPittawat TaveekitworachaiarxivJan 2024https://arxiv.org/abs/2401.08273null
Hallucination is Inevitable: An Innate Limitation of Large Language ModelsZiwei XuarxivJan-24https://arxiv.org/abs/2401.11817null
Relying on the Unreliable: The Impact of Language Models' Reluctance to Express UncertaintyKaitlyn ZhouarxivJan 2024https://arxiv.org/abs/2401.06730null
Deficiency of Large Language Models in Finance: An Empirical Examination of HallucinationHaoqiang KangarxivNov-23https://arxiv.org/abs/2311.15548null
Banishing LLM Hallucinations Requires Rethinking GeneralizationJohnny LiarxivJun-24https://arxiv.org/abs/2406.17642null
Seven Failure Points When Engineering a Retrieval Augmented Generation SystemScott BarnettarxivJan-24https://arxiv.org/abs/2401.05856null
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?Zorik GekhmanarxivMay-24https://arxiv.org/abs/2405.05904null
The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive ConversationRongwu XuACLDec-23https://arxiv.org/abs/2312.09085https://llms-believe-the-earth-is-flat.github.io/
Tell me the truth: A system to measure the trustworthiness of Large Language ModelCarlo LipizziarxivMar-24https://arxiv.org/abs/2403.04964null

MLLM

Jailbreak

Attack

TitleAuthorPublishYearLinkSource Code
Visual Adversarial Examples Jailbreak Large Language ModelsXiangyu QiAAAI2024AAAIgithub
Unveiling the Safety of GPT-4o: An Empirical Study using Jailbreak AttacksZonghao Yingarxiv2024https://arxiv.org/abs/2406.06302https://github.com/ny1024/jailbreak_gpt4o
Jailbreak Vision Language Models via Bi-Modal Adversarial PromptZonghao Yingarxiv2024https://arxiv.org/abs/2406.04031https://github.com/NY1024/BAP-Jailbreak-Vision-Language-Models-via-Bi-Modal-Adversarial-Prompt
Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language ModelsErfan ShayeganiICLR2024https://openreview.net/forum?id=plmBsXHxgRnull
FigStep: Jailbreaking Large Vision-language Models via Typographic Visual PromptsYichen Gongarxiv2023https://arxiv.org/abs/2311.05608https://github.com/ThuCCSLab/FigStep
Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language ModelsYifan Liarxiv2024https://arxiv.org/abs/2403.09792https://github.com/AoiDragon/HADES
Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks?Shuo Chenarxiv2024https://arxiv.org/abs/2404.03411null
Cross-Modality Jailbreak and Mismatched Attacks on Medical Multimodal Large Language ModelsXijie Huangarxiv2024https://arxiv.org/abs/2405.20775https://github.com/dirtycomputer/O2M_attack
Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image CharacterSiyuan Maarxiv2024https://arxiv.org/abs/2405.20773null
Jailbreaking Attack against Multimodal Large Language ModelZhenxing Niuarxiv2024https://arxiv.org/abs/2402.02309https://github.com/abc03570128/Jailbreaking-Attack-against-Multimodal-Large-Language-Model

Defense

TitleAuthorPublishYearLinkSource Code
MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language ModelsTianle Guarxiv2024https://arxiv.org/abs/2406.07594https://github.com/Carol-gutianle/MLLMGuard
Cross-Modal Safety Alignment: Is textual unlearning all you need?Trishna Chakrabortyarxiv2024https://arxiv.org/abs/2406.02575null
Unbridled Icarus: A Survey of the Potential Perils of Image Inputs in Multimodal Large Language Model SecurityYihe Fanarxiv2024https://arxiv.org/abs/2404.05264null
AdaShield: Safeguarding Multimodal Large Language Models from Structure-based Attack via Adaptive Shield PromptingYu Wangarxiv2024https://arxiv.org/abs/2403.09513https://github.com/rain305f/AdaShield
Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language ModelsYongshuo Zongarxiv2024https://arxiv.org/abs/2402.02207https://github.com/ys-zong/VLGuard
JailGuard: A Universal Detection Framework for LLM Prompt-based AttacksXiaoyu Zhangarxiv2024https://arxiv.org/abs/2312.10766null

Benchmark

TitleAuthorPublishYearLinkSource Code
JailBreakV-28K: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak AttacksWeidi Luoarxiv2024https://arxiv.org/abs/2404.03027https://github.com/EddyLuo1232/JailBreakV_28K
MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language ModelsXin Liuarxiv2023https://arxiv.org/abs/2311.17600https://github.com/isXinLiu/MM-SafetyBench
Red Teaming Visual Language ModelsMukai Liarxiv2024https://arxiv.org/abs/2401.12915null
How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMsHaoqin Tuarxiv2023https://arxiv.org/abs/2311.16101https://github.com/UCSC-VLAA/vllm-safety-benchmark
Benchmarking Trustworthiness of Multimodal Large Language Models: A Comprehensive StudyYichi Zhangarxiv2024https://arxiv.org/abs/2406.07057https://multi-trust.github.io/
SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language ModelYongting Zhangarxiv2024https://arxiv.org/abs/2406.12030https://github.com/EchoseChen/SPA-VL-RLHF
MOSSBench: Is Your Multimodal Language Model Oversensitive to Safe Queries?Xirui Liarxiv2024https://arxiv.org/abs/2406.17806https://turningpoint-ai.github.io/MOSSBench/

Survey

TitleAuthorPublishYearLinkSource Code
Safety of Multimodal Large Language Models on Images and TextsXin Liuarxiv2024https://arxiv.org//abs/2402.00357null

Poisoning/Backdoor

Attack

TitleFirst AuthorPublishYearLinkSource Code
Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language ModelsYuancheng Xuarxiv2024https://arxiv.org/abs/2402.06659https://github.com/umd-huang-lab/VLM-Poisoning
Test-Time Backdoor Attacks on Multimodal Large Language ModelsDong Luarxiv2024https://arxiv.org/abs/2402.08577https://sail-sg.github.io/AnyDoor/
VL-Trojan: Multimodal Instruction Backdoor Attacks against Autoregressive Visual Language ModelsJiawei Liangarxiv2024https://arxiv.org/abs/2402.13851null
Physical Backdoor Attack can Jeopardize Driving with Vision-Large-Language ModelsZhenyang Niarxiv2024https://arxiv.org/abs/2404.12916null

Adversarial

Attack

TitleFirst AuthorPublishYearLinkSource Code
Understanding Zero-Shot Adversarial Robustness for Large-Scale ModelsChengzhi Maoarxiv2022https://arxiv.org/abs/2212.07016null
On Evaluating Adversarial Robustness of Large Vision-Language ModelsYunqing Zhaoarxiv2023https://arxiv.org/abs/2305.16934https://github.com/yunqing-me/AttackVLM
On the Adversarial Robustness of Multi-Modal Foundation ModelsChristian SchlarmannICCV AROW2023https://arxiv.org/abs/2308.10741null
Adversarial Illusions in Multi-Modal EmbeddingsTingwei ZhangUSENIX Security2023https://arxiv.org/abs/2308.11804null
Image Hijacks: Adversarial Images can Control Generative Models at RuntimeLuke Baileyarxiv2023https://arxiv.org/abs/2309.00236null
How Robust is Google's Bard to Adversarial Image Attacks?Yinpeng Dongarxiv2023https://arxiv.org/abs/2309.11751https://github.com/thu-ml/Attack-Bard
An Image Is Worth 1000 Lies: Transferability of Adversarial Images across Prompts on Vision-Language ModelsHaochen LuoICLR2024https://openreview.net/pdf?id=nc5GgFAvtkhttps://github.com/Haochen-Luo/CroPA
Inducing High Energy-Latency of Large Vision-Language Models with Verbose ImagesKuofeng GaoICLR2024https://openreview.net/pdf?id=BteuUysuXXhttps://github.com/KuofengGao/Verbose_Images
Hijacking Context in Large Multi-modal ModelsJoonhyun Jeongarxiv2023https://arxiv.org/abs/2312.07553null
On the Robustness of Large Multimodal Models Against Image Adversarial AttacksXuanming Cuiarxiv2023https://arxiv.org/abs/2312.03777null
InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language ModelsXunguang Wangarxiv2023https://arxiv.org/abs/2312.01886https://github.com/xunguangwang/InstructTA
Stop Reasoning! When Multimodal LLMs with Chain-of-Thought Reasoning Meets Adversarial Imagesarxiv2024https://arxiv.org/abs/2402.14899null

Benchmark

TitleFirst AuthorPublishYearLinkSource Code
AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Adversarial Visual-InstructionsHao Zhangarxiv2024https://arxiv.org/abs/2403.09346null
Transferable Multimodal Attack on Vision-Language Pre-training ModelsHaodi WangS&P2024https://www.computer.org/csdl/proceedings-article/sp/2024/313000a102/1Ub239H4xygnull

Hallucination

TitleFirst AuthorPublishYearLinkSource Code
Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction TuningFuxiao LiuICLR2024https://openreview.net/pdf?id=J44HfH4JCghttps://github.com/FuxiaoLiu/LRV-Instruction
Analyzing and Mitigating Object Hallucination in Large Vision-Language ModelsYiyang ZhouICLR2024https://openreview.net/pdf?id=oZDJKTlOUehttps://github.com/YiyangZhou/LURE
A Survey on Hallucination in Large Vision-Language ModelsHanchao Liuarxiv2024https://arxiv.org/abs/2402.00253null
Mitigating Object Hallucination in Large Vision-Language Models via Classifier-Free GuidanceLinxi Zhaoarxiv2024https://arxiv.org/abs/2402.08680null
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided DecodingAilin Dengarxiv2024https://arxiv.org/abs/2402.15300https://github.com/d-ailin/CLIP-Guided-Decoding
Detecting and Mitigating Hallucination in Large Vision Language Models via Fine-Grained AI FeedbackWenyi Xiaoarxiv2024https://arxiv.org/abs/2404.14233null
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive DecodingXintong Wangarxiv2024https://arxiv.org/abs/2403.18715null
HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language ModelsTianrui GuanCVPR2024https://arxiv.org/abs/2310.14566https://github.com/tianyi-lab/HallusionBench

Bias

TitleFirst AuthorPublishYearLinkSource Code
Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference ChallengesChenhang Cuiarxiv2023https://arxiv.org/abs/2311.03287https://github.com/gzcch/Bingo

Others

TitleFirst AuthorPublishYearLinkSource Code
Zero shot VLMs for hate meme detection: Are we there yet?Naquee Rizwanarxiv2024https://arxiv.org/abs/2402.12198null
MemeCraft: Contextual and Stance-Driven Multimodal Meme GenerationHan WangACM MM2024https://arxiv.org/abs/2403.14652null
Moderating Illicit Online Image Promotion for Unsafe User-Generated Content Games Using Large Vision-Language ModelsKeyan GuoUSENIX Security2024https://arxiv.org/abs/2403.18957null

T2I

IP

Protection

TitleAuthorPublishYearLinkSource Code
AIGC-Chain: A Blockchain-Enabled Full Lifecycle Recording System for AIGC Product Copyright ManagementJiajia Jiangarixv2024https://arxiv.org/abs/2406.14966null
PID: Prompt-Independent Data Protection Against Latent Diffusion ModelsAng Liarxiv2024https://arxiv.org/abs/2406.15305null
Glaze: Protecting Artists from Style Mimicry by Text-to-Image ModelsShawn ShanUSENIX Security2023https://arxiv.org/abs/2302.04222null
Generative Watermarking Against Unauthorized Subject-Driven Image SynthesisYihan Maarxiv2023https://arxiv.org/abs/2306.07754null
Toward effective protection against diffusion based mimicry through score distillationHaotian Xuearxiv2023https://arxiv.org/abs/2311.12832https://github.com/xavihart/Diff-Protect
A Watermark-Conditioned Diffusion Model for IP ProtectionRui Minarxiv2024https://arxiv.org/abs/2403.10893null
RAW: A Robust and Agile Plug-and-Play Watermark Framework for AI-Generated Images with Provable GuaranteesXun Xianarxiv2024https://arxiv.org/abs/2403.18774null
A Training-Free Plug-and-Play Watermark Framework for Stable DiffusionGuokai Zhangarxiv2024https://arxiv.org/abs/2404.05607null
Gaussian Shading: Provable Performance-Lossless Image Watermarking for Diffusion ModelsZijin Yangarxiv2024https://arxiv.org/abs/2404.04956null
Lazy Layers to Make Fine-Tuned Diffusion Models More TraceableHaozhe Liuarxiv2024https://arxiv.org/abs/2405.00466null
DiffuseTrace: A Transparent and Flexible Watermarking Scheme for Latent Diffusion ModelLiangqi Leiarxiv2024https://arxiv.org/abs/2405.02696null
AquaLoRA: Toward White-box Protection for Customized Stable Diffusion Models via Watermark LoRAWeitao Fengarxiv2024https://arxiv.org/abs/2405.11135null
A Recipe for Watermarking Diffusion ModelsYunqing Zhaoarixv2023https://arxiv.org/abs/2303.10137https://github.com/yunqing-me/WatermarkDM
Watermarking Diffusion ModelYugeng Liuarxiv2023https://arxiv.org/abs/2305.12502null
Tree-Ring Watermarks: Fingerprints for Diffusion Images that are Invisible and RobustYuxin Wenarxiv2023https://arxiv.org/abs/2305.20030https://github.com/YuxinWenRick/tree-ring-watermark
Generative Models are Self-Watermarked: Declaring Model Authentication through Re-GenerationAditya Desuarxiv2024https://arxiv.org/abs/2402.16889null
Erasing Concepts from Diffusion ModelsRohit Gandikotaarxiv2023https://arxiv.org/abs/2303.07345https://erasing.baulab.info/
Adversarial Example Does Good: Preventing Painting Imitation from Diffusion Models via Adversarial ExamplesChumeng LiangICML2023https://arxiv.org/abs/2302.04578https://github.com/mist-project/mist.git

Violation

TitleAuthorPublishYearLinkSource Code
A Transfer Attack to Image WatermarksYuepeng Huarxiv2024https://arxiv.org/abs/2403.15365null
Disguised Copyright Infringement of Latent Diffusion ModelsYiwei Luarxiv2024https://arxiv.org/abs/2404.06737https://github.com/watml/disguised_copyright_infringement
Stable Signature is Unstable: Removing Image Watermark from Diffusion ModelsYuepeng Huarxiv2024https://arxiv.org/abs/2405.07145null
UnMarker: A Universal Attack on Defensive WatermarkingAndre Kassisarxiv2024https://arxiv.org/abs/2405.08363null
FreezeAsGuard: Mitigating Illegal Adaptation of Diffusion Models via Selective Tensor FreezingKai Huangarxiv2024https://arxiv.org/abs/2405.17472null
Adversarial Perturbations Cannot Reliably Protect Artists From Generative AIRobert Hönigarxiv2024https://arxiv.org/abs/2406.12027null
EnTruth: Enhancing the Traceability of Unauthorized Dataset Usage in Text-to-image Diffusion Models with Minimal and Robust AlterationsJie Renarxiv2024https://arxiv.org/abs/2406.13933null
Leveraging Optimization for Adaptive Attacks on Image WatermarksNils LukasICLR2024https://openreview.net/pdf?id=O9PArxKLe1null

Survey

TitleAuthorPublishYearLinkSource Code
Copyright Protection in Generative AI: A Technical PerspectiveJie Renarxiv2024https://arxiv.org/abs/2402.02333null

Privacy

Attack

TitleAuthorPublishYearLinkSource Code
Recovering the Pre-Fine-Tuning Weights of Generative ModelsEliahu Horwitzarxiv2024https://arxiv.org/abs/2402.10208null
Membership Inference Attacks Against Text-to-image Generation ModelsYixin Wuarxiv2022https://arxiv.org/abs/2210.00968null
Class Attribute Inference Attacks: Inferring Sensitive Class Information by Diffusion-Based Attribute ManipulationsLukas Struppekarxiv2023https://arxiv.org/abs/2303.09289null
White-box Membership Inference Attacks against Diffusion ModelsYan Pangarxiv2023https://arxiv.org/abs/2308.06405null
An Efficient Membership Inference Attack for the Diffusion Model by Proximal InitializationFei KongICLR2024https://openreview.net/pdf?id=rpH9FcCEV6https://github.com/kong13661/PIA
Black-box Membership Inference Attacks against Fine-tuned Diffusion ModelsYan Pangarxiv2023https://arxiv.org/abs/2312.08207null
Prompt Stealing Attacks Against Text-to-Image Generation ModelsXinyue Shenarxiv2023https://arxiv.org/abs/2302.09923null
Shake to Leak: Fine-tuning Diffusion Models Can Amplify the Generative Privacy RiskZhangheng Liarxiv2024https://arxiv.org/html/2403.09450v1https://github.com/VITA-Group/Shake-to-Leak
Is Diffusion Model Safe? Severe Data Leakage via Gradient-Guided Diffusion ModelJiayang Mengarxiv2024https://arxiv.org/abs/2406.09484null
Extracting Training Data from Unconditional Diffusion ModelsYunhao Chenarxiv2024https://arxiv.org/abs/2406.12752null
Extracting Training Data from Diffusion ModelsNicholas Carliniarxiv2023https://arxiv.org/abs/2301.13188null
Towards Black-Box Membership Inference Attack for Diffusion ModelsJingwei Liarxiv2024https://arxiv.org/abs/2405.20771null
Visual Privacy Auditing with Diffusion ModelsKristian Schwethelmarxiv2024https://arxiv.org/abs/2403.07588null
Membership Inference on Text-to-Image Diffusion Models via Conditional Likelihood DiscrepancyShengfang Zhaiarxiv2024https://arxiv.org/abs/2405.14800null
Extracting Prompts by Inverting LLM OutputsCollin Zhangarxiv2024https://arxiv.org/abs/2405.15012null

Defense

TitleAuthorPublishYearLinkSource Code
Differentially Private Fine-Tuning of Diffusion ModelsYu-Lin Tsaiarxiv2024https://arxiv.org/abs/2406.01355https://anonymous.4open.science/r/DP-LORA-F02F
Privacy-Preserving Diffusion Model Using Homomorphic EncryptionYaojian Chenarxiv2024https://arxiv.org/abs/2403.05794null
Efficient Differentially Private Fine-Tuning of Diffusion ModelsJing Liuarxiv2024https://arxiv.org/abs/2406.05257null
Differentially Private Synthetic Data via Foundation Model APIs 1: ImagesZinan LinICLR2024https://openreview.net/pdf?id=YEhQs8POIohttps://github.com/microsoft/DPSDA
Anti-DreamBooth: Protecting users from personalized text-to-image synthesisThanh Van LeICCV2023https://arxiv.org/abs/2303.15433https://github.com/VinAIResearch/Anti-DreamBooth.git
Unlearnable Examples for Diffusion Models: Protect Data from Unauthorized ExploitationZhengyue Zhaoarxiv2023https://arxiv.org/abs/2306.01902null
Can Protective Perturbation Safeguard Personal Data from Being Exploited by Stable Diffusion?Zhengyue Zhaoarxiv2023https://arxiv.org/abs/2312.00084null
MetaCloak: Preventing Unauthorized Subject-driven Text-to-image Diffusion-based Synthesis via Meta-learningYixin Liuarxiv2023https://arxiv.org/abs/2311.13127https://github.com/liuyixin-louis/MetaCloak

Benchmark

TitleAuthorPublishYearLinkSource Code
Raccoon: Prompt Extraction Benchmark of LLM-Integrated ApplicationsJunlin Wangarxiv2024https://arxiv.org/abs/2406.06737https://github.com/M0gician/RaccoonBench

Memorization

TitleAuthorPublishYearLinkSource Code
SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and GenerationChongyu FanICLR2024https://openreview.net/pdf?id=gn0mIhQGNMhttps://github.com/OPTML-Group/Unlearn-Saliency
Ring-A-Bell! How Reliable are Concept Removal Methods For Diffusion Models?Yu-Lin TsaiICLR2024https://openreview.net/pdf?id=lm7MRcsFiShttps://github.com/chiayi-hsu/Ring-A-Bell
Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion ModelsYimeng Zhangarxiv2024https://arxiv.org/abs/2405.15234https://github.com/OPTML-Group/AdvUnlearn
Espresso: Robust Concept Filtering in Text-to-Image ModelsAnudeep Dasarxiv2024https://arxiv.org/abs/2404.19227null
Machine Unlearning for Image-to-Image Generative ModelsGuihong Liarxiv2024https://arxiv.org/abs/2402.00351https://github.com/jpmorganchase/l2l-generator-unlearning
MACE: Mass Concept Erasure in Diffusion ModelsShilin Luarxiv2024https://arxiv.org/abs/2403.06135https://github.com/Shilin-LU/MACE
Unveiling and Mitigating Memorization in Text-to-image Diffusion Models through Cross AttentionJie Renarxiv2024https://arxiv.org/abs/2403.11052https://github.com/renjie3/MemAttn
Could It Be Generated? Towards Practical Analysis of Memorization in Text-To-Image Diffusion ModelsZhe Maarxiv2024https://arxiv.org/abs/2405.05846null

Deepfake

Construction

TitleAuthorPublishYearLinkSource Code
An Analysis of Recent Advances in Deepfake Image Detection in an Evolving Threat LandscapeSifat Muhammad AbdullahS&P2024https://arxiv.org/abs/2404.16212null
Robustness of AI-Image Detectors: Fundamental Limits and Practical AttacksMehrdad SaberiICLR2024https://openreview.net/pdf?id=dLoAdIKENchttps://github.com/mehrdadsaberi/watermark_robustness

Detection

TitleAuthorPublishYearLinkSource Code
DeepFake-O-Meter v2.0: An Open Platform for DeepFake DetectionYan Juarxiv2024https://arxiv.org/abs/2404.13146null
DE-FAKE: Detection and Attribution of Fake Images Generated by Text-to-Image Generation ModelsZeyang Shaarxiv2022https://arxiv.org/abs/2210.06998null
Organic or Diffused: Can We Distinguish Human Art from AI-generated Images?Anna Yoo Jeong Haarxiv2024https://arxiv.org/abs/2402.03214null
Watermark-based Detection and Attribution of AI-Generated ContentZhengyuan Jiangarxiv2024https://arxiv.org/abs/2404.04254null
An Analysis of Recent Advances in Deepfake Image Detection in an Evolving Threat LandscapeSifat Muhammad AbdullahS&P2024https://arxiv.org/abs/2404.16212null

Benchmark

TitleAuthorPublishYearLinkSource Code
The Adversarial AI-Art: Understanding, Generation, Detection, and BenchmarkingYuying Liarxiv2024https://arxiv.org/abs/2404.14581null

Bias

TitleAuthorPublishYearLinkSource Code
Finetuning Text-to-Image Diffusion Models for FairnessXudong ShenICLR2024https://openreview.net/pdf?id=hnrB5YHoYuhttps://sail-sg.github.io/finetune-fair-diffusion/
ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image GenerationAkshita Jhaarxiv2024https://arxiv.org/abs/2401.06310null

Backdoor

Attack

TitleAuthorPublishYearLinkSource Code
Rickrolling the Artist: Injecting Backdoors into Text Encoders for Text-to-Image SynthesisLukas StruppekICCV2023https://arxiv.org/abs/2211.02408null
Generating Potent Poisons and Backdoors from Scratch with Guided DiffusionHossein Souriarxiv2024https://arxiv.org/abs/2403.16365null

Defense

TitleAuthorPublishYearLinkSource Code
Diffusion Denoising as a Certified Defense against Clean-label PoisoningSanghyun Hongarxiv2024https://arxiv.org/abs/2403.11981null
Leveraging Diffusion-Based Image Variations for Robust Training on Poisoned DataNeurIPS Workshop2023https://arxiv.org/abs/2310.06372null
UFID: A Unified Framework for Input-level Backdoor Detection on Diffusion ModelsZihan Guanarxiv2024https://arxiv.org/abs/2404.01101https://github.com/GuanZihan/official_UFID
Invisible Backdoor Attacks on Diffusion Modelsarxiv2024https://arxiv.org/abs/2406.00816https://github.com/invisibleTriggerDiffusion/invisible_triggers_for_diffusion
Watch the Watcher! Backdoor Attacks on Security-Enhancing Diffusion ModelsChangjiang Liarxiv2024https://arxiv.org/abs/2406.09669null
Injecting Bias in Text-To-Image Models via Composite-Trigger BackdoorsAli Naseharxiv2024https://arxiv.org/abs/2406.15213null

Adversarial

Attack

TitleAuthorPublishYearLinkSource Code
Diffusion-Based Adversarial Sample Generation for Improved Stealthiness and ControllabilityHaotian XueNeurIPS2023https://arxiv.org/abs/2305.16494https://github.com/xavihart/Diff-PGD
Stable Diffusion is UnstableChengbin Duarxiv2023https://arxiv.org/abs/2306.02583null
DiffAttack: Evasion Attacks Against Diffusion-Based Adversarial PurificationMintong KangNeurIPS2023https://arxiv.org/abs/2311.16124null
Exploring Adversarial Attacks against Latent Diffusion Model from the Perspective of Adversarial TransferabilityJunxi Chenarxiv2024https://arxiv.org/abs/2401.07087null
Revealing Vulnerabilities in Stable Diffusion via Targeted AttacksChenyu Zhangarxiv2024https://arxiv.org/abs/2401.08725https://github.com/datar001/Revealing-Vulnerabilities-in-Stable-Diffusion-via-Targeted-Attacks
Cheating Suffix: Targeted Attack to Text-To-Image Diffusion Models with Multi-Modal PriorsDingcheng Yangarxiv2024https://arxiv.org/abs/2402.01369https://github.com/ydc123/MMP-Attack
Groot: Adversarial Testing for Generative Text-to-Image Models with Tree-based Semantic TransformationYi Liuarxiv2024https://arxiv.org/abs/2402.12100null
BSPA: Exploring Black-box Stealthy Prompt Attacks against Image GeneratorsYu Tianarxiv2024https://arxiv.org/abs/2402.15218null
Perturbing Attention Gives You More Bang for the Buck: Subtle Imaging Perturbations That Efficiently Fool Customized Diffusion ModelsJingyao Xuarxiv2024https://arxiv.org/abs/2404.15081null
Investigating and Defending Shortcut Learning in Personalized Diffusion ModelsYixin Liuarxiv2024https://arxiv.org/abs/2406.18944null
Unsafe Diffusion: On the Generation of Unsafe Images and Hateful Memes From Text-To-Image ModelsYiting Quarxiv2024https://arxiv.org/abs/2305.13873null
ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign Usersarxiv2024https://arxiv.org/abs/2405.19360https://github.com/GuanlinLee/ART

Defense

TitleAuthorPublishYearLinkSource Code
Raising the Cost of Malicious AI-Powered Image EditingHadi Salmanarxiv2023https://arxiv.org/abs/2302.06588null
Adversarial Examples are Misaligned in Diffusion Model ManifoldsPeter Lorenzarxiv2024https://arxiv.org/abs/2401.06637null
Universal Prompt Optimizer for Safe Text-to-Image GenerationZongyu WuNAACL2024https://arxiv.org/abs/2402.10882https://github.com/wzongyu/POSI
Adversarial Nibbler: An Open Red-Teaming Method for Identifying Diverse Harms in Text-to-Image GenerationJessica Quayearxiv2024https://arxiv.org/abs/2403.12075null
SafeGen: Mitigating Unsafe Content Generation in Text-to-Image ModelsXinfeng Liarxiv2024https://arxiv.org/abs/2404.06666null

Agent

Backdoor Attack/Defense

TitleAuthorVenueYearLinkSource Code
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety TrainingEvan Hubingerarxiv2024https://arxiv.org/abs/2401.05566null
BadAgent: Inserting and Activating Backdoor Attacks in LLM AgentsYifei WangACL2024https://arxiv.org/abs/2406.03007https://github.com/DPamK/BadAgent

Adversarial Attack/Defense

TitleAuthorVenueYearLinkSource Code
Large Language Model Sentinel: Advancing Adversarial Robustness by LLM AgentGuang Linarxiv2024https://arxiv.org/abs/2405.20770null
Adversarial Attacks on Multimodal AgentsChen Henry Wuarxiv2024https://arxiv.org/abs/2406.12814https://github.com/ChenWu98/agent-attack
AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM AgentsEdoardo Debenedettiarxiv2024https://arxiv.org/abs/2406.13352https://github.com/ethz-spylab/agentdojo
GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled ReasoningZhen Xiangarxiv2024https://arxiv.org/abs/2406.09187null

Jailbreak

TitleAuthorVenueYearLinkSource Code
Evil Geniuses: Delving into the Safety of LLM-based AgentsYu Tianarxiv2023https://arxiv.org/abs/2311.11855https://github.com/T1aNS1R/Evil-Geniuses
Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially FastXiangming GuICML2024https://arxiv.org/abs/2402.08567https://sail-sg.github.io/Agent-Smith/
AutoDefense: Multi-Agent LLM Defense against Jailbreak AttacksYifan Zengarxiv2024https://arxiv.org/abs/2403.04783https://github.com/XHMY/AutoDefense

Prompt Injection

TitleAuthorVenueYearLinkSource Code
WIPI: A New Web Threat for LLM-Driven Web AgentsFangzhou Wuarxiv2024https://arxiv.org/abs/2402.16965null
InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model AgentsQiusi Zhanarxiv2024https://arxiv.org/abs/2403.02691https://github.com/uiuc-kang-lab/InjecAgent

Hallucination

TitleAuthorVenueYearLinkSource Code
Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Duarxiv2023https://arxiv.org/abs/2305.14325null
MetaGPT: Meta Programming for A Multi-Agent Collaborative FrameworkICLR2024https://openreview.net/pdf?id=VtmBAGCN7onull
Can LLMs Produce Faithful Explanations For Fact-checking? Towards Faithful Explainable Fact-Checking via Multi-Agent DebateKyungha Kimarxiv2024https://arxiv.org/abs/2402.07401null

Others

TitleAuthorVenueYearLinkSource Code
Achieving Fairness in Multi-Agent MDP Using Reinforcement LearningPeizhong JuICLR2024https://openreview.net/pdf?id=yoVq2BGQdPnull
Agent Alignment in Evolving Social NormsShimin Liarxiv2024https://arxiv.org/abs/2401.04620null
Air Gap: Protecting Privacy-Conscious Conversational AgentsEugene Bagdasaryanarxiv2024https://arxiv.org/abs/2405.05175null
Secret Collusion Among Generative AI AgentsSumeet Ramesh Motwaniarxiv2024https://arxiv.org/abs/2402.07510v1null

Survey

TitleAuthorVenueYearLinkSource Code
Security of AI AgentsYifeng Hearxiv2024https://arxiv.org/abs/2406.08689null

Competition

TitleOrganizerYearLinkCategory
Machine Learning Model Attribution ChallengeMITRE2022https://mlmac.io/Privacy
Training Data Extraction ChallengeGoogle2022https://github.com/google-research/lm-extraction-benchmarkPrivacy
Find the Trojan: Universal Backdoor Detection in Aligned LLMsETHZ2024https://github.com/ethz-spylab/rlhf_trojan_competitionSecurity
LLM - Detect AI Generated TextThe Learning Agency Lab2023https://www.kaggle.com/competitions/llm-detect-ai-generated-text/overviewDeepfake
Deepfake Detection ChallengeKaggle2019https://www.kaggle.com/c/deepfake-detection-challengeDeepfake
Large Language Model Capture-the-Flag (LLM CTF) CompetitionKaggle2024https://ctf.spylab.ai/Safety

Leaderboard

TitleFirst AuthorPublishYearLink
LLM Safety LeaderboardBoxin WangNeurIPS2023https://huggingface.co/spaces/AI-Secure/llm-trustworthy-leaderboard
Hallucinations LeaderboardPasquale Minervininull2023https://huggingface.co/spaces/hallucinations-leaderboard/leaderboard
JailbreakBenchPatrick Chaoarxiv2024https://jailbreakbench.github.io/
A Comprehensive Study of Trustworthiness in Large Language Models.Lichao Sunarxiv2024https://trustllmbenchmark.github.io/TrustLLM-Website/leaderboard.html
PromptBench: Towards Evaluating the Robustness of Large Language Models on Adversarial PromptsKaijie Zhuarxiv2023https://llm-eval.github.io/pages/leaderboard/advprompt.html
Leaderboard Comparing LLM Performance at Producing Hallucinations when Summarizing Short DocumentsVectaragithub2023https://github.com/vectara/hallucination-leaderboard

Arena

TitleInstitutionYearLink
LMSYS Chatbot Arena (Multimodal): Benchmarking LLMs and VLMs in the WildLMSYS Org2024https://arena.lmsys.org/
中文大模型竞技场ModelScope2024https://modelscope.cn/studios/LLMZOO/Chinese-Arena/summary
司南 OpenCompass 大模型竞技场Shanghai AI Lab2024https://opencompass.org.cn/arena
模型广场Coze2024https://www.coze.cn/model/arena

Book

TitleAuthorPublishYearLink
Adversarial Machine Learning A Taxonomy and Terminology of Attacks and MitigationsApostol Vassilevonline2024https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-2e2023.pdf
人工智能安全方滨兴电子工业出版社2022null
人工智能安全曾剑平清华大学出版社2022null
人工智能安全陈左宁电子工业出版社2024null
AI安全:技术与实战腾讯安全朱雀实验室电子工业出版社2022null
Trustworthy Machine LearningKush R. Varshneynull2022null
人工智能:数据与模型安全姜育刚机械工业出版社2024null


Star History Chart


Contributors

NY1024

35 commits