A resource repository for machine unlearning in large language models
626
184 commits
updated Aug 6, 2026
A curated collection of papers, surveys, benchmarks, frameworks, and blog posts for machine unlearning in large language models.
As of the last commit, there are 616 papers, 18 surveys and position papers, 3 frameworks, and 2 blog posts.
If you believe your paper on LLM unlearning is not included, or if you find a mistake, typo, or information that is not up to date, please open an issue or submit a pull request, and I will be happy to update the list.
UMU-Bench: Closing the Modality Gap in Multimodal Unlearning Evaluation
Elastic Robust Unlearning of Specific Knowledge in Large Language Models
MPSelectTune: Prompt-type Selection for Fine-tuning improves Concept Unlearning in LLMs
Identifying Unlearned Data in LLMs via Membership Inference Attacks
Unlearners Can Lie: Evaluating "Honesty" in LLM Unlearning
The Role of Learning and Memorization in Relabeling-based Unlearning for LLMs
On the Fragility of Latent Knowledge: Layer-wise Influence under Unlearning in Large Language Model
Lifelong Unlearning for Multimodal Large Language Models
MOUCHI: Mitigating Over-forgetting in Unlearning Copyrighted Information
Provably Continual Unlearning for Large Language Model
SELU: Energy-based Targeted Unlearning in LLMs
Refusal Is Not an Option: Unlearning Safety Alignment of Large Language Models
Rethinking Unlearning for Large Reasoning Models
Unlearning in Large Language Models: We Are Not There Yet
Investigating Model Editing for Unlearning in Large Language Models
The Erasure Illusion: Stress-Testing the Generalization of LLM Forgetting Evaluation
Towards Benchmarking Privacy Vulnerabilities in Selective Forgetting with Large Language Models
Towards Reasoning-Preserving Unlearning in Multimodal Large Language Models
Feature-Selective Representation Misdirection for Machine Unlearning
FAME: Fictional Actors for Multilingual Erasure
Explainable reinforcement learning from human feedback to improve alignment
FROC: A Unified Framework with Risk-Optimized Control for Machine Unlearning in LLMs
MLLM Machine Unlearning via Visual Knowledge Distillation
MedForget: Hierarchy-Aware Multimodal Unlearning Testbed for Medical AI
LUNE: Efficient LLM Unlearning via LoRA Fine-Tuning with Negative Examples
Recover-to-Forget: Gradient Reconstruction from LoRA for Efficient LLM Unlearning
Delete and Retain: Efficient Unlearning for Document Classification
Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
When Forgetting Builds Reliability: LLM Unlearning for Reliable Hardware Code Generation
RapidUn: Influence-Driven Parameter Reweighting for Efficient Large Language Model Unlearning
RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories
Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning
SineProject: Machine Unlearning for Stable Vision Language Alignment
From Narrow Unlearning to Emergent Misalignment: Causes, Consequences, and Containment in LLMs
Forgetting-MarI: LLM Unlearning via Marginal Information Regularization
AUVIC: Adversarial Unlearning of Visual Concepts for Multi-modal Large Language Models
Unlearning Imperative: Securing Trustworthy and Responsible LLMs through Engineered Forgetting
Cross-Modal Unlearning via Influential Neuron Path Editing in Multimodal Large Language Models
DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
Leak@$k$: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding
REMIND: Input Loss Landscapes Reveal Residual Memorization in Post-Unlearning LLMs
The Realignment Problem: When Right becomes Wrong in LLMs
From Memorization to Reasoning in the Spectrum of Loss Curvature
Uncovering the Potential Risks in Unlearning: Danger of English-only Unlearning in Multilingual LLMs
Probing Knowledge Holes in Unlearned LLMs
The Evaluation of Retrieval-Based Unlearning Mechanisms on Large Language Models
OFFSIDE: Benchmarking Unlearning Misinformation in Multimodal Large Language Models
Label Smoothing Improves Gradient Ascent in LLM Unlearning
Leverage Unlearning to Sanitize LLMs
Hubble: a Model Suite to Advance the Study of LLM Memorization
LLM Unlearning with LLM Beliefs
Forget to Know, Remember to Use: Context-Aware Unlearning for Large Language Models
Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting
Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
Hierarchical Federated Unlearning for Large Language Models
On the Impossibility of Retrain Equivalence in Machine Unlearning
Reference-Specific Unlearning Metrics Can Hide the Truth: A Reality Check
LLM Unlearning on Noisy Forget Sets: A Study of Incomplete, Rewritten, and Watermarked Data
Approximate Domain Unlearning for Vision-Language Models
SIMU: Selective Influence Machine Unlearning
LLM Unlearning Under the Microscope: A Full-Stack View on Methods and Metrics
Cross-Modal Attention Guided Unlearning in Vision-Language Models
(Token-Level) InfoRMIA: Stronger Membership Inference and Memorization Assessment for LLMs
Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning
Machine Unlearning Meets Adversarial Robustness via Constrained Interventions on LLMs
Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning
KnowledgeSmith: Uncovering Knowledge Updating in LLMs with Model Editing and Unlearning
Direct Token Optimization: A Self-contained Approach to Large Language Model Unlearning
Scalable and Robust LLM Unlearning by Correcting Responses with Retrieved Exclusions
Mitigating Biases in Language Models via Bias Unlearning
Understanding the Dilemma of Unlearning for Large Language Models
Stable Forgetting: Bounded Parameter-Efficient Unlearning in LLMs
Dual-Space Smoothness for Robust and Balanced LLM Unlearning
OFMU: Optimization-Driven Framework for Machine Unlearning
Erase or Hide? Suppressing Spurious Unlearning Neurons for Robust Unlearning
CLUE: Conflict-guided Localization for LLM Unlearning Framework
Beyond Sharp Minima: Robust LLM Unlearning via Feedback-Guided Multi-Point Optimization
Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models
Concept Unlearning in Large Language Models via Self-Constructed Knowledge Triplets
Reveal and Release: Iterative LLM Unlearning with Self-generated Data
Scrub It Out! Erasing Sensitive Memorization in Code Language Models via Machine Unlearning
Collapse of Irrelevant Representations (CIR) Ensures Robust and Non-Disruptive LLM Unlearning
Customized Retrieval-Augmented Generation with LLM for Debiasing Recommendation Unlearning
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs
Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?
Unlearning That Lasts: Utility-Preserving, Robust, and Almost Irreversible Forgetting in LLMs
Standard vs. Modular Sampling: Best Practices for Reliable LLM Unlearning
CURE: A Unified Framework for Class and Concept Unlearning via Retraining Emulation
Improving Fisher Information Estimation and Efficiency for LoRA-based LLM Unlearning
Unlearning as Ablation: Toward a Falsifiable Benchmark for Generative Scientific Discovery
Reliable Unlearning Harmful Information in LLMs with Metamorphosis Representation Projection
SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
Unlearning at Scale: Implementing the Right to be Forgotten in Large Language Models
Oblivionis: A Lightweight Learning and Unlearning Framework for Federated Large Language Models
Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
LLM Unlearning using Gradient Ratio-Based Influence Estimation and Noise Injection
LLM Unlearning Without an Expert Curated Dataset
Analyzing and Mitigating Object Hallucination: A Training Bias Perspective
From Learning to Unlearning: Biomedical Security Protection in Multimodal Large Language Models
DUP: Detection-guided Unlearning for Backdoor Purification in Language Models
Towards Evaluation for Real-World LLM Unlearning
Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
What Should LLMs Forget? Quantifying Personal Data in LLMs for Right-to-Be-Forgotten Requests
Automating Evaluation of Diffusion Model Unlearning with (Vision-) Language Model World Knowledge
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs
Unlearning the Noisy Correspondence Makes CLIP More Robust
PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning
Agents Are All You Need for LLM Unlearning
SoK: Semantic Privacy in Large Language Models
Model State Arithmetic for Machine Unlearning
Step-by-Step Reasoning Attack: Revealing 'Erased' Knowledge in Large Language Models
Does Multimodal Large Language Model Truly Unlearn? Stealthy MLLM Unlearning Attack
Large Language Model Unlearning for Source Code
Mr. Snuffleupagus at SemEval-2025 Task 4: Unlearning Factual Knowledge from LLMs Using Adaptive RMU
BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap
Learning-Time Encoding Shapes Unlearning in LLMs
Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs
Align-then-Unlearn: Embedding Alignment for LLM Unlearning
Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps
Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills
OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization
UCD: Unlearning in LLMs via Contrastive Decoding
Lifting Data-Tracing Machine Unlearning to Knowledge-Tracing for Foundation Models
GUARD: Guided Unlearning and Retention via Data Attribution for Large Language Models
SoK: Machine Unlearning for Large Language Models
BLUR: A Bi-Level Optimization Approach for LLM Unlearning
LLM Unlearning Should Be Form-Independent
RULE: Reinforcement UnLEarning Achieves Forget-Retain Pareto Optimality
Distillation Robustifies Unlearning
Towards Lifecycle Unlearning Commitment Management: Measuring Sample-level Unlearning Completeness
Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence Awareness
Constrained Entropic Unlearning: A Primal-Dual Framework for Large Language Models
Quantifying Cross-Modality Memorization in Vision-Language Models
Lacuna Inc. at SemEval-2025 Task 4: LoRA-Enhanced Influence-Based Unlearning for LLMs
Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning
Not All Tokens Are Meant to Be Forgotten
Rethinking Post-Unlearning Behavior of Large Vision-Language Models
Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-Tuning
Existing Large Language Model Unlearning Evaluations Are Inconclusive
Keeping an Eye on LLM Unlearning: The Hidden Risk and Remedy
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
Model Unlearning via Sparse Autoencoder Subspace Guided Projections
Does Machine Unlearning Truly Remove Model Knowledge? A Framework for Auditing Unlearning in LLMs
From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space Regularization
Graceful Forgetting in Generative Language Models
Safety Alignment via Constrained Knowledge Unlearning
T2VUnlearning: A Concept Erasing Method for Text-to-Video Diffusion Models
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
Losing is for Cherishing: Data Valuation Based on Machine Unlearning and Shapley Value
UniErase: Unlearning Token as a Universal Erasure Primitive for Language Models
R-TOFU: Unlearning in Large Reasoning Models
DUSK: Do Not Unlearn Shared Knowledge
SEPS: A Separability Measure for Robust Unlearning in LLMs
GUARD: Generation-time LLM Unlearning via Adaptive Restriction and Detection
Exploring Criteria of Loss Reweighting to Enhance LLM Unlearning
Unilogit: Robust Machine Unlearning for LLMs Using Uniform-Target Self-Distillation
OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models
Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?
Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
Token-Level Knowledge Unlearning for Large Language Model Security
AegisLLM: Scaling Agentic Systems for Self-Reflective Defense in LLM Security
DualOptim: Enhancing Efficacy and Stability in Machine Unlearning with Dual Optimizers
ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
DP2Unlearning: An Efficient and Guaranteed Unlearning Framework for LLMs
A mean teacher algorithm for unlearning of language models
GRAIL: Gradient-Based Adaptive Unlearning for Privacy and Copyright in LLMs
LLM Unlearning Reveals a Stronger-Than-Expected Coreset Effect in Current Benchmarks
Bridging the Gap Between Preference Alignment and Machine Unlearning
Exact Unlearning of Finetuning Data via Model Merging at Scale
SUV: Scalable Large Language Model Copyright Compliance with Regularized Selective Unlearning
Effective Skill Unlearning through Intervention and Abstention
ZJUKLAB at SemEval-2025 Task 4: Unlearning via Model Merging
Deep Contrastive Unlearning for Language Models
SAUCE: Selective Concept Unlearning in Vision-Language Models with Sparse Autoencoders
Atyaephyra at SemEval-2025 Task 4: Low-Rank NPO
PEBench: A Fictitious Dataset to Benchmark Machine Unlearning for Multimodal Large Language Models
Hyperbolic Safety-Aware Vision-Language Models
Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-tuning
Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
GRU: Mitigating the Trade-off between Unlearning and Retention for Large Language Models
UIPE: Enhancing LLM Unlearning by Removing Knowledge Related to Forgetting Targets
Improving LLM Safety Alignment with Dual-Objective Optimization
CE-U: Cross Entropy Unlearning
Rectifying Belief Space via Unlearning to Harness LLMs' Reasoning
Erasing Without Remembering: Safeguarding Knowledge Forgetting in Large Language Models
Rethinking LLM Unlearning Objectives: A Gradient Perspective and Go Beyond
A General Framework to Enhance Fine-tuning-based LLM Unlearning
Modality-Aware Neuron Pruning for Unlearning in Multimodal Large Language Models
Soft Token Attacks Cannot Reliably Audit Unlearning in Large Language Models
CoME: An Unlearning-based Approach to Conflict-free Model Editing
LUME: LLM Unlearning with Multitask Evaluations
UPCORE: Utility-Preserving Coreset Selection for Balanced Unlearning
Obliviate: Efficient Unmemorization for Protecting Intellectual Property in Large Language Models
Beyond Single-Value Metrics: Evaluating and Enhancing LLM Unlearning with Cognitive Diagnosis
Which Retain Set Matters for LLM Unlearning? A Case Study on Entity Unlearning
LUNAR: LLM Unlearning via Neural Activation Redirection
Mitigating Sensitive Information Leakage in LLMs4Code through Machine Unlearning
A Lightweight Method to Disrupt Memorized Sequences in LLM
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
Improving the Robustness of Representation Misdirection for Large Language Model Unlearning
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
Precise In-Parameter Concept Erasure in Large Language Models
A Fully Probabilistic Perspective on Large Language Model Unlearning: Evaluation and Optimization
Machine Unlearning of Personally Identifiable Information in Large Language Models
Adaptive Localization of Knowledge Negation for Continual LLM Unlearning
LUSB: Formalizing and Benchmarking Unlearning Attacks and Defenses against Large Language Models
White-Box Auditing of Large Language Model Unlearning
UUE: Untargeted Language Model Unlearning via Null-Space-Guided Editing with Lightweight Adapters
UnRe: Zero-Shot LLM Unlearning via Dynamic Contextual Retrieval
Effective Unlearning in LLMs Relies on the Right Data Retention Strategy
Decoupling Memories, Muting Neurons: Towards Practical Machine Unlearning for Large Language Models
LLM-Eraser: Optimizing Large Language Model Unlearning through Selective Pruning
SemEval-2025 Task 4: Unlearning Sensitive Content from Large Language Models
NEKO at SemEval-2025 Task 4: A Gradient Ascent Based Machine Unlearning Strategy
NeuroReset: LLM Unlearning via Dual Phase Mixed Methodology
MALTO at SemEval-2025 Task 4: Dual Teachers for Unlearning Sensitive Content in LLMs
YNU at SemEval-2025 Task 4: Synthetic Token Alternative Training for LLM Unlearning
JU-CSE-NLP'25 at SemEval-2025 Task 4: Learning to Unlearn LLMs
NLPART at SemEval-2025 Task 4: Forgetting is Harder than Learning
Orthogonal Gradient Projection for Continual LLM Unlearning
Position: The Term "Machine Unlearning" Is Overused in LLMs
Machine Unlearning in Large Language Models: A Survey of Challenges and Methods
Is your algorithm unlearning or untraining?
Unlearning in LLMs: Methods, Evaluation, and Open Challenges
A Survey on Unlearning in Large Language Models
A Comprehensive Survey of Machine Unlearning Techniques for Large Language Models
Open Problems in Machine Unlearning for AI Safety
Position: LLM Unlearning Benchmarks are Weak Measures of Progress
Preserving Privacy in Large Language Models: A Survey on Current Threats and Solutions
Machine Unlearning in Generative AI: A Survey
Digital Forgetting in Large Language Models: A Survey of Unlearning Methods
Machine Unlearning for Traditional Models and Large Language Models: A Short Survey
The Frontier of Data Erasure: Machine Unlearning for Large Language Models
Rethinking Machine Unlearning for Large Language Models
Eight Methods to Evaluate Robust Unlearning in LLMs
Knowledge Unlearning for LLMs: Tasks, Methods, and Challenges
Right to be Forgotten in the Era of Large Language Models: Implications, Challenges, and Solutions
Contributions welcome! Please open a PR if you know of papers, benchmarks, or tools related to LLM unlearning.
- [Paper Title](url)
- Author(s): Name1, Name2, ...
- Date: YYYY-MM
- Venue: VenueName Year (or - if preprint)
- Code: [](url) (or - if none)
If you find this repository useful, please consider citing it:
@software{awesome-llm-unlearning,
title = {{Awesome Large Language Model Unlearning}},
author = {Liu, Chris Yuhao and others},
year = {2024},
doi = {10.5281/zenodo.19411433},
url = {https://github.com/chrisliu298/awesome-llm-unlearning},
version = {v1.0.0}
}
A resource repository for machine unlearning in large language models
626
184 commits
updated Aug 6, 2026
A curated collection of papers, surveys, benchmarks, frameworks, and blog posts for machine unlearning in large language models.
As of the last commit, there are 616 papers, 18 surveys and position papers, 3 frameworks, and 2 blog posts.
If you believe your paper on LLM unlearning is not included, or if you find a mistake, typo, or information that is not up to date, please open an issue or submit a pull request, and I will be happy to update the list.
UMU-Bench: Closing the Modality Gap in Multimodal Unlearning Evaluation
Elastic Robust Unlearning of Specific Knowledge in Large Language Models
MPSelectTune: Prompt-type Selection for Fine-tuning improves Concept Unlearning in LLMs
Identifying Unlearned Data in LLMs via Membership Inference Attacks
Unlearners Can Lie: Evaluating "Honesty" in LLM Unlearning
The Role of Learning and Memorization in Relabeling-based Unlearning for LLMs
On the Fragility of Latent Knowledge: Layer-wise Influence under Unlearning in Large Language Model
Lifelong Unlearning for Multimodal Large Language Models
MOUCHI: Mitigating Over-forgetting in Unlearning Copyrighted Information
Provably Continual Unlearning for Large Language Model
SELU: Energy-based Targeted Unlearning in LLMs
Refusal Is Not an Option: Unlearning Safety Alignment of Large Language Models
Rethinking Unlearning for Large Reasoning Models
Unlearning in Large Language Models: We Are Not There Yet
Investigating Model Editing for Unlearning in Large Language Models
The Erasure Illusion: Stress-Testing the Generalization of LLM Forgetting Evaluation
Towards Benchmarking Privacy Vulnerabilities in Selective Forgetting with Large Language Models
Towards Reasoning-Preserving Unlearning in Multimodal Large Language Models
Feature-Selective Representation Misdirection for Machine Unlearning
FAME: Fictional Actors for Multilingual Erasure
Explainable reinforcement learning from human feedback to improve alignment
FROC: A Unified Framework with Risk-Optimized Control for Machine Unlearning in LLMs
MLLM Machine Unlearning via Visual Knowledge Distillation
MedForget: Hierarchy-Aware Multimodal Unlearning Testbed for Medical AI
LUNE: Efficient LLM Unlearning via LoRA Fine-Tuning with Negative Examples
Recover-to-Forget: Gradient Reconstruction from LoRA for Efficient LLM Unlearning
Delete and Retain: Efficient Unlearning for Document Classification
Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
When Forgetting Builds Reliability: LLM Unlearning for Reliable Hardware Code Generation
RapidUn: Influence-Driven Parameter Reweighting for Efficient Large Language Model Unlearning
RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories
Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning
SineProject: Machine Unlearning for Stable Vision Language Alignment
From Narrow Unlearning to Emergent Misalignment: Causes, Consequences, and Containment in LLMs
Forgetting-MarI: LLM Unlearning via Marginal Information Regularization
AUVIC: Adversarial Unlearning of Visual Concepts for Multi-modal Large Language Models
Unlearning Imperative: Securing Trustworthy and Responsible LLMs through Engineered Forgetting
Cross-Modal Unlearning via Influential Neuron Path Editing in Multimodal Large Language Models
DRAGON: Guard LLM Unlearning in Context via Negative Detection and Reasoning
Leak@$k$: Unlearning Does Not Make LLMs Forget Under Probabilistic Decoding
REMIND: Input Loss Landscapes Reveal Residual Memorization in Post-Unlearning LLMs
The Realignment Problem: When Right becomes Wrong in LLMs
From Memorization to Reasoning in the Spectrum of Loss Curvature
Uncovering the Potential Risks in Unlearning: Danger of English-only Unlearning in Multilingual LLMs
Probing Knowledge Holes in Unlearned LLMs
The Evaluation of Retrieval-Based Unlearning Mechanisms on Large Language Models
OFFSIDE: Benchmarking Unlearning Misinformation in Multimodal Large Language Models
Label Smoothing Improves Gradient Ascent in LLM Unlearning
Leverage Unlearning to Sanitize LLMs
Hubble: a Model Suite to Advance the Study of LLM Memorization
LLM Unlearning with LLM Beliefs
Forget to Know, Remember to Use: Context-Aware Unlearning for Large Language Models
Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting
Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
Hierarchical Federated Unlearning for Large Language Models
On the Impossibility of Retrain Equivalence in Machine Unlearning
Reference-Specific Unlearning Metrics Can Hide the Truth: A Reality Check
LLM Unlearning on Noisy Forget Sets: A Study of Incomplete, Rewritten, and Watermarked Data
Approximate Domain Unlearning for Vision-Language Models
SIMU: Selective Influence Machine Unlearning
LLM Unlearning Under the Microscope: A Full-Stack View on Methods and Metrics
Cross-Modal Attention Guided Unlearning in Vision-Language Models
(Token-Level) InfoRMIA: Stronger Membership Inference and Memorization Assessment for LLMs
Distribution Preference Optimization: A Fine-grained Perspective for LLM Unlearning
Machine Unlearning Meets Adversarial Robustness via Constrained Interventions on LLMs
Downgrade to Upgrade: Optimizer Simplification Enhances Robustness in LLM Unlearning
KnowledgeSmith: Uncovering Knowledge Updating in LLMs with Model Editing and Unlearning
Direct Token Optimization: A Self-contained Approach to Large Language Model Unlearning
Scalable and Robust LLM Unlearning by Correcting Responses with Retrieved Exclusions
Mitigating Biases in Language Models via Bias Unlearning
Understanding the Dilemma of Unlearning for Large Language Models
Stable Forgetting: Bounded Parameter-Efficient Unlearning in LLMs
Dual-Space Smoothness for Robust and Balanced LLM Unlearning
OFMU: Optimization-Driven Framework for Machine Unlearning
Erase or Hide? Suppressing Spurious Unlearning Neurons for Robust Unlearning
CLUE: Conflict-guided Localization for LLM Unlearning Framework
Beyond Sharp Minima: Robust LLM Unlearning via Feedback-Guided Multi-Point Optimization
Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models
Concept Unlearning in Large Language Models via Self-Constructed Knowledge Triplets
Reveal and Release: Iterative LLM Unlearning with Self-generated Data
Scrub It Out! Erasing Sensitive Memorization in Code Language Models via Machine Unlearning
Collapse of Irrelevant Representations (CIR) Ensures Robust and Non-Disruptive LLM Unlearning
Customized Retrieval-Augmented Generation with LLM for Debiasing Recommendation Unlearning
AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs
Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?
Unlearning That Lasts: Utility-Preserving, Robust, and Almost Irreversible Forgetting in LLMs
Standard vs. Modular Sampling: Best Practices for Reliable LLM Unlearning
CURE: A Unified Framework for Class and Concept Unlearning via Retraining Emulation
Improving Fisher Information Estimation and Efficiency for LoRA-based LLM Unlearning
Unlearning as Ablation: Toward a Falsifiable Benchmark for Generative Scientific Discovery
Reliable Unlearning Harmful Information in LLMs with Metamorphosis Representation Projection
SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
Unlearning at Scale: Implementing the Right to be Forgotten in Large Language Models
Oblivionis: A Lightweight Learning and Unlearning Framework for Federated Large Language Models
Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs
LLM Unlearning using Gradient Ratio-Based Influence Estimation and Noise Injection
LLM Unlearning Without an Expert Curated Dataset
Analyzing and Mitigating Object Hallucination: A Training Bias Perspective
From Learning to Unlearning: Biomedical Security Protection in Multimodal Large Language Models
DUP: Detection-guided Unlearning for Backdoor Purification in Language Models
Towards Evaluation for Real-World LLM Unlearning
Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
What Should LLMs Forget? Quantifying Personal Data in LLMs for Right-to-Be-Forgotten Requests
Automating Evaluation of Diffusion Model Unlearning with (Vision-) Language Model World Knowledge
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs
Unlearning the Noisy Correspondence Makes CLIP More Robust
PULSE: Practical Evaluation Scenarios for Large Multimodal Model Unlearning
Agents Are All You Need for LLM Unlearning
SoK: Semantic Privacy in Large Language Models
Model State Arithmetic for Machine Unlearning
Step-by-Step Reasoning Attack: Revealing 'Erased' Knowledge in Large Language Models
Does Multimodal Large Language Model Truly Unlearn? Stealthy MLLM Unlearning Attack
Large Language Model Unlearning for Source Code
Mr. Snuffleupagus at SemEval-2025 Task 4: Unlearning Factual Knowledge from LLMs Using Adaptive RMU
BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap
Learning-Time Encoding Shapes Unlearning in LLMs
Unlearning Isn't Invisible: Detecting Unlearning Traces in LLMs from Model Outputs
Align-then-Unlearn: Embedding Alignment for LLM Unlearning
Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps
Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills
OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization
UCD: Unlearning in LLMs via Contrastive Decoding
Lifting Data-Tracing Machine Unlearning to Knowledge-Tracing for Foundation Models
GUARD: Guided Unlearning and Retention via Data Attribution for Large Language Models
SoK: Machine Unlearning for Large Language Models
BLUR: A Bi-Level Optimization Approach for LLM Unlearning
LLM Unlearning Should Be Form-Independent
RULE: Reinforcement UnLEarning Achieves Forget-Retain Pareto Optimality
Distillation Robustifies Unlearning
Towards Lifecycle Unlearning Commitment Management: Measuring Sample-level Unlearning Completeness
Do LLMs Really Forget? Evaluating Unlearning with Knowledge Correlation and Confidence Awareness
Constrained Entropic Unlearning: A Primal-Dual Framework for Large Language Models
Quantifying Cross-Modality Memorization in Vision-Language Models
Lacuna Inc. at SemEval-2025 Task 4: LoRA-Enhanced Influence-Based Unlearning for LLMs
Vulnerability-Aware Alignment: Mitigating Uneven Forgetting in Harmful Fine-Tuning
Not All Tokens Are Meant to Be Forgotten
Rethinking Post-Unlearning Behavior of Large Vision-Language Models
Invariance Makes LLM Unlearning Resilient Even to Unanticipated Downstream Fine-Tuning
Existing Large Language Model Unlearning Evaluations Are Inconclusive
Keeping an Eye on LLM Unlearning: The Hidden Risk and Remedy
Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race
Model Unlearning via Sparse Autoencoder Subspace Guided Projections
Does Machine Unlearning Truly Remove Model Knowledge? A Framework for Auditing Unlearning in LLMs
From Dormant to Deleted: Tamper-Resistant Unlearning Through Weight-Space Regularization
Graceful Forgetting in Generative Language Models
Safety Alignment via Constrained Knowledge Unlearning
T2VUnlearning: A Concept Erasing Method for Text-to-Video Diffusion Models
Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning
Losing is for Cherishing: Data Valuation Based on Machine Unlearning and Shapley Value
UniErase: Unlearning Token as a Universal Erasure Primitive for Language Models
R-TOFU: Unlearning in Large Reasoning Models
DUSK: Do Not Unlearn Shared Knowledge
SEPS: A Separability Measure for Robust Unlearning in LLMs
GUARD: Generation-time LLM Unlearning via Adaptive Restriction and Detection
Exploring Criteria of Loss Reweighting to Enhance LLM Unlearning
Unilogit: Robust Machine Unlearning for LLMs Using Uniform-Target Self-Distillation
OBLIVIATE: Robust and Practical Machine Unlearning for Large Language Models
Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?
Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
Token-Level Knowledge Unlearning for Large Language Model Security
AegisLLM: Scaling Agentic Systems for Self-Reflective Defense in LLM Security
DualOptim: Enhancing Efficacy and Stability in Machine Unlearning with Dual Optimizers
ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
DP2Unlearning: An Efficient and Guaranteed Unlearning Framework for LLMs
A mean teacher algorithm for unlearning of language models
GRAIL: Gradient-Based Adaptive Unlearning for Privacy and Copyright in LLMs
LLM Unlearning Reveals a Stronger-Than-Expected Coreset Effect in Current Benchmarks
Bridging the Gap Between Preference Alignment and Machine Unlearning
Exact Unlearning of Finetuning Data via Model Merging at Scale
SUV: Scalable Large Language Model Copyright Compliance with Regularized Selective Unlearning
Effective Skill Unlearning through Intervention and Abstention
ZJUKLAB at SemEval-2025 Task 4: Unlearning via Model Merging
Deep Contrastive Unlearning for Language Models
SAUCE: Selective Concept Unlearning in Vision-Language Models with Sparse Autoencoders
Atyaephyra at SemEval-2025 Task 4: Low-Rank NPO
PEBench: A Fictitious Dataset to Benchmark Machine Unlearning for Multimodal Large Language Models
Hyperbolic Safety-Aware Vision-Language Models
Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-tuning
Don't Forget It! Conditional Sparse Autoencoder Clamping Works for Unlearning
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
GRU: Mitigating the Trade-off between Unlearning and Retention for Large Language Models
UIPE: Enhancing LLM Unlearning by Removing Knowledge Related to Forgetting Targets
Improving LLM Safety Alignment with Dual-Objective Optimization
CE-U: Cross Entropy Unlearning
Rectifying Belief Space via Unlearning to Harness LLMs' Reasoning
Erasing Without Remembering: Safeguarding Knowledge Forgetting in Large Language Models
Rethinking LLM Unlearning Objectives: A Gradient Perspective and Go Beyond
A General Framework to Enhance Fine-tuning-based LLM Unlearning
Modality-Aware Neuron Pruning for Unlearning in Multimodal Large Language Models
Soft Token Attacks Cannot Reliably Audit Unlearning in Large Language Models
CoME: An Unlearning-based Approach to Conflict-free Model Editing
LUME: LLM Unlearning with Multitask Evaluations
UPCORE: Utility-Preserving Coreset Selection for Balanced Unlearning
Obliviate: Efficient Unmemorization for Protecting Intellectual Property in Large Language Models
Beyond Single-Value Metrics: Evaluating and Enhancing LLM Unlearning with Cognitive Diagnosis
Which Retain Set Matters for LLM Unlearning? A Case Study on Entity Unlearning
LUNAR: LLM Unlearning via Neural Activation Redirection
Mitigating Sensitive Information Leakage in LLMs4Code through Machine Unlearning
A Lightweight Method to Disrupt Memorized Sequences in LLM
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
Improving the Robustness of Representation Misdirection for Large Language Model Unlearning
Backdoor Token Unlearning: Exposing and Defending Backdoors in Pretrained Language Models
Precise In-Parameter Concept Erasure in Large Language Models
A Fully Probabilistic Perspective on Large Language Model Unlearning: Evaluation and Optimization
Machine Unlearning of Personally Identifiable Information in Large Language Models
Adaptive Localization of Knowledge Negation for Continual LLM Unlearning
LUSB: Formalizing and Benchmarking Unlearning Attacks and Defenses against Large Language Models
White-Box Auditing of Large Language Model Unlearning
UUE: Untargeted Language Model Unlearning via Null-Space-Guided Editing with Lightweight Adapters
UnRe: Zero-Shot LLM Unlearning via Dynamic Contextual Retrieval
Effective Unlearning in LLMs Relies on the Right Data Retention Strategy
Decoupling Memories, Muting Neurons: Towards Practical Machine Unlearning for Large Language Models
LLM-Eraser: Optimizing Large Language Model Unlearning through Selective Pruning
SemEval-2025 Task 4: Unlearning Sensitive Content from Large Language Models
NEKO at SemEval-2025 Task 4: A Gradient Ascent Based Machine Unlearning Strategy
NeuroReset: LLM Unlearning via Dual Phase Mixed Methodology
MALTO at SemEval-2025 Task 4: Dual Teachers for Unlearning Sensitive Content in LLMs
YNU at SemEval-2025 Task 4: Synthetic Token Alternative Training for LLM Unlearning
JU-CSE-NLP'25 at SemEval-2025 Task 4: Learning to Unlearn LLMs
NLPART at SemEval-2025 Task 4: Forgetting is Harder than Learning
Orthogonal Gradient Projection for Continual LLM Unlearning
Position: The Term "Machine Unlearning" Is Overused in LLMs
Machine Unlearning in Large Language Models: A Survey of Challenges and Methods
Is your algorithm unlearning or untraining?
Unlearning in LLMs: Methods, Evaluation, and Open Challenges
A Survey on Unlearning in Large Language Models
A Comprehensive Survey of Machine Unlearning Techniques for Large Language Models
Open Problems in Machine Unlearning for AI Safety
Position: LLM Unlearning Benchmarks are Weak Measures of Progress
Preserving Privacy in Large Language Models: A Survey on Current Threats and Solutions
Machine Unlearning in Generative AI: A Survey
Digital Forgetting in Large Language Models: A Survey of Unlearning Methods
Machine Unlearning for Traditional Models and Large Language Models: A Short Survey
The Frontier of Data Erasure: Machine Unlearning for Large Language Models
Rethinking Machine Unlearning for Large Language Models
Eight Methods to Evaluate Robust Unlearning in LLMs
Knowledge Unlearning for LLMs: Tasks, Methods, and Challenges
Right to be Forgotten in the Era of Large Language Models: Implications, Challenges, and Solutions
Contributions welcome! Please open a PR if you know of papers, benchmarks, or tools related to LLM unlearning.
- [Paper Title](url)
- Author(s): Name1, Name2, ...
- Date: YYYY-MM
- Venue: VenueName Year (or - if preprint)
- Code: [](url) (or - if none)
If you find this repository useful, please consider citing it:
@software{awesome-llm-unlearning,
title = {{Awesome Large Language Model Unlearning}},
author = {Liu, Chris Yuhao and others},
year = {2024},
doi = {10.5281/zenodo.19411433},
url = {https://github.com/chrisliu298/awesome-llm-unlearning},
version = {v1.0.0}
}