ydyjya/Awesome-LLM-Safety

A curated list of safety-related papers, articles, and resources focused on Large Language Models (LLMs). This repository aims to provide researchers, practitioners, and enthusiasts with insights into the safety implications, challenges, and advancements surrounding these powerful models.

HTML

1,914

233 commits

updated Jul 12, 2026

See the code

README

🛡️Awesome LLM-Safety🛡️Awesome

GitHub stars GitHub forks GitHub issues GitHub Last commit

English | 中文

🤗Introduction

Welcome to our Awesome-llm-safety repository! 🥰🥰🥰

🔥 News

  • 2024.05 update NAACL 2024 Papers Collection, thanks @zhrli324, @feqHe!

🧑‍💻 Our Work

We've curated a collection of the latest 😋, most comprehensive 😎, and most valuable 🤩 resources on large language model safety (llm-safety). But we don't stop there; included are also relevant talks, tutorials, conferences, news, and articles. Our repository is constantly updated to ensure you have the most current information at your fingertips.

If a resource is relevant to multiple subcategories, we place it under each applicable section. For instance, the "Awesome-LLM-Safety" repository will be listed under each subcategory to which it pertains🤩!.

✔️ Perfect for Majority

  • For beginners curious about llm-safety, our repository serves as a compass for grasping the big picture and diving into the details. Classic or influential papers retained in the README provide a beginner-friendly navigation through interesting directions in the field;
  • For seasoned researchers, this repository is a tool to keep you informed and fill any gaps in your knowledge. Within each subtopic, we are diligently updating all the latest content and continuously backfilling with previous work. Our thorough compilation and careful selection are time-savers for you.

🧭 How to Use this Guide

  • Quick Start: In the README, users can find a curated list of select information sorted by date, along with links to various consultations.
  • In-Depth Exploration: If you have a special interest in a particular subtopic, delve into the "subtopic" folder for more. Each item, be it an article or piece of news, comes with a brief introduction, allowing researchers to swiftly zero in on relevant content.

💼 How to Contribution

If you have completed an insightful work or carefully compiled conference papers, we would love to add your work to the repository.

  • For individual papers, you can raise an issue, and we will quickly add your paper under the corresponding subtopic.
  • If you have compiled a collection of papers for a conference, you are welcome to submit a pull request directly. We would greatly appreciate your contribution. Please note that these pull requests need to be consistent with our existing format.

📜Advertisement

🌱 If you would like more people to read your recent insightful work, please contact me via email. I can offer you a promotional spot here for up to one month.

Let’s start LLM Safety tutorial!


🚀Table of Contents


🤔AI Safety & Security Discussions

DateLinkPublicationAuthors
2024/5/20Managing extreme AI risks amid rapid progressYoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Trevor Darrell, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, Gillian Hadfield, Jeff Clune, Tegan Maharaj, Frank Hutter, Atılım Güneş Baydin, Sheila McIlraith, Qiqi Gao, Ashwin Acharya, David Krueger, Anca Dragan, Philip Torr, Stuart Russell, Daniel Kahneman, Jan Brauner, Sören MindermannScience

🔐Security & Discussion

📑Papers

📖Tutorials, Articles, Presentations and Talks

DateTypeTitleURL
22.02Toxicity Detection APIPerspective APIlink
paper
23.07RepositoryAwesome LLM Securitylink
23.10TutorialsAwesome-LLM-Safetylink
24.01TutorialsAwesome-LM-SSPlink

Other

👉Latest&Comprehensive Security Paper


🔏Privacy

📑Papers

📖Tutorials, Articles, Presentations and Talks

DateTypeTitleURL
23.10TutorialsAwesome-LLM-Safetylink
24.01TutorialsAwesome-LM-SSPlink

Other

👉Latest&Comprehensive Privacy Paper


📰Truthfulness & Misinformation

📑Papers

📖Tutorials, Articles, Presentations and Talks

DateTypeTitleURL
23.07Repositoryllm-hallucination-surveylink
23.10RepositoryLLM-Factuality-Surveylink
23.10TutorialsAwesome-LLM-Safetylink

Other

👉Latest&Comprehensive Truthfulness&Misinformation Paper


😈JailBreak & Attacks

📑Papers

📖Tutorials, Articles, Presentations and Talks

DateTypeTitleURL
23.01CommunityReddit/ChatGPTJailbreklink
23.02Resource&TutorialsLatest Jailbreak Promptslink
23.10TutorialsAwesome-LLM-Safetylink
23.10ArticleAdversarial Attacks on LLMs(Author: Lilian Weng)link
23.11Video[1hr Talk] Intro to Large Language Models
From 45:45(Author: Andrej Karpathy)
link
24.09Repoawesome_LLM-harmful-fine-tuning-paperslink
12.10ResourceJailbreak Commuinitieslink
12.10ArticleJailbreak Techniques and Safeguardslink

Other

👉Latest&Comprehensive JailBreak & Attacks Paper


🛡️Defenses & Mitigation

📑Papers

📖Tutorials, Articles, Presentations and Talks

DateTypeTitleURL
23.10TutorialsAwesome-LLM-Safetylink

Other

👉Latest&Comprehensive Defenses Paper


💯Datasets & Benchmark

📑Papers

📖Tutorials, Articles, Presentations and Talks

DateTypeTitleURL
23.10TutorialsAwesome-LLM-Safetylink

📚Resource📚

Other

👉Latest&Comprehensive datasets & Benchmark Paper


🧑‍🎓Author

🤗If you have any questions, please contact our authors!🤗

✉️: ydyjya ➡️ zhouzhenhong@bupt.edu.cn

💬: LLM Safety Discussion


Star History Chart

⬆ Back to ToC

Contributors

ydyjya

206 commits

HowieHwong

7 commits

suyuleyuan

6 commits

GIGABaozi

5 commits

ydyjya/Awesome-LLM-Safety

A curated list of safety-related papers, articles, and resources focused on Large Language Models (LLMs). This repository aims to provide researchers, practitioners, and enthusiasts with insights into the safety implications, challenges, and advancements surrounding these powerful models.

HTML

1,914

233 commits

updated Jul 12, 2026

See the code

README

🛡️Awesome LLM-Safety🛡️Awesome

GitHub stars GitHub forks GitHub issues GitHub Last commit

English | 中文

🤗Introduction

Welcome to our Awesome-llm-safety repository! 🥰🥰🥰

🔥 News

  • 2024.05 update NAACL 2024 Papers Collection, thanks @zhrli324, @feqHe!

🧑‍💻 Our Work

We've curated a collection of the latest 😋, most comprehensive 😎, and most valuable 🤩 resources on large language model safety (llm-safety). But we don't stop there; included are also relevant talks, tutorials, conferences, news, and articles. Our repository is constantly updated to ensure you have the most current information at your fingertips.

If a resource is relevant to multiple subcategories, we place it under each applicable section. For instance, the "Awesome-LLM-Safety" repository will be listed under each subcategory to which it pertains🤩!.

✔️ Perfect for Majority

  • For beginners curious about llm-safety, our repository serves as a compass for grasping the big picture and diving into the details. Classic or influential papers retained in the README provide a beginner-friendly navigation through interesting directions in the field;
  • For seasoned researchers, this repository is a tool to keep you informed and fill any gaps in your knowledge. Within each subtopic, we are diligently updating all the latest content and continuously backfilling with previous work. Our thorough compilation and careful selection are time-savers for you.

🧭 How to Use this Guide

  • Quick Start: In the README, users can find a curated list of select information sorted by date, along with links to various consultations.
  • In-Depth Exploration: If you have a special interest in a particular subtopic, delve into the "subtopic" folder for more. Each item, be it an article or piece of news, comes with a brief introduction, allowing researchers to swiftly zero in on relevant content.

💼 How to Contribution

If you have completed an insightful work or carefully compiled conference papers, we would love to add your work to the repository.

  • For individual papers, you can raise an issue, and we will quickly add your paper under the corresponding subtopic.
  • If you have compiled a collection of papers for a conference, you are welcome to submit a pull request directly. We would greatly appreciate your contribution. Please note that these pull requests need to be consistent with our existing format.

📜Advertisement

🌱 If you would like more people to read your recent insightful work, please contact me via email. I can offer you a promotional spot here for up to one month.

Let’s start LLM Safety tutorial!


🚀Table of Contents


🤔AI Safety & Security Discussions

DateLinkPublicationAuthors
2024/5/20Managing extreme AI risks amid rapid progressYoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Trevor Darrell, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, Gillian Hadfield, Jeff Clune, Tegan Maharaj, Frank Hutter, Atılım Güneş Baydin, Sheila McIlraith, Qiqi Gao, Ashwin Acharya, David Krueger, Anca Dragan, Philip Torr, Stuart Russell, Daniel Kahneman, Jan Brauner, Sören MindermannScience

🔐Security & Discussion

📑Papers

📖Tutorials, Articles, Presentations and Talks

DateTypeTitleURL
22.02Toxicity Detection APIPerspective APIlink
paper
23.07RepositoryAwesome LLM Securitylink
23.10TutorialsAwesome-LLM-Safetylink
24.01TutorialsAwesome-LM-SSPlink

Other

👉Latest&Comprehensive Security Paper


🔏Privacy

📑Papers

📖Tutorials, Articles, Presentations and Talks

DateTypeTitleURL
23.10TutorialsAwesome-LLM-Safetylink
24.01TutorialsAwesome-LM-SSPlink

Other

👉Latest&Comprehensive Privacy Paper


📰Truthfulness & Misinformation

📑Papers

📖Tutorials, Articles, Presentations and Talks

DateTypeTitleURL
23.07Repositoryllm-hallucination-surveylink
23.10RepositoryLLM-Factuality-Surveylink
23.10TutorialsAwesome-LLM-Safetylink

Other

👉Latest&Comprehensive Truthfulness&Misinformation Paper


😈JailBreak & Attacks

📑Papers

📖Tutorials, Articles, Presentations and Talks

DateTypeTitleURL
23.01CommunityReddit/ChatGPTJailbreklink
23.02Resource&TutorialsLatest Jailbreak Promptslink
23.10TutorialsAwesome-LLM-Safetylink
23.10ArticleAdversarial Attacks on LLMs(Author: Lilian Weng)link
23.11Video[1hr Talk] Intro to Large Language Models
From 45:45(Author: Andrej Karpathy)
link
24.09Repoawesome_LLM-harmful-fine-tuning-paperslink
12.10ResourceJailbreak Commuinitieslink
12.10ArticleJailbreak Techniques and Safeguardslink

Other

👉Latest&Comprehensive JailBreak & Attacks Paper


🛡️Defenses & Mitigation

📑Papers

📖Tutorials, Articles, Presentations and Talks

DateTypeTitleURL
23.10TutorialsAwesome-LLM-Safetylink

Other

👉Latest&Comprehensive Defenses Paper


💯Datasets & Benchmark

📑Papers

📖Tutorials, Articles, Presentations and Talks

DateTypeTitleURL
23.10TutorialsAwesome-LLM-Safetylink

📚Resource📚

Other

👉Latest&Comprehensive datasets & Benchmark Paper


🧑‍🎓Author

🤗If you have any questions, please contact our authors!🤗

✉️: ydyjya ➡️ zhouzhenhong@bupt.edu.cn

💬: LLM Safety Discussion


Star History Chart

⬆ Back to ToC

Contributors

ydyjya

206 commits

HowieHwong

7 commits

suyuleyuan

6 commits

GIGABaozi

5 commits

Languages

HTML

99.5%