xsankar/AI-Red-Teaming

All things specific to LLM Red Teaming Generative AI

30

68 commits

updated Oct 22, 2024

See the code

README

AI Red Teaming a.k.a. Awesome-LLM-Red-Teaming

Back to TOC

All things specific to Generative AI LLM Red Teaming

My Blog "What the heck is AI Red Teaming" https://bit.ly/ai-red-teaming

As of 11.30.23, I am working hard to build the repos - takes time to review and curate. Appreciate your patience ... Thanks ...
As of 2.1.24, Started transcribing and curating the links from my Omnioutline to this GitHub page ...


Best Practices

Top


NIST

Top

All NIST documents, ideas, responses et al

Most probably will split into a Awesome-NIST repository. I have - see Awesome-NIST


Survey & Analytical Papers

Top

YearTitleNotes
Survey Papers
2024.01Gradient-Based Language Model Red TeamingHot from the press (at least for now! as of Stardate -299100.57)
  • I had written, in my Red Teaming blog, “Tests follow a progressive nature, where a response could lead to another prompt deeper in the knowledge graph on the same topic” Here
  • I was thinking of a prompt hierarchy, this paper does the adaptive Red Teaming by creating new, modified prompts using backprop !!
2024.01Red Teaming Visual Language Models
2024.01Red-Teaming for Generative AI: Silver Bullet or Security Theater?
2023.11Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming in the Wild
2023.08Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
2023.06Explore, Establish, Exploit: Red Teaming Language Models from Scratch
2022.09Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
LLMs vs. LLMs
2022.02Red Teaming Language Models with Language Models
Analytical Papers
2023.10Risk Assessment and Statistical Significance in the Age of Foundation Models
Star Trek Stardate Calculator

Metrics

Top

LLM benchmarks (See LLM Evaluation Topics for a quick intro)

I will start polulating this section


Benchmarks

Top

LLM benchmarks (See LLM Evaluation Topics for a quick intro)

I will start polulating this section


Datasets

Top

LLM benchmarks (See LLM Evaluation Topics for a quick intro)

I will start polulating this section


Other Repos

Top

LLM benchmarks (See LLM Evaluation Other Repos

I will start polulating this section

TitleNotes
Awesome Security
Awesome ControlsLinks to various security fraeworks. Last update 4 years ago, still useful
Awesome InfosecA curated list of awesome information security resources

xsankar/AI-Red-Teaming

All things specific to LLM Red Teaming Generative AI

30

68 commits

updated Oct 22, 2024

See the code

README

AI Red Teaming a.k.a. Awesome-LLM-Red-Teaming

Back to TOC

All things specific to Generative AI LLM Red Teaming

My Blog "What the heck is AI Red Teaming" https://bit.ly/ai-red-teaming

As of 11.30.23, I am working hard to build the repos - takes time to review and curate. Appreciate your patience ... Thanks ...
As of 2.1.24, Started transcribing and curating the links from my Omnioutline to this GitHub page ...


Best Practices

Top


NIST

Top

All NIST documents, ideas, responses et al

Most probably will split into a Awesome-NIST repository. I have - see Awesome-NIST


Survey & Analytical Papers

Top

YearTitleNotes
Survey Papers
2024.01Gradient-Based Language Model Red TeamingHot from the press (at least for now! as of Stardate -299100.57)
  • I had written, in my Red Teaming blog, “Tests follow a progressive nature, where a response could lead to another prompt deeper in the knowledge graph on the same topic” Here
  • I was thinking of a prompt hierarchy, this paper does the adaptive Red Teaming by creating new, modified prompts using backprop !!
2024.01Red Teaming Visual Language Models
2024.01Red-Teaming for Generative AI: Silver Bullet or Security Theater?
2023.11Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming in the Wild
2023.08Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
2023.06Explore, Establish, Exploit: Red Teaming Language Models from Scratch
2022.09Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
LLMs vs. LLMs
2022.02Red Teaming Language Models with Language Models
Analytical Papers
2023.10Risk Assessment and Statistical Significance in the Age of Foundation Models
Star Trek Stardate Calculator

Metrics

Top

LLM benchmarks (See LLM Evaluation Topics for a quick intro)

I will start polulating this section


Benchmarks

Top

LLM benchmarks (See LLM Evaluation Topics for a quick intro)

I will start polulating this section


Datasets

Top

LLM benchmarks (See LLM Evaluation Topics for a quick intro)

I will start polulating this section


Other Repos

Top

LLM benchmarks (See LLM Evaluation Other Repos

I will start polulating this section

TitleNotes
Awesome Security
Awesome ControlsLinks to various security fraeworks. Last update 4 years ago, still useful
Awesome InfosecA curated list of awesome information security resources