LLM Jailbreaking and Safety

10 repos

Methods, frameworks, and benchmarks for testing and understanding adversarial attacks against large language models, with a focus on jailbreak techniques that circumvent safety constraints. The cluster centers on discrete optimization approaches to generate effective adversarial prompts, alongside tools and evaluation frameworks for assessing LLM robustness. Researchers here explore both attack generation (including automated methods like TAP and EasyJailbreak) and defensive evaluation to improve model safety and alignment.

Python · 9
Jupyter Notebook · 1
discrete-optimization ·910
jailbreak ·910
jailbreak-framework ·910
large-language-model ·910
llm-safety-benchmark ·910
llm-security ·910