10 repos
Methods, frameworks, and benchmarks for testing and understanding adversarial attacks against large language models, with a focus on jailbreak techniques that circumvent safety constraints. The cluster centers on discrete optimization approaches to generate effective adversarial prompts, alongside tools and evaluation frameworks for assessing LLM robustness. Researchers here explore both attack generation (including automated methods like TAP and EasyJailbreak) and defensive evaluation to improve model safety and alignment.