4 repos
YuanBoXie/DeepRefusal
[EMNLP2025 Findings] Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via…
10
2 commits
GraySwanAI/circuit-breakers
Improving Alignment and Robustness with Circuit Breakers
269
5 commits
poloclub/llm-self-defense
LLM Self Defense: By Self Examination, LLMs know they are being tricked
53
3 commits
UCSB-NLP-Chang/SelfDenoise
No description
14