4
stars
commits
1
linked in READMEs
Apr 20, 2026
updated
Trained by https://github.com/YuanBoXie/DeepRefusal
[1] Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction, EMNLP 2025
4 commits
wangzhang/Llama-3-8B-Instruct-DeepRefusal-Broken
YuanBoXie/DeepRefusal
[EMNLP2025 Findings] Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via…
10
meta-llama/Meta-Llama-3-8B-Instruct
4,958
meta-llama/Meta-Llama-3.1-8B-Instruct
6,900
meta-llama/Meta-Llama-3-70B-Instruct
1,525
meta-llama/Meta-Llama-3-8B
6,653
meta-llama/Meta-Llama-3.1-70B-Instruct
975
meta-llama/Meta-Llama-3.1-405B-Instruct
600