5 repos
GeorgeMLP/reasoning-probing
Official code for ICML 2026 paper "Do Sparse Autoencoders Identify Reasoning Features in Language…
5
103 commits
caiovicentino1/FabricationGuard-linearprobe-qwen36-27b
No description
0
10 commits
Yusen-Peng/CE-Bench
[EMNLP-W 2025] CE-Bench: Towards a Reliable Contrastive Evaluation Benchmark of Interpretability of…
4
124 commits
caiovicentino1/ReasoningGuard-linearprobe-qwen36-27b
4 commits
crate-lm/crate-lm
Official Repository of the Paper "Improving Neuron-level Interpretability with White-box Language…
6
15 commits