A curated list of LLM Interpretability related material - Tutorial, Library, Survey, Paper, Blog, etc..
306
69 commits
updated Jan 22, 2026
A curated list of LLM Interpretability related material.
Note: These Alignment surveys discuss the relation between Interpretability and LLM Alignment.
Large Language Model Alignment: A Survey [arxiv 2309]
AI Alignment: A Comprehensive Survey [arxiv 2310] [github] [website]
๐ICML 2024 Workshop on Mechanistic Interpretability [openreview]
๐Transformer Circuits Thread [blog]
BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP [workshop]
AI Alignment Forum [forum]
Lesswrong [forum]
Neel Nanda [blog] [google scholar]
Mor Geva [google scholar]
David Bau [google scholar]
Jacob Steinhardt [google scholar]
Yonatan Belinkov [google scholar]

๐interpreting GPT: the logit lens [Lesswrong 2020]
๐Analyzing Transformers in Embedding Space [ACL 2023]
Eliciting Latent Predictions from Transformers with the Tuned Lens [arxiv 2303]
An Adversarial Example for Direct Logit Attribution: Memory Management in gelu-4l arxiv 2310
Future Lens: Anticipating Subsequent Tokens from a Single Hidden State [CoNLL 2023]
SelfIE: Self-Interpretation of Large Language Model Embeddings [arxiv 2403]
InversionView: A General-Purpose Method for Reading Information from Neural Activations [ICML 2024 MI Workshop]
๐Awesome-Attention-Heads [github]
๐In-context learning and induction heads [Transformer Circuits Thread]
On the Expressivity Role of LayerNorm in Transformers' Attention [ACL 2023 Findings]
On the Role of Attention in Prompt-tuning [ICML 2023]
Copy Suppression: Comprehensively Understanding an Attention Head [ICLR 2024]
Successor Heads: Recurring, Interpretable Attention Heads In The Wild [ICLR 2024]
A phase transition between positional and semantic learning in a solvable model of dot-product attention [arxiv 2024]
Retrieval Head Mechanistically Explains Long-Context Factuality [arxiv 2404]
Iteration Head: A Mechanistic Study of Chain-of-Thought [arxiv 2406]
When Attention Sink Emerges in Language Models: An Empirical View [arxiv 2410]
66 commits
3 commits
A curated list of LLM Interpretability related material - Tutorial, Library, Survey, Paper, Blog, etc..
306
69 commits
updated Jan 22, 2026
A curated list of LLM Interpretability related material.
Note: These Alignment surveys discuss the relation between Interpretability and LLM Alignment.
Large Language Model Alignment: A Survey [arxiv 2309]
AI Alignment: A Comprehensive Survey [arxiv 2310] [github] [website]
๐ICML 2024 Workshop on Mechanistic Interpretability [openreview]
๐Transformer Circuits Thread [blog]
BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP [workshop]
AI Alignment Forum [forum]
Lesswrong [forum]
Neel Nanda [blog] [google scholar]
Mor Geva [google scholar]
David Bau [google scholar]
Jacob Steinhardt [google scholar]
Yonatan Belinkov [google scholar]

๐interpreting GPT: the logit lens [Lesswrong 2020]
๐Analyzing Transformers in Embedding Space [ACL 2023]
Eliciting Latent Predictions from Transformers with the Tuned Lens [arxiv 2303]
An Adversarial Example for Direct Logit Attribution: Memory Management in gelu-4l arxiv 2310
Future Lens: Anticipating Subsequent Tokens from a Single Hidden State [CoNLL 2023]
SelfIE: Self-Interpretation of Large Language Model Embeddings [arxiv 2403]
InversionView: A General-Purpose Method for Reading Information from Neural Activations [ICML 2024 MI Workshop]
๐Awesome-Attention-Heads [github]
๐In-context learning and induction heads [Transformer Circuits Thread]
On the Expressivity Role of LayerNorm in Transformers' Attention [ACL 2023 Findings]
On the Role of Attention in Prompt-tuning [ICML 2023]
Copy Suppression: Comprehensively Understanding an Attention Head [ICLR 2024]
Successor Heads: Recurring, Interpretable Attention Heads In The Wild [ICLR 2024]
A phase transition between positional and semantic learning in a solvable model of dot-product attention [arxiv 2024]
Retrieval Head Mechanistically Explains Long-Context Factuality [arxiv 2404]
Iteration Head: A Mechanistic Study of Chain-of-Thought [arxiv 2406]
When Attention Sink Emerges in Language Models: An Empirical View [arxiv 2410]
66 commits
3 commits