Neural Network Mechanistic Interpretability

14 repos

Methods and tools for understanding how neural networks process information through activation intervention, patching, and lens-based analysis techniques. This cluster focuses on mechanistic interpretability—reverse-engineering the internal computations and learned representations that drive model behavior—using approaches like activation patching, intervention-based probing, and visualization tools that expose model internals. Researchers here build both foundational analysis frameworks and domain-specific applications (vision-language models, language models) to uncover what individual neurons, attention heads, and pathways actually compute.

Python · 9
HTML · 1
interpretability ·1,273
mechanistic-interpretability ·1,264
activation-patching ·899
intervention ·899
activation-intervention ·899
llm ·254
transformers ·252
qwen ·246
pytorch ·245
visualization ·242