14 repos
Methods and tools for understanding how neural networks process information through activation intervention, patching, and lens-based analysis techniques. This cluster focuses on mechanistic interpretability—reverse-engineering the internal computations and learned representations that drive model behavior—using approaches like activation patching, intervention-based probing, and visualization tools that expose model internals. Researchers here build both foundational analysis frameworks and domain-specific applications (vision-language models, language models) to uncover what individual neurons, attention heads, and pathways actually compute.