12 repos
Tools and research for understanding neural network internals through sparse autoencoders (SAEs) and mechanistic interpretability analysis. SAELens, the most central project, provides a comprehensive framework for training, evaluating, and analyzing sparse autoencoders to decompose learned representations into interpretable features. The cluster includes benchmarking systems (SAEBench), theoretical foundations in integral superposition, and circuit analysis methods that help researchers reverse-engineer model behavior and understand neural computation at a granular level.