2 repos
Mechanistic interpretability research focused on sparse autoencoders (SAEs) applied to vision transformer models, particularly CLIP. The cluster contains trained SAE checkpoints and supporting code for extracting and analyzing learned features from different transformer layers and activation hooks. This represents work in neural network interpretability aimed at decomposing model internals into interpretable sparse components.