26 repos
Sparse autoencoders (SAEs) applied to interpretability research on vision-language models, particularly CLIP. This cluster focuses on mechanistic interpretability—understanding how neural networks process and represent information by decomposing activations into sparse, interpretable features. The central repositories contain trained SAE models across different layers and activation hooks of CLIP-B32, enabling researchers to probe what concepts these models learn and how they make predictions, with applications to concept-based explanations and domain-specific analysis like dermatology.