This repository contains the following SAEs:
Model described in the paper I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders. Code available at https://github.com/AIRI-Institute/SAE-Reasoning
Load these SAEs using SAELens as below:
from sae_lens import SAE
sae, cfg_dict, sparsity = SAE.from_pretrained("andreuka18/deepseek-r1-distill-llama-8b-lmsys-openthoughts", "<sae_id>")
7 commits
1 commits
This repository contains the following SAEs:
Model described in the paper I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders. Code available at https://github.com/AIRI-Institute/SAE-Reasoning
Load these SAEs using SAELens as below:
from sae_lens import SAE
sae, cfg_dict, sparsity = SAE.from_pretrained("andreuka18/deepseek-r1-distill-llama-8b-lmsys-openthoughts", "<sae_id>")
7 commits
1 commits