Low-rank sparse attention decomposition for LLM interpretability; active development continues in Llamascopium
See the codeLorsa is a novel attention decomposition method designed to tackle attention superposition, extracting tens of thousands of interpretable attention units from the attention layers of large language models.
The complete implementation of Lorsa has been migrated to the Language-Model-SAEs repository.
This repository is a comprehensive, fully-distributed framework for training, analyzing, and visualizing Sparse Autoencoders (SAEs) and their frontier variants, including:
Lorsa employs low-rank sparse decomposition to decompose attention layer outputs into interpretable feature units, effectively addressing the feature superposition problem in attention mechanisms. This enables a deeper understanding of how attention mechanisms work in large language models.
Please visit the Language-Model-SAEs repository for:
If you use Lorsa in your research, please cite:
@article{He2025Lorsa,
author = {Zhengfu He and Junxuan Wang and Rui Lin and Xuyang Ge and
Wentao Shu and Qiong Tang and Junping Zhang and Xipeng Qiu},
title = {Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition},
journal = {CoRR},
volume = {abs/2504.20938},
year = {2025},
url = {https://arxiv.org/abs/2504.20938},
eprint = {2504.20938},
eprinttype = {arXiv}
}
For any questions or suggestions, please submit an Issue in the Language-Model-SAEs repository!
Python
56.1%
TypeScript
42.0%
JavaScript
1.2%
Low-rank sparse attention decomposition for LLM interpretability; active development continues in Llamascopium
See the codeLorsa is a novel attention decomposition method designed to tackle attention superposition, extracting tens of thousands of interpretable attention units from the attention layers of large language models.
The complete implementation of Lorsa has been migrated to the Language-Model-SAEs repository.
This repository is a comprehensive, fully-distributed framework for training, analyzing, and visualizing Sparse Autoencoders (SAEs) and their frontier variants, including:
Lorsa employs low-rank sparse decomposition to decompose attention layer outputs into interpretable feature units, effectively addressing the feature superposition problem in attention mechanisms. This enables a deeper understanding of how attention mechanisms work in large language models.
Please visit the Language-Model-SAEs repository for:
If you use Lorsa in your research, please cite:
@article{He2025Lorsa,
author = {Zhengfu He and Junxuan Wang and Rui Lin and Xuyang Ge and
Wentao Shu and Qiong Tang and Junping Zhang and Xipeng Qiu},
title = {Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition},
journal = {CoRR},
volume = {abs/2504.20938},
year = {2025},
url = {https://arxiv.org/abs/2504.20938},
eprint = {2504.20938},
eprinttype = {arXiv}
}
For any questions or suggestions, please submit an Issue in the Language-Model-SAEs repository!
Python
56.1%
TypeScript
42.0%
JavaScript
1.2%