Neural Network Activation Interpretability

13 repos

Techniques and models for interpreting neural network internals through activation analysis and decoding. This cluster focuses on understanding what happens inside large language models by examining their learned representations, with particular emphasis on activation-level interpretability methods. The central repositories represent various model scales (from 7B to 70B parameters) across different architectures (Llama, Gemma, Qwen) that have been instrumented or analyzed for activation decoding and similar interpretability research.

Python · 2
activation-decoding ·26
interpretability ·26
nla ·26
safetensors ·26
qwen2 ·15
gemma3_text ·11
gpt-oss-20b ·0
hidden-states ·0
vector-quantization ·0
llama ·0