This project is under heavy active development. We are planning to release a stable version of the framework in the coming weeks.
In the meantime, join our Slack Community
Learn more about our approach and early results:
This project trains Sparse Autoencoders (SAEs) on the internal activations of an AI agent to understand why it selects specific tools. Given a user request like "Search our docs for API limits," the agent must choose between tools (e.g., internal_search vs web_search). We extract the model's hidden representations at the moment of that decision, decompose them into interpretable features using a JumpReLU SAE, and validate the resulting explanations through automated fuzzing and causal ablation experiments.
The key insight: train the SAE on raw activations (not difference vectors), then use contrastive pairs post-hoc to identify which learned features correspond to specific tool-selection decisions. This preserves the SAE's general feature dictionary while enabling targeted analysis of decision-relevant features.
For loading and running pretrained SAEs:
pip install kiji-inspector
For the HuggingFace-based training and analysis extras (accelerate etc.):
pip install 'kiji-inspector[full]'
The vLLM extraction path is not covered by any extra: upstream vLLM wheels do
not ship the hidden-states connector, so it requires the Docker image built
from the repository Dockerfile, which compiles vLLM from the
Davidnet/vllm fork.
from kiji_inspector import SAE
sae, feature_descriptions = SAE.from_pretrained(
repo_id="575-lab/kiji-inspector-google-gemma-4-E4B-it",
layer=30,
)
# encode/decode operate in the SAE's normalized space
features = sae.encode(sae.normalize_input(activations))
reconstruction = sae.denormalize_output(sae.decode(features))
normalize_input applies the exact transform the SAE was trained under —
(x - mean_vec) / rms_scale, a single global mean vector and RMS constant
computed over the training set. Raw activations must go through it or the
JumpReLU thresholds are meaningless; denormalize_output inverts it when a
reconstruction is written back into the model.
Training and data-generation entrypoints live under the package namespace:
python -m kiji_inspector.generate_pairs 1300
python -m kiji_inspector.pipeline --layers 10 20 30
Activation extraction uses vLLM's native extract_hidden_states speculator
method together with the ExampleHiddenStatesConnector, which writes captured
hidden states to safetensors files that the extractor loads and cleans up per
request. This capability ships in the 575lab/kiji-inspector:dev image, which
builds the Davidnet/vllm fork
(branch hidden-states-inline-return-squashed); the public v0.19.0 wheel does
not contain the connector.
Run all extraction, tests, and API checks inside that image:
docker pull 575lab/kiji-inspector:dev
# Smoke test the connector (Qwen3-8B):
samples/run_hidden_states_test.sh 575lab/kiji-inspector:dev
# Iterate on the checked-out source with a bind mount:
docker run --rm --gpus all \
-v "$PWD:/workspace" \
-v "${HF_CACHE:-$HOME/.cache/huggingface}:/root/.cache/huggingface" \
-e HF_HOME=/root/.cache/huggingface \
-e PYTHONPATH=/workspace/src \
-w /workspace \
575lab/kiji-inspector:dev \
python -m pytest tests/test_activation_extractors.py
The historical patches/ directory (applied against a stock v0.19.0 wheel)
predates the fork-based image and is retained for reference only; the current
image bakes those changes into the fork and does not run apply-patch.sh. See
patches/README_PATCH.md for the legacy workflow.
To run the pipeline against a Qwen3.6 subject model (a reasoning model), pass --no-thinking so the decision token sits at the final-answer position:
python -m kiji_inspector.pipeline --subject-model Qwen/Qwen3.6-35B-A3B --no-thinking
--subject-model also accepts a local model directory (paths must start with
/, ./, ../, or ~; anything else is treated as a hub ID). Inside Docker
the directory must be volume-mounted, and you pass the container path:
docker run --rm --gpus all \
-v "$PWD:/workspace" \
-v /home/user/models:/models:ro \
-v "${HF_CACHE:-$HOME/.cache/huggingface}:/root/.cache/huggingface" \
-e HF_HOME=/root/.cache/huggingface \
-e PYTHONPATH=/workspace/src \
-w /workspace \
575lab/kiji-inspector:dev \
python -m kiji_inspector.pipeline \
--subject-model /models/programs/downloaded-model
Note: SAE.from_pretrained(base_model=...) resolves SAE repos through an
exact-match registry of hub IDs, so it won't match a local path — pass
repo_id directly instead.
Two end-to-end notebooks demonstrate the library. Both run on Colab — the first needs an L4 (24 GB), the second an A100 high-RAM runtime.
| Notebook | What it shows | Open |
|---|---|---|
quickstart_colab.ipynb | Minimal walkthrough: capture the decision-token activation of google/gemma-4-E4B-it via a forward pre-hook, load the layer-30 SAE with SAE.from_pretrained, normalize with sae.normalize_input, and describe the top features firing on a single prompt. | |
home_repair_colab.ipynb | Full agent demo: a Nemotron-3-Nano-30B home repair advisor calls four tools across three appliance problems, the residual stream is captured at every decision point, and a trained JumpReLU SAE decomposes those activations into themed features. Includes the interactive index.html viewer served from Colab. |
We welcome contributions! Whether you're fixing a bug, improving documentation, or proposing a new feature, your help is appreciated.
Copyright (c) 2026 Dataiku SAS
This project is licensed under the Apache 2.0 License - see the LICENSE file for details.
HTML
89.7%
Python
7.8%
TeX
2.0%
This project is under heavy active development. We are planning to release a stable version of the framework in the coming weeks.
In the meantime, join our Slack Community
Learn more about our approach and early results:
This project trains Sparse Autoencoders (SAEs) on the internal activations of an AI agent to understand why it selects specific tools. Given a user request like "Search our docs for API limits," the agent must choose between tools (e.g., internal_search vs web_search). We extract the model's hidden representations at the moment of that decision, decompose them into interpretable features using a JumpReLU SAE, and validate the resulting explanations through automated fuzzing and causal ablation experiments.
The key insight: train the SAE on raw activations (not difference vectors), then use contrastive pairs post-hoc to identify which learned features correspond to specific tool-selection decisions. This preserves the SAE's general feature dictionary while enabling targeted analysis of decision-relevant features.
For loading and running pretrained SAEs:
pip install kiji-inspector
For the HuggingFace-based training and analysis extras (accelerate etc.):
pip install 'kiji-inspector[full]'
The vLLM extraction path is not covered by any extra: upstream vLLM wheels do
not ship the hidden-states connector, so it requires the Docker image built
from the repository Dockerfile, which compiles vLLM from the
Davidnet/vllm fork.
from kiji_inspector import SAE
sae, feature_descriptions = SAE.from_pretrained(
repo_id="575-lab/kiji-inspector-google-gemma-4-E4B-it",
layer=30,
)
# encode/decode operate in the SAE's normalized space
features = sae.encode(sae.normalize_input(activations))
reconstruction = sae.denormalize_output(sae.decode(features))
normalize_input applies the exact transform the SAE was trained under —
(x - mean_vec) / rms_scale, a single global mean vector and RMS constant
computed over the training set. Raw activations must go through it or the
JumpReLU thresholds are meaningless; denormalize_output inverts it when a
reconstruction is written back into the model.
Training and data-generation entrypoints live under the package namespace:
python -m kiji_inspector.generate_pairs 1300
python -m kiji_inspector.pipeline --layers 10 20 30
Activation extraction uses vLLM's native extract_hidden_states speculator
method together with the ExampleHiddenStatesConnector, which writes captured
hidden states to safetensors files that the extractor loads and cleans up per
request. This capability ships in the 575lab/kiji-inspector:dev image, which
builds the Davidnet/vllm fork
(branch hidden-states-inline-return-squashed); the public v0.19.0 wheel does
not contain the connector.
Run all extraction, tests, and API checks inside that image:
docker pull 575lab/kiji-inspector:dev
# Smoke test the connector (Qwen3-8B):
samples/run_hidden_states_test.sh 575lab/kiji-inspector:dev
# Iterate on the checked-out source with a bind mount:
docker run --rm --gpus all \
-v "$PWD:/workspace" \
-v "${HF_CACHE:-$HOME/.cache/huggingface}:/root/.cache/huggingface" \
-e HF_HOME=/root/.cache/huggingface \
-e PYTHONPATH=/workspace/src \
-w /workspace \
575lab/kiji-inspector:dev \
python -m pytest tests/test_activation_extractors.py
The historical patches/ directory (applied against a stock v0.19.0 wheel)
predates the fork-based image and is retained for reference only; the current
image bakes those changes into the fork and does not run apply-patch.sh. See
patches/README_PATCH.md for the legacy workflow.
To run the pipeline against a Qwen3.6 subject model (a reasoning model), pass --no-thinking so the decision token sits at the final-answer position:
python -m kiji_inspector.pipeline --subject-model Qwen/Qwen3.6-35B-A3B --no-thinking
--subject-model also accepts a local model directory (paths must start with
/, ./, ../, or ~; anything else is treated as a hub ID). Inside Docker
the directory must be volume-mounted, and you pass the container path:
docker run --rm --gpus all \
-v "$PWD:/workspace" \
-v /home/user/models:/models:ro \
-v "${HF_CACHE:-$HOME/.cache/huggingface}:/root/.cache/huggingface" \
-e HF_HOME=/root/.cache/huggingface \
-e PYTHONPATH=/workspace/src \
-w /workspace \
575lab/kiji-inspector:dev \
python -m kiji_inspector.pipeline \
--subject-model /models/programs/downloaded-model
Note: SAE.from_pretrained(base_model=...) resolves SAE repos through an
exact-match registry of hub IDs, so it won't match a local path — pass
repo_id directly instead.
Two end-to-end notebooks demonstrate the library. Both run on Colab — the first needs an L4 (24 GB), the second an A100 high-RAM runtime.
| Notebook | What it shows | Open |
|---|---|---|
quickstart_colab.ipynb | Minimal walkthrough: capture the decision-token activation of google/gemma-4-E4B-it via a forward pre-hook, load the layer-30 SAE with SAE.from_pretrained, normalize with sae.normalize_input, and describe the top features firing on a single prompt. | |
home_repair_colab.ipynb | Full agent demo: a Nemotron-3-Nano-30B home repair advisor calls four tools across three appliance problems, the residual stream is captured at every decision point, and a trained JumpReLU SAE decomposes those activations into themed features. Includes the interactive index.html viewer served from Colab. |
We welcome contributions! Whether you're fixing a bug, improving documentation, or proposing a new feature, your help is appreciated.
Copyright (c) 2026 Dataiku SAS
This project is licensed under the Apache 2.0 License - see the LICENSE file for details.
HTML
89.7%
Python
7.8%
TeX
2.0%