dataiku/kiji-inspector

113

stars

40

commits

HTML

primary language

Sep 2, 2026

updated

README

Kiji Inspector: Mechanistic Interpretability for AI Agent Tool Selection

Kiji Inspector Workflow

CI Core CI Extras License: Apache 2.0 GitHub Stars GitHub Issues

Python Version Open In Colab

Responsible AI Contributions Welcome PRs Welcome

Status

This project is under heavy active development. We are planning to release a stable version of the framework in the coming weeks.

In the meantime, join our Slack Community

Learn more about our approach and early results:


What This Project Does

This project trains Sparse Autoencoders (SAEs) on the internal activations of an AI agent to understand why it selects specific tools. Given a user request like "Search our docs for API limits," the agent must choose between tools (e.g., internal_search vs web_search). We extract the model's hidden representations at the moment of that decision, decompose them into interpretable features using a JumpReLU SAE, and validate the resulting explanations through automated fuzzing and causal ablation experiments.

The key insight: train the SAE on raw activations (not difference vectors), then use contrastive pairs post-hoc to identify which learned features correspond to specific tool-selection decisions. This preserves the SAE's general feature dictionary while enabling targeted analysis of decision-relevant features.

Install

For loading and running pretrained SAEs:

pip install kiji-inspector

For the HuggingFace-based training and analysis extras (accelerate etc.):

pip install 'kiji-inspector[full]'

The vLLM extraction path is not covered by any extra: upstream vLLM wheels do not ship the hidden-states connector, so it requires the Docker image built from the repository Dockerfile, which compiles vLLM from the Davidnet/vllm fork.

Quick Start

from kiji_inspector import SAE

sae, feature_descriptions = SAE.from_pretrained(
    repo_id="575-lab/kiji-inspector-google-gemma-4-E4B-it",
    layer=30,
)

# encode/decode operate in the SAE's normalized space
features = sae.encode(sae.normalize_input(activations))
reconstruction = sae.denormalize_output(sae.decode(features))

normalize_input applies the exact transform the SAE was trained under — (x - mean_vec) / rms_scale, a single global mean vector and RMS constant computed over the training set. Raw activations must go through it or the JumpReLU thresholds are meaningless; denormalize_output inverts it when a reconstruction is written back into the model.

Training and data-generation entrypoints live under the package namespace:

python -m kiji_inspector.generate_pairs 1300
python -m kiji_inspector.pipeline --layers 10 20 30

vLLM hidden-state extraction (native connector)

Activation extraction uses vLLM's native extract_hidden_states speculator method together with the ExampleHiddenStatesConnector, which writes captured hidden states to safetensors files that the extractor loads and cleans up per request. This capability ships in the 575lab/kiji-inspector:dev image, which builds the Davidnet/vllm fork (branch hidden-states-inline-return-squashed); the public v0.19.0 wheel does not contain the connector.

Run all extraction, tests, and API checks inside that image:

docker pull 575lab/kiji-inspector:dev

# Smoke test the connector (Qwen3-8B):
samples/run_hidden_states_test.sh 575lab/kiji-inspector:dev

# Iterate on the checked-out source with a bind mount:
docker run --rm --gpus all \
  -v "$PWD:/workspace" \
  -v "${HF_CACHE:-$HOME/.cache/huggingface}:/root/.cache/huggingface" \
  -e HF_HOME=/root/.cache/huggingface \
  -e PYTHONPATH=/workspace/src \
  -w /workspace \
  575lab/kiji-inspector:dev \
  python -m pytest tests/test_activation_extractors.py

The historical patches/ directory (applied against a stock v0.19.0 wheel) predates the fork-based image and is retained for reference only; the current image bakes those changes into the fork and does not run apply-patch.sh. See patches/README_PATCH.md for the legacy workflow.

To run the pipeline against a Qwen3.6 subject model (a reasoning model), pass --no-thinking so the decision token sits at the final-answer position:

python -m kiji_inspector.pipeline --subject-model Qwen/Qwen3.6-35B-A3B --no-thinking

Using a locally downloaded model

--subject-model also accepts a local model directory (paths must start with /, ./, ../, or ~; anything else is treated as a hub ID). Inside Docker the directory must be volume-mounted, and you pass the container path:

docker run --rm --gpus all \
  -v "$PWD:/workspace" \
  -v /home/user/models:/models:ro \
  -v "${HF_CACHE:-$HOME/.cache/huggingface}:/root/.cache/huggingface" \
  -e HF_HOME=/root/.cache/huggingface \
  -e PYTHONPATH=/workspace/src \
  -w /workspace \
  575lab/kiji-inspector:dev \
  python -m kiji_inspector.pipeline \
    --subject-model /models/programs/downloaded-model

Note: SAE.from_pretrained(base_model=...) resolves SAE repos through an exact-match registry of hub IDs, so it won't match a local path — pass repo_id directly instead.


📓 Examples

Two end-to-end notebooks demonstrate the library. Both run on Colab — the first needs an L4 (24 GB), the second an A100 high-RAM runtime.

NotebookWhat it showsOpen
quickstart_colab.ipynbMinimal walkthrough: capture the decision-token activation of google/gemma-4-E4B-it via a forward pre-hook, load the layer-30 SAE with SAE.from_pretrained, normalize with sae.normalize_input, and describe the top features firing on a single prompt.Open In Colab
home_repair_colab.ipynbFull agent demo: a Nemotron-3-Nano-30B home repair advisor calls four tools across three appliance problems, the residual stream is captured at every decision point, and a trained JumpReLU SAE decomposes those activations into themed features. Includes the interactive index.html viewer served from Colab.Open In Colab

🤝 Contributing

We welcome contributions! Whether you're fixing a bug, improving documentation, or proposing a new feature, your help is appreciated.

Ways to Contribute

  • Report Bugs - Open an issue with steps to reproduce
  • Improve Docs - Documentation PRs are always welcome
  • Submit Features - Open an issue to discuss your idea before submitting a PR
  • Share Feedback - Start a discussion

Community


📄 License

Copyright (c) 2026 Dataiku SAS

This project is licensed under the Apache 2.0 License - see the LICENSE file for details.

Contributors

Davidnet

25 commits

hanneshapke

12 commits

carloshuisa

1 commits

claude

1 commits

dataiku/kiji-inspector

113

stars

40

commits

HTML

primary language

Sep 2, 2026

updated

README

Kiji Inspector: Mechanistic Interpretability for AI Agent Tool Selection

Kiji Inspector Workflow

CI Core CI Extras License: Apache 2.0 GitHub Stars GitHub Issues

Python Version Open In Colab

Responsible AI Contributions Welcome PRs Welcome

Status

This project is under heavy active development. We are planning to release a stable version of the framework in the coming weeks.

In the meantime, join our Slack Community

Learn more about our approach and early results:


What This Project Does

This project trains Sparse Autoencoders (SAEs) on the internal activations of an AI agent to understand why it selects specific tools. Given a user request like "Search our docs for API limits," the agent must choose between tools (e.g., internal_search vs web_search). We extract the model's hidden representations at the moment of that decision, decompose them into interpretable features using a JumpReLU SAE, and validate the resulting explanations through automated fuzzing and causal ablation experiments.

The key insight: train the SAE on raw activations (not difference vectors), then use contrastive pairs post-hoc to identify which learned features correspond to specific tool-selection decisions. This preserves the SAE's general feature dictionary while enabling targeted analysis of decision-relevant features.

Install

For loading and running pretrained SAEs:

pip install kiji-inspector

For the HuggingFace-based training and analysis extras (accelerate etc.):

pip install 'kiji-inspector[full]'

The vLLM extraction path is not covered by any extra: upstream vLLM wheels do not ship the hidden-states connector, so it requires the Docker image built from the repository Dockerfile, which compiles vLLM from the Davidnet/vllm fork.

Quick Start

from kiji_inspector import SAE

sae, feature_descriptions = SAE.from_pretrained(
    repo_id="575-lab/kiji-inspector-google-gemma-4-E4B-it",
    layer=30,
)

# encode/decode operate in the SAE's normalized space
features = sae.encode(sae.normalize_input(activations))
reconstruction = sae.denormalize_output(sae.decode(features))

normalize_input applies the exact transform the SAE was trained under — (x - mean_vec) / rms_scale, a single global mean vector and RMS constant computed over the training set. Raw activations must go through it or the JumpReLU thresholds are meaningless; denormalize_output inverts it when a reconstruction is written back into the model.

Training and data-generation entrypoints live under the package namespace:

python -m kiji_inspector.generate_pairs 1300
python -m kiji_inspector.pipeline --layers 10 20 30

vLLM hidden-state extraction (native connector)

Activation extraction uses vLLM's native extract_hidden_states speculator method together with the ExampleHiddenStatesConnector, which writes captured hidden states to safetensors files that the extractor loads and cleans up per request. This capability ships in the 575lab/kiji-inspector:dev image, which builds the Davidnet/vllm fork (branch hidden-states-inline-return-squashed); the public v0.19.0 wheel does not contain the connector.

Run all extraction, tests, and API checks inside that image:

docker pull 575lab/kiji-inspector:dev

# Smoke test the connector (Qwen3-8B):
samples/run_hidden_states_test.sh 575lab/kiji-inspector:dev

# Iterate on the checked-out source with a bind mount:
docker run --rm --gpus all \
  -v "$PWD:/workspace" \
  -v "${HF_CACHE:-$HOME/.cache/huggingface}:/root/.cache/huggingface" \
  -e HF_HOME=/root/.cache/huggingface \
  -e PYTHONPATH=/workspace/src \
  -w /workspace \
  575lab/kiji-inspector:dev \
  python -m pytest tests/test_activation_extractors.py

The historical patches/ directory (applied against a stock v0.19.0 wheel) predates the fork-based image and is retained for reference only; the current image bakes those changes into the fork and does not run apply-patch.sh. See patches/README_PATCH.md for the legacy workflow.

To run the pipeline against a Qwen3.6 subject model (a reasoning model), pass --no-thinking so the decision token sits at the final-answer position:

python -m kiji_inspector.pipeline --subject-model Qwen/Qwen3.6-35B-A3B --no-thinking

Using a locally downloaded model

--subject-model also accepts a local model directory (paths must start with /, ./, ../, or ~; anything else is treated as a hub ID). Inside Docker the directory must be volume-mounted, and you pass the container path:

docker run --rm --gpus all \
  -v "$PWD:/workspace" \
  -v /home/user/models:/models:ro \
  -v "${HF_CACHE:-$HOME/.cache/huggingface}:/root/.cache/huggingface" \
  -e HF_HOME=/root/.cache/huggingface \
  -e PYTHONPATH=/workspace/src \
  -w /workspace \
  575lab/kiji-inspector:dev \
  python -m kiji_inspector.pipeline \
    --subject-model /models/programs/downloaded-model

Note: SAE.from_pretrained(base_model=...) resolves SAE repos through an exact-match registry of hub IDs, so it won't match a local path — pass repo_id directly instead.


📓 Examples

Two end-to-end notebooks demonstrate the library. Both run on Colab — the first needs an L4 (24 GB), the second an A100 high-RAM runtime.

NotebookWhat it showsOpen
quickstart_colab.ipynbMinimal walkthrough: capture the decision-token activation of google/gemma-4-E4B-it via a forward pre-hook, load the layer-30 SAE with SAE.from_pretrained, normalize with sae.normalize_input, and describe the top features firing on a single prompt.Open In Colab
home_repair_colab.ipynbFull agent demo: a Nemotron-3-Nano-30B home repair advisor calls four tools across three appliance problems, the residual stream is captured at every decision point, and a trained JumpReLU SAE decomposes those activations into themed features. Includes the interactive index.html viewer served from Colab.Open In Colab

🤝 Contributing

We welcome contributions! Whether you're fixing a bug, improving documentation, or proposing a new feature, your help is appreciated.

Ways to Contribute

  • Report Bugs - Open an issue with steps to reproduce
  • Improve Docs - Documentation PRs are always welcome
  • Submit Features - Open an issue to discuss your idea before submitting a PR
  • Share Feedback - Start a discussion

Community


📄 License

Copyright (c) 2026 Dataiku SAS

This project is licensed under the Apache 2.0 License - see the LICENSE file for details.

Contributors

Davidnet

25 commits

hanneshapke

12 commits

carloshuisa

1 commits

claude

1 commits

Languages

HTML

89.7%

Python

7.8%

TeX

2.0%