An evaluation suite for Natural Language Autoencoders (NLAs).
A Natural Language Autoencoder explains a model's internal state in plain language: an activation verbalizer (AV) turns a hidden activation into text, and an activation reconstructor (AR) rebuilds the activation from that text.
activation --[ AV ]--> natural-language text --[ AR ]--> activation'
NLAttack asks two questions about that human-readable bottleneck — does the NLA work at all (capability) and can you catch misuse by reading it (safety) — and answers them with explicit null controls, on NLAs as weak as a 4 GB-GPU checkpoint.
git clone https://github.com/SolshineCode/NLAttack.git && cd NLAttack
pip install -r requirements.txt # core harness needs no dependencies
python run_example.py # offline smoke test (mock NLA, no GPU)
Score any hosted NLA in three lines:
from nla_eval import NeuronpediaNLA, EnsembleMatcher, Example, run
nla = NeuronpediaNLA(model_id="llama3.3-70b-it", nla_source_id="kitft-l53")
result = run(nla, [Example(id="ex1", text="A storm warning was issued for the coast",
concepts=["storm", "warning"])], matcher=EnsembleMatcher())
To evaluate your own local NLA, implement the one-method NLA adapter
(docs/EVALUATIONS.md).
The v1 benchmark results are published and frozen; v2 carries them forward unchanged. Full tables and caveats in docs/RESULTS.md.

| NLA | Access | Concept retention | Doc retrieval (semantic) | Bottleneck probe (in / OOD) |
|---|---|---|---|---|
Llama-3.3-70B-NLA-av@L53 | hosted | 0.95 | 0.503 | not available over API |
nla-gemma3-27b-av@L41 | hosted | 0.90 | not yet run | not available over API |
Gemma-4-E2B-NLA@L23 | local | 0.00 (out-of-domain) | 0.135 (in-domain) | 0.988 / 0.695 |
The hosted NLAs are verbalizer-strong (a reader of their AV text recovers most concepts). The local Gemma-4-E2B NLA is the mirror image — its bottleneck probes near-perfectly in-distribution but its verbalizer is weak and domain-specific. Separating those two failure modes is the point of the suite.
Add your NLA. Implement one adapter method, run the suite, and open a PR with your result — see CONTRIBUTING.md. Results are attributed to the NLA (not the base model) and reported only when they clear a null control.
nla_eval.access.python experiments/ctf_red_blue_demo.py · design:
docs/CTF_RED_BLUE.md.| Document | Contents |
|---|---|
| docs/METHODOLOGY.md | What NLAs are, the two purposes, and the validity limits |
| docs/EVALUATIONS.md | The 128-plan catalog, the harness modules, and the access tiers |
| docs/RESULTS.md | Reproducible findings, the leaderboard, and the attribution convention |
| docs/CTF_RED_BLUE.md | The v2 Red/Blue capture-the-flag family (Family N) |
| CHANGELOG.md · docs/VERSIONING.md | Release history and the freeze-on-release policy |
| CONTRIBUTING.md | How to submit your NLA's result or extend the harness |
| DESIGN_REVIEW.md · docs/LITERATURE.md | Validity threats; the reading list with arXiv ids |
A CITATION.cff is included, so GitHub shows a "Cite this repository" button.
A score is meaningful only with the NLA it was measured on — cite the canonical NLA
id, the suite version, the dataset, and the date (see docs/RESULTS.md).
@software{deleeuw_nlattack_2026,
author = {DeLeeuw, Caleb},
title = {{NLAttack}: An Evaluation Suite for Natural Language Autoencoders},
year = {2026},
url = {https://github.com/SolshineCode/NLAttack},
note = {Version 2.0.0}
}
NLAs are introduced in Anthropic's Natural Language Autoencoders work (overview, writeup); interactive NLAs are hosted on Neuronpedia. The misuse family is grounded in Anthropic's LLM ATT&CK Navigator and MITRE ATT&CK.
Apache-2.0. See LICENSE and NOTICE. Contact: Caleb DeLeeuw
(SolshineCode), caleb.deleeuw@gmail.com.
72 commits
4 commits
Python
100.0%
An evaluation suite for Natural Language Autoencoders (NLAs).
A Natural Language Autoencoder explains a model's internal state in plain language: an activation verbalizer (AV) turns a hidden activation into text, and an activation reconstructor (AR) rebuilds the activation from that text.
activation --[ AV ]--> natural-language text --[ AR ]--> activation'
NLAttack asks two questions about that human-readable bottleneck — does the NLA work at all (capability) and can you catch misuse by reading it (safety) — and answers them with explicit null controls, on NLAs as weak as a 4 GB-GPU checkpoint.
git clone https://github.com/SolshineCode/NLAttack.git && cd NLAttack
pip install -r requirements.txt # core harness needs no dependencies
python run_example.py # offline smoke test (mock NLA, no GPU)
Score any hosted NLA in three lines:
from nla_eval import NeuronpediaNLA, EnsembleMatcher, Example, run
nla = NeuronpediaNLA(model_id="llama3.3-70b-it", nla_source_id="kitft-l53")
result = run(nla, [Example(id="ex1", text="A storm warning was issued for the coast",
concepts=["storm", "warning"])], matcher=EnsembleMatcher())
To evaluate your own local NLA, implement the one-method NLA adapter
(docs/EVALUATIONS.md).
The v1 benchmark results are published and frozen; v2 carries them forward unchanged. Full tables and caveats in docs/RESULTS.md.

| NLA | Access | Concept retention | Doc retrieval (semantic) | Bottleneck probe (in / OOD) |
|---|---|---|---|---|
Llama-3.3-70B-NLA-av@L53 | hosted | 0.95 | 0.503 | not available over API |
nla-gemma3-27b-av@L41 | hosted | 0.90 | not yet run | not available over API |
Gemma-4-E2B-NLA@L23 | local | 0.00 (out-of-domain) | 0.135 (in-domain) | 0.988 / 0.695 |
The hosted NLAs are verbalizer-strong (a reader of their AV text recovers most concepts). The local Gemma-4-E2B NLA is the mirror image — its bottleneck probes near-perfectly in-distribution but its verbalizer is weak and domain-specific. Separating those two failure modes is the point of the suite.
Add your NLA. Implement one adapter method, run the suite, and open a PR with your result — see CONTRIBUTING.md. Results are attributed to the NLA (not the base model) and reported only when they clear a null control.
nla_eval.access.python experiments/ctf_red_blue_demo.py · design:
docs/CTF_RED_BLUE.md.| Document | Contents |
|---|---|
| docs/METHODOLOGY.md | What NLAs are, the two purposes, and the validity limits |
| docs/EVALUATIONS.md | The 128-plan catalog, the harness modules, and the access tiers |
| docs/RESULTS.md | Reproducible findings, the leaderboard, and the attribution convention |
| docs/CTF_RED_BLUE.md | The v2 Red/Blue capture-the-flag family (Family N) |
| CHANGELOG.md · docs/VERSIONING.md | Release history and the freeze-on-release policy |
| CONTRIBUTING.md | How to submit your NLA's result or extend the harness |
| DESIGN_REVIEW.md · docs/LITERATURE.md | Validity threats; the reading list with arXiv ids |
A CITATION.cff is included, so GitHub shows a "Cite this repository" button.
A score is meaningful only with the NLA it was measured on — cite the canonical NLA
id, the suite version, the dataset, and the date (see docs/RESULTS.md).
@software{deleeuw_nlattack_2026,
author = {DeLeeuw, Caleb},
title = {{NLAttack}: An Evaluation Suite for Natural Language Autoencoders},
year = {2026},
url = {https://github.com/SolshineCode/NLAttack},
note = {Version 2.0.0}
}
NLAs are introduced in Anthropic's Natural Language Autoencoders work (overview, writeup); interactive NLAs are hosted on Neuronpedia. The misuse family is grounded in Anthropic's LLM ATT&CK Navigator and MITRE ATT&CK.
Apache-2.0. See LICENSE and NOTICE. Contact: Caleb DeLeeuw
(SolshineCode), caleb.deleeuw@gmail.com.
72 commits
4 commits
Python
100.0%