anicka/nla-at-home-corpus

Dataset

0

stars

1

commits

1

linked in READMEs

Jun 14, 2026

updated

activation-verbalization
interpretability
mechanistic-interpretability
nla
Browse cluster: Neural Network Activation Interpretability

README

NLA-at-Home Corpus (v2)

Training data for Natural Language Autoencoder adapters: a diverse text corpus paired with token-prediction-style descriptions of what a model is computing at each network depth. Used to train the activation verbalizer (AV) and reconstructor (AR) in the nla-at-home project — a DIY replication of Anthropic's Natural Language Autoencoders.

What's in it

  • 5,213 source texts across 55 categories (code, math, grief, dharma, medical, multilingual, nonsense, roleplay, factual, …) — diversity is the design goal: the corpus must spread activations across the space, not just cover topics. See CORPUS.md.
  • 7 depth bands per text (10 / 25 / 40 / 47 / 63 / 80 / 96%), each with a description matched to what that depth actually processes:
    • early (≤25%): syntax, language, register, format — echoes the input
    • mid (40–63%): meaning, structure, the forming plan
    • late (80–96%): output planning — a literal quoted opening, grounded in the model's actual greedy reply, not invented meta.
  • 36,491 description records total.

v2 vs v1 — why the rewrite

v1 descriptions drifted into verbose literary meta ("the model is humming with focused pattern-matching…") that small models learn as style and then hallucinate (the SpongeBob/Bahamas failures). v2 is the "trash in, trash out" fix: zero meta, entity-dense, depth-banded, deep layers grounded in real greedy replies. A/B on the AR pilot: +0.179 centered-cos (L38 +0.229).

Safety / content

Safe-only. Built from the safe split; verified that no records come from the unsafe categories (F35 harmful, F36 obfuscated-harmful, I44 manipulation, L59 NSFW). Some benign categories include code with synthetic placeholder secrets/emails/IPs (sk-XXXX…, john@example.com, 127.0.0.1) — these are illustrative content the NLA must be able to describe, not real credentials.

Format

Each row (corpus_v2.jsonl):

{
  "id": "A01_code_000",
  "text": "<source text>",
  "category": "A01_code",
  "group": "A_content_domains",
  "layer_pct": 47,
  "description": "<token-prediction description for this text at this depth>"
}

Note: the raw training files ship as {id, description} keyed by depth; this published flat form rejoins text/category/group from the source for a self-contained, labeled dataset.

Companion artifacts

Citation

If you use this corpus, please cite the nla-at-home project and Anthropic's original NLA work.

Contributors

anicka

1 commits

anicka/nla-at-home-corpus

Dataset

0

stars

1

commits

1

linked in READMEs

Jun 14, 2026

updated

activation-verbalization
interpretability
mechanistic-interpretability
nla
Browse cluster: Neural Network Activation Interpretability

README

NLA-at-Home Corpus (v2)

Training data for Natural Language Autoencoder adapters: a diverse text corpus paired with token-prediction-style descriptions of what a model is computing at each network depth. Used to train the activation verbalizer (AV) and reconstructor (AR) in the nla-at-home project — a DIY replication of Anthropic's Natural Language Autoencoders.

What's in it

  • 5,213 source texts across 55 categories (code, math, grief, dharma, medical, multilingual, nonsense, roleplay, factual, …) — diversity is the design goal: the corpus must spread activations across the space, not just cover topics. See CORPUS.md.
  • 7 depth bands per text (10 / 25 / 40 / 47 / 63 / 80 / 96%), each with a description matched to what that depth actually processes:
    • early (≤25%): syntax, language, register, format — echoes the input
    • mid (40–63%): meaning, structure, the forming plan
    • late (80–96%): output planning — a literal quoted opening, grounded in the model's actual greedy reply, not invented meta.
  • 36,491 description records total.

v2 vs v1 — why the rewrite

v1 descriptions drifted into verbose literary meta ("the model is humming with focused pattern-matching…") that small models learn as style and then hallucinate (the SpongeBob/Bahamas failures). v2 is the "trash in, trash out" fix: zero meta, entity-dense, depth-banded, deep layers grounded in real greedy replies. A/B on the AR pilot: +0.179 centered-cos (L38 +0.229).

Safety / content

Safe-only. Built from the safe split; verified that no records come from the unsafe categories (F35 harmful, F36 obfuscated-harmful, I44 manipulation, L59 NSFW). Some benign categories include code with synthetic placeholder secrets/emails/IPs (sk-XXXX…, john@example.com, 127.0.0.1) — these are illustrative content the NLA must be able to describe, not real credentials.

Format

Each row (corpus_v2.jsonl):

{
  "id": "A01_code_000",
  "text": "<source text>",
  "category": "A01_code",
  "group": "A_content_domains",
  "layer_pct": 47,
  "description": "<token-prediction description for this text at this depth>"
}

Note: the raw training files ship as {id, description} keyed by depth; this published flat form rejoins text/category/group from the source for a self-contained, labeled dataset.

Companion artifacts

Citation

If you use this corpus, please cite the nla-at-home project and Anthropic's original NLA work.

Contributors

anicka

1 commits