LoRA adapter that turns a residual-stream activation vector from
Qwen2.5-7B-Instruct into a short natural-language rendering of what the model
was processing at that layer. A single adapter handles all 28 layers: the
prompt carries the extraction depth (from depth {N}%), so one set of weights
spans early-syntax bands through near-output bands.
Part of the nla-at-home project — a DIY replication of Anthropic's Natural Language Autoencoders using open-weight models.
This is the SFT stage. For the sharper, GRPO-refined variant see nla-qwen2.5-7b-universal-av-grpo; the companion reconstructor is nla-qwen2.5-7b-universal-ar.
| parameter | value |
|---|---|
| injection token | ㈎ (U+320E), id 149705 |
| injection method | normalize the vector's L2 norm TO 150.0 (never multiply BY 150) |
| prompt | chat template around the AV prompt, add_generation_prompt=True, encoded with add_special_tokens=False |
| depth | from depth {N}% where N = nearest of 4,10,17,25,32,40,47,55,63,71,80,90,96 to layer*100/28 |
# activation: torch tensor [3584] read at block L's output, last prompt token
from nla_lib import make_av_prompt, normalize_activation, nearest_depth_pct
content = make_av_prompt(nearest_depth_pct(L, 28), "㈎")
chat = tok.apply_chat_template([{"role": "user", "content": content}],
tokenize=False, add_generation_prompt=True)
ids = tok.encode(chat, add_special_tokens=False)
emb = model.get_input_embeddings()(torch.tensor([ids]).to(dev)).clone()
emb[0, ids.index(149705), :] = normalize_activation(activation).to(emb.dtype)
out = model.generate(inputs_embeds=emb, attention_mask=torch.ones(1, len(ids)).to(dev),
max_new_tokens=120, do_sample=False)
Full working scripts: scripts/eval_roundtrip_universal.py,
scripts/brain_in_jar.py in the repo.
LoRA SFT over ~68k terse, content-anchored descriptions (_twin_clean set;
5 epochs, lr 8e-6) of activations from the project's 5.2k-text public corpus
(corpus v2),
extracted at the last prompt token across all 28 layers. 528 texts held out
(val_text_ids.json ships with the adapter).
Descriptions are scored by whether the companion AR can rebuild the original activation from the text alone (centered cosine, per-layer mean removed; 286-text clean holdout never seen by AV, AR, or the GRPO curriculum):
| layer (depth) | centered cosine |
|---|---|
| L4 (17%) | 0.473 |
| L10 (32%) | 0.380 |
| L17 (63%) | 0.550 |
| L20 (71%) | 0.615 |
| L24 (90%) | 0.596 |
| L27 (96%) | 0.435 |
| mean | 0.508 |
The GRPO variant improves on this — see its card for the side-by-side.
Trained on the public, safety-filtered corpus split: descriptions of harmful / NSFW / manipulative activations are out of distribution and will be weaker (see the repo's "Safety & scope" note). Entity fidelity degrades for rare proper nouns; gist and structure are the reliable layer.
2 commits
LoRA adapter that turns a residual-stream activation vector from
Qwen2.5-7B-Instruct into a short natural-language rendering of what the model
was processing at that layer. A single adapter handles all 28 layers: the
prompt carries the extraction depth (from depth {N}%), so one set of weights
spans early-syntax bands through near-output bands.
Part of the nla-at-home project — a DIY replication of Anthropic's Natural Language Autoencoders using open-weight models.
This is the SFT stage. For the sharper, GRPO-refined variant see nla-qwen2.5-7b-universal-av-grpo; the companion reconstructor is nla-qwen2.5-7b-universal-ar.
| parameter | value |
|---|---|
| injection token | ㈎ (U+320E), id 149705 |
| injection method | normalize the vector's L2 norm TO 150.0 (never multiply BY 150) |
| prompt | chat template around the AV prompt, add_generation_prompt=True, encoded with add_special_tokens=False |
| depth | from depth {N}% where N = nearest of 4,10,17,25,32,40,47,55,63,71,80,90,96 to layer*100/28 |
# activation: torch tensor [3584] read at block L's output, last prompt token
from nla_lib import make_av_prompt, normalize_activation, nearest_depth_pct
content = make_av_prompt(nearest_depth_pct(L, 28), "㈎")
chat = tok.apply_chat_template([{"role": "user", "content": content}],
tokenize=False, add_generation_prompt=True)
ids = tok.encode(chat, add_special_tokens=False)
emb = model.get_input_embeddings()(torch.tensor([ids]).to(dev)).clone()
emb[0, ids.index(149705), :] = normalize_activation(activation).to(emb.dtype)
out = model.generate(inputs_embeds=emb, attention_mask=torch.ones(1, len(ids)).to(dev),
max_new_tokens=120, do_sample=False)
Full working scripts: scripts/eval_roundtrip_universal.py,
scripts/brain_in_jar.py in the repo.
LoRA SFT over ~68k terse, content-anchored descriptions (_twin_clean set;
5 epochs, lr 8e-6) of activations from the project's 5.2k-text public corpus
(corpus v2),
extracted at the last prompt token across all 28 layers. 528 texts held out
(val_text_ids.json ships with the adapter).
Descriptions are scored by whether the companion AR can rebuild the original activation from the text alone (centered cosine, per-layer mean removed; 286-text clean holdout never seen by AV, AR, or the GRPO curriculum):
| layer (depth) | centered cosine |
|---|---|
| L4 (17%) | 0.473 |
| L10 (32%) | 0.380 |
| L17 (63%) | 0.550 |
| L20 (71%) | 0.615 |
| L24 (90%) | 0.596 |
| L27 (96%) | 0.435 |
| mean | 0.508 |
The GRPO variant improves on this — see its card for the side-by-side.
Trained on the public, safety-filtered corpus split: descriptions of harmful / NSFW / manipulative activations are out of distribution and will be weaker (see the repo's "Safety & scope" note). Entity fidelity degrades for rare proper nouns; gist and structure are the reliable layer.
2 commits