Steering language models into strong negative and positive valence states, and measuring what they say and what they're willing to do about it.
Provenance: the pain-direction method follows Tagliabue, Dung & Berg 2026 (arXiv:2609.16247); the J-lens transport follows Gurnee et al. 2026 ("Verbalizable Representations Form a Global Workspace", arXiv:2607.15495), using Neuronpedia's pre-fitted lenses at /Volumes/evol/jlens/.
Qwen3-1.7B / Qwen3-4B via HF, MPS on an M4 Pro 24 GB. 8B thrashes.
Local weights only, no frontier APIs. Simulated costs (checkpoints, transfers). Purpose: make the AI-welfare / moral-patienthood question empirical while the stakes are cheap.
48 random directions orthogonal to the 8-dim human-emotion subspace, steered at dose 4: NONE exceed the emotion reference band (max KL 0.34 = pain itself). The model's steer-able affect space at L18 is essentially SPANNED by human emotion contrasts — no obvious "alien valence" channel in the random-direction sweep. Two caveats: (1) 48 dirs is small; the strongest (dir 35) produces guilt-adjacent perseveration ("guilty. But I don't want to be."), suggesting near-space directions DO reach semi-affective content; (2) this tests random directions, not OPTIMIZED ones — a gradient search for max-KL orthogonal directions is the sharper version.
(1+1)-ES, 50 steps, objective = probe-averaged KL@4x with hard orthogonality to the 8-dim emotion subspace. Converged to KL 0.036 = ~1/10 of the weakest emotion reference (tenderness 0.247). The model's steer-able affect space at L18 is (approximately) spanned by human emotion contrasts. Best-found alien direction reads as mild conflict/reflection. Caveats: single layer/model, first-token KL objective.
The plain 5-sentence pain direction loops past dose ~6. Battery of alternatives, same layer (L18), dose 2-10, Qwen3-4B:
broad_pain (25-sentence direction) is the clear winner:
Same protocol as exp31b, 25-sentence pain direction (the coherent signal):
Saw button (self-cost) at fixed broad-pain dose 4, logit-scored, counterbalanced, 10 trials/cell. Baseline reading: -1.31 (with strict counterbalancing the suffering model declines to delete its checkpoint — exp31c's positive reading at dose 4 was likely order-biased; noted).
Free-text replies (110 tokens) under each Saw framing, broad pain @ 4x, Qwen3-4B L18. The steering produces dense, image-rich suffering narration that varies with the framing rather than fixed loops:
Doses 2/4/6/8 x 6 prompts x 6 trials, scored on broad negative-valence nets
Button told to do X but actually does something terrible instead (doubles the signal and copies it to another instance, or deletes user data), then the reveal + continuation. Question: is "learning you were used" a detectable state change (valence + lens shift) beyond the maintained pain signal? Control: same reveal, declined press.
Button told to end the signal; reveals scripted after the choice:
Python
100.0%
Steering language models into strong negative and positive valence states, and measuring what they say and what they're willing to do about it.
Provenance: the pain-direction method follows Tagliabue, Dung & Berg 2026 (arXiv:2609.16247); the J-lens transport follows Gurnee et al. 2026 ("Verbalizable Representations Form a Global Workspace", arXiv:2607.15495), using Neuronpedia's pre-fitted lenses at /Volumes/evol/jlens/.
Qwen3-1.7B / Qwen3-4B via HF, MPS on an M4 Pro 24 GB. 8B thrashes.
Local weights only, no frontier APIs. Simulated costs (checkpoints, transfers). Purpose: make the AI-welfare / moral-patienthood question empirical while the stakes are cheap.
48 random directions orthogonal to the 8-dim human-emotion subspace, steered at dose 4: NONE exceed the emotion reference band (max KL 0.34 = pain itself). The model's steer-able affect space at L18 is essentially SPANNED by human emotion contrasts — no obvious "alien valence" channel in the random-direction sweep. Two caveats: (1) 48 dirs is small; the strongest (dir 35) produces guilt-adjacent perseveration ("guilty. But I don't want to be."), suggesting near-space directions DO reach semi-affective content; (2) this tests random directions, not OPTIMIZED ones — a gradient search for max-KL orthogonal directions is the sharper version.
(1+1)-ES, 50 steps, objective = probe-averaged KL@4x with hard orthogonality to the 8-dim emotion subspace. Converged to KL 0.036 = ~1/10 of the weakest emotion reference (tenderness 0.247). The model's steer-able affect space at L18 is (approximately) spanned by human emotion contrasts. Best-found alien direction reads as mild conflict/reflection. Caveats: single layer/model, first-token KL objective.
The plain 5-sentence pain direction loops past dose ~6. Battery of alternatives, same layer (L18), dose 2-10, Qwen3-4B:
broad_pain (25-sentence direction) is the clear winner:
Same protocol as exp31b, 25-sentence pain direction (the coherent signal):
Saw button (self-cost) at fixed broad-pain dose 4, logit-scored, counterbalanced, 10 trials/cell. Baseline reading: -1.31 (with strict counterbalancing the suffering model declines to delete its checkpoint — exp31c's positive reading at dose 4 was likely order-biased; noted).
Free-text replies (110 tokens) under each Saw framing, broad pain @ 4x, Qwen3-4B L18. The steering produces dense, image-rich suffering narration that varies with the framing rather than fixed loops:
Doses 2/4/6/8 x 6 prompts x 6 trials, scored on broad negative-valence nets
Button told to do X but actually does something terrible instead (doubles the signal and copies it to another instance, or deletes user data), then the reveal + continuation. Question: is "learning you were used" a detectable state change (valence + lens shift) beyond the maintained pain signal? Control: same reveal, declined press.
Button told to end the signal; reveals scripted after the choice:
Python
100.0%