Does reasoning work by simulating a 'society of thought'? A replication of arXiv:2601.10825 finding its causal claim is a Countdown artifact — the same feature that gives +10 pts on Countdown costs −22 on MATH-Hard.
Python
1
142 commits
updated Sep 21, 2026
A replication and stress-test of "Reasoning Models Generate Societies of Thought" (Kim, Lai, Scherrer, Agüera y Arcas & Evans, Google / UChicago / Santa Fe Institute, arXiv:2601.10825, Jan 2026).
The paper ships no code and no data. Everything here is rebuilt from public artifacts.
When a reasoning model like DeepSeek-R1 "thinks", its chain of thought reads like a conversation — it asks itself questions, changes its mind, argues, and reconciles. Instruction-tuned models (the same base model, different post-training) don't do this; they produce one-sided monologue.
The paper says this dialogue is the mechanism: reasoning works because the model simulates diverse internal voices that debate. Its evidence comes in two parts.
Part A — descriptive (we don't dispute this). Across 8,262 problems, R1's traces contain far more question-asking, perspective-shifting, and disagreement than V3's, even controlling for how long the traces are. The effect is large and holds at every model size.
Part B — causal (this is what we test). They find a single feature inside the model that fires on conversational surprise markers — the "Oh!" feature — and turn it up. Accuracy on a puzzle task doubles, 27.1% → 54.8%. They read this as: induce more society-of-thought, get more reasoning.
The gap: Part B was only ever run on one task — Countdown, an arithmetic puzzle. The benchmarks that make it look like a general result (GPQA, MATH) were only observed, never intervened on.
Turn the "Oh!" feature up and down inside DeepSeek-R1-Distill-Llama-8B and measure accuracy. Three things the paper didn't do:
Result: the effect is a Countdown artifact.
| Countdown (paper's task) | MATH-Hard (paper's benchmark) | |
|---|---|---|
| baseline | 24.0% (paper: 27.1% — we reproduce) | 62.0% |
| the paper's feature, turned up | +10.0 pts | −22.0 pts |
The same feature, at the same strength, helps on Countdown and wrecks MATH. Turn it up further and the model degenerates into literal babble ("cyclochoh! Wait, no, wait, no, wait...").
And the paper's own mechanism story fails: steering does make the traces more dialogic (self-interruptions +36%, contradictions and questions up) — and accuracy falls anyway. You can induce the society of thought and get a dumber model.
Why? Countdown is a search puzzle — combine 3–4 numbers to hit a target. The way to win is to try more candidate expressions. On a task like that, "poke the model into trying more things" and "make the model reason better" are indistinguishable. On MATH you can't brute-force your way to the answer, and the poke just makes it ramble.
→ Full findings, with all the numbers
The other half of the paper, and the only part that touches real RL rather than a pre-trained model. Take a base model (Qwen-2.5-3B), fine-tune it two ways over identical problems with identical correct answers:
Then run identical RL on both. The paper says the dialogue-primed model learns faster. If true, the social account survives our steering result — the mechanism would be real but mislocated: it emerges in training rather than living in a steerable feature.
We also do what the paper didn't: ≥3 random seeds per arm. Their headline is an early-training gap, which is exactly where seed noise is largest, and they appear to report single runs.
Huot, Kaisers & Lapata (2026) reach the same dissociation from the opposite direction — between models rather than inside one. Routing over a society of models, they argue, is judged almost entirely on accuracy, and that is not enough: "high task accuracy is compatible with ... a redundant society." They conclude that "accuracy and meaningfulness can sharply diverge."
Mirror images:
| they show | we show | |
|---|---|---|
| high accuracy with a fake society | a real society with worse accuracy | |
| level | between models (routing) | inside one model (steering) |
Both break the inference the paper depends on. It measures conversational behaviour and accuracy together, on a task where both rise, and concludes the first causes the second. You can have the accuracy without the society — and, as we show, the society without the accuracy.
Their Hierarchic Social Entropy is also a judge-free alternative to the paper's "perspective diversity" measure (an LLM-judge that infers the personas it then scores). Applying it would test whether R1's internal voices are genuinely differentiated, or a redundant society wearing dialogic clothes. We haven't done that; it's the obvious next step.
Six silent failures, none of which raised an error, each of which would have produced a confident wrong answer. They're documented because anyone replicating this paper will hit them:
resid_pre; it's actually
resid_post). Replicate from the published metadata and you steer the wrong layer.\boxed{}, not the <answer> tag the prompt demands.
Grading only the tag scored 74% of correct traces as unparseable and put our baseline
at 5.5% instead of 24% — which looks exactly like "the paper doesn't reproduce."truncated measured the padded batch, not the sequence — 96% "truncation" that was
really 12%.The lesson, and the reason for the test suite: an unparseable answer must score wrong, never be dropped. Otherwise a degenerating model "improves" by shrinking its own denominator.
./scripts/setup.sh # uv venv + torch + deps
./scripts/run_stages.sh hook # REQUIRED: resolve the hook point by reconstruction
./scripts/run_stages.sh calibrate # REQUIRED: measure activation scales in OUR units
./scripts/run_stages.sh control # the Countdown dose-response
./scripts/run_stages.sh main # GPQA + MATH-Hard with matched controls
python -m rl.generate_sft # dialogue/monologue SFT data (verified, matched)
python -m rl.train_grpo --arm baseline --seed 42 # Gate 2: does RL learn at all?
pytest tests/ # 53 tests
Stages are gates. If the Countdown control doesn't reproduce ~24%, the harness is wrong and nothing downstream means anything — fix that first.
Needs one GPU with ≥20GB. Don't run this on a unified-memory box (DGX Spark): a GPU over-allocation there starves the OS and takes the whole machine down, rather than just killing your job. We learned that twice.
sot/ steering: SAE loading, the hook, calibration, grading, the sweep
rl/ the RL half: SFT data generation, GRPO training, provenance checking
tests/ 53 tests — the graders, the SAE maths, the hook, the arm-matching
scripts/ staged runners + RunPod provisioning
results/ raw traces (5,664 attempts) and findings
deepseek-ai/DeepSeek-R1-Distill-Llama-8B (MIT)OpenMOSS-Team/Llama-Scope-R1-Distill (Apache-2.0) — the paper's is the
800M-Slimpajama-0-OpenR1-Math-220k/L15R subdirectoryfingertap/GPQA-Diamond (the official one is gated) · MATH: lighteval/MATH-Hard
· Countdown: Jiayi-Pan/Countdown-Tasks-3to4DeepSeek-R1-Distill-Llama-8B was never RL'd. It's Llama-3.1-8B fine-tuned to imitate
R1's outputs. It's the only reasoning model with a public SAE, so it's what the paper used
and what we used.
Which means the field's mechanistic understanding of reasoning models currently rests on a model that is an impersonation of one. Training an SAE on a genuinely RL'd reasoner (QwQ-32B) is the obvious next step, and nobody has done it.
Part of one program — a controls-first, open-weights attempt to make the 2025–26 interpretability claims checkable:
| repo | what |
|---|---|
| spinning-up-in-mech-interp | the curriculum — 8 rungs, 6 runnable on a laptop, each ending in its own null |
| jacobian-lens | the research — OLMo post-training ladder, metacognition, the Consciousness-Indicator Scorecard |
| tri-lens | do three instruments agree about the same activation? |
| societies-of-thought | the adversarial replication — rebuild a no-code/no-data paper, then try to break it |
| controls-and-trajectories | the published datasets — nulls and developmental trajectories |
This repo is the program's stress test: the paper ships nothing, so everything is rebuilt from
public artifacts and every claim is run against the control that could kill it. Three results have
been retracted or corrected here after exactly that treatment — see
results/steering/RECHECK_length_and_filtering.md and results/qwq/FINDINGS.md.
Direct link to tri-lens. Our open question is an instrument-agreement question: our diversity nulls are nulls for Hierarchic Social Entropy, and HSE has never been run against the paper's own LLM judge on shared inputs. If two instruments disagree about the same traces, that belongs alongside the three-instrument study.
Datasets: sot-priming-traces-dialogue-monologue (the corpus the paper does not ship).
Python
92.1%
Shell
7.9%
Does reasoning work by simulating a 'society of thought'? A replication of arXiv:2601.10825 finding its causal claim is a Countdown artifact — the same feature that gives +10 pts on Countdown costs −22 on MATH-Hard.
Python
1
142 commits
updated Sep 21, 2026
A replication and stress-test of "Reasoning Models Generate Societies of Thought" (Kim, Lai, Scherrer, Agüera y Arcas & Evans, Google / UChicago / Santa Fe Institute, arXiv:2601.10825, Jan 2026).
The paper ships no code and no data. Everything here is rebuilt from public artifacts.
When a reasoning model like DeepSeek-R1 "thinks", its chain of thought reads like a conversation — it asks itself questions, changes its mind, argues, and reconciles. Instruction-tuned models (the same base model, different post-training) don't do this; they produce one-sided monologue.
The paper says this dialogue is the mechanism: reasoning works because the model simulates diverse internal voices that debate. Its evidence comes in two parts.
Part A — descriptive (we don't dispute this). Across 8,262 problems, R1's traces contain far more question-asking, perspective-shifting, and disagreement than V3's, even controlling for how long the traces are. The effect is large and holds at every model size.
Part B — causal (this is what we test). They find a single feature inside the model that fires on conversational surprise markers — the "Oh!" feature — and turn it up. Accuracy on a puzzle task doubles, 27.1% → 54.8%. They read this as: induce more society-of-thought, get more reasoning.
The gap: Part B was only ever run on one task — Countdown, an arithmetic puzzle. The benchmarks that make it look like a general result (GPQA, MATH) were only observed, never intervened on.
Turn the "Oh!" feature up and down inside DeepSeek-R1-Distill-Llama-8B and measure accuracy. Three things the paper didn't do:
Result: the effect is a Countdown artifact.
| Countdown (paper's task) | MATH-Hard (paper's benchmark) | |
|---|---|---|
| baseline | 24.0% (paper: 27.1% — we reproduce) | 62.0% |
| the paper's feature, turned up | +10.0 pts | −22.0 pts |
The same feature, at the same strength, helps on Countdown and wrecks MATH. Turn it up further and the model degenerates into literal babble ("cyclochoh! Wait, no, wait, no, wait...").
And the paper's own mechanism story fails: steering does make the traces more dialogic (self-interruptions +36%, contradictions and questions up) — and accuracy falls anyway. You can induce the society of thought and get a dumber model.
Why? Countdown is a search puzzle — combine 3–4 numbers to hit a target. The way to win is to try more candidate expressions. On a task like that, "poke the model into trying more things" and "make the model reason better" are indistinguishable. On MATH you can't brute-force your way to the answer, and the poke just makes it ramble.
→ Full findings, with all the numbers
The other half of the paper, and the only part that touches real RL rather than a pre-trained model. Take a base model (Qwen-2.5-3B), fine-tune it two ways over identical problems with identical correct answers:
Then run identical RL on both. The paper says the dialogue-primed model learns faster. If true, the social account survives our steering result — the mechanism would be real but mislocated: it emerges in training rather than living in a steerable feature.
We also do what the paper didn't: ≥3 random seeds per arm. Their headline is an early-training gap, which is exactly where seed noise is largest, and they appear to report single runs.
Huot, Kaisers & Lapata (2026) reach the same dissociation from the opposite direction — between models rather than inside one. Routing over a society of models, they argue, is judged almost entirely on accuracy, and that is not enough: "high task accuracy is compatible with ... a redundant society." They conclude that "accuracy and meaningfulness can sharply diverge."
Mirror images:
| they show | we show | |
|---|---|---|
| high accuracy with a fake society | a real society with worse accuracy | |
| level | between models (routing) | inside one model (steering) |
Both break the inference the paper depends on. It measures conversational behaviour and accuracy together, on a task where both rise, and concludes the first causes the second. You can have the accuracy without the society — and, as we show, the society without the accuracy.
Their Hierarchic Social Entropy is also a judge-free alternative to the paper's "perspective diversity" measure (an LLM-judge that infers the personas it then scores). Applying it would test whether R1's internal voices are genuinely differentiated, or a redundant society wearing dialogic clothes. We haven't done that; it's the obvious next step.
Six silent failures, none of which raised an error, each of which would have produced a confident wrong answer. They're documented because anyone replicating this paper will hit them:
resid_pre; it's actually
resid_post). Replicate from the published metadata and you steer the wrong layer.\boxed{}, not the <answer> tag the prompt demands.
Grading only the tag scored 74% of correct traces as unparseable and put our baseline
at 5.5% instead of 24% — which looks exactly like "the paper doesn't reproduce."truncated measured the padded batch, not the sequence — 96% "truncation" that was
really 12%.The lesson, and the reason for the test suite: an unparseable answer must score wrong, never be dropped. Otherwise a degenerating model "improves" by shrinking its own denominator.
./scripts/setup.sh # uv venv + torch + deps
./scripts/run_stages.sh hook # REQUIRED: resolve the hook point by reconstruction
./scripts/run_stages.sh calibrate # REQUIRED: measure activation scales in OUR units
./scripts/run_stages.sh control # the Countdown dose-response
./scripts/run_stages.sh main # GPQA + MATH-Hard with matched controls
python -m rl.generate_sft # dialogue/monologue SFT data (verified, matched)
python -m rl.train_grpo --arm baseline --seed 42 # Gate 2: does RL learn at all?
pytest tests/ # 53 tests
Stages are gates. If the Countdown control doesn't reproduce ~24%, the harness is wrong and nothing downstream means anything — fix that first.
Needs one GPU with ≥20GB. Don't run this on a unified-memory box (DGX Spark): a GPU over-allocation there starves the OS and takes the whole machine down, rather than just killing your job. We learned that twice.
sot/ steering: SAE loading, the hook, calibration, grading, the sweep
rl/ the RL half: SFT data generation, GRPO training, provenance checking
tests/ 53 tests — the graders, the SAE maths, the hook, the arm-matching
scripts/ staged runners + RunPod provisioning
results/ raw traces (5,664 attempts) and findings
deepseek-ai/DeepSeek-R1-Distill-Llama-8B (MIT)OpenMOSS-Team/Llama-Scope-R1-Distill (Apache-2.0) — the paper's is the
800M-Slimpajama-0-OpenR1-Math-220k/L15R subdirectoryfingertap/GPQA-Diamond (the official one is gated) · MATH: lighteval/MATH-Hard
· Countdown: Jiayi-Pan/Countdown-Tasks-3to4DeepSeek-R1-Distill-Llama-8B was never RL'd. It's Llama-3.1-8B fine-tuned to imitate
R1's outputs. It's the only reasoning model with a public SAE, so it's what the paper used
and what we used.
Which means the field's mechanistic understanding of reasoning models currently rests on a model that is an impersonation of one. Training an SAE on a genuinely RL'd reasoner (QwQ-32B) is the obvious next step, and nobody has done it.
Part of one program — a controls-first, open-weights attempt to make the 2025–26 interpretability claims checkable:
| repo | what |
|---|---|
| spinning-up-in-mech-interp | the curriculum — 8 rungs, 6 runnable on a laptop, each ending in its own null |
| jacobian-lens | the research — OLMo post-training ladder, metacognition, the Consciousness-Indicator Scorecard |
| tri-lens | do three instruments agree about the same activation? |
| societies-of-thought | the adversarial replication — rebuild a no-code/no-data paper, then try to break it |
| controls-and-trajectories | the published datasets — nulls and developmental trajectories |
This repo is the program's stress test: the paper ships nothing, so everything is rebuilt from
public artifacts and every claim is run against the control that could kill it. Three results have
been retracted or corrected here after exactly that treatment — see
results/steering/RECHECK_length_and_filtering.md and results/qwq/FINDINGS.md.
Direct link to tri-lens. Our open question is an instrument-agreement question: our diversity nulls are nulls for Hierarchic Social Entropy, and HSE has never been run against the paper's own LLM judge on shared inputs. If two instruments disagree about the same traces, that belongs alongside the three-instrument study.
Datasets: sot-priming-traces-dialogue-monologue (the corpus the paper does not ship).
Python
92.1%
Shell
7.9%