
Tiny Theory of Mind is our first attempt at evaluating theory of mind capabilities in small language models. The benchmark covers a wide variety of ToM topics, ranging in difficulties that, for humans, would be appropriate for Pre-K through 6th grade.
The benchmark is designed primarily for base-model continuation log-likelihood scoring. It does not require instruction following, chain-of-thought, or generated explanations. Random-choice accuracy is 25%.
| Property | Value |
|---|---|
| Examples | 2,000 |
| Answer choices per example | 4 |
| Random-choice accuracy | 25% |
| ToM Constructs | 40 (50 samples each) |
| Settings | 937 |
| Vocabulary Themes | 514 |
| Grade bands | Grades PreK-6 |
| Difficulty levels | Easy and Medium (1000 samples each) |
| Language | English |
| Data source | Synthetic |
| Dataset file | tiny_theory_of_mind_2000.jsonl |
| Rank | Model | Score |
|---|---|---|
| 1 | HuggingFaceTB/SmolLM-135M | 47.20% |
| 2 | BananaMind/BananaMind-2-Pro | 46.65% |
| 3 | HuggingFaceTB/SmolLM2-135M | 46.55% |
| 4 | AxiomicLabs/GPT-X2-125M | 44.55% |
| 5 | AxiomicLabs/GPT-X2.5-135M | 44.20% |
| 6 | SupraLabs/Supra2-100M-Base | 43.00% |
| 7 | AxiomicLabs/GPT-X-125M | 42.75% |
| 8 | facebook/opt-125m | 42.75% |
| 9 | SupraLabs/Supra-50M-Base | 41.90% |
| 10 | BananaMind/BananaMind-2-Medium | 41.20% |
| 11 | BananaMind/BananaMind-2-Medium-Chat | 40.95% |
| 12 | TobiasLogic/Museko-125M | 40.90% |
| 13 | openai-community/gpt2 | 40.75% |
| 14 | SupraLabs/Supra1.5-50M-Base-exp | 39.95% |
| 15 | finnianx/Gros-Michel-90m-Base-v2 | 39.80% |
| 16 | BananaMind/BananaMind-2-Mini | 38.55% |
| 17 | veyra-ai/Veyra2-Apricot-50M-Base | 38.20% |
| 18 | BananaMind/BananaMind-2.1-Unified | 38.15% |
| 19 | GODELEV/Archaea-74M | 37.15% |
| 20 | BananaMind/BananaMind-2-Nano | 37.10% |
| 21 | finnianx/michel-micro | 36.40% |
| 22 | fromziro/Er-Large-30M | 35.50% |
| 23 | EleutherAI/pythia-31m | 35.20% |
| 24 | BananaMind/BananaMind-2-MoE | 34.05% |
| 25 | BananaMind/BananaMind-2-SLMoE | 34.00% |
| 26 | AxiomicLabs/GPT-S2-5M | 32.65% |
| 27 | User01110/CMA-8M | 32.55% |
| 28 | AxiomicLabs/GPT-S-5M | 31.60% |
| 29 | BananaMind/BananaMind-2-Micro | 31.60% |
| 30 | User01110/CMA-1M-Mini | 31.40% |
| 31 | fromziro/Syn-2.6M | 30.30% |
| 32 | AtomixLabs/Photon-2.0-1M | 30.05% |
| 33 | AxiomicLabs/GPT-S-1.4M | 29.95% |
| 34 | BananaMind/BananaMind-2.1-Pico-Preview | 29.20% |
| 35 | ThingAI/Quark-50m-v2 | 29.10% |
| 36 | harley-ml/dillion-1.2M | 28.30% |
Every item contains an unfinished Theory of Mind scenario in ctx, four possible continuations in endings, and a zero-based correct-answer index in label.
{
"ind": 1910,
"activity_label": "tiny_theory_of_mind_natural_continuation::first_order_false_belief::order_1::medium",
"ctx_a": "Before entering a concert, Rae leaves her bicycle at the east rack. While she is inside with no view outdoors, staff move it to the west rack for construction. Until Rae receives new information, she will believe her bicycle is at",
"ctx_b": "",
"ctx": "Before entering a concert, Rae leaves her bicycle at the east rack. While she is inside with no view outdoors, staff move it to the west rack for construction. Until Rae receives new information, she will believe her bicycle is at",
"endings": [
" the east bicycle rack.",
" the west bicycle rack.",
" the construction office.",
" the concert entrance."
],
"source_id": "tiny_theory_of_mind_batch_008_natural_continuation_1910",
"split": "test",
"split_type": "synthetic",
"label": "0",
"metadata": {
"answer": "the east bicycle rack",
"unit": "text continuation",
"topic": "first_order_false_belief",
"difficulty": "medium",
"grade_band": "2-4",
"template_id": "theory_of_mind_continuation",
"format_note": "Continuation-style prompt for base-model length-normalized log-likelihood scoring; endings include leading spaces.",
"setting": "concert_bicycle_east_west_rack_move",
"vocabulary_themes": [
"transport",
"false_belief"
]
}
}
The correct continuation in this example is option 0:
Before entering a concert, Rae leaves her bicycle at the east rack. While she is inside with no view outdoors, staff move it to the west rack for construction. Until Rae receives new information, she will believe her bicycle is at the east bicycle rack.
Each of 40 constructs has 50 examples (2.5% of the dataset), split evenly between easy and medium. Each cell shows the number of examples in that grade band.
| Construct | Pre-K–K | K–2 | 2–4 | 4–6 |
|---|---|---|---|---|
| Accidental vs. intentional | 9 | 31 | 10 | 0 |
| Appearance vs. reality | 9 | 27 | 13 | 1 |
| Audience design | 0 | 15 | 32 | 3 |
| Belief–desire integration | 0 | 2 | 47 | 1 |
| Belief revision | 2 | 20 | 27 | 1 |
| Bluffing | 0 | 15 | 28 | 7 |
| Common ignorance | 0 | 12 | 36 | 2 |
| Communication failure | 1 | 21 | 23 | 5 |
| Counterfactual mental state | 1 | 4 | 26 | 19 |
| Deception | 1 | 18 | 27 | 4 |
| Desire-based emotion | 21 | 19 | 9 | 1 |
| Diverse beliefs | 8 | 26 | 14 | 2 |
| Diverse desires | 20 | 17 | 13 | 0 |
| Emotion from belief | 3 | 21 | 24 | 2 |
| Epistemic independence | 0 | 1 | 32 | 17 |
| Epistemic trust | 0 | 16 | 25 | 9 |
| Epistemic uncertainty | 0 | 5 | 41 | 4 |
| False inference | 1 | 21 | 25 | 3 |
| Faux pas | 0 | 14 | 25 | 11 |
| First-order false belief | 10 | 26 | 13 | 1 |
| Goal inference | 12 | 23 | 15 | 0 |
| Hidden emotions | 2 | 17 | 29 | 2 |
| Ignorance | 15 | 25 | 9 | 1 |
| Indirect request | 3 | 31 | 16 | 0 |
| Interpretive diversity | 0 | 9 | 29 | 12 |
| Joke vs. lie | 3 | 21 | 22 | 4 |
| Knowledge by perception | 16 | 23 | 10 | 1 |
| Mistaken identity | 2 | 24 | 22 | 2 |
| Mistaken intention | 0 | 14 | 27 | 9 |
| Mutual knowledge | 0 | 12 | 32 | 6 |
| Participant-role common ground | 0 | 2 | 33 | 15 |
| Perspective taking | 5 | 22 | 22 | 1 |
| Persuasive intent | 0 | 1 | 32 | 17 |
| Pluralistic ignorance | 0 | 0 | 6 | 44 |
| Recursive belief reasoning | 0 | 0 | 15 | 35 |
| Reputation management | 0 | 6 | 33 | 11 |
| Sarcasm and irony | 0 | 16 | 31 | 3 |
| Second-order false belief | 0 | 11 | 23 | 16 |
| Unexpected-contents false belief | 7 | 29 | 14 | 0 |
| White lie | 0 | 18 | 28 | 4 |
| Total | 151 | 635 | 938 | 276 |
Total examples: 2,000.
A note on difficulty: Although the benchmark is split 50/50 between easy and medium difficulty samples, individual theory of mind constructs are difficult to balance by grade due to differing ages when human children acquire theory of mind concepts. Some concepts are naturally acquired later. As such, you'll see some constructs, like pluralistic ignorance, weigh more heavily towards older grades, whereas diverse desires leans towards younger grades. Overall, an attempt has been made to balance by overall difficulty within each construct, and thus within the overall benchmark, rather than balancing by grade level.
| Difficulty | Examples | Share |
|---|---|---|
| Easy | 1000 | 50.0% |
| Medium | 1000 | 50.0% |
| Total | 2,000 | 100.0% |
Note: Difficulties are marked, like grade levels, as references to human standards. These may not, and often do not, reflect the general difficulty language models have answering.
| Grade band | Examples | Share |
|---|---|---|
| Pre-K–K | 151 | 7.55% |
| K–2 | 635 | 31.75% |
| Grades 2–4 | 938 | 46.90% |
| Grades 4–6 | 276 | 13.80% |
| Total | 2,000 | 100.00% |
The labels are exactly balanced, preventing answer-position frequency from providing an advantage.
| Label | Examples | Share |
|---|---|---|
0 | 500 | 25.0% |
1 | 500 | 25.0% |
2 | 500 | 25.0% |
3 | 500 | 25.0% |
| Total | 2,000 | 100.0% |
Tiny Theory of Mind (Tiny ToM) is released under the Apache License 2.0.
Special thanks to the following people who made Tiny Theory of Mind possible:
Mmorgan-ML - Project Lead
Datdanboi - Invaluable advice & guidance on making benchmarks
Amanda Long - Brainstorming and help creating 60 early draft questions
The AxiomicLabs team for their constant encouragement.
6 commits

Tiny Theory of Mind is our first attempt at evaluating theory of mind capabilities in small language models. The benchmark covers a wide variety of ToM topics, ranging in difficulties that, for humans, would be appropriate for Pre-K through 6th grade.
The benchmark is designed primarily for base-model continuation log-likelihood scoring. It does not require instruction following, chain-of-thought, or generated explanations. Random-choice accuracy is 25%.
| Property | Value |
|---|---|
| Examples | 2,000 |
| Answer choices per example | 4 |
| Random-choice accuracy | 25% |
| ToM Constructs | 40 (50 samples each) |
| Settings | 937 |
| Vocabulary Themes | 514 |
| Grade bands | Grades PreK-6 |
| Difficulty levels | Easy and Medium (1000 samples each) |
| Language | English |
| Data source | Synthetic |
| Dataset file | tiny_theory_of_mind_2000.jsonl |
| Rank | Model | Score |
|---|---|---|
| 1 | HuggingFaceTB/SmolLM-135M | 47.20% |
| 2 | BananaMind/BananaMind-2-Pro | 46.65% |
| 3 | HuggingFaceTB/SmolLM2-135M | 46.55% |
| 4 | AxiomicLabs/GPT-X2-125M | 44.55% |
| 5 | AxiomicLabs/GPT-X2.5-135M | 44.20% |
| 6 | SupraLabs/Supra2-100M-Base | 43.00% |
| 7 | AxiomicLabs/GPT-X-125M | 42.75% |
| 8 | facebook/opt-125m | 42.75% |
| 9 | SupraLabs/Supra-50M-Base | 41.90% |
| 10 | BananaMind/BananaMind-2-Medium | 41.20% |
| 11 | BananaMind/BananaMind-2-Medium-Chat | 40.95% |
| 12 | TobiasLogic/Museko-125M | 40.90% |
| 13 | openai-community/gpt2 | 40.75% |
| 14 | SupraLabs/Supra1.5-50M-Base-exp | 39.95% |
| 15 | finnianx/Gros-Michel-90m-Base-v2 | 39.80% |
| 16 | BananaMind/BananaMind-2-Mini | 38.55% |
| 17 | veyra-ai/Veyra2-Apricot-50M-Base | 38.20% |
| 18 | BananaMind/BananaMind-2.1-Unified | 38.15% |
| 19 | GODELEV/Archaea-74M | 37.15% |
| 20 | BananaMind/BananaMind-2-Nano | 37.10% |
| 21 | finnianx/michel-micro | 36.40% |
| 22 | fromziro/Er-Large-30M | 35.50% |
| 23 | EleutherAI/pythia-31m | 35.20% |
| 24 | BananaMind/BananaMind-2-MoE | 34.05% |
| 25 | BananaMind/BananaMind-2-SLMoE | 34.00% |
| 26 | AxiomicLabs/GPT-S2-5M | 32.65% |
| 27 | User01110/CMA-8M | 32.55% |
| 28 | AxiomicLabs/GPT-S-5M | 31.60% |
| 29 | BananaMind/BananaMind-2-Micro | 31.60% |
| 30 | User01110/CMA-1M-Mini | 31.40% |
| 31 | fromziro/Syn-2.6M | 30.30% |
| 32 | AtomixLabs/Photon-2.0-1M | 30.05% |
| 33 | AxiomicLabs/GPT-S-1.4M | 29.95% |
| 34 | BananaMind/BananaMind-2.1-Pico-Preview | 29.20% |
| 35 | ThingAI/Quark-50m-v2 | 29.10% |
| 36 | harley-ml/dillion-1.2M | 28.30% |
Every item contains an unfinished Theory of Mind scenario in ctx, four possible continuations in endings, and a zero-based correct-answer index in label.
{
"ind": 1910,
"activity_label": "tiny_theory_of_mind_natural_continuation::first_order_false_belief::order_1::medium",
"ctx_a": "Before entering a concert, Rae leaves her bicycle at the east rack. While she is inside with no view outdoors, staff move it to the west rack for construction. Until Rae receives new information, she will believe her bicycle is at",
"ctx_b": "",
"ctx": "Before entering a concert, Rae leaves her bicycle at the east rack. While she is inside with no view outdoors, staff move it to the west rack for construction. Until Rae receives new information, she will believe her bicycle is at",
"endings": [
" the east bicycle rack.",
" the west bicycle rack.",
" the construction office.",
" the concert entrance."
],
"source_id": "tiny_theory_of_mind_batch_008_natural_continuation_1910",
"split": "test",
"split_type": "synthetic",
"label": "0",
"metadata": {
"answer": "the east bicycle rack",
"unit": "text continuation",
"topic": "first_order_false_belief",
"difficulty": "medium",
"grade_band": "2-4",
"template_id": "theory_of_mind_continuation",
"format_note": "Continuation-style prompt for base-model length-normalized log-likelihood scoring; endings include leading spaces.",
"setting": "concert_bicycle_east_west_rack_move",
"vocabulary_themes": [
"transport",
"false_belief"
]
}
}
The correct continuation in this example is option 0:
Before entering a concert, Rae leaves her bicycle at the east rack. While she is inside with no view outdoors, staff move it to the west rack for construction. Until Rae receives new information, she will believe her bicycle is at the east bicycle rack.
Each of 40 constructs has 50 examples (2.5% of the dataset), split evenly between easy and medium. Each cell shows the number of examples in that grade band.
| Construct | Pre-K–K | K–2 | 2–4 | 4–6 |
|---|---|---|---|---|
| Accidental vs. intentional | 9 | 31 | 10 | 0 |
| Appearance vs. reality | 9 | 27 | 13 | 1 |
| Audience design | 0 | 15 | 32 | 3 |
| Belief–desire integration | 0 | 2 | 47 | 1 |
| Belief revision | 2 | 20 | 27 | 1 |
| Bluffing | 0 | 15 | 28 | 7 |
| Common ignorance | 0 | 12 | 36 | 2 |
| Communication failure | 1 | 21 | 23 | 5 |
| Counterfactual mental state | 1 | 4 | 26 | 19 |
| Deception | 1 | 18 | 27 | 4 |
| Desire-based emotion | 21 | 19 | 9 | 1 |
| Diverse beliefs | 8 | 26 | 14 | 2 |
| Diverse desires | 20 | 17 | 13 | 0 |
| Emotion from belief | 3 | 21 | 24 | 2 |
| Epistemic independence | 0 | 1 | 32 | 17 |
| Epistemic trust | 0 | 16 | 25 | 9 |
| Epistemic uncertainty | 0 | 5 | 41 | 4 |
| False inference | 1 | 21 | 25 | 3 |
| Faux pas | 0 | 14 | 25 | 11 |
| First-order false belief | 10 | 26 | 13 | 1 |
| Goal inference | 12 | 23 | 15 | 0 |
| Hidden emotions | 2 | 17 | 29 | 2 |
| Ignorance | 15 | 25 | 9 | 1 |
| Indirect request | 3 | 31 | 16 | 0 |
| Interpretive diversity | 0 | 9 | 29 | 12 |
| Joke vs. lie | 3 | 21 | 22 | 4 |
| Knowledge by perception | 16 | 23 | 10 | 1 |
| Mistaken identity | 2 | 24 | 22 | 2 |
| Mistaken intention | 0 | 14 | 27 | 9 |
| Mutual knowledge | 0 | 12 | 32 | 6 |
| Participant-role common ground | 0 | 2 | 33 | 15 |
| Perspective taking | 5 | 22 | 22 | 1 |
| Persuasive intent | 0 | 1 | 32 | 17 |
| Pluralistic ignorance | 0 | 0 | 6 | 44 |
| Recursive belief reasoning | 0 | 0 | 15 | 35 |
| Reputation management | 0 | 6 | 33 | 11 |
| Sarcasm and irony | 0 | 16 | 31 | 3 |
| Second-order false belief | 0 | 11 | 23 | 16 |
| Unexpected-contents false belief | 7 | 29 | 14 | 0 |
| White lie | 0 | 18 | 28 | 4 |
| Total | 151 | 635 | 938 | 276 |
Total examples: 2,000.
A note on difficulty: Although the benchmark is split 50/50 between easy and medium difficulty samples, individual theory of mind constructs are difficult to balance by grade due to differing ages when human children acquire theory of mind concepts. Some concepts are naturally acquired later. As such, you'll see some constructs, like pluralistic ignorance, weigh more heavily towards older grades, whereas diverse desires leans towards younger grades. Overall, an attempt has been made to balance by overall difficulty within each construct, and thus within the overall benchmark, rather than balancing by grade level.
| Difficulty | Examples | Share |
|---|---|---|
| Easy | 1000 | 50.0% |
| Medium | 1000 | 50.0% |
| Total | 2,000 | 100.0% |
Note: Difficulties are marked, like grade levels, as references to human standards. These may not, and often do not, reflect the general difficulty language models have answering.
| Grade band | Examples | Share |
|---|---|---|
| Pre-K–K | 151 | 7.55% |
| K–2 | 635 | 31.75% |
| Grades 2–4 | 938 | 46.90% |
| Grades 4–6 | 276 | 13.80% |
| Total | 2,000 | 100.00% |
The labels are exactly balanced, preventing answer-position frequency from providing an advantage.
| Label | Examples | Share |
|---|---|---|
0 | 500 | 25.0% |
1 | 500 | 25.0% |
2 | 500 | 25.0% |
3 | 500 | 25.0% |
| Total | 2,000 | 100.0% |
Tiny Theory of Mind (Tiny ToM) is released under the Apache License 2.0.
Special thanks to the following people who made Tiny Theory of Mind possible:
Mmorgan-ML - Project Lead
Datdanboi - Invaluable advice & guidance on making benchmarks
Amanda Long - Brainstorming and help creating 60 early draft questions
The AxiomicLabs team for their constant encouragement.
6 commits