AxiomicLabs/Tiny_Theory_of_Mind

Dataset

Tiny Theory of Mind

12

6 commits

updated Sep 23, 2026

See the code

README

axiomic banner

Tiny Theory of Mind

Tiny Theory of Mind is our first attempt at evaluating theory of mind capabilities in small language models. The benchmark covers a wide variety of ToM topics, ranging in difficulties that, for humans, would be appropriate for Pre-K through 6th grade.

The benchmark is designed primarily for base-model continuation log-likelihood scoring. It does not require instruction following, chain-of-thought, or generated explanations. Random-choice accuracy is 25%.

At a glance

PropertyValue
Examples2,000
Answer choices per example4
Random-choice accuracy25%
ToM Constructs40 (50 samples each)
Settings937
Vocabulary Themes514
Grade bandsGrades PreK-6
Difficulty levelsEasy and Medium (1000 samples each)
LanguageEnglish
Data sourceSynthetic
Dataset filetiny_theory_of_mind_2000.jsonl

Baseline Results

View all 36 baseline results
RankModelScore
1HuggingFaceTB/SmolLM-135M47.20%
2BananaMind/BananaMind-2-Pro46.65%
3HuggingFaceTB/SmolLM2-135M46.55%
4AxiomicLabs/GPT-X2-125M44.55%
5AxiomicLabs/GPT-X2.5-135M44.20%
6SupraLabs/Supra2-100M-Base43.00%
7AxiomicLabs/GPT-X-125M42.75%
8facebook/opt-125m42.75%
9SupraLabs/Supra-50M-Base41.90%
10BananaMind/BananaMind-2-Medium41.20%
11BananaMind/BananaMind-2-Medium-Chat40.95%
12TobiasLogic/Museko-125M40.90%
13openai-community/gpt240.75%
14SupraLabs/Supra1.5-50M-Base-exp39.95%
15finnianx/Gros-Michel-90m-Base-v239.80%
16BananaMind/BananaMind-2-Mini38.55%
17veyra-ai/Veyra2-Apricot-50M-Base38.20%
18BananaMind/BananaMind-2.1-Unified38.15%
19GODELEV/Archaea-74M37.15%
20BananaMind/BananaMind-2-Nano37.10%
21finnianx/michel-micro36.40%
22fromziro/Er-Large-30M35.50%
23EleutherAI/pythia-31m35.20%
24BananaMind/BananaMind-2-MoE34.05%
25BananaMind/BananaMind-2-SLMoE34.00%
26AxiomicLabs/GPT-S2-5M32.65%
27User01110/CMA-8M32.55%
28AxiomicLabs/GPT-S-5M31.60%
29BananaMind/BananaMind-2-Micro31.60%
30User01110/CMA-1M-Mini31.40%
31fromziro/Syn-2.6M30.30%
32AtomixLabs/Photon-2.0-1M30.05%
33AxiomicLabs/GPT-S-1.4M29.95%
34BananaMind/BananaMind-2.1-Pico-Preview29.20%
35ThingAI/Quark-50m-v229.10%
36harley-ml/dillion-1.2M28.30%

Task Format

Every item contains an unfinished Theory of Mind scenario in ctx, four possible continuations in endings, and a zero-based correct-answer index in label.

{
  "ind": 1910,
  "activity_label": "tiny_theory_of_mind_natural_continuation::first_order_false_belief::order_1::medium",
  "ctx_a": "Before entering a concert, Rae leaves her bicycle at the east rack. While she is inside with no view outdoors, staff move it to the west rack for construction. Until Rae receives new information, she will believe her bicycle is at",
  "ctx_b": "",
  "ctx": "Before entering a concert, Rae leaves her bicycle at the east rack. While she is inside with no view outdoors, staff move it to the west rack for construction. Until Rae receives new information, she will believe her bicycle is at",
  "endings": [
    " the east bicycle rack.",
    " the west bicycle rack.",
    " the construction office.",
    " the concert entrance."
  ],
  "source_id": "tiny_theory_of_mind_batch_008_natural_continuation_1910",
  "split": "test",
  "split_type": "synthetic",
  "label": "0",
  "metadata": {
    "answer": "the east bicycle rack",
    "unit": "text continuation",
    "topic": "first_order_false_belief",
    "difficulty": "medium",
    "grade_band": "2-4",
    "template_id": "theory_of_mind_continuation",
    "format_note": "Continuation-style prompt for base-model length-normalized log-likelihood scoring; endings include leading spaces.",
    "setting": "concert_bicycle_east_west_rack_move",
    "vocabulary_themes": [
      "transport",
      "false_belief"
    ]
  }
}

The correct continuation in this example is option 0:

Before entering a concert, Rae leaves her bicycle at the east rack. While she is inside with no view outdoors, staff move it to the west rack for construction. Until Rae receives new information, she will believe her bicycle is at the east bicycle rack.

Dataset composition

Theory of Mind Constructs

Each of 40 constructs has 50 examples (2.5% of the dataset), split evenly between easy and medium. Each cell shows the number of examples in that grade band.

ConstructPre-K–KK–22–44–6
Accidental vs. intentional931100
Appearance vs. reality927131
Audience design015323
Belief–desire integration02471
Belief revision220271
Bluffing015287
Common ignorance012362
Communication failure121235
Counterfactual mental state142619
Deception118274
Desire-based emotion211991
Diverse beliefs826142
Diverse desires2017130
Emotion from belief321242
Epistemic independence013217
Epistemic trust016259
Epistemic uncertainty05414
False inference121253
Faux pas0142511
First-order false belief1026131
Goal inference1223150
Hidden emotions217292
Ignorance152591
Indirect request331160
Interpretive diversity092912
Joke vs. lie321224
Knowledge by perception1623101
Mistaken identity224222
Mistaken intention014279
Mutual knowledge012326
Participant-role common ground023315
Perspective taking522221
Persuasive intent013217
Pluralistic ignorance00644
Recursive belief reasoning001535
Reputation management063311
Sarcasm and irony016313
Second-order false belief0112316
Unexpected-contents false belief729140
White lie018284
Total151635938276

Total examples: 2,000.

A note on difficulty: Although the benchmark is split 50/50 between easy and medium difficulty samples, individual theory of mind constructs are difficult to balance by grade due to differing ages when human children acquire theory of mind concepts. Some concepts are naturally acquired later. As such, you'll see some constructs, like pluralistic ignorance, weigh more heavily towards older grades, whereas diverse desires leans towards younger grades. Overall, an attempt has been made to balance by overall difficulty within each construct, and thus within the overall benchmark, rather than balancing by grade level.

Difficulty

DifficultyExamplesShare
Easy100050.0%
Medium100050.0%
Total2,000100.0%

Note: Difficulties are marked, like grade levels, as references to human standards. These may not, and often do not, reflect the general difficulty language models have answering.

Grade bands

Grade bandExamplesShare
Pre-K–K1517.55%
K–263531.75%
Grades 2–493846.90%
Grades 4–627613.80%
Total2,000100.00%

Correct answer positions

The labels are exactly balanced, preventing answer-position frequency from providing an advantage.

LabelExamplesShare
050025.0%
150025.0%
250025.0%
350025.0%
Total2,000100.0%

License

Tiny Theory of Mind (Tiny ToM) is released under the Apache License 2.0.

Acknowledgments

Special thanks to the following people who made Tiny Theory of Mind possible:

Mmorgan-ML - Project Lead

Datdanboi - Invaluable advice & guidance on making benchmarks

Amanda Long - Brainstorming and help creating 60 early draft questions

The AxiomicLabs team for their constant encouragement.

benchmark
language model evaluation
multiple choice
synthetic
theory of mind

Contributors

Mmorgan-ML

6 commits

AxiomicLabs/Tiny_Theory_of_Mind

Dataset

Tiny Theory of Mind

12

6 commits

updated Sep 23, 2026

See the code

README

axiomic banner

Tiny Theory of Mind

Tiny Theory of Mind is our first attempt at evaluating theory of mind capabilities in small language models. The benchmark covers a wide variety of ToM topics, ranging in difficulties that, for humans, would be appropriate for Pre-K through 6th grade.

The benchmark is designed primarily for base-model continuation log-likelihood scoring. It does not require instruction following, chain-of-thought, or generated explanations. Random-choice accuracy is 25%.

At a glance

PropertyValue
Examples2,000
Answer choices per example4
Random-choice accuracy25%
ToM Constructs40 (50 samples each)
Settings937
Vocabulary Themes514
Grade bandsGrades PreK-6
Difficulty levelsEasy and Medium (1000 samples each)
LanguageEnglish
Data sourceSynthetic
Dataset filetiny_theory_of_mind_2000.jsonl

Baseline Results

View all 36 baseline results
RankModelScore
1HuggingFaceTB/SmolLM-135M47.20%
2BananaMind/BananaMind-2-Pro46.65%
3HuggingFaceTB/SmolLM2-135M46.55%
4AxiomicLabs/GPT-X2-125M44.55%
5AxiomicLabs/GPT-X2.5-135M44.20%
6SupraLabs/Supra2-100M-Base43.00%
7AxiomicLabs/GPT-X-125M42.75%
8facebook/opt-125m42.75%
9SupraLabs/Supra-50M-Base41.90%
10BananaMind/BananaMind-2-Medium41.20%
11BananaMind/BananaMind-2-Medium-Chat40.95%
12TobiasLogic/Museko-125M40.90%
13openai-community/gpt240.75%
14SupraLabs/Supra1.5-50M-Base-exp39.95%
15finnianx/Gros-Michel-90m-Base-v239.80%
16BananaMind/BananaMind-2-Mini38.55%
17veyra-ai/Veyra2-Apricot-50M-Base38.20%
18BananaMind/BananaMind-2.1-Unified38.15%
19GODELEV/Archaea-74M37.15%
20BananaMind/BananaMind-2-Nano37.10%
21finnianx/michel-micro36.40%
22fromziro/Er-Large-30M35.50%
23EleutherAI/pythia-31m35.20%
24BananaMind/BananaMind-2-MoE34.05%
25BananaMind/BananaMind-2-SLMoE34.00%
26AxiomicLabs/GPT-S2-5M32.65%
27User01110/CMA-8M32.55%
28AxiomicLabs/GPT-S-5M31.60%
29BananaMind/BananaMind-2-Micro31.60%
30User01110/CMA-1M-Mini31.40%
31fromziro/Syn-2.6M30.30%
32AtomixLabs/Photon-2.0-1M30.05%
33AxiomicLabs/GPT-S-1.4M29.95%
34BananaMind/BananaMind-2.1-Pico-Preview29.20%
35ThingAI/Quark-50m-v229.10%
36harley-ml/dillion-1.2M28.30%

Task Format

Every item contains an unfinished Theory of Mind scenario in ctx, four possible continuations in endings, and a zero-based correct-answer index in label.

{
  "ind": 1910,
  "activity_label": "tiny_theory_of_mind_natural_continuation::first_order_false_belief::order_1::medium",
  "ctx_a": "Before entering a concert, Rae leaves her bicycle at the east rack. While she is inside with no view outdoors, staff move it to the west rack for construction. Until Rae receives new information, she will believe her bicycle is at",
  "ctx_b": "",
  "ctx": "Before entering a concert, Rae leaves her bicycle at the east rack. While she is inside with no view outdoors, staff move it to the west rack for construction. Until Rae receives new information, she will believe her bicycle is at",
  "endings": [
    " the east bicycle rack.",
    " the west bicycle rack.",
    " the construction office.",
    " the concert entrance."
  ],
  "source_id": "tiny_theory_of_mind_batch_008_natural_continuation_1910",
  "split": "test",
  "split_type": "synthetic",
  "label": "0",
  "metadata": {
    "answer": "the east bicycle rack",
    "unit": "text continuation",
    "topic": "first_order_false_belief",
    "difficulty": "medium",
    "grade_band": "2-4",
    "template_id": "theory_of_mind_continuation",
    "format_note": "Continuation-style prompt for base-model length-normalized log-likelihood scoring; endings include leading spaces.",
    "setting": "concert_bicycle_east_west_rack_move",
    "vocabulary_themes": [
      "transport",
      "false_belief"
    ]
  }
}

The correct continuation in this example is option 0:

Before entering a concert, Rae leaves her bicycle at the east rack. While she is inside with no view outdoors, staff move it to the west rack for construction. Until Rae receives new information, she will believe her bicycle is at the east bicycle rack.

Dataset composition

Theory of Mind Constructs

Each of 40 constructs has 50 examples (2.5% of the dataset), split evenly between easy and medium. Each cell shows the number of examples in that grade band.

ConstructPre-K–KK–22–44–6
Accidental vs. intentional931100
Appearance vs. reality927131
Audience design015323
Belief–desire integration02471
Belief revision220271
Bluffing015287
Common ignorance012362
Communication failure121235
Counterfactual mental state142619
Deception118274
Desire-based emotion211991
Diverse beliefs826142
Diverse desires2017130
Emotion from belief321242
Epistemic independence013217
Epistemic trust016259
Epistemic uncertainty05414
False inference121253
Faux pas0142511
First-order false belief1026131
Goal inference1223150
Hidden emotions217292
Ignorance152591
Indirect request331160
Interpretive diversity092912
Joke vs. lie321224
Knowledge by perception1623101
Mistaken identity224222
Mistaken intention014279
Mutual knowledge012326
Participant-role common ground023315
Perspective taking522221
Persuasive intent013217
Pluralistic ignorance00644
Recursive belief reasoning001535
Reputation management063311
Sarcasm and irony016313
Second-order false belief0112316
Unexpected-contents false belief729140
White lie018284
Total151635938276

Total examples: 2,000.

A note on difficulty: Although the benchmark is split 50/50 between easy and medium difficulty samples, individual theory of mind constructs are difficult to balance by grade due to differing ages when human children acquire theory of mind concepts. Some concepts are naturally acquired later. As such, you'll see some constructs, like pluralistic ignorance, weigh more heavily towards older grades, whereas diverse desires leans towards younger grades. Overall, an attempt has been made to balance by overall difficulty within each construct, and thus within the overall benchmark, rather than balancing by grade level.

Difficulty

DifficultyExamplesShare
Easy100050.0%
Medium100050.0%
Total2,000100.0%

Note: Difficulties are marked, like grade levels, as references to human standards. These may not, and often do not, reflect the general difficulty language models have answering.

Grade bands

Grade bandExamplesShare
Pre-K–K1517.55%
K–263531.75%
Grades 2–493846.90%
Grades 4–627613.80%
Total2,000100.00%

Correct answer positions

The labels are exactly balanced, preventing answer-position frequency from providing an advantage.

LabelExamplesShare
050025.0%
150025.0%
250025.0%
350025.0%
Total2,000100.0%

License

Tiny Theory of Mind (Tiny ToM) is released under the Apache License 2.0.

Acknowledgments

Special thanks to the following people who made Tiny Theory of Mind possible:

Mmorgan-ML - Project Lead

Datdanboi - Invaluable advice & guidance on making benchmarks

Amanda Long - Brainstorming and help creating 60 early draft questions

The AxiomicLabs team for their constant encouragement.

benchmark
language model evaluation
multiple choice
synthetic
theory of mind

Contributors

Mmorgan-ML

6 commits