ying-hui-he/Hi-ToM_dataset

[EMNLP 2023] Hi-ToM benchmark

Jupyter Notebook

21

65 commits

updated Oct 11, 2025

See the code

See what people are saying

SourceMessageScoreDate

Difficult reasoning competition for LLMs - Hi-ToM (r/ClaudeAI)

Source: [https://github.com/ying-hui-he/Hi-ToM\_dataset](https://github.com/ying-hui-he/Hi-ToM_dataset) I asked several LLMs to solve this Hi-ToM puzzle: The following story happens in chronological order. You will be given a multiple-choice question and a note at the end. Directly output the…

7

Oct 1, 2026

README

Hi-ToM Dataset

This is the dataset for the paper "Hi-ToM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models".

<img src=media/Picture1.png height=430>

Hi-ToM dataset

Hi-ToM_data/Hi-ToM_data.json contains 1.2k higher-order ToM question-answer pairs.

Generate new data and prompts

Please run the script generate_tomh.sh to automatically generate new stories along questions and answers.

ying-hui-he/Hi-ToM_dataset

[EMNLP 2023] Hi-ToM benchmark

Jupyter Notebook

21

65 commits

updated Oct 11, 2025

See the code

See what people are saying

SourceMessageScoreDate

Difficult reasoning competition for LLMs - Hi-ToM (r/ClaudeAI)

Source: [https://github.com/ying-hui-he/Hi-ToM\_dataset](https://github.com/ying-hui-he/Hi-ToM_dataset) I asked several LLMs to solve this Hi-ToM puzzle: The following story happens in chronological order. You will be given a multiple-choice question and a note at the end. Directly output the…

7

Oct 1, 2026

README

Hi-ToM Dataset

This is the dataset for the paper "Hi-ToM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models".

<img src=media/Picture1.png height=430>

Hi-ToM dataset

Hi-ToM_data/Hi-ToM_data.json contains 1.2k higher-order ToM question-answer pairs.

Generate new data and prompts

Please run the script generate_tomh.sh to automatically generate new stories along questions and answers.

Languages

Jupyter Notebook

97.9%

Python

2.0%