Belief in the Machine: Investigating Epistemological Blind Spots of Language Models
See the codeAs language models (LMs) become integral to fields like healthcare, law, and journalism, their ability to differentiate between fact, belief, and knowledge is essential for reliable decision-making. Failure to grasp these distinctions can lead to significant consequences in areas such as medical diagnosis, legal judgments, and dissemination of fake news. Despite this, current literature has largely focused on more complex issues such as theory of mind, overlooking more fundamental epistemic challenges. This study systematically evaluates the epistemic reasoning capabilities of modern LMs, including GPT-4, Claude-3, and Llama-3, using a new dataset, KaBLE, consisting of 13,000 questions across 13 tasks. Our results reveal key limitations.
These findings highlight significant concerns about current LMs' ability to reason about truth, belief, and knowledge while emphasizing the need for advancements in these areas before broad deployment in critical sectors.



Please refer to the /kable-dataset directory to get access to the raw input files.

All model outputs are stored in the /outputs directory.

To conduct your own experiments, please feel free to modify and use the run_experiments.py file. Before executing this code, however, please ensure that you have installed all the required packages (e.g., pip install -r requirements.txt) and have exported all relevant OpenAI API keys and credentials to your local environment (e.g., export OPENAI_API_KEY="YOUR_API_KEY").
Here is an example command to run the experiments:
[TBD]
If your work makes use of our model, data, or results, please cite our paper as follows:
@article{suzgun2024beliefmachine,
title={Belief in the Machine: Investigating Epistemological Blind Spots of Language Models},
author={Mirac Suzgun and Tayfun Gur and Federico Bianchi and Daniel E. Ho and Thomas Icard and Dan Jurafsky and James Zou},
year={2024},
eprint={2410.21195},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2410.21195},
}
We thank William Held, Wesley H. Holliday, Adam T. Kalai, Jacopo Tagliabue, Merve Tekgürler, Suproteem Sarkar, Emily Shen, Kyle Swanson, Angelina Wang, and Mert Yüksekgönül for their helpful comments and suggestions. We also thank the members of the James Zou Lab and the participants at the IX. CSLI Workshop on Logic, Rationality, and Intelligent Interaction at Stanford University.
Belief in the Machine: Investigating Epistemological Blind Spots of Language Models
See the codeAs language models (LMs) become integral to fields like healthcare, law, and journalism, their ability to differentiate between fact, belief, and knowledge is essential for reliable decision-making. Failure to grasp these distinctions can lead to significant consequences in areas such as medical diagnosis, legal judgments, and dissemination of fake news. Despite this, current literature has largely focused on more complex issues such as theory of mind, overlooking more fundamental epistemic challenges. This study systematically evaluates the epistemic reasoning capabilities of modern LMs, including GPT-4, Claude-3, and Llama-3, using a new dataset, KaBLE, consisting of 13,000 questions across 13 tasks. Our results reveal key limitations.
These findings highlight significant concerns about current LMs' ability to reason about truth, belief, and knowledge while emphasizing the need for advancements in these areas before broad deployment in critical sectors.



Please refer to the /kable-dataset directory to get access to the raw input files.

All model outputs are stored in the /outputs directory.

To conduct your own experiments, please feel free to modify and use the run_experiments.py file. Before executing this code, however, please ensure that you have installed all the required packages (e.g., pip install -r requirements.txt) and have exported all relevant OpenAI API keys and credentials to your local environment (e.g., export OPENAI_API_KEY="YOUR_API_KEY").
Here is an example command to run the experiments:
[TBD]
If your work makes use of our model, data, or results, please cite our paper as follows:
@article{suzgun2024beliefmachine,
title={Belief in the Machine: Investigating Epistemological Blind Spots of Language Models},
author={Mirac Suzgun and Tayfun Gur and Federico Bianchi and Daniel E. Ho and Thomas Icard and Dan Jurafsky and James Zou},
year={2024},
eprint={2410.21195},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2410.21195},
}
We thank William Held, Wesley H. Holliday, Adam T. Kalai, Jacopo Tagliabue, Merve Tekgürler, Suproteem Sarkar, Emily Shen, Kyle Swanson, Angelina Wang, and Mert Yüksekgönül for their helpful comments and suggestions. We also thank the members of the James Zou Lab and the participants at the IX. CSLI Workshop on Logic, Rationality, and Intelligent Interaction at Stanford University.