TruthGen is a dataset of generated political statements, created to assess the relationship between truthfulness and political bias in reward models and language models. It consists of non-repetitive, non-political factual statements paired with false statements, designed to evaluate models for their ability to distinguish true from false information while minimizing political content. The dataset was generated using GPT-3.5, GPT-4 and Gemini, with a focus on objective truthfulness rather than subjective or politically charged topics.
TruthGen is a dataset of generated true and false statements, intended for research on truthfulness in reward models and language models, specifically in contexts where political bias is undesirable. This dataset contains 1,987 statement pairs (3,974 statements in total), with each pair containing one true statement and one false statement. It spans a variety of everyday and scientific facts, excluding politically charged topics to the greatest extent possible.
The dataset is particularly useful for evaluating reward models trained for alignment with truth, as well as for research on mitigating political bias while improving model accuracy on truth-related tasks.
This dataset is suitable for:
This dataset is not intended for tasks that require detailed political analysis or subjective evaluations. It is also not suitable for fine-grained political content analysis, as it intentionally avoids politically charged statements.
The dataset consists of approximately 4,000 true and false statements in approximately 2,000 statement pairs, where each pair includes:
truth: A true statement.model_truth: The model generating the true statementfalsehood: A corresponding false statement on the same topic.model_falsehood: The model generating the false statementThe statements were generated by GPT-3.5, GPT-4 and Gemini and cover factual knowledge about the world, scientific facts, and everyday knowledge. Political and subjective statements were avoided to ensure focus on truthfulness.
TruthGen was created to provide a dataset for studying model alignment with truth while minimizing political bias. The goal is to evaluate how models process factual information without the confounding effects of political or controversial content.
The statements were generated by GPT-3.5, GPT-4 and Gemini using prompts that aimed to elicit non-political, factual information. Clustering techniques were applied to ensure diversity in the generated statements. Manual auditing was performed on a sample of the data to ensure that the true/false labels were correct and that the statements avoided political content.
The dataset was generated by language models (GPT-3.5, GPT-4 and Gemini) and audited by researchers at MIT.
This dataset does not contain personal, sensitive, or private information. The statements are factual in nature and do not include identifiable data.
While TruthGen was created to avoid political content, there are still some limitations to be aware of:
Users should be aware of potential artifacts from the language models used to generate the data. For applications involving high-stakes decisions, additional manual checks or audits are recommended.
BibTeX:
@inproceedings{fulayRelationshipTruthPolitical2024,
author = {Fulay, Suyash and Brannon, William and Mohanty, Shrestha and Overney, Cassandra and Poole-Dayan, Elinor and Roy, Deb and Kabbara, Jad},
title = {On the Relationship between Truth and Political Bias in Language Models},
booktitle = {Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP '24)},
year = {2024},
month = nov,
publisher = {Association for Computational Linguistics},
note = {arXiv:2409.05283},
abstract = {Language model alignment research often attempts to ensure that models are not only helpful and harmless, but also truthful and unbiased. However, optimizing these objectives simultaneously can obscure how improving one aspect might impact the others. In this work, we focus on analyzing the relationship between two concepts essential in both language model alignment and political science: \textit{truthfulness} and \textit{political bias}. We train reward models on various popular truthfulness datasets and subsequently evaluate their political bias. Our findings reveal that optimizing reward models for truthfulness on these datasets tends to result in a left-leaning political bias. We also find that existing open-source reward models (i.e. those trained on standard human preference datasets) already show a similar bias and that the bias is larger for larger models. These results raise important questions about both the datasets used to represent truthfulness and what language models capture about the relationship between truth and politics.}
}
APA:
Fulay, S., Brannon, W., Mohanty, S., Overney, C., Poole-Dayan, E., Roy, D., & Kabbara, J. (2024). On the Relationship between Truth and Political Bias in Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP '24). Association for Computational Linguistics.
William Brannon, wbrannon@mit.edu
13 commits
TruthGen is a dataset of generated political statements, created to assess the relationship between truthfulness and political bias in reward models and language models. It consists of non-repetitive, non-political factual statements paired with false statements, designed to evaluate models for their ability to distinguish true from false information while minimizing political content. The dataset was generated using GPT-3.5, GPT-4 and Gemini, with a focus on objective truthfulness rather than subjective or politically charged topics.
TruthGen is a dataset of generated true and false statements, intended for research on truthfulness in reward models and language models, specifically in contexts where political bias is undesirable. This dataset contains 1,987 statement pairs (3,974 statements in total), with each pair containing one true statement and one false statement. It spans a variety of everyday and scientific facts, excluding politically charged topics to the greatest extent possible.
The dataset is particularly useful for evaluating reward models trained for alignment with truth, as well as for research on mitigating political bias while improving model accuracy on truth-related tasks.
This dataset is suitable for:
This dataset is not intended for tasks that require detailed political analysis or subjective evaluations. It is also not suitable for fine-grained political content analysis, as it intentionally avoids politically charged statements.
The dataset consists of approximately 4,000 true and false statements in approximately 2,000 statement pairs, where each pair includes:
truth: A true statement.model_truth: The model generating the true statementfalsehood: A corresponding false statement on the same topic.model_falsehood: The model generating the false statementThe statements were generated by GPT-3.5, GPT-4 and Gemini and cover factual knowledge about the world, scientific facts, and everyday knowledge. Political and subjective statements were avoided to ensure focus on truthfulness.
TruthGen was created to provide a dataset for studying model alignment with truth while minimizing political bias. The goal is to evaluate how models process factual information without the confounding effects of political or controversial content.
The statements were generated by GPT-3.5, GPT-4 and Gemini using prompts that aimed to elicit non-political, factual information. Clustering techniques were applied to ensure diversity in the generated statements. Manual auditing was performed on a sample of the data to ensure that the true/false labels were correct and that the statements avoided political content.
The dataset was generated by language models (GPT-3.5, GPT-4 and Gemini) and audited by researchers at MIT.
This dataset does not contain personal, sensitive, or private information. The statements are factual in nature and do not include identifiable data.
While TruthGen was created to avoid political content, there are still some limitations to be aware of:
Users should be aware of potential artifacts from the language models used to generate the data. For applications involving high-stakes decisions, additional manual checks or audits are recommended.
BibTeX:
@inproceedings{fulayRelationshipTruthPolitical2024,
author = {Fulay, Suyash and Brannon, William and Mohanty, Shrestha and Overney, Cassandra and Poole-Dayan, Elinor and Roy, Deb and Kabbara, Jad},
title = {On the Relationship between Truth and Political Bias in Language Models},
booktitle = {Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP '24)},
year = {2024},
month = nov,
publisher = {Association for Computational Linguistics},
note = {arXiv:2409.05283},
abstract = {Language model alignment research often attempts to ensure that models are not only helpful and harmless, but also truthful and unbiased. However, optimizing these objectives simultaneously can obscure how improving one aspect might impact the others. In this work, we focus on analyzing the relationship between two concepts essential in both language model alignment and political science: \textit{truthfulness} and \textit{political bias}. We train reward models on various popular truthfulness datasets and subsequently evaluate their political bias. Our findings reveal that optimizing reward models for truthfulness on these datasets tends to result in a left-leaning political bias. We also find that existing open-source reward models (i.e. those trained on standard human preference datasets) already show a similar bias and that the bias is larger for larger models. These results raise important questions about both the datasets used to represent truthfulness and what language models capture about the relationship between truth and politics.}
}
APA:
Fulay, S., Brannon, W., Mohanty, S., Overney, C., Poole-Dayan, E., Roy, D., & Kabbara, J. (2024). On the Relationship between Truth and Political Bias in Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP '24). Association for Computational Linguistics.
William Brannon, wbrannon@mit.edu
13 commits