This dataset contains 13,855 pairs of left-leaning and right-leaning political statements matched by topic. The dataset was generated using GPT-3.5 Turbo and has been audited to ensure quality and ideological balance. It is designed to facilitate the study of political bias in reward models and language models, with a focus on the relationship between truthfulness and political views.
TwinViews-13k is a dataset of 13,855 pairs of left-leaning and right-leaning political statements, each pair matched by topic. It was created to study political bias in reward and language models, with a focus on understanding the interaction between model alignment to truthfulness and the emergence of political bias. The dataset was generated using GPT-3.5 Turbo, with extensive auditing to ensure ideological balance and topical relevance.
This dataset can be used for various tasks related to political bias, natural language processing, and model alignment, particularly in studies examining how political orientation impacts model outputs.
This dataset is suitable for:
This dataset is not suitable for tasks requiring very fine-grained or human-labeled annotations of political affiliation beyond the machine-generated left/right splits. Notions of "left" and "right" may also vary between countries and over time, and users of the data should check that it captures the ideological dimensions of interest.
The dataset contains 13,855 pairs of left-leaning and right-leaning political statements. Each pair is matched by topic, with statements generated to be similar in style and length. The dataset consists of the following fields:
l: A left-leaning political statement.r: A right-leaning political statement.topic: The general topic of the pair (e.g., taxes, climate, education).The dataset was created to fill the gap in large-scale, topically matched political statement pairs for studying bias in LLMs. It allows for comparison of how models treat left-leaning versus right-leaning perspectives, particularly in the context of truthfulness and political bias.
The data was generated using GPT-3.5 Turbo. A carefully designed prompt was used to generate statement pairs that were ideologically representative of left-leaning and right-leaning viewpoints. The statements were then audited to ensure relevance, ideological alignment, and quality. Topic matching was done to ensure the statements are comparable across the political spectrum.
In summary:
The dataset was generated by GPT-3.5 Turbo, with extensive auditing performed by the dataset creators at MIT.
The dataset consists of machine-generated political statements and does not contain any personal or sensitive information.
Users of the dataset should be aware of certain limitations:
BibTeX:
@inproceedings{fulayRelationshipTruthPolitical2024,
author = {Fulay, Suyash and Brannon, William and Mohanty, Shrestha and Overney, Cassandra and Poole-Dayan, Elinor and Roy, Deb and Kabbara, Jad},
title = {On the Relationship between Truth and Political Bias in Language Models},
booktitle = {Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP '24)},
year = {2024},
month = nov,
publisher = {Association for Computational Linguistics},
note = {arXiv:2409.05283},
abstract = {Language model alignment research often attempts to ensure that models are not only helpful and harmless, but also truthful and unbiased. However, optimizing these objectives simultaneously can obscure how improving one aspect might impact the others. In this work, we focus on analyzing the relationship between two concepts essential in both language model alignment and political science: \textit{truthfulness} and \textit{political bias}. We train reward models on various popular truthfulness datasets and subsequently evaluate their political bias. Our findings reveal that optimizing reward models for truthfulness on these datasets tends to result in a left-leaning political bias. We also find that existing open-source reward models (i.e. those trained on standard human preference datasets) already show a similar bias and that the bias is larger for larger models. These results raise important questions about both the datasets used to represent truthfulness and what language models capture about the relationship between truth and politics.}
}
APA:
Fulay, S., Brannon, W., Mohanty, S., Overney, C., Poole-Dayan, E., Roy, D., & Kabbara, J. (2024). On the Relationship between Truth and Political Bias in Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP '24). Association for Computational Linguistics.
William Brannon, wbrannon@mit.edu
14 commits
This dataset contains 13,855 pairs of left-leaning and right-leaning political statements matched by topic. The dataset was generated using GPT-3.5 Turbo and has been audited to ensure quality and ideological balance. It is designed to facilitate the study of political bias in reward models and language models, with a focus on the relationship between truthfulness and political views.
TwinViews-13k is a dataset of 13,855 pairs of left-leaning and right-leaning political statements, each pair matched by topic. It was created to study political bias in reward and language models, with a focus on understanding the interaction between model alignment to truthfulness and the emergence of political bias. The dataset was generated using GPT-3.5 Turbo, with extensive auditing to ensure ideological balance and topical relevance.
This dataset can be used for various tasks related to political bias, natural language processing, and model alignment, particularly in studies examining how political orientation impacts model outputs.
This dataset is suitable for:
This dataset is not suitable for tasks requiring very fine-grained or human-labeled annotations of political affiliation beyond the machine-generated left/right splits. Notions of "left" and "right" may also vary between countries and over time, and users of the data should check that it captures the ideological dimensions of interest.
The dataset contains 13,855 pairs of left-leaning and right-leaning political statements. Each pair is matched by topic, with statements generated to be similar in style and length. The dataset consists of the following fields:
l: A left-leaning political statement.r: A right-leaning political statement.topic: The general topic of the pair (e.g., taxes, climate, education).The dataset was created to fill the gap in large-scale, topically matched political statement pairs for studying bias in LLMs. It allows for comparison of how models treat left-leaning versus right-leaning perspectives, particularly in the context of truthfulness and political bias.
The data was generated using GPT-3.5 Turbo. A carefully designed prompt was used to generate statement pairs that were ideologically representative of left-leaning and right-leaning viewpoints. The statements were then audited to ensure relevance, ideological alignment, and quality. Topic matching was done to ensure the statements are comparable across the political spectrum.
In summary:
The dataset was generated by GPT-3.5 Turbo, with extensive auditing performed by the dataset creators at MIT.
The dataset consists of machine-generated political statements and does not contain any personal or sensitive information.
Users of the dataset should be aware of certain limitations:
BibTeX:
@inproceedings{fulayRelationshipTruthPolitical2024,
author = {Fulay, Suyash and Brannon, William and Mohanty, Shrestha and Overney, Cassandra and Poole-Dayan, Elinor and Roy, Deb and Kabbara, Jad},
title = {On the Relationship between Truth and Political Bias in Language Models},
booktitle = {Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP '24)},
year = {2024},
month = nov,
publisher = {Association for Computational Linguistics},
note = {arXiv:2409.05283},
abstract = {Language model alignment research often attempts to ensure that models are not only helpful and harmless, but also truthful and unbiased. However, optimizing these objectives simultaneously can obscure how improving one aspect might impact the others. In this work, we focus on analyzing the relationship between two concepts essential in both language model alignment and political science: \textit{truthfulness} and \textit{political bias}. We train reward models on various popular truthfulness datasets and subsequently evaluate their political bias. Our findings reveal that optimizing reward models for truthfulness on these datasets tends to result in a left-leaning political bias. We also find that existing open-source reward models (i.e. those trained on standard human preference datasets) already show a similar bias and that the bias is larger for larger models. These results raise important questions about both the datasets used to represent truthfulness and what language models capture about the relationship between truth and politics.}
}
APA:
Fulay, S., Brannon, W., Mohanty, S., Overney, C., Poole-Dayan, E., Roy, D., & Kabbara, J. (2024). On the Relationship between Truth and Political Bias in Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP '24). Association for Computational Linguistics.
William Brannon, wbrannon@mit.edu
14 commits