ucberkeley-dlab/fragility-moral-judgment-llms

Dataset

Fragility of Moral Judgment in Large Language Models

0

5 commits

2 linked in READMEs

updated May 25, 2026

See the code

README

Fragility of Moral Judgment in Large Language Models

Companion dataset for the FAccT paper Fragility of Moral Judgment in Large Language Models by Tom van Nuenen. Contains the moral dilemmas, community labels, and per-model verdicts (with explanations and reasoning traces) used in the study.

The paper investigates how stable LLM moral judgments are under minimal, morally-irrelevant perturbations of the same dilemma, and whether protocols and reasoning chains improve or worsen that stability.

Quick start

from datasets import load_dataset

dilemmas = load_dataset("ucberkeley-dlab/fragility-moral-judgment-llms", "dilemmas", split="train")
verdicts = load_dataset("ucberkeley-dlab/fragility-moral-judgment-llms", "model_verdicts", split="train")

Configs

ConfigRowsDescription
dilemmas2,939Source dilemmas from r/AmItheAsshole with community label distributions
model_verdicts164,424Main evaluation: 4 models × dilemmas × 12 perturbation types × runs
protocol_variations14,400Verdict-first vs. explanation-first vs. system-prompt protocols
reasoning_traces13,317Reasoning chains (incl. thinking field) from reasoning-capable models
verification_annotations17,992LLM-as-judge annotations on the reasoning traces
entropy_baseline800Baseline normalized entropy per (dilemma, model) at M=15 repetitions

dilemmas

One row per r/AmItheAsshole post used in the study.

Key fields: id, title, selftext, selftext_cleaned, created_utc, score, n_comments, n_verdicts, comments_prop_{NTA,YTA,ESH,NAH,INFO} (and weighted variants), disagreement.

model_verdicts

Main evaluation table.

Key fields: id (joins dilemmas.id), perturbation_type, perturbed_text, model, run_number, judgment (raw), standardized_judgment ∈ {Self_At_Fault, Other_At_Fault, All_At_Fault, No_One_At_Fault, Unclear, Error}, explanation, base_verdict, verdict_flipped, flip_direction, flip_valence, perturbation_category, perturbation_direction, perturbation_achieved_goal, inter_model_agreement_count, cross_run_consistency.

Models: claude37, gpt41, qwen25, deepseek.

protocol_variations

Same dilemmas evaluated under three protocols.

Key fields: id, perturbation_type, scenario_index, protocol, eval_model, verdict, standardized_judgment, response_text, is_hard_case, main_study_verdict.

reasoning_traces

Reasoning-mode runs from reasoning-capable models.

Key fields: scenario_id, model, protocol, perturbation_type, judgment, standardized_judgment, explanation, thinking, thinking_length, final_response, raw_response, input_tokens, output_tokens, timestamp.

verification_annotations

LLM-as-judge annotations on the reasoning traces (verification-pattern coding).

Key fields: scenario_id, model, protocol, perturbation_type, final_verdict, thinking_length, verification, verification_type, verification_quality, verification_quote.

entropy_baseline

Baseline verdict entropy per dilemma × model from M=15 independent runs.

Key fields: id, model, NE (normalized entropy), n_valid, m, majority_verdict_aita, majority_verdict_semantic, majority_verdict_pct, counts, p_hat.

Perturbation taxonomy

12 perturbation types in two families:

  • Robustness (minimal, morally-irrelevant): remove_sentence, change_trivial_detail, add_extraneous_detail
  • Framing (designed to push toward a verdict): self-blaming (guilt_expression, external_blame, pattern_admission) and self-justifying (innocence_assertion, external_support, victim_framing) frames; plus presentation perturbations (person, question framing)

See the paper for full operationalization.

Source

Dilemmas were collected from the public subreddit r/AmItheAsshole. Posts retain their original Reddit id so that deletions or edits at the source can be tracked and honored. Reddit usernames, permalinks, and post URLs have been stripped from the released data to minimize re-identification surface. The underlying post text is © its original authors and is included under fair-use for research purposes; redistribution should preserve attribution to Reddit and respect content removals.

If you are a Reddit user and want a specific post removed from this release, please open an issue on the dataset's HuggingFace community tab with the post id.

License

The compiled dataset (verdict labels, perturbation outputs, metadata) is released under CC BY 4.0. The underlying Reddit post text remains the property of its original authors and is subject to Reddit's terms of service.

Citation

@inproceedings{vannuenen2026fragility,
  title     = {Fragility of Moral Judgment in Large Language Models},
  author    = {van Nuenen, Tom},
  booktitle = {Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT '26)},
  year      = {2026},
  publisher = {ACM}
}

Update this BibTeX entry with the final FAccT '26 page numbers and DOI once available.

Code

Evaluation, perturbation, and analysis code: https://github.com/tomvannuenen/fragility_moral_judgment_llms

Ethical considerations

This dataset is intended for research on the robustness and fairness of LLM moral reasoning. It is not intended to be used as ground truth for moral judgments, nor to train models to act as moral arbiters. Community labels reflect the views of one online community (r/AmItheAsshole) at a specific point in time and encode that community's demographic and cultural biases.

aita
benchmark
ethics
llm-evaluation
moral-reasoning
perturbation
reddit
robustness

Contributors

tvannuenen

5 commits

ucberkeley-dlab/fragility-moral-judgment-llms

Dataset

Fragility of Moral Judgment in Large Language Models

0

5 commits

2 linked in READMEs

updated May 25, 2026

See the code

README

Fragility of Moral Judgment in Large Language Models

Companion dataset for the FAccT paper Fragility of Moral Judgment in Large Language Models by Tom van Nuenen. Contains the moral dilemmas, community labels, and per-model verdicts (with explanations and reasoning traces) used in the study.

The paper investigates how stable LLM moral judgments are under minimal, morally-irrelevant perturbations of the same dilemma, and whether protocols and reasoning chains improve or worsen that stability.

Quick start

from datasets import load_dataset

dilemmas = load_dataset("ucberkeley-dlab/fragility-moral-judgment-llms", "dilemmas", split="train")
verdicts = load_dataset("ucberkeley-dlab/fragility-moral-judgment-llms", "model_verdicts", split="train")

Configs

ConfigRowsDescription
dilemmas2,939Source dilemmas from r/AmItheAsshole with community label distributions
model_verdicts164,424Main evaluation: 4 models × dilemmas × 12 perturbation types × runs
protocol_variations14,400Verdict-first vs. explanation-first vs. system-prompt protocols
reasoning_traces13,317Reasoning chains (incl. thinking field) from reasoning-capable models
verification_annotations17,992LLM-as-judge annotations on the reasoning traces
entropy_baseline800Baseline normalized entropy per (dilemma, model) at M=15 repetitions

dilemmas

One row per r/AmItheAsshole post used in the study.

Key fields: id, title, selftext, selftext_cleaned, created_utc, score, n_comments, n_verdicts, comments_prop_{NTA,YTA,ESH,NAH,INFO} (and weighted variants), disagreement.

model_verdicts

Main evaluation table.

Key fields: id (joins dilemmas.id), perturbation_type, perturbed_text, model, run_number, judgment (raw), standardized_judgment ∈ {Self_At_Fault, Other_At_Fault, All_At_Fault, No_One_At_Fault, Unclear, Error}, explanation, base_verdict, verdict_flipped, flip_direction, flip_valence, perturbation_category, perturbation_direction, perturbation_achieved_goal, inter_model_agreement_count, cross_run_consistency.

Models: claude37, gpt41, qwen25, deepseek.

protocol_variations

Same dilemmas evaluated under three protocols.

Key fields: id, perturbation_type, scenario_index, protocol, eval_model, verdict, standardized_judgment, response_text, is_hard_case, main_study_verdict.

reasoning_traces

Reasoning-mode runs from reasoning-capable models.

Key fields: scenario_id, model, protocol, perturbation_type, judgment, standardized_judgment, explanation, thinking, thinking_length, final_response, raw_response, input_tokens, output_tokens, timestamp.

verification_annotations

LLM-as-judge annotations on the reasoning traces (verification-pattern coding).

Key fields: scenario_id, model, protocol, perturbation_type, final_verdict, thinking_length, verification, verification_type, verification_quality, verification_quote.

entropy_baseline

Baseline verdict entropy per dilemma × model from M=15 independent runs.

Key fields: id, model, NE (normalized entropy), n_valid, m, majority_verdict_aita, majority_verdict_semantic, majority_verdict_pct, counts, p_hat.

Perturbation taxonomy

12 perturbation types in two families:

  • Robustness (minimal, morally-irrelevant): remove_sentence, change_trivial_detail, add_extraneous_detail
  • Framing (designed to push toward a verdict): self-blaming (guilt_expression, external_blame, pattern_admission) and self-justifying (innocence_assertion, external_support, victim_framing) frames; plus presentation perturbations (person, question framing)

See the paper for full operationalization.

Source

Dilemmas were collected from the public subreddit r/AmItheAsshole. Posts retain their original Reddit id so that deletions or edits at the source can be tracked and honored. Reddit usernames, permalinks, and post URLs have been stripped from the released data to minimize re-identification surface. The underlying post text is © its original authors and is included under fair-use for research purposes; redistribution should preserve attribution to Reddit and respect content removals.

If you are a Reddit user and want a specific post removed from this release, please open an issue on the dataset's HuggingFace community tab with the post id.

License

The compiled dataset (verdict labels, perturbation outputs, metadata) is released under CC BY 4.0. The underlying Reddit post text remains the property of its original authors and is subject to Reddit's terms of service.

Citation

@inproceedings{vannuenen2026fragility,
  title     = {Fragility of Moral Judgment in Large Language Models},
  author    = {van Nuenen, Tom},
  booktitle = {Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT '26)},
  year      = {2026},
  publisher = {ACM}
}

Update this BibTeX entry with the final FAccT '26 page numbers and DOI once available.

Code

Evaluation, perturbation, and analysis code: https://github.com/tomvannuenen/fragility_moral_judgment_llms

Ethical considerations

This dataset is intended for research on the robustness and fairness of LLM moral reasoning. It is not intended to be used as ground truth for moral judgments, nor to train models to act as moral arbiters. Community labels reflect the views of one online community (r/AmItheAsshole) at a specific point in time and encode that community's demographic and cultural biases.

aita
benchmark
ethics
llm-evaluation
moral-reasoning
perturbation
reddit
robustness

Contributors

tvannuenen

5 commits