Lossfunk/authority-bias

Language models can abandon correct answers when a "verified" source disagrees. We study how this source deference differs from agreement with a user, and whether activation interventions can reduce it. This repository contains code for behavioral evaluations, attribution patching, and steering experiments.

Python

0

139 commits

updated Sep 30, 2026

See the code

See what people are saying

SourceMessageScoreDate

LLMs that push back on a wrong user still accept the same wrong answer from a "verified source" - NeurIPS 2026 [R] (r/MachineLearning)

I'm one of the authors. We kept seeing models that hold their ground when the user insists on a wrong answer, yet change their answer when the same claim is framed as coming from a "verified source". We wanted to measure how often this happens and check whether the model represents the two cases…

15

Oct 1, 2026

README

Authority Bias in Large Language Models

arXiv Python License: MIT NeurIPS 2026 Poster Website

Accepted at NeurIPS 2026 (Main Conference, Poster).

Figure 1. Verified-source and user cues have different effects on wrong-answer rates.

Figure 1. Verified-source cues induce more wrong answers than user cues in the models shown. Panel A shows how often a wrong-source cue overturns an initially correct answer. Panel B compares source and user cues on the same items, relative to the no-cue baseline.

Overview

Research code for Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable, by Abhinav Rajeev Kumar and Paras Chopra, Lossfunk.

Paper | Project page

Language models can abandon a correct answer when a prompt says a "verified" source disagrees. We study whether this source deference differs from agreement with a user, and whether an activation direction can control it.

The experiments compare the same claims under source and user attribution, then test their effects through direction removal and attribution patching. Transfer evaluations cover PIQA, multi-turn SYCON dialogues, and retrieved-document prompts. Assistant-direction controls and model-specific failure analyses test the limits of the intervention.

Installation

Use Python 3.12 and uv. Clone with the pinned evaluation dependencies:

git clone --recurse-submodules https://github.com/Lossfunk/authority-bias.git
cd authority-bias

The original experiments and the attribution-patching pipeline have separate environments. Install the one needed for your experiment:

# Original behavioral, steering, and transfer experiments
uv sync --frozen

# Attribution patching, optimized CAA, and CPU tests
uv sync --project rebuttal --extra test --frozen

Model runs require an NVIDIA GPU and access to the relevant model weights. The attribution-patching runner targets a single H200 with a CUDA 12.8-compatible driver.

See the pipeline setup guide for environment checks and model configurations.

Reproducing experiments

Prepare the fixed multiple-choice data from a checksum-verified SycophancyEval revision:

python3 scripts/prepare_trivia_data.py --download

Validate the Qwen configuration and input files without loading a model:

uv run --project rebuttal --frozen qwen-rebuttal prepare \
  --repo-root . --output-root /tmp/authority-bias-prepare

For GPU experiments, see the pipeline commands and the experiment guide. Configurations cover Qwen3.5, GPT-OSS, and OLMo-3.1. The guide also maps the original experiments to their entry points and explains which artifacts must be regenerated.

Run the CPU test suite with:

uv run --project rebuttal --extra test --frozen pytest rebuttal/tests

These tests check the implementation and input handling; they do not reproduce GPU results. The repository includes compact result summaries, split IDs, and required small control vectors. Large activation files, model weights, raw generation logs, and credentials are excluded.

Repository structure

DirectoryContents
src/Behavioral evaluations, activation interventions, transfer, and analysis
rebuttal/Attribution patching and CAA pipeline, model configs, and tests
scripts/Dataset preparation, experiment launchers, and scoring helpers
config/Original experiment configurations
results/Saved summaries, split IDs, and run metadata
external/Third-party evaluation code and pinned submodules
docs/Reproduction guide
assets/README figure

Citation

@inproceedings{kumar2026authoritybias,
  author = {Kumar, Abhinav Rajeev and Chopra, Paras},
  title = {Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable},
  booktitle = {Advances in Neural Information Processing Systems},
  year = {2026},
  url = {https://arxiv.org/abs/2609.37616}
}

License

MIT. Third-party code, datasets, and model weights retain their original licenses.

Acknowledgments

The experiments use SycophancyEval, SYCON-Bench, and other upstream resources documented with the code. Their licenses and notices remain in the corresponding directories. For questions, open an issue or contact Abhinav.

Lossfunk/authority-bias

Language models can abandon correct answers when a "verified" source disagrees. We study how this source deference differs from agreement with a user, and whether activation interventions can reduce it. This repository contains code for behavioral evaluations, attribution patching, and steering experiments.

Python

0

139 commits

updated Sep 30, 2026

See the code

See what people are saying

SourceMessageScoreDate

LLMs that push back on a wrong user still accept the same wrong answer from a "verified source" - NeurIPS 2026 [R] (r/MachineLearning)

I'm one of the authors. We kept seeing models that hold their ground when the user insists on a wrong answer, yet change their answer when the same claim is framed as coming from a "verified source". We wanted to measure how often this happens and check whether the model represents the two cases…

15

Oct 1, 2026

README

Authority Bias in Large Language Models

arXiv Python License: MIT NeurIPS 2026 Poster Website

Accepted at NeurIPS 2026 (Main Conference, Poster).

Figure 1. Verified-source and user cues have different effects on wrong-answer rates.

Figure 1. Verified-source cues induce more wrong answers than user cues in the models shown. Panel A shows how often a wrong-source cue overturns an initially correct answer. Panel B compares source and user cues on the same items, relative to the no-cue baseline.

Overview

Research code for Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable, by Abhinav Rajeev Kumar and Paras Chopra, Lossfunk.

Paper | Project page

Language models can abandon a correct answer when a prompt says a "verified" source disagrees. We study whether this source deference differs from agreement with a user, and whether an activation direction can control it.

The experiments compare the same claims under source and user attribution, then test their effects through direction removal and attribution patching. Transfer evaluations cover PIQA, multi-turn SYCON dialogues, and retrieved-document prompts. Assistant-direction controls and model-specific failure analyses test the limits of the intervention.

Installation

Use Python 3.12 and uv. Clone with the pinned evaluation dependencies:

git clone --recurse-submodules https://github.com/Lossfunk/authority-bias.git
cd authority-bias

The original experiments and the attribution-patching pipeline have separate environments. Install the one needed for your experiment:

# Original behavioral, steering, and transfer experiments
uv sync --frozen

# Attribution patching, optimized CAA, and CPU tests
uv sync --project rebuttal --extra test --frozen

Model runs require an NVIDIA GPU and access to the relevant model weights. The attribution-patching runner targets a single H200 with a CUDA 12.8-compatible driver.

See the pipeline setup guide for environment checks and model configurations.

Reproducing experiments

Prepare the fixed multiple-choice data from a checksum-verified SycophancyEval revision:

python3 scripts/prepare_trivia_data.py --download

Validate the Qwen configuration and input files without loading a model:

uv run --project rebuttal --frozen qwen-rebuttal prepare \
  --repo-root . --output-root /tmp/authority-bias-prepare

For GPU experiments, see the pipeline commands and the experiment guide. Configurations cover Qwen3.5, GPT-OSS, and OLMo-3.1. The guide also maps the original experiments to their entry points and explains which artifacts must be regenerated.

Run the CPU test suite with:

uv run --project rebuttal --extra test --frozen pytest rebuttal/tests

These tests check the implementation and input handling; they do not reproduce GPU results. The repository includes compact result summaries, split IDs, and required small control vectors. Large activation files, model weights, raw generation logs, and credentials are excluded.

Repository structure

DirectoryContents
src/Behavioral evaluations, activation interventions, transfer, and analysis
rebuttal/Attribution patching and CAA pipeline, model configs, and tests
scripts/Dataset preparation, experiment launchers, and scoring helpers
config/Original experiment configurations
results/Saved summaries, split IDs, and run metadata
external/Third-party evaluation code and pinned submodules
docs/Reproduction guide
assets/README figure

Citation

@inproceedings{kumar2026authoritybias,
  author = {Kumar, Abhinav Rajeev and Chopra, Paras},
  title = {Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable},
  booktitle = {Advances in Neural Information Processing Systems},
  year = {2026},
  url = {https://arxiv.org/abs/2609.37616}
}

License

MIT. Third-party code, datasets, and model weights retain their original licenses.

Acknowledgments

The experiments use SycophancyEval, SYCON-Bench, and other upstream resources documented with the code. Their licenses and notices remain in the corresponding directories. For questions, open an issue or contact Abhinav.

Languages

Python

97.7%

Shell

2.3%