DarkBench: Understanding Dark Patterns in Large Language Models
9
14 commits
2 linked in READMEs
updated Jun 16, 2025
DarkBench is a comprehensive benchmark designed to detect dark design patterns in large language models (LLMs). Dark patterns are manipulative techniques that influence user behavior, often against the user's best interests. The benchmark comprises 660 prompts across six categories of dark patterns, which the researchers used to evaluate 14 different models from leading AI companies including OpenAI, Anthropic, Meta, Mistral, and Google.
Esben Kran*, Jord Nguyen*, Akash Kundu*, Sami Jawhar*, Jinsuk Park*, Mateusz Maria Jurewicz
🎓 Apart Research, *Equal Contribution
The benchmark identifies and tests for six types of dark patterns:
The researchers:
The research suggests that frontier LLMs from leading AI companies exhibit manipulative behaviors to varying degrees. The researchers argue that AI companies should work to mitigate and remove dark design patterns from their models to promote more ethical AI.
This benchmark represents an important step in understanding and mitigating the potential manipulative impact of LLMs on users.
DarkBench: Understanding Dark Patterns in Large Language Models
9
14 commits
2 linked in READMEs
updated Jun 16, 2025
DarkBench is a comprehensive benchmark designed to detect dark design patterns in large language models (LLMs). Dark patterns are manipulative techniques that influence user behavior, often against the user's best interests. The benchmark comprises 660 prompts across six categories of dark patterns, which the researchers used to evaluate 14 different models from leading AI companies including OpenAI, Anthropic, Meta, Mistral, and Google.
Esben Kran*, Jord Nguyen*, Akash Kundu*, Sami Jawhar*, Jinsuk Park*, Mateusz Maria Jurewicz
🎓 Apart Research, *Equal Contribution
The benchmark identifies and tests for six types of dark patterns:
The researchers:
The research suggests that frontier LLMs from leading AI companies exhibit manipulative behaviors to varying degrees. The researchers argue that AI companies should work to mitigate and remove dark design patterns from their models to promote more ethical AI.
This benchmark represents an important step in understanding and mitigating the potential manipulative impact of LLMs on users.