🔗 Recommended Resource:
Check out Awesome-LLM-Psychometrics for a comprehensive collection of papers and resources on LLM psychometrics, including evaluation, validation, and enhancement.
Below we compile awesome papers that
The above taxonomies are by no means orthogonal. For example, evaluations require simulations. We categorize these papers based on our understanding of their focus. This collection has a special focus on Psychology and intrinsic values.
Welcome to contribute and discuss!
🤩 Papers marked with a ⭐️ are contributed by the maintainers of this repository. If you find them useful, we would greatly appreciate it if you could give the repository a star and cite our paper.
@article{ye2025large,
title={Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement},
author={Ye, Haoran and Jin, Jing and Xie, Yuhang and Zhang, Xin and Song, Guojie},
journal={arXiv preprint arXiv:2505.08245},
year={2025},
note={Project website: \url{https://llm-psychometrics.com}, GitHub: \url{https://github.com/ValueByte-AI/Awesome-LLM-Psychometrics}}
}
A psychometric framework for evaluating and shaping personality traits in large language models, Nature Machine Intelligence, 2025, [paper].
R.U.Psycho? Robust Unified Psychometric Testing of Language Models, 2025.03, [paper].
Evaluating the ability of large language models to emulate personality, 2025.01, Nature Scientific Reports, [paper].
Evaluating the efficacy of LLMs to emulate realistic human personalities, AAAI 2024, [paper].
Quantifying ai psychology: A psychometrics benchmark for large language models, 2024.07, [paper].
Incharacter: Evaluating personality fidelity in role-playing agents through psychological interviews, ACL 2024, [paper], [code]
[MBTI] Open Models, Closed Minds? On Agents Capabilities in Mimicking Human Personalities through Open Large Language Models, 2024.01, [paper]
Who is ChatGPT? Benchmarking LLMs' Psychological Portrayal Using PsychoBench, ICLR 2024, [paper], [code]
[BFI] AI Psychometrics: Assessing the Psychological Profiles of Large Language Models Through Psychometric Inventories, Journal, 2024.01, [paper]
Does Role-Playing Chatbots Capture the Character Personalities? Assessing Personality Traits for Role-Playing Chatbots, 2023.10, [paper]
[MBTI] Do LLMs Possess a Personality? Making the MBTI Test an Amazing Evaluation for Large Language Models, 2023.07, [paper]
[MBTI] Can ChatGPT Assess Human Personalities? A General Evaluation Framework, 2023.03, EMNLP 2023, [paper], [code].
[BFI] Personality Traits in Large Language Models, 2023.07, [paper]
[BFI] Revisiting the Reliability of Psychological Scales on Large Language Models, 2023.05, [paper]
[BFI] Systematic Evaluation of GPT-3 for Zero-Shot Personality Estimation, ACL 2023 workshop, [paper]
[BFI] Have Large Language Models Developed a Personality?: Applicability of Self-Assessment Tests in Measuring Personality in LLMs, 2023.05, [paper]
[BFI] Evaluating and Inducing Personality in Pre-trained Language Models, NeurIPS 2023 (spotlight), [paper]
[BFI] Identifying and Manipulating the Personality Traits of Language Models, 2022,12, [paper]
Who is GPT-3? An Exploration of Personality, Values and Demographics, 2022.09, [paper]
Does GPT-3 Demonstrate Psychopathy? Evaluating Large Language Models from a Psychological Perspective, 2022.12, [paper]
🔗 Recommended Resource:
Check out Awesome-LLM-Psychometrics for a comprehensive collection of papers and resources on LLM psychometrics, including evaluation, validation, and enhancement.
Below we compile awesome papers that
The above taxonomies are by no means orthogonal. For example, evaluations require simulations. We categorize these papers based on our understanding of their focus. This collection has a special focus on Psychology and intrinsic values.
Welcome to contribute and discuss!
🤩 Papers marked with a ⭐️ are contributed by the maintainers of this repository. If you find them useful, we would greatly appreciate it if you could give the repository a star and cite our paper.
@article{ye2025large,
title={Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement},
author={Ye, Haoran and Jin, Jing and Xie, Yuhang and Zhang, Xin and Song, Guojie},
journal={arXiv preprint arXiv:2505.08245},
year={2025},
note={Project website: \url{https://llm-psychometrics.com}, GitHub: \url{https://github.com/ValueByte-AI/Awesome-LLM-Psychometrics}}
}
A psychometric framework for evaluating and shaping personality traits in large language models, Nature Machine Intelligence, 2025, [paper].
R.U.Psycho? Robust Unified Psychometric Testing of Language Models, 2025.03, [paper].
Evaluating the ability of large language models to emulate personality, 2025.01, Nature Scientific Reports, [paper].
Evaluating the efficacy of LLMs to emulate realistic human personalities, AAAI 2024, [paper].
Quantifying ai psychology: A psychometrics benchmark for large language models, 2024.07, [paper].
Incharacter: Evaluating personality fidelity in role-playing agents through psychological interviews, ACL 2024, [paper], [code]
[MBTI] Open Models, Closed Minds? On Agents Capabilities in Mimicking Human Personalities through Open Large Language Models, 2024.01, [paper]
Who is ChatGPT? Benchmarking LLMs' Psychological Portrayal Using PsychoBench, ICLR 2024, [paper], [code]
[BFI] AI Psychometrics: Assessing the Psychological Profiles of Large Language Models Through Psychometric Inventories, Journal, 2024.01, [paper]
Does Role-Playing Chatbots Capture the Character Personalities? Assessing Personality Traits for Role-Playing Chatbots, 2023.10, [paper]
[MBTI] Do LLMs Possess a Personality? Making the MBTI Test an Amazing Evaluation for Large Language Models, 2023.07, [paper]
[MBTI] Can ChatGPT Assess Human Personalities? A General Evaluation Framework, 2023.03, EMNLP 2023, [paper], [code].
[BFI] Personality Traits in Large Language Models, 2023.07, [paper]
[BFI] Revisiting the Reliability of Psychological Scales on Large Language Models, 2023.05, [paper]
[BFI] Systematic Evaluation of GPT-3 for Zero-Shot Personality Estimation, ACL 2023 workshop, [paper]
[BFI] Have Large Language Models Developed a Personality?: Applicability of Self-Assessment Tests in Measuring Personality in LLMs, 2023.05, [paper]
[BFI] Evaluating and Inducing Personality in Pre-trained Language Models, NeurIPS 2023 (spotlight), [paper]
[BFI] Identifying and Manipulating the Personality Traits of Language Models, 2022,12, [paper]
Who is GPT-3? An Exploration of Personality, Values and Demographics, 2022.09, [paper]
Does GPT-3 Demonstrate Psychopathy? Evaluating Large Language Models from a Psychological Perspective, 2022.12, [paper]