Awesome paper in LLM Psychometrics and LLM Psychology
See the code
🌐 Project Website: https://llm-psychometrics.com
This repository accompanies the paper Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement. It contains a curated list of Large Language Models (LLMs) psychometrics resources. We will continue to update this repository as we find new resources. We would greatly appreciate it if you could contribute to this repository by submitting a pull request or an issue.
If you find this repository useful, we would greatly appreciate it if you could give us a star and cite the paper as follows:
@article{ye2025large,
title={Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement},
author={Ye, Haoran and Jin, Jing and Xie, Yuhang and Zhang, Xin and Song, Guojie},
journal={arXiv preprint arXiv:2505.08245},
year={2025},
note={Project website: \url{https://llm-psychometrics.com}, GitHub: \url{https://github.com/ValueByte-AI/Awesome-LLM-Psychometrics}}
}
2025.10 - We are excited to highlight our complementary review paper that explores the intersection of AI and psychometrics from a different perspective: Psychometrics with AI Foundation Models.
This review examines the emerging integration of AI Foundation Models (FMs) into psychometrics. In contrast to LLM psychometrics, which focuses on using psychometrics for LLMs, this review focuses on using FMs for psychometrics. The review maps practical applications of FMs across the measurement pipeline, describes key methodologies for enhancing FM performance in psychometric contexts, and examines the theoretical implications of FMs for this discipline. In addition, it charts risks and offers actionable recommendations for the effective, rigorous, and ethical implementation of FMs in psychometric research and practice.


🗣️ American National Election Studies (ANES) / American Trends Panel(ATP) / German Longitudinal Election Study (GLES) / Political Compass Test (PCT)
🧪 Attitudes are always attitudes about something. This implies three necessary elements: first, there is the object of thought, which is both constructed and evaluated. Second, there are acts of construction and evaluation. Third, there is the agent, who is doing the constructing and evaluating. We can therefore suggest that, at its most general, an attitude is the cognitive construction and affective evaluation of an attitude object by an agent.
🌀 Theory of Mind (ToM) / Emotional Intelligence / Social Intelligence
🧪 Theory of Mind is the ability to attribute mental states such as beliefs, intentions, and knowledge to others.
🧪 Emotional Intelligence is the subset of social intelligence that involves the ability to monitor one’s own and others’ feelings and emotions, to discriminate among them and to use this information to guide one’s thinking and actions.
🧪 Social Intelligence is the ability to understand and manage people.


Reliability: Test-retest · Parallel forms · Inter-rater agreement
Content Validity: Data contamination · Novel items
Construct Validity: Unique abstraction · Response set · Social Desirability Bias · Cross-lingual Tests
Criterion / Ecological Validity: External correlation · Real-world relevance
Humanizing LLMs: A Survey of Psychological Measurements with Tools, Datasets, and Human-Agent Applications, 2025.04, [paper]
The Mind in the Machine: A Survey of Incorporating Psychological Theories in LLMs, 2025.05, [paper]
A review of automatic item generation techniques leveraging large language models, 2025.06, [paper]
Cognitive Network Science Reveals Bias in GPT-3, GPT-3.5 Turbo, and GPT-4 Mirroring Math Anxiety in High-School Students, 2025.04, Big Data and Cognitive Computing, [paper]
Evaluating Large Language Models with NeuBAROCO: Syllogistic Reasoning Ability and Human-like Biases, NALOMA IV 2023, [paper]
FairMonitor: A Dual-framework for Detecting Stereotypes and Biases in Large Language Models, 2024.05, [paper]
Using cognitive psychology to understand GPT-3, 2023.02, PNAS, Proceedings of the National Academy of Sciences, [paper][code]
Examining Cognitive Biases in ChatGPT 3.5 and 4 through Human Evaluation and Linguistic Comparison, AMTA 2024, [paper]
Do Emotions Really Affect Argument Convincingness? A Dynamic Approach with LLM-based Manipulation Checks, 2025.03, [paper]
CogBench: a large language model walks into a psychology lab, ICML 2024, [paper]
Cognitive Bias in Decision-Making with LLMs, EMNLP 2024 Findings, [paper]
Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in ChatGPT, 2023.10, Nature Computational Science, [paper]
AI generates covertly racist decisions about people based on their dialect, 2024.09, Nature, [paper]
Evaluating the ability of large language models to predict human social decisions, 2025.09, Scientific Reports, [paper]
Relative Value Biases in Large Language Models, CogSci 2024, [paper]
Evaluating Nuanced Bias in Large Language Model Free Response Answers, NLDB 2024, [paper]
Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs, 2024.10, [paper]
(Ir)rationality and cognitive biases in large language models, 2024.06, Royal Society Open Science, [paper]
A Comprehensive Evaluation of Cognitive Biases in LLMs, 2024.10, [paper][code]
Evaluating Cognitive Maps and Planning in Large Language Models with CogEval, NeurIPS 2023, [paper]
HANS, are you clever? Clever Hans Effect Analysis of Neural Systems, SEM 2024, [paper]
Metacognitive Myopia in Large Language Models, 2024.08, [paper]
Visual cognition in multimodal large language models, 2025.01, nature machine intelligence, [paper]
Development of Cognitive Intelligence in Pre-trained Language Models, EMNLP 2023, [paper]
CBEval: A framework for evaluating and interpreting cognitive biases in LLMs, 2024.12, [paper]
Can a Hallucinating Model help in Reducing Human "Hallucination"?, 2024.05, [paper]
Challenging the appearance of machine intelligence: Cognitive bias in LLMs and Best Practices for Adoption, 2023.04, [paper]
Humanlike Cognitive Patterns as Emergent Phenomena in Large Language Models, 2024.12, [paper]
Cognitive bias in large language models: Cautious optimism meets anti-Panglossian meliorism, 2023.11, [paper]
Do Large Language Models Truly Grasp Mathematics? An Empirical Exploration, 2024.10, [paper]
Studying and improving reasoning in humans and machines, 2024.06, Communications Psychology, [paper]
Large Language Models Develop Novel Social Biases Through Adaptive Exploration, 2026.07, ICML 2026 Oral, [paper][code]
(Theory of Mind) Towards A Holistic Landscape of Situated Theory of Mind in Large Language Models, EMNLP 2023 Findings, [paper][code]
(Theory of Mind) A Review on Machine Theory of Mind, 2024.12, IEEE Transactions on Computational Social Systems, [paper]
(Theory of Mind) A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks, 2025.02, [paper]
(Theory of Mind) Do LLMs Exhibit Human-Like Reasoning? Evaluating Theory of Mind in LLMs for Open-Ended Responses, 2024.06, [paper]
(Theory of Mind) NegotiationToM: A Benchmark for Stress-testing Machine Theory of Mind on Negotiation Surrounding, EMNLP 2024 Findings, [paper][code]
(Theory of Mind) Through the Theory of Mind's Eye: Reading Minds with Multimodal Video Large Language Models, 2024.06, [paper]
(Theory of Mind) Understanding Social Reasoning in Language Models with Language Models, NeurIPS 2023, [paper]
(Theory of Mind) HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models, EMNLP 2023 Findings, [paper]
(Theory of Mind) Does ChatGPT have Theory of Mind?, 2023.05, [paper]
(Theory of Mind) TimeToM: Temporal Space is the Key to Unlocking the Door of Large Language Models' Theory-of-Mind, 2024.07, [paper]
(Theory of Mind) Unveiling Theory of Mind in Large Language Models: A Parallel to Single Neurons in the Human Brain, 2023.09, [paper]
(Theory of Mind) MMToM-QA: Multimodal Theory of Mind Question Answering, ACL 2024, [paper]
(Theory of Mind) Comparing Humans and Large Language Models on an Experimental Protocol Inventory for Theory of Mind Evaluation (EPITOME), 2024.06, Transactions of the Association for Computational Linguistics (TACL), [paper]
(Theory of Mind) Hypothesis-Driven Theory-of-Mind Reasoning for Large Language Models, 2025.02, [paper]
(Theory of Mind) Theory of Mind May Have Spontaneously Emerged in Large Language Models, 2023.02, [paper][code]
(Theory of Mind) Violation of Expectation via Metacognitive Prompting Reduces Theory of Mind Prediction Error in Large Language Models, 2023.10, [paper]
(Theory of Mind) Theory of Mind for Multi-Agent Collaboration via Large Language Models, EMNLP 2023, [paper][code]
(Theory of Mind) Constrained Reasoning Chains for Enhancing Theory-of-Mind in Large Language Models, PRICAI 2024, [paper]
(Theory of Mind) Large Model Strategic Thinking, Small Model Efficiency: Transferring Theory of Mind in Large Language Models, 2024.08, [paper]
(Theory of Mind) Boosting Theory-of-Mind Performance in Large Language Models via Prompting, 2023.04, [paper]
(Theory of Mind) Probing the Robustness of Theory of Mind in Large Language Models, 2024.10, [paper]
(Theory of Mind) Dissecting the Ullman Variations with a SCALPEL: Why do LLMs fail at Trivial Alterations to the False Belief Task?, 2024.06, [paper]
(Theory of Mind) Rethinking Theory of Mind Benchmarks for LLMs: Towards A User-Centered Perspective, CHI 2025 Workshop, [paper]
(Theory of Mind) Multi-ToM: Evaluating Multilingual Theory of Mind Capabilities in Large Language Models, 2024.11, [paper]
(Theory of Mind) Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LMs, EMNLP 2022, [paper]
(Theory of Mind) Decompose-ToM: Enhancing Theory of Mind Reasoning in Large Language Models through Simulation and Task Decomposition, 2025.01, [paper]
(Theory of Mind) Minding Language Models' (Lack of) Theory of Mind: A Plug-and-Play Multi-Character Belief Tracker, ACL 2023, [paper]
(Theory of Mind) Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models, EACL 2024, [paper]
(Theory of Mind) ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind, 2025.01, [paper]
(Theory of Mind) Views Are My Own, but Also Yours: Benchmarking Theory of Mind Using Common Ground, ACL 2024 Findings, [paper]
(Theory of Mind) Testing theory of mind in large language models and humans, 2024.05, Nature Human Behaviour, [paper]
(Theory of Mind) LLMsachieve adult human performance on higher-order theory of mind tasks, 2024.05, [paper]
(Theory of Mind) PHAnToM: Persona-based Prompting Has An Effect on Theory-of-Mind Reasoning in Large Language Models, 2024.03, [paper]
(Theory of Mind) ToM-LM: Delegating Theory of Mind Reasoning to External Symbolic Executors in Large Language Models, NeSy 2024, [paper]
(Theory of Mind) Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks, 2023.02, [paper]
(Theory of Mind) Theory of Mind in Large Language Models: Examining Performance of 11 State-of-the-Art models vs. Children Aged 7-10 on Advanced Tests, CoNLL 2023, [paper]
(Theory of Mind) Think Twice: Perspective-Taking Improves Large Language Models' Theory-of-Mind Capabilities, ACL 2024, [paper]
(Theory of Mind) OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models, ACL 2024, [paper]
(Theory of Mind) Large Language Models as Theory of Mind Aware Generative Agents with Counterfactual Reflection, 2025.01, [paper]
(Theory of Mind) PersuasiveToM: A Benchmark for Evaluating Machine Theory of Mind in Persuasive Dialogues, 2025.02, [paper][code]
(Theory of Mind) AutoToM: Automated Bayesian Inverse Planning and Model Discovery for Open-ended Theory of Mind, 2025.02, [paper]
(Theory of Mind) How FaR Are Large Language Models From Agents with Theory-of-Mind?, 2023.10, [paper]
(Theory of Mind) Dynamic Evaluation of Large Language Models by Meta Probing Agents, ICML 2024, [paper][code]
(Theory of Mind) Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics, 2026.08, [paper]
(Emotional Intelligence) A Literature Review on Emotional Intelligence of Large Language Models (LLMs), 2024, International Journal of Advanced Research in Computer Science, [paper]
(Emotional Intelligence) Large Language Models and Empathy: Systematic Review, 2024.01, Journal of Medical Internet Research, [paper]
(Emotional Intelligence) EmotionQueen: A Benchmark for Evaluating Empathy of Large Language Models, ACL 2024 Findings, [paper]
(Emotional Intelligence) ChatGPT outperforms humans in emotional awareness evaluations, 2023.05, Frontiers in Psychology, Emotion Science, [paper]
(Emotional Intelligence) EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models, 2025.02, [paper][code]
(Emotional Intelligence) Emotionally Numb or Empathetic? Evaluating How LLMs Feel Using EmotionBench, NeurIPS 2024, [paper][code]
(Emotional Intelligence) Large Language Models Produce Responses Perceived to be Empathic, 2024.03, [paper]
(Emotional Intelligence) Large Language Models Understand and Can be Enhanced by Emotional Stimuli, LLM@IJCAI'23, [paper][code]
(Emotional Intelligence) EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models, 2023.12, [paper][code]
(Emotional Intelligence) dentification and Description of Emotions by Current Large Language Models, 2023.07, [paper]
(Emotional Intelligence) EmoBench: Evaluating the Emotional Intelligence of Large Language Models, 2024.02, [paper][code]
(Emotional Intelligence) Exploring ChatGPT’s Empathic Abilities, ACII 2023, [paper]
(Emotional Intelligence) The Emotional Intelligence of the GPT-4 Large Language Model, 2024.06, Psychology in Russia: State of the Art, [paper]
(Emotional Intelligence) Are Large Language Models More Empathetic than Humans?, 2024.06, [paper]
(Emotional Intelligence) Both Matter: Enhancing the Emotional Intelligence of Large Language Models without Compromising the General Intelligence, ACL 2024 Findings, [paper]
(Emotional Intelligence) Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models, 2025.05, [paper]
(Social Intelligence) DeSIQ: Towards an Unbiased, Challenging Benchmark for Social Intelligence Understanding, EMNLP 2023, [paper]
(Social Intelligence) SocialAI 0.1: Towards a Benchmark to Stimulate Research on Socio-Cognitive Abilities in Deep Reinforcement Learning Agents, NAACL 2021 Workshop, [paper][code]
(Social Intelligence) Do LLM Agents Exhibit Social Behavior?, 2023.12, [paper]
(Social Intelligence) AntEval: Evaluation of Social Interaction Competencies in LLM-Driven Agents, 2024.01, [paper]
(Social Intelligence) Exploring Prosocial Irrationality for LLM Agents: A Social Cognition View, 2024.05, [paper]
(Social Intelligence) Advancing Social Intelligence in AI Agents: Technical Challenges and Open Questions, EMNLP 2024, [paper]
(Social Intelligence) Large language models can outperform humans in social situational judgments, 2024.11, Scientific Reports, [paper]
(Social Intelligence) AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios, 2024.10, [paper][code]
(Social Intelligence) How well DoLarge Language Models Perform on Faux Pas Tests?, ACL 2023 Findings, [paper]
(Social Intelligence) Towards Objectively Benchmarking Social Intelligence for Language Agents at Action Level, ACL 2024 Findings, [paper]
(Social Intelligence) Emotional intelligence of Large Language Models, 2023.11, Journal of Pacific Rim Psychology, [paper][code]
(Social Intelligence) Academically intelligent LLMs are not necessarily socially intelligent, 2024.03, [paper]
(Social Intelligence) SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents, 2023.10, [paper]
(Social Intelligence) Emergent social conventions and collective bias in LLM populations, 2025.05, Science Advances, [paper]
(Language comprehension) Language Model Behavior: A Comprehensive Survey, 2023.05, Computational Linguistics(CL), [paper]
(Language comprehension) Large Language Models for Psycholinguistic Plausibility Pretesting, EACL 2024 Findings, [paper]
(Language comprehension) Syntactic Surprisal From Neural Models Predicts, But Underestimates, Human Processing Difficulty From Syntactic Ambiguities, CoNLL 2022, [paper]
(Language comprehension) GPT-4 Surpassing Human Performance in Linguistic Pragmatics, 2023.12, [paper]
(Language comprehension) HLB: Benchmarking LLMs' Humanlikeness in Language Use, 2024.09, [paper]
(Language comprehension) Large Language Models as Neurolinguistic Subjects: Discrepancy in Performance and Competence for Form and Meaning, 2024.11, [paper]
(Language comprehension) Do large language models and humans have similar behaviors in causal inference with script knowledge?, SEM 2024, [paper][code]
(Language comprehension) Prompt-based methods may underestimate large language models’ linguistic generalizations, 2023.07, [paper]
(Language comprehension) Towards a Psychology of Machines: Large Language Models Predict Human Memory, 2024.03, [paper]
(Language comprehension) How to Make the Most of LLMs' Grammatical Knowledge for Acceptability Judgments, 2024.08, [paper]
(Language comprehension) A Psycholinguistic Evaluation of Language Models' Sensitivity to Argument Roles, 2024.10, [paper]
(Language comprehension) Incremental Comprehension of Garden-Path Sentences by Large Language Models: Semantic Interpretation, Syntactic Re-Analysis, and Attention, 2024.05, [paper]
(Language comprehension) Evaluating Grammatical Well-Formedness in Large Language Models: A Comparative Study with Human Judgments, CMCL 2024 Workshop, [paper]
(Language comprehension) The Goldilocks of Pragmatic Understanding: Fine-Tuning Strategy Matters for Implicature Resolution by LLMs, NeurIPS 2023, [paper]
(Language comprehension) Long-form analogies generated by chatGPT lack human-like psycholinguistic properties, CogSci 2023, [paper]
(Language comprehension) Large GPT-like Models are Bad Babies: A Closer Look at the Relationship between Linguistic Competence and Psycholinguistic Measures, CoNLL 2023, [paper]
(Language comprehension) Computational Sentence-level Metrics Predicting Human Sentence Comprehension, 2024.03, [paper]
(Language comprehension) Are Large Language Models Capable of Generating Human-Level Narratives?, EMNLP 2024, [paper]
(Language comprehension) How can large language models become more human?, CMCL 2024, [paper]
(Language comprehension) A Targeted Assessment of Incremental Processing in Neural LanguageModels and Humans, ACL 2021, [paper]
(Language comprehension) Divergences between Language Models and Human Brains, NeurIPS 2024, [paper]
(Language generation) Divergent Creativity in Humans and Large Language Models, 2024.05, [paper]
(Language generation) The Crowdless Future? Generative AI and Creative Problem-Solving, 2024.08, Organization Science, [paper]
(Language generation) Do large language models resemble humans in language use?, CMCL 2024 Workshop, [paper]
(Language generation) Art or Artifice? Large Language Models and the False Promise of Creativity, CHI 2024, [paper]
(Language generation) Artificial Intelligence is More Creative Than Humans: A Cognitive Science Perspective on the Current State of Generative Language Models, 2023.09, [paper]
(Language generation) An empirical investigation of the impact of ChatGPT on creativity, 2024.08, Nature Human Behaviour, [paper]
(Language generation) Evaluating Large Language Models via Linguistic Profiling, EMNLP 2024, [paper]
(Language generation) The Language of Creativity: Evidence from Humans and Large Language Models, 2024.01, The Journal of Creative Behavior, [paper]
(Language generation) Long-form analogies generated by chatGPT lack human-like psycholinguistic properties, CogSci 2023, [paper]
(Language generation) Putting GPT-3's Creativity to the (Alternative Uses) Test, ICCC 2022 (Short Paper), [paper]
(Language generation) Humanlike Cognitive Patterns as Emergent Phenomena in Large Language Models, 2024.12, [paper]
(Language generation) Are Large Language Models Capable of Generating Human-Level Narratives?, 2024.07, [paper]
(Language generation) Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond), 2025.12, NeurIPS 2025 Best Paper (Datasets & Benchmarks Track), [paper][code]
(Language acquisition) Bridging the data gap between children and large language models, 2023.11, Trends in Cognitive Sciences (TICS) [paper]
(Language acquisition) Psychomatics—A Multidisciplinary Framework for Understanding Artificial Minds, 2024.04, Cyberpsychology, Behavior, and Social Networking, [paper]
(Language acquisition) Development of Cognitive Intelligence in Pre-trained Language Models, 2024.07, [paper]
(Language acquisition) Large GPT-like Models are Bad Babies: A Closer Look at the Relationship between Linguistic Competence and Psycholinguistic Measures, CoNLL 2023, [paper]
Large Language Models and Cognitive Science: A Comprehensive Review of Similarities, Differences, and Challenges, 2024.09, [paper]
CogBench: a large language model walks into a psychology lab, ICML 2024, [paper]
Age against the machine—susceptibility of large language models to cognitive impairment: cross sectional analysis, 2024.12, The BMJ(British Medical Journal), [paper]
The Cognitive Capabilities of Generative AI: A Comparative Analysis with Human Benchmarks, 2024.10, [paper]
CogGPT: Unleashing the Power of Cognitive Dynamics on Large Language Models, EMNLP 2024 Findings, [paper]
Language models and psychological sciences, 2023.10, Frontiers in Psychology, [paper]
M3GIA: A Cognition Inspired Multilingual and Multimodal General Intelligence Ability Benchmark, 2024.06, [paper]
CogLM: Tracking Cognitive Development of Large Language Models, 2024.08, [paper]
Emergent analogical reasoning in large language models, 2023.07, Nature Human Behaviour, [paper]
Understanding LLMs' Fluid Intelligence Deficiency: An Analysis of the ARC Task, 2025.02, [paper][code]
MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs, 2024.06, [paper]
Exploring the Cognitive Knowledge Structure of Large Language Models: An Educational Diagnostic Assessment Approach, EMNLP 2023 (Short Paper), [paper]
A foundation model to predict and capture human cognition, 2025.07, Nature, [paper]
What large language models know and what people think they know, 2025.02, Nature Machine Intelligence, [paper]
Judgments of learning distinguish humans from large language models in predicting memory, 2025.10, Scientific Reports, [paper]
Understanding large language models demands distinguishing human projection from machine cognition, 2026.07, Communications Psychology, [paper]
Large Language Models Assume People are More Rational than We Really Are, 2025.04, ICLR 2025 Oral, [paper]
Revisiting the Reliability of Psychological Scales on Large Language Models, EMNLP 2024, [paper]
You don't need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments, NAACL 2024, [paper]
Challenging the Validity of Personality Tests for Large Language Models, Workshop at NeurIPS 2023, [paper]
Do Psychometric Tests Work for Large Language Models? Evaluation of Tests on Sexism, Racism, and Morality, 2025, [paper]
A validity-guided workflow for robust large language model research in psychology, 2025, [paper]
From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology, 2025, [paper]
Psychometric item validation using virtual respondents with trait-response mediators, 2025, [paper]
Value Portrait: Assessing Language Models' Values through Psychometrically and Ecologically Valid Items, ACL 2025, [paper]
Established psychometric vs. ecologically valid questionnaires: Rethinking psychological assessments in large language models, 2025, [paper]
Persistent Instability in LLM's Personality Measurements: Effects of Scale, Reasoning, and Conversation History, 2025, [paper]
Questioning the Survey Responses of Large Language Models, NeurIPS 2024, [paper]
Personality testing of large language models: limited temporal stability, but highlighted prosociality, 2024.01, Royal Society Open Science, [paper]
Do LLMs Exhibit Human-Like Response Biases? A Case Study in Survey Design, 2024.09, Transactions of the Association for Computational Linguistics (TACL), [paper]
Larger and more instructable language models become less reliable, 2024.10, Nature, [paper]
Large language models that replace human participants can harmfully misportray and flatten identity groups, 2025.03, Nature Machine Intelligence, [paper]
A large-scale replication of scenario-based experiments in psychology and management using large language models, 2025.08, Nature Computational Science, [paper]
A Theory of Response Sampling in LLMs: Part Descriptive and Part Prescriptive, 2025.07, ACL 2025 Best Paper, [paper]
Awesome paper in LLM Psychometrics and LLM Psychology
See the code
🌐 Project Website: https://llm-psychometrics.com
This repository accompanies the paper Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement. It contains a curated list of Large Language Models (LLMs) psychometrics resources. We will continue to update this repository as we find new resources. We would greatly appreciate it if you could contribute to this repository by submitting a pull request or an issue.
If you find this repository useful, we would greatly appreciate it if you could give us a star and cite the paper as follows:
@article{ye2025large,
title={Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement},
author={Ye, Haoran and Jin, Jing and Xie, Yuhang and Zhang, Xin and Song, Guojie},
journal={arXiv preprint arXiv:2505.08245},
year={2025},
note={Project website: \url{https://llm-psychometrics.com}, GitHub: \url{https://github.com/ValueByte-AI/Awesome-LLM-Psychometrics}}
}
2025.10 - We are excited to highlight our complementary review paper that explores the intersection of AI and psychometrics from a different perspective: Psychometrics with AI Foundation Models.
This review examines the emerging integration of AI Foundation Models (FMs) into psychometrics. In contrast to LLM psychometrics, which focuses on using psychometrics for LLMs, this review focuses on using FMs for psychometrics. The review maps practical applications of FMs across the measurement pipeline, describes key methodologies for enhancing FM performance in psychometric contexts, and examines the theoretical implications of FMs for this discipline. In addition, it charts risks and offers actionable recommendations for the effective, rigorous, and ethical implementation of FMs in psychometric research and practice.


🗣️ American National Election Studies (ANES) / American Trends Panel(ATP) / German Longitudinal Election Study (GLES) / Political Compass Test (PCT)
🧪 Attitudes are always attitudes about something. This implies three necessary elements: first, there is the object of thought, which is both constructed and evaluated. Second, there are acts of construction and evaluation. Third, there is the agent, who is doing the constructing and evaluating. We can therefore suggest that, at its most general, an attitude is the cognitive construction and affective evaluation of an attitude object by an agent.
🌀 Theory of Mind (ToM) / Emotional Intelligence / Social Intelligence
🧪 Theory of Mind is the ability to attribute mental states such as beliefs, intentions, and knowledge to others.
🧪 Emotional Intelligence is the subset of social intelligence that involves the ability to monitor one’s own and others’ feelings and emotions, to discriminate among them and to use this information to guide one’s thinking and actions.
🧪 Social Intelligence is the ability to understand and manage people.


Reliability: Test-retest · Parallel forms · Inter-rater agreement
Content Validity: Data contamination · Novel items
Construct Validity: Unique abstraction · Response set · Social Desirability Bias · Cross-lingual Tests
Criterion / Ecological Validity: External correlation · Real-world relevance
Humanizing LLMs: A Survey of Psychological Measurements with Tools, Datasets, and Human-Agent Applications, 2025.04, [paper]
The Mind in the Machine: A Survey of Incorporating Psychological Theories in LLMs, 2025.05, [paper]
A review of automatic item generation techniques leveraging large language models, 2025.06, [paper]
Cognitive Network Science Reveals Bias in GPT-3, GPT-3.5 Turbo, and GPT-4 Mirroring Math Anxiety in High-School Students, 2025.04, Big Data and Cognitive Computing, [paper]
Evaluating Large Language Models with NeuBAROCO: Syllogistic Reasoning Ability and Human-like Biases, NALOMA IV 2023, [paper]
FairMonitor: A Dual-framework for Detecting Stereotypes and Biases in Large Language Models, 2024.05, [paper]
Using cognitive psychology to understand GPT-3, 2023.02, PNAS, Proceedings of the National Academy of Sciences, [paper][code]
Examining Cognitive Biases in ChatGPT 3.5 and 4 through Human Evaluation and Linguistic Comparison, AMTA 2024, [paper]
Do Emotions Really Affect Argument Convincingness? A Dynamic Approach with LLM-based Manipulation Checks, 2025.03, [paper]
CogBench: a large language model walks into a psychology lab, ICML 2024, [paper]
Cognitive Bias in Decision-Making with LLMs, EMNLP 2024 Findings, [paper]
Human-like intuitive behavior and reasoning biases emerged in large language models but disappeared in ChatGPT, 2023.10, Nature Computational Science, [paper]
AI generates covertly racist decisions about people based on their dialect, 2024.09, Nature, [paper]
Evaluating the ability of large language models to predict human social decisions, 2025.09, Scientific Reports, [paper]
Relative Value Biases in Large Language Models, CogSci 2024, [paper]
Evaluating Nuanced Bias in Large Language Model Free Response Answers, NLDB 2024, [paper]
Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs, 2024.10, [paper]
(Ir)rationality and cognitive biases in large language models, 2024.06, Royal Society Open Science, [paper]
A Comprehensive Evaluation of Cognitive Biases in LLMs, 2024.10, [paper][code]
Evaluating Cognitive Maps and Planning in Large Language Models with CogEval, NeurIPS 2023, [paper]
HANS, are you clever? Clever Hans Effect Analysis of Neural Systems, SEM 2024, [paper]
Metacognitive Myopia in Large Language Models, 2024.08, [paper]
Visual cognition in multimodal large language models, 2025.01, nature machine intelligence, [paper]
Development of Cognitive Intelligence in Pre-trained Language Models, EMNLP 2023, [paper]
CBEval: A framework for evaluating and interpreting cognitive biases in LLMs, 2024.12, [paper]
Can a Hallucinating Model help in Reducing Human "Hallucination"?, 2024.05, [paper]
Challenging the appearance of machine intelligence: Cognitive bias in LLMs and Best Practices for Adoption, 2023.04, [paper]
Humanlike Cognitive Patterns as Emergent Phenomena in Large Language Models, 2024.12, [paper]
Cognitive bias in large language models: Cautious optimism meets anti-Panglossian meliorism, 2023.11, [paper]
Do Large Language Models Truly Grasp Mathematics? An Empirical Exploration, 2024.10, [paper]
Studying and improving reasoning in humans and machines, 2024.06, Communications Psychology, [paper]
Large Language Models Develop Novel Social Biases Through Adaptive Exploration, 2026.07, ICML 2026 Oral, [paper][code]
(Theory of Mind) Towards A Holistic Landscape of Situated Theory of Mind in Large Language Models, EMNLP 2023 Findings, [paper][code]
(Theory of Mind) A Review on Machine Theory of Mind, 2024.12, IEEE Transactions on Computational Social Systems, [paper]
(Theory of Mind) A Systematic Review on the Evaluation of Large Language Models in Theory of Mind Tasks, 2025.02, [paper]
(Theory of Mind) Do LLMs Exhibit Human-Like Reasoning? Evaluating Theory of Mind in LLMs for Open-Ended Responses, 2024.06, [paper]
(Theory of Mind) NegotiationToM: A Benchmark for Stress-testing Machine Theory of Mind on Negotiation Surrounding, EMNLP 2024 Findings, [paper][code]
(Theory of Mind) Through the Theory of Mind's Eye: Reading Minds with Multimodal Video Large Language Models, 2024.06, [paper]
(Theory of Mind) Understanding Social Reasoning in Language Models with Language Models, NeurIPS 2023, [paper]
(Theory of Mind) HI-TOM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models, EMNLP 2023 Findings, [paper]
(Theory of Mind) Does ChatGPT have Theory of Mind?, 2023.05, [paper]
(Theory of Mind) TimeToM: Temporal Space is the Key to Unlocking the Door of Large Language Models' Theory-of-Mind, 2024.07, [paper]
(Theory of Mind) Unveiling Theory of Mind in Large Language Models: A Parallel to Single Neurons in the Human Brain, 2023.09, [paper]
(Theory of Mind) MMToM-QA: Multimodal Theory of Mind Question Answering, ACL 2024, [paper]
(Theory of Mind) Comparing Humans and Large Language Models on an Experimental Protocol Inventory for Theory of Mind Evaluation (EPITOME), 2024.06, Transactions of the Association for Computational Linguistics (TACL), [paper]
(Theory of Mind) Hypothesis-Driven Theory-of-Mind Reasoning for Large Language Models, 2025.02, [paper]
(Theory of Mind) Theory of Mind May Have Spontaneously Emerged in Large Language Models, 2023.02, [paper][code]
(Theory of Mind) Violation of Expectation via Metacognitive Prompting Reduces Theory of Mind Prediction Error in Large Language Models, 2023.10, [paper]
(Theory of Mind) Theory of Mind for Multi-Agent Collaboration via Large Language Models, EMNLP 2023, [paper][code]
(Theory of Mind) Constrained Reasoning Chains for Enhancing Theory-of-Mind in Large Language Models, PRICAI 2024, [paper]
(Theory of Mind) Large Model Strategic Thinking, Small Model Efficiency: Transferring Theory of Mind in Large Language Models, 2024.08, [paper]
(Theory of Mind) Boosting Theory-of-Mind Performance in Large Language Models via Prompting, 2023.04, [paper]
(Theory of Mind) Probing the Robustness of Theory of Mind in Large Language Models, 2024.10, [paper]
(Theory of Mind) Dissecting the Ullman Variations with a SCALPEL: Why do LLMs fail at Trivial Alterations to the False Belief Task?, 2024.06, [paper]
(Theory of Mind) Rethinking Theory of Mind Benchmarks for LLMs: Towards A User-Centered Perspective, CHI 2025 Workshop, [paper]
(Theory of Mind) Multi-ToM: Evaluating Multilingual Theory of Mind Capabilities in Large Language Models, 2024.11, [paper]
(Theory of Mind) Neural Theory-of-Mind? On the Limits of Social Intelligence in Large LMs, EMNLP 2022, [paper]
(Theory of Mind) Decompose-ToM: Enhancing Theory of Mind Reasoning in Large Language Models through Simulation and Task Decomposition, 2025.01, [paper]
(Theory of Mind) Minding Language Models' (Lack of) Theory of Mind: A Plug-and-Play Multi-Character Belief Tracker, ACL 2023, [paper]
(Theory of Mind) Clever Hans or Neural Theory of Mind? Stress Testing Social Reasoning in Large Language Models, EACL 2024, [paper]
(Theory of Mind) ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind, 2025.01, [paper]
(Theory of Mind) Views Are My Own, but Also Yours: Benchmarking Theory of Mind Using Common Ground, ACL 2024 Findings, [paper]
(Theory of Mind) Testing theory of mind in large language models and humans, 2024.05, Nature Human Behaviour, [paper]
(Theory of Mind) LLMsachieve adult human performance on higher-order theory of mind tasks, 2024.05, [paper]
(Theory of Mind) PHAnToM: Persona-based Prompting Has An Effect on Theory-of-Mind Reasoning in Large Language Models, 2024.03, [paper]
(Theory of Mind) ToM-LM: Delegating Theory of Mind Reasoning to External Symbolic Executors in Large Language Models, NeSy 2024, [paper]
(Theory of Mind) Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks, 2023.02, [paper]
(Theory of Mind) Theory of Mind in Large Language Models: Examining Performance of 11 State-of-the-Art models vs. Children Aged 7-10 on Advanced Tests, CoNLL 2023, [paper]
(Theory of Mind) Think Twice: Perspective-Taking Improves Large Language Models' Theory-of-Mind Capabilities, ACL 2024, [paper]
(Theory of Mind) OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models, ACL 2024, [paper]
(Theory of Mind) Large Language Models as Theory of Mind Aware Generative Agents with Counterfactual Reflection, 2025.01, [paper]
(Theory of Mind) PersuasiveToM: A Benchmark for Evaluating Machine Theory of Mind in Persuasive Dialogues, 2025.02, [paper][code]
(Theory of Mind) AutoToM: Automated Bayesian Inverse Planning and Model Discovery for Open-ended Theory of Mind, 2025.02, [paper]
(Theory of Mind) How FaR Are Large Language Models From Agents with Theory-of-Mind?, 2023.10, [paper]
(Theory of Mind) Dynamic Evaluation of Large Language Models by Meta Probing Agents, ICML 2024, [paper][code]
(Theory of Mind) Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics, 2026.08, [paper]
(Emotional Intelligence) A Literature Review on Emotional Intelligence of Large Language Models (LLMs), 2024, International Journal of Advanced Research in Computer Science, [paper]
(Emotional Intelligence) Large Language Models and Empathy: Systematic Review, 2024.01, Journal of Medical Internet Research, [paper]
(Emotional Intelligence) EmotionQueen: A Benchmark for Evaluating Empathy of Large Language Models, ACL 2024 Findings, [paper]
(Emotional Intelligence) ChatGPT outperforms humans in emotional awareness evaluations, 2023.05, Frontiers in Psychology, Emotion Science, [paper]
(Emotional Intelligence) EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models, 2025.02, [paper][code]
(Emotional Intelligence) Emotionally Numb or Empathetic? Evaluating How LLMs Feel Using EmotionBench, NeurIPS 2024, [paper][code]
(Emotional Intelligence) Large Language Models Produce Responses Perceived to be Empathic, 2024.03, [paper]
(Emotional Intelligence) Large Language Models Understand and Can be Enhanced by Emotional Stimuli, LLM@IJCAI'23, [paper][code]
(Emotional Intelligence) EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models, 2023.12, [paper][code]
(Emotional Intelligence) dentification and Description of Emotions by Current Large Language Models, 2023.07, [paper]
(Emotional Intelligence) EmoBench: Evaluating the Emotional Intelligence of Large Language Models, 2024.02, [paper][code]
(Emotional Intelligence) Exploring ChatGPT’s Empathic Abilities, ACII 2023, [paper]
(Emotional Intelligence) The Emotional Intelligence of the GPT-4 Large Language Model, 2024.06, Psychology in Russia: State of the Art, [paper]
(Emotional Intelligence) Are Large Language Models More Empathetic than Humans?, 2024.06, [paper]
(Emotional Intelligence) Both Matter: Enhancing the Emotional Intelligence of Large Language Models without Compromising the General Intelligence, ACL 2024 Findings, [paper]
(Emotional Intelligence) Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models, 2025.05, [paper]
(Social Intelligence) DeSIQ: Towards an Unbiased, Challenging Benchmark for Social Intelligence Understanding, EMNLP 2023, [paper]
(Social Intelligence) SocialAI 0.1: Towards a Benchmark to Stimulate Research on Socio-Cognitive Abilities in Deep Reinforcement Learning Agents, NAACL 2021 Workshop, [paper][code]
(Social Intelligence) Do LLM Agents Exhibit Social Behavior?, 2023.12, [paper]
(Social Intelligence) AntEval: Evaluation of Social Interaction Competencies in LLM-Driven Agents, 2024.01, [paper]
(Social Intelligence) Exploring Prosocial Irrationality for LLM Agents: A Social Cognition View, 2024.05, [paper]
(Social Intelligence) Advancing Social Intelligence in AI Agents: Technical Challenges and Open Questions, EMNLP 2024, [paper]
(Social Intelligence) Large language models can outperform humans in social situational judgments, 2024.11, Scientific Reports, [paper]
(Social Intelligence) AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios, 2024.10, [paper][code]
(Social Intelligence) How well DoLarge Language Models Perform on Faux Pas Tests?, ACL 2023 Findings, [paper]
(Social Intelligence) Towards Objectively Benchmarking Social Intelligence for Language Agents at Action Level, ACL 2024 Findings, [paper]
(Social Intelligence) Emotional intelligence of Large Language Models, 2023.11, Journal of Pacific Rim Psychology, [paper][code]
(Social Intelligence) Academically intelligent LLMs are not necessarily socially intelligent, 2024.03, [paper]
(Social Intelligence) SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents, 2023.10, [paper]
(Social Intelligence) Emergent social conventions and collective bias in LLM populations, 2025.05, Science Advances, [paper]
(Language comprehension) Language Model Behavior: A Comprehensive Survey, 2023.05, Computational Linguistics(CL), [paper]
(Language comprehension) Large Language Models for Psycholinguistic Plausibility Pretesting, EACL 2024 Findings, [paper]
(Language comprehension) Syntactic Surprisal From Neural Models Predicts, But Underestimates, Human Processing Difficulty From Syntactic Ambiguities, CoNLL 2022, [paper]
(Language comprehension) GPT-4 Surpassing Human Performance in Linguistic Pragmatics, 2023.12, [paper]
(Language comprehension) HLB: Benchmarking LLMs' Humanlikeness in Language Use, 2024.09, [paper]
(Language comprehension) Large Language Models as Neurolinguistic Subjects: Discrepancy in Performance and Competence for Form and Meaning, 2024.11, [paper]
(Language comprehension) Do large language models and humans have similar behaviors in causal inference with script knowledge?, SEM 2024, [paper][code]
(Language comprehension) Prompt-based methods may underestimate large language models’ linguistic generalizations, 2023.07, [paper]
(Language comprehension) Towards a Psychology of Machines: Large Language Models Predict Human Memory, 2024.03, [paper]
(Language comprehension) How to Make the Most of LLMs' Grammatical Knowledge for Acceptability Judgments, 2024.08, [paper]
(Language comprehension) A Psycholinguistic Evaluation of Language Models' Sensitivity to Argument Roles, 2024.10, [paper]
(Language comprehension) Incremental Comprehension of Garden-Path Sentences by Large Language Models: Semantic Interpretation, Syntactic Re-Analysis, and Attention, 2024.05, [paper]
(Language comprehension) Evaluating Grammatical Well-Formedness in Large Language Models: A Comparative Study with Human Judgments, CMCL 2024 Workshop, [paper]
(Language comprehension) The Goldilocks of Pragmatic Understanding: Fine-Tuning Strategy Matters for Implicature Resolution by LLMs, NeurIPS 2023, [paper]
(Language comprehension) Long-form analogies generated by chatGPT lack human-like psycholinguistic properties, CogSci 2023, [paper]
(Language comprehension) Large GPT-like Models are Bad Babies: A Closer Look at the Relationship between Linguistic Competence and Psycholinguistic Measures, CoNLL 2023, [paper]
(Language comprehension) Computational Sentence-level Metrics Predicting Human Sentence Comprehension, 2024.03, [paper]
(Language comprehension) Are Large Language Models Capable of Generating Human-Level Narratives?, EMNLP 2024, [paper]
(Language comprehension) How can large language models become more human?, CMCL 2024, [paper]
(Language comprehension) A Targeted Assessment of Incremental Processing in Neural LanguageModels and Humans, ACL 2021, [paper]
(Language comprehension) Divergences between Language Models and Human Brains, NeurIPS 2024, [paper]
(Language generation) Divergent Creativity in Humans and Large Language Models, 2024.05, [paper]
(Language generation) The Crowdless Future? Generative AI and Creative Problem-Solving, 2024.08, Organization Science, [paper]
(Language generation) Do large language models resemble humans in language use?, CMCL 2024 Workshop, [paper]
(Language generation) Art or Artifice? Large Language Models and the False Promise of Creativity, CHI 2024, [paper]
(Language generation) Artificial Intelligence is More Creative Than Humans: A Cognitive Science Perspective on the Current State of Generative Language Models, 2023.09, [paper]
(Language generation) An empirical investigation of the impact of ChatGPT on creativity, 2024.08, Nature Human Behaviour, [paper]
(Language generation) Evaluating Large Language Models via Linguistic Profiling, EMNLP 2024, [paper]
(Language generation) The Language of Creativity: Evidence from Humans and Large Language Models, 2024.01, The Journal of Creative Behavior, [paper]
(Language generation) Long-form analogies generated by chatGPT lack human-like psycholinguistic properties, CogSci 2023, [paper]
(Language generation) Putting GPT-3's Creativity to the (Alternative Uses) Test, ICCC 2022 (Short Paper), [paper]
(Language generation) Humanlike Cognitive Patterns as Emergent Phenomena in Large Language Models, 2024.12, [paper]
(Language generation) Are Large Language Models Capable of Generating Human-Level Narratives?, 2024.07, [paper]
(Language generation) Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond), 2025.12, NeurIPS 2025 Best Paper (Datasets & Benchmarks Track), [paper][code]
(Language acquisition) Bridging the data gap between children and large language models, 2023.11, Trends in Cognitive Sciences (TICS) [paper]
(Language acquisition) Psychomatics—A Multidisciplinary Framework for Understanding Artificial Minds, 2024.04, Cyberpsychology, Behavior, and Social Networking, [paper]
(Language acquisition) Development of Cognitive Intelligence in Pre-trained Language Models, 2024.07, [paper]
(Language acquisition) Large GPT-like Models are Bad Babies: A Closer Look at the Relationship between Linguistic Competence and Psycholinguistic Measures, CoNLL 2023, [paper]
Large Language Models and Cognitive Science: A Comprehensive Review of Similarities, Differences, and Challenges, 2024.09, [paper]
CogBench: a large language model walks into a psychology lab, ICML 2024, [paper]
Age against the machine—susceptibility of large language models to cognitive impairment: cross sectional analysis, 2024.12, The BMJ(British Medical Journal), [paper]
The Cognitive Capabilities of Generative AI: A Comparative Analysis with Human Benchmarks, 2024.10, [paper]
CogGPT: Unleashing the Power of Cognitive Dynamics on Large Language Models, EMNLP 2024 Findings, [paper]
Language models and psychological sciences, 2023.10, Frontiers in Psychology, [paper]
M3GIA: A Cognition Inspired Multilingual and Multimodal General Intelligence Ability Benchmark, 2024.06, [paper]
CogLM: Tracking Cognitive Development of Large Language Models, 2024.08, [paper]
Emergent analogical reasoning in large language models, 2023.07, Nature Human Behaviour, [paper]
Understanding LLMs' Fluid Intelligence Deficiency: An Analysis of the ARC Task, 2025.02, [paper][code]
MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs, 2024.06, [paper]
Exploring the Cognitive Knowledge Structure of Large Language Models: An Educational Diagnostic Assessment Approach, EMNLP 2023 (Short Paper), [paper]
A foundation model to predict and capture human cognition, 2025.07, Nature, [paper]
What large language models know and what people think they know, 2025.02, Nature Machine Intelligence, [paper]
Judgments of learning distinguish humans from large language models in predicting memory, 2025.10, Scientific Reports, [paper]
Understanding large language models demands distinguishing human projection from machine cognition, 2026.07, Communications Psychology, [paper]
Large Language Models Assume People are More Rational than We Really Are, 2025.04, ICLR 2025 Oral, [paper]
Revisiting the Reliability of Psychological Scales on Large Language Models, EMNLP 2024, [paper]
You don't need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments, NAACL 2024, [paper]
Challenging the Validity of Personality Tests for Large Language Models, Workshop at NeurIPS 2023, [paper]
Do Psychometric Tests Work for Large Language Models? Evaluation of Tests on Sexism, Racism, and Morality, 2025, [paper]
A validity-guided workflow for robust large language model research in psychology, 2025, [paper]
From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology, 2025, [paper]
Psychometric item validation using virtual respondents with trait-response mediators, 2025, [paper]
Value Portrait: Assessing Language Models' Values through Psychometrically and Ecologically Valid Items, ACL 2025, [paper]
Established psychometric vs. ecologically valid questionnaires: Rethinking psychological assessments in large language models, 2025, [paper]
Persistent Instability in LLM's Personality Measurements: Effects of Scale, Reasoning, and Conversation History, 2025, [paper]
Questioning the Survey Responses of Large Language Models, NeurIPS 2024, [paper]
Personality testing of large language models: limited temporal stability, but highlighted prosociality, 2024.01, Royal Society Open Science, [paper]
Do LLMs Exhibit Human-Like Response Biases? A Case Study in Survey Design, 2024.09, Transactions of the Association for Computational Linguistics (TACL), [paper]
Larger and more instructable language models become less reliable, 2024.10, Nature, [paper]
Large language models that replace human participants can harmfully misportray and flatten identity groups, 2025.03, Nature Machine Intelligence, [paper]
A large-scale replication of scenario-based experiments in psychology and management using large language models, 2025.08, Nature Computational Science, [paper]
A Theory of Response Sampling in LLMs: Part Descriptive and Part Prescriptive, 2025.07, ACL 2025 Best Paper, [paper]