This repo include the papers discussed in our latest survey paper on Awesome-LLM-as-a-judge.
Our official websit: Website
:book: Read the full paper here: Paper Link
2025-5 We update our paper list and include papers of LLM-as-a-judge in April & May 2025!2025-4 We update our paper list and include papers of LLM-as-a-judge in March 2025!2025-3 Want to learn more about risks and safety problems of using LLM-based annotation? Check out our new paper list on AI supervision risk!2025-3 We update our paper list and include papers of LLM-as-a-judge in February 2025, together with papers about thinking LLM as a judge!2025-2 We update our paper list and include papers of LLM-as-a-judge in January 2025!2025-2 Check our new paper about preference leakage in LLM-as-a-judge!2024-12 Also check our paper list and survey on LLM-based data annotation and synthesis!2024-12 We update our paper list and include papers of LLM-as-a-judge in December 2024!2024-12 We update the slides, talk and report of our paper, check them in our Website!If our survey is useful for your research, please kindly cite our paper:
@article{li2024llmasajudge,
title = {From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge},
author = {Dawei Li and Bohan Jiang and Liangjie Huang and Alimohammad Beigi and Chengshuai Zhao and Zhen Tan and Amrita Bhattacharjee and Yuxuan Jiang and Canyu Chen and Tianhao Wu and Kai Shu and Lu Cheng and Huan Liu},
year = {2024},
journal = {arXiv preprint arXiv: 2411.16594}
}

CaseSumm: A Large-Scale Dataset for Long-Context Summarization from U.S. Supreme Court Opinions. ArXiv preprint (2024) [Paper]
Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models Fan Zhang, Shulin Tian, Ziqi Huang, Yu Qiao, Ziwei Liu. ArXiv preprint (2024) [Paper]
Engineering AI Judge Systems Jiahuei Lin (Justina), Dayi Lin, Sky Zhang, Ahmed E. Hassan. ArXiv preprint (2024) [Paper]
LMUnit: Fine-grained Evaluation with Natural Language Unit Tests Jon Saad-Falcon, Rajan Vivek, William Berrios, Nandita Shankar Naik, Matija Franklin, Bertie Vidgen, Amanpreet Singh, Douwe Kiela, Shikib Mehri. ArXiv preprint (2024) [Paper]
ACE-M3: Automatic Capability Evaluator for Multimodal Medical Models Xiechi Zhang, Shunfan Zheng, Linlin Wang, Gerard de Melo, Zhu Cao, Xiaoling Wang, Liang He. ArXiv preprint (2024) [Paper]
IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering Ruosen Li, Ruochen Li, Barry Wang, Xinya Du. ArXiv preprint (2024) [Paper]
StrategyLLM: Large Language Models as Strategy Generators, Executors, Optimizers, and Evaluators for Problem Solving Chang Gao, Haiyun Jiang, Deng Cai, Shuming Shi, Wai Lam. ArXiv preprint (2024) [Paper]
ALI-Agent: Assessing LLMs' Alignment with Human Values via Agent-based Evaluation Jingnan Zheng, Han Wang, An Zhang, Tai D. Nguyen, Jun Sun, Tat-Seng Chua. ArXiv preprint (2024) [Paper]
On scalable oversight with weak LLMs judging strong LLMs Zachary Kenton, Noah Y. Siegel, János Kramár, Jonah Brown-Cohen, Samuel Albanie, Jannis Bulian, Rishabh Agarwal, David Lindner, Yunhao Tang, Noah D. Goodman, Rohin Shah. ArXiv preprint (2024) [Paper]
Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision Zhiqing Sun, Longhui Yu, Yikang Shen, Weiyang Liu, Yiming Yang, Sean Welleck, Chuang Gan. ArXiv preprint (2024) [Paper]
CriticEval: Evaluating Large Language Model as Critic Tian Lan, Wenwei Zhang, Chen Xu, Heyan Huang, Dahua Lin, Kai Chen, Xian-ling Mao. ArXiv preprint (2024) [Paper]
Adaptive Image Quality Assessment via Teaching Large Multimodal Model to Compare Hanwei Zhu, Haoning Wu, Yixuan Li, Zicheng Zhang, Baoliang Chen, Lingyu Zhu, Yuming Fang, Guangtao Zhai, Weisi Lin, Shiqi Wang. ArXiv preprint (2024) [Paper]
Beyond Guilt: Legal Judgment Prediction with Trichotomous Reasoning Kepu Zhang, Haoyue Yang, Xu Tang, Weijie Yu, Jun Xu. ArXiv preprint (2024) [Paper]
Let your LLM generate a few tokens and you will reduce the need for retrieval Hervé Déjean. ArXiv preprint (2024) [Paper]
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Kayla Schroeder, Zach Wood-Doughty. ArXiv preprint (2024) [Paper]
Using LLM-Generated Draft Replies to Support Human Experts in Responding to Stakeholder Inquiries in Maritime Industry: A Real-World Case Study of Industrial AI Tita Alissa Bach, Aleksandar Babic, Narae Park, Tor Sporsem, Rasmus Ulfsnes, Henrik Smith-Meyer, Torkel Skeie. ArXiv preprint (2024) [Paper]
An Exploratory Study of ML Sketches and Visual Code Assistants Luís F. Gomes, Vincent J. Hellendoorn, Jonathan Aldrich, Rui Abreu. ArXiv preprint (2024) [Paper]
Steering Large Language Models to Evaluate and Amplify Creativity Matthew Lyle Olson, Neale Ratzlaff, Musashi Hinck, Shao-yen Tseng, Vasudev Lal. ArXiv preprint (2024) [Paper]
Optimizing Alignment with Less: Leveraging Data Augmentation for Personalized Evaluation Javad Seraj, Mohammad Mahdi Mohajeri, Mohammad Javad Dousti, Majid Nili Ahmadabadi. ArXiv preprint (2024) [Paper]
Exploring Large Language Models on Cross-Cultural Values in Connection with Training Methodology Minsang Kim, Seungjun Baek. ArXiv preprint (2024) [Paper]
GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking Darshan Deshpande, Selvan Sunitha Ravi, Sky CH-Wang, Bartosz Mielczarek, Anand Kannappan, Rebecca Qian. ArXiv preprint (2024) [Paper]
CharacterBench: Benchmarking Character Customization of Large Language Models Jinfeng Zhou, Yongkang Huang, Bosi Wen, Guanqun Bi, Yuxuan Chen, Pei Ke, Zhuang Chen, Xiyao Xiao, Libiao Peng, Kuntian Tang, Rongsheng Zhang, Le Zhang, Tangjie Lv, Zhipeng Hu, Hongning Wang, Minlie Huang. ArXiv preprint (2024) [Paper]
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xiaozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Lei Hou, Yuxiao Dong, Jie Tang, Juanzi Li. ArXiv preprint (2024) [Paper]
JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment Hossein A. Rahmani, Emine Yilmaz, Nick Craswell, Bhaskar Mitra. ArXiv preprint (2024) [Paper]
Assessing the Impact of Conspiracy Theories Using Large Language Models Bohan Jiang, Dawei Li, Zhen Tan, Xinyi Zhou, Ashwin Rao, Kristina Lerman, H. Russell Bernard, Huan Liu. ArXiv preprint (2024) [Paper]
Outcome-Refining Process Supervision for Code Generation Zhuohao Yu, Weizheng Gu, Yidong Wang, Zhengran Zeng, Jindong Wang, Wei Ye, Shikun Zhang. ArXiv preprint (2024) [Paper]
RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment Zhuoran Jin, Hongbang Yuan, Tianyi Men, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao. ArXiv preprint (2024) [Paper]
JuStRank: Benchmarking LLM Judges for System Ranking Ariel Gera, Odellia Boni, Yotam Perlitz, Roy Bar-Haim, Lilach Eden, Asaf Yehudai. ArXiv preprint (2024) [Paper]
INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness Hung Le, Yingbo Zhou, Caiming Xiong, Silvio Savarese, Doyen Sahoo. ArXiv preprint (2024) [Paper]
LLM Evaluators Recognize and Favor Their Own Generations Arjun Panickssery, Samuel R. Bowman, Shi Feng. ArXiv preprint (2024) [Paper]
Training Language Models to Critique With Multi-agent Feedback Tian Lan, Wenwei Zhang, Chengqi Lyu, Shuaibin Li, Chen Xu, Heyan Huang, Dahua Lin, Xian-Ling Mao, Kai Chen. ArXiv preprint (2024) [Paper]
Beyond Exact Match: Semantically Reassessing Event Extraction by Large Language Models Yi-Fan Lu, Xian-Ling Mao, Tian Lan, Chen Xu, Heyan Huang. ArXiv preprint (2024) [Paper]
Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark Rong-Cheng Tu, Zi-Ao Ma, Tian Lan, Yuehao Zhao, Heyan Huang, Xian-Ling Mao. ArXiv preprint (2024) [Paper]
Multi-modal Retrieval Augmented Multi-modal Generation: A Benchmark, Evaluate Metrics and Strong Baselines Zi-Ao Ma, Tian Lan, Rong-Cheng Tu, Yong Hu, Heyan Huang, Xian-Ling Mao. ArXiv preprint (2024) [Paper]
LLM-AS-AN-INTERVIEWER: Beyond Static Testing Through Dynamic LLM Evaluation Eunsu Kim, Juyoung Suk, Seungone Kim, Niklas Muennighoff, Dongkwan Kim, Alice Oh. ArXiv preprint (2024) [Paper]
MAG-V: A Multi-Agent Framework for Synthetic Data Generation and Verification Saptarshi Sengupta, Kristal Curtis, Akshay Mallipeddi, Abhinav Mathur, Joseph Ross, Liang Gou. ArXiv preprint (2024) [Paper]
Law of the Weakest Link: Cross Capabilities of Large Language Models Ming Zhong, Aston Zhang, Xuewei Wang, Rui Hou, Wenhan Xiong, Chenguang Zhu, Zhengxing Chen, Liang Tan, Chloe Bi, Mike Lewis, Sravya Popuri, Sharan Narang, Melanie Kambadur, Dhruv Mahajan, Sergey Edunov, Jiawei Han, Laurens van der Maaten. ArXiv preprint (2024) [Paper]
MCQG-SRefine: Multiple Choice Question Generation and Evaluation with Iterative Self-Critique, Correction, and Comparison Feedback Zonghai Yao, Aditya Parashar, Huixue Zhou, Won Seok Jang, Feiyun Ouyang, Zhichao Yang, Hong Yu. ArXiv preprint (2024) [Paper]
ALMA: Alignment with Minimal Annotation Michihiro Yasunaga, Leonid Shamis, Chunting Zhou, Andrew Cohen, Jason Weston, Luke Zettlemoyer, Marjan Ghazvininejad. ArXiv preprint (2024) [Paper]
ConQRet: Benchmarking Fine-Grained Evaluation of Retrieval Augmented Argumentation with LLM Judges Kaustubh D. Dhole, Kai Shu, Eugene Agichtein. ArXiv preprint (2024) [Paper]
Evaluating and Aligning CodeLLMs on Human Preference Jian Yang, Jiaxi Yang, Ke Jin, Yibo Miao, Lei Zhang, Liqun Yang, Zeyu Cui, Yichang Zhang, Binyuan Hui, Junyang Lin. ArXiv preprint (2024) [Paper]
Benchmarking LLMs' Judgments with No Gold Standard Shengwei Xu, Yuxuan Lu, Grant Schoenebeck, Yuqing Kong. ArXiv preprint (2024) [Paper]
The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance? Sourav Banerjee, Ayushi Agarwal, Eishkaran Singh. ArXiv preprint (2024) [Paper]
Breaking Event Rumor Detection via Stance-Separated Multi-Agent Debate Mingqing Zhang, Haisong Gong, Qiang Liu, Shu Wu, Liang Wang. ArXiv preprint (2024) [Paper]
Can Large Language Models Serve as Evaluators for Code Summarization? Yang Wu, Yao Wan, Zhaoyang Chu, Wenting Zhao, Ye Liu, Hongyu Zhang, Xuanhua Shi, Philip S. Yu. ArXiv preprint (2024) [Paper]
FutureAGI agent-opt - Open-source library for automated optimization of AI agent workflows with evaluation-driven prompt and config tuning, including LLM-as-judge support.
Oceangpt: A large language model for ocean science tasks. Bi, Zhen, Zhang, Ningyu, Xue, Yida, Ou, Yixin, Ji, Daxiong, Zheng, Guozhou, and Chen, Huajun. ArXiv preprint (2023) [Paper]
Lawbench: Benchmarking legal knowledge of large language models. Fei, Zhiwei, Shen, Xiaoyu, Zhu, Dawei, Zhou, Fengzhe, Han, Zhuo, Zhang, Songyang, Chen, Kai, Shen, Zongwen, and Ge, Jidong. ArXiv preprint (2023) [Paper]
Sotopia: Interactive evaluation for social intelligence in language agents. Zhou, Xuhui, Zhu, Hao, Mathur, Leena, Zhang, Ruohong, Yu, Haofei, Qi, Zhengyang, Morency, Louis-Philippe, Bisk, Yonatan, Fried, Daniel, Neubig, Graham, and others. ArXiv preprint (2023) [Paper]
Can {C}hat{GPT} Defend its Belief in Truth? Evaluating {LLM} Reasoning via Debate. Wang, Boshi , Yue, Xiang , and Sun, Huan. Findings of the Association for Computational Linguistics: EMNLP 2023 (2023) [Paper]
On Evaluating the Integration of Reasoning and Action in {LLM} Agents with Database Question Answering. Nan, Linyong , Zhang, Ellen , Zou, Weijin , Zhao, Yilun , Zhou, Wenfei , and Cohan, Arman. Findings of the Association for Computational Linguistics: NAACL 2024 (2024) [Paper]
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Lianmin Zheng, Wei{-}Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 (2023) [Paper]
Human-like summarization evaluation with chatgpt. Gao, Mingqi, Ruan, Jie, Sun, Renliang, Yin, Xunjian, Yang, Shiping, and Wan, Xiaojun. ArXiv preprint (2023) [Paper]
Large language models are diverse role-players for summarization evaluation. Wu, Ning, Gong, Ming, Shou, Linjun, Liang, Shining, and Jiang, Daxin. CCF International Conference on Natural Language Processing and Chinese Computing (2023) [Paper]
Evaluating hallucinations in chinese large language models. Cheng, Qinyuan, Sun, Tianxiang, Zhang, Wenwei, Wang, Siyin, Liu, Xiangyang, Zhang, Mozhi, He, Junliang, Huang, Mianqiu, Yin, Zhangyue, Chen, Kai, and others. ArXiv preprint (2023) [Paper]
LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models. Lin, Yen-Ting , and Chen, Yun-Nung. Proceedings of the 5th Workshop on NLP for Conversational AI (NLP4ConvAI 2023) (2023) [Paper]
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models--A Survey. Mondorf, Philipp, and Plank, Barbara. ArXiv preprint (2024) [Paper]
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form Text. Badshah, Sher, and Sajjad, Hassan. ArXiv preprint (2024) [Paper]
Benchmarking Foundation Models with Language-Model-as-an-Examiner. Yushi Bai, Jiahao Ying, Yixin Cao, Xin Lv, Yuze He, Xiaozhi Wang, Jifan Yu, Kaisheng Zeng, Yijia Xiao, Haozhe Lyu, Jiayin Zhang, Juanzi Li, and Lei Hou. Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 (2023) [Paper]
Decoding biases: Automated methods and llm judges for gender bias detection in language models. Kumar, Shachi H, Sahay, Saurav, Mazumder, Sahisnu, Okur, Eda, Manuvinakurike, Ramesh, Beckage, Nicole, Su, Hsuan, Lee, Hung-yi, and Nachman, Lama. ArXiv preprint (2024) [Paper]
Halu-J: Critique-Based Hallucination Judge. Wang, Binjie, Chern, Steffi, Chern, Ethan, and Liu, Pengfei. ArXiv preprint (2024) [Paper]
Salad-bench: A hierarchical and comprehensive safety benchmark for large language models. Li, Lijun, Dong, Bowen, Wang, Ruohui, Hu, Xuhao, Zuo, Wangmeng, Lin, Dahua, Qiao, Yu, and Shao, Jing. ArXiv preprint (2024) [Paper]
Sorry-bench: Systematically evaluating large language model safety refusal behaviors. Xie, Tinghao, Qi, Xiangyu, Zeng, Yi, Huang, Yangsibo, Sehwag, Udari Madhushani, Huang, Kaixuan, He, Luxi, Wei, Boyi, Li, Dacheng, Sheng, Ying, and others. ArXiv preprint (2024) [Paper]
ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate. Chan, Chi-Min, Chen, Weize, Su, Yusheng, Yu, Jianxuan, Xue, Wei, Zhang, Shanghang, Fu, Jie, and Liu, Zhiyuan. The Twelfth International Conference on Learning Representations (2023) [Paper]
Evaluating the Performance of Large Language Models via Debates. Moniri, Behrad, Hassani, Hamed, and Dobriban, Edgar. ArXiv preprint (2024) [Paper]
Evaluating Mathematical Reasoning Beyond Accuracy. Xia, Shijie, Li, Xuefeng, Liu, Yixin, Wu, Tongshuang, and Liu, Pengfei. ArXiv preprint (2024) [Paper]
Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning. Fatemi, Bahare, Kazemi, Mehran, Tsitsulin, Anton, Malkan, Karishma, Yim, Jinyeong, Palowitch, John, Seo, Sungyong, Halcrow, Jonathan, and Perozzi, Bryan. ArXiv preprint (2024) [Paper]
LogicBench: Towards systematic evaluation of logical reasoning ability of large language models. Parmar, Mihir, Patel, Nisarg, Varshney, Neeraj, Nakamura, Mutsumi, Luo, Man, Mashetty, Santosh, Mitra, Arindam, and Baral, Chitta. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (2024) [Paper]
Academically intelligent LLMs are not necessarily socially intelligent. Xu, Ruoxi, Lin, Hongyu, Han, Xianpei, Sun, Le, and Sun, Yingfei. ArXiv preprint (2024) [Paper]
LLaVA-Critic: Learning to Evaluate Multimodal Models. Xiong, Tianyi, Wang, Xiyao, Guo, Dong, Ye, Qinghao, Fan, Haoqi, Gu, Quanquan, Huang, Heng, and Li, Chunyuan. ArXiv preprint (2024) [Paper]
Automated evaluation of large vision-language models on self-driving corner cases. Chen, Kai, Li, Yanze, Zhang, Wenhua, Liu, Yanxin, Li, Pengxiang, Gao, Ruiyuan, Hong, Lanqing, Tian, Meng, Zhao, Xinhai, Li, Zhenguo, and others. ArXiv preprint (2024) [Paper]
CodeJudge-Eval: A Benchmark for Evaluating Code Generation. Zhao, John, and others. ArXiv preprint (2024) [Paper]
Prompt-Gaming: A Pilot Study on LLM-Evaluating Agent in a Meaningful Energy Game. Isaza-Giraldo, Andr{'e}s, Bala, Paulo, Campos, Pedro F, and Pereira, Lucas. Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (2024) [Paper]
HealthQ: Unveiling Questioning Capabilities of LLM Chains in Healthcare Conversations. Wang, Ziyu, Li, Hao, Huang, Di, and Rahmani, Amir M. ArXiv preprint (2024) [Paper]
This repo include the papers discussed in our latest survey paper on Awesome-LLM-as-a-judge.
Our official websit: Website
:book: Read the full paper here: Paper Link
2025-5 We update our paper list and include papers of LLM-as-a-judge in April & May 2025!2025-4 We update our paper list and include papers of LLM-as-a-judge in March 2025!2025-3 Want to learn more about risks and safety problems of using LLM-based annotation? Check out our new paper list on AI supervision risk!2025-3 We update our paper list and include papers of LLM-as-a-judge in February 2025, together with papers about thinking LLM as a judge!2025-2 We update our paper list and include papers of LLM-as-a-judge in January 2025!2025-2 Check our new paper about preference leakage in LLM-as-a-judge!2024-12 Also check our paper list and survey on LLM-based data annotation and synthesis!2024-12 We update our paper list and include papers of LLM-as-a-judge in December 2024!2024-12 We update the slides, talk and report of our paper, check them in our Website!If our survey is useful for your research, please kindly cite our paper:
@article{li2024llmasajudge,
title = {From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge},
author = {Dawei Li and Bohan Jiang and Liangjie Huang and Alimohammad Beigi and Chengshuai Zhao and Zhen Tan and Amrita Bhattacharjee and Yuxuan Jiang and Canyu Chen and Tianhao Wu and Kai Shu and Lu Cheng and Huan Liu},
year = {2024},
journal = {arXiv preprint arXiv: 2411.16594}
}

CaseSumm: A Large-Scale Dataset for Long-Context Summarization from U.S. Supreme Court Opinions. ArXiv preprint (2024) [Paper]
Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models Fan Zhang, Shulin Tian, Ziqi Huang, Yu Qiao, Ziwei Liu. ArXiv preprint (2024) [Paper]
Engineering AI Judge Systems Jiahuei Lin (Justina), Dayi Lin, Sky Zhang, Ahmed E. Hassan. ArXiv preprint (2024) [Paper]
LMUnit: Fine-grained Evaluation with Natural Language Unit Tests Jon Saad-Falcon, Rajan Vivek, William Berrios, Nandita Shankar Naik, Matija Franklin, Bertie Vidgen, Amanpreet Singh, Douwe Kiela, Shikib Mehri. ArXiv preprint (2024) [Paper]
ACE-M3: Automatic Capability Evaluator for Multimodal Medical Models Xiechi Zhang, Shunfan Zheng, Linlin Wang, Gerard de Melo, Zhu Cao, Xiaoling Wang, Liang He. ArXiv preprint (2024) [Paper]
IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering Ruosen Li, Ruochen Li, Barry Wang, Xinya Du. ArXiv preprint (2024) [Paper]
StrategyLLM: Large Language Models as Strategy Generators, Executors, Optimizers, and Evaluators for Problem Solving Chang Gao, Haiyun Jiang, Deng Cai, Shuming Shi, Wai Lam. ArXiv preprint (2024) [Paper]
ALI-Agent: Assessing LLMs' Alignment with Human Values via Agent-based Evaluation Jingnan Zheng, Han Wang, An Zhang, Tai D. Nguyen, Jun Sun, Tat-Seng Chua. ArXiv preprint (2024) [Paper]
On scalable oversight with weak LLMs judging strong LLMs Zachary Kenton, Noah Y. Siegel, János Kramár, Jonah Brown-Cohen, Samuel Albanie, Jannis Bulian, Rishabh Agarwal, David Lindner, Yunhao Tang, Noah D. Goodman, Rohin Shah. ArXiv preprint (2024) [Paper]
Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision Zhiqing Sun, Longhui Yu, Yikang Shen, Weiyang Liu, Yiming Yang, Sean Welleck, Chuang Gan. ArXiv preprint (2024) [Paper]
CriticEval: Evaluating Large Language Model as Critic Tian Lan, Wenwei Zhang, Chen Xu, Heyan Huang, Dahua Lin, Kai Chen, Xian-ling Mao. ArXiv preprint (2024) [Paper]
Adaptive Image Quality Assessment via Teaching Large Multimodal Model to Compare Hanwei Zhu, Haoning Wu, Yixuan Li, Zicheng Zhang, Baoliang Chen, Lingyu Zhu, Yuming Fang, Guangtao Zhai, Weisi Lin, Shiqi Wang. ArXiv preprint (2024) [Paper]
Beyond Guilt: Legal Judgment Prediction with Trichotomous Reasoning Kepu Zhang, Haoyue Yang, Xu Tang, Weijie Yu, Jun Xu. ArXiv preprint (2024) [Paper]
Let your LLM generate a few tokens and you will reduce the need for retrieval Hervé Déjean. ArXiv preprint (2024) [Paper]
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Kayla Schroeder, Zach Wood-Doughty. ArXiv preprint (2024) [Paper]
Using LLM-Generated Draft Replies to Support Human Experts in Responding to Stakeholder Inquiries in Maritime Industry: A Real-World Case Study of Industrial AI Tita Alissa Bach, Aleksandar Babic, Narae Park, Tor Sporsem, Rasmus Ulfsnes, Henrik Smith-Meyer, Torkel Skeie. ArXiv preprint (2024) [Paper]
An Exploratory Study of ML Sketches and Visual Code Assistants Luís F. Gomes, Vincent J. Hellendoorn, Jonathan Aldrich, Rui Abreu. ArXiv preprint (2024) [Paper]
Steering Large Language Models to Evaluate and Amplify Creativity Matthew Lyle Olson, Neale Ratzlaff, Musashi Hinck, Shao-yen Tseng, Vasudev Lal. ArXiv preprint (2024) [Paper]
Optimizing Alignment with Less: Leveraging Data Augmentation for Personalized Evaluation Javad Seraj, Mohammad Mahdi Mohajeri, Mohammad Javad Dousti, Majid Nili Ahmadabadi. ArXiv preprint (2024) [Paper]
Exploring Large Language Models on Cross-Cultural Values in Connection with Training Methodology Minsang Kim, Seungjun Baek. ArXiv preprint (2024) [Paper]
GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking Darshan Deshpande, Selvan Sunitha Ravi, Sky CH-Wang, Bartosz Mielczarek, Anand Kannappan, Rebecca Qian. ArXiv preprint (2024) [Paper]
CharacterBench: Benchmarking Character Customization of Large Language Models Jinfeng Zhou, Yongkang Huang, Bosi Wen, Guanqun Bi, Yuxuan Chen, Pei Ke, Zhuang Chen, Xiyao Xiao, Libiao Peng, Kuntian Tang, Rongsheng Zhang, Le Zhang, Tangjie Lv, Zhipeng Hu, Hongning Wang, Minlie Huang. ArXiv preprint (2024) [Paper]
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xiaozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Lei Hou, Yuxiao Dong, Jie Tang, Juanzi Li. ArXiv preprint (2024) [Paper]
JudgeBlender: Ensembling Judgments for Automatic Relevance Assessment Hossein A. Rahmani, Emine Yilmaz, Nick Craswell, Bhaskar Mitra. ArXiv preprint (2024) [Paper]
Assessing the Impact of Conspiracy Theories Using Large Language Models Bohan Jiang, Dawei Li, Zhen Tan, Xinyi Zhou, Ashwin Rao, Kristina Lerman, H. Russell Bernard, Huan Liu. ArXiv preprint (2024) [Paper]
Outcome-Refining Process Supervision for Code Generation Zhuohao Yu, Weizheng Gu, Yidong Wang, Zhengran Zeng, Jindong Wang, Wei Ye, Shikun Zhang. ArXiv preprint (2024) [Paper]
RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment Zhuoran Jin, Hongbang Yuan, Tianyi Men, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao. ArXiv preprint (2024) [Paper]
JuStRank: Benchmarking LLM Judges for System Ranking Ariel Gera, Odellia Boni, Yotam Perlitz, Roy Bar-Haim, Lilach Eden, Asaf Yehudai. ArXiv preprint (2024) [Paper]
INDICT: Code Generation with Internal Dialogues of Critiques for Both Security and Helpfulness Hung Le, Yingbo Zhou, Caiming Xiong, Silvio Savarese, Doyen Sahoo. ArXiv preprint (2024) [Paper]
LLM Evaluators Recognize and Favor Their Own Generations Arjun Panickssery, Samuel R. Bowman, Shi Feng. ArXiv preprint (2024) [Paper]
Training Language Models to Critique With Multi-agent Feedback Tian Lan, Wenwei Zhang, Chengqi Lyu, Shuaibin Li, Chen Xu, Heyan Huang, Dahua Lin, Xian-Ling Mao, Kai Chen. ArXiv preprint (2024) [Paper]
Beyond Exact Match: Semantically Reassessing Event Extraction by Large Language Models Yi-Fan Lu, Xian-Ling Mao, Tian Lan, Chen Xu, Heyan Huang. ArXiv preprint (2024) [Paper]
Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark Rong-Cheng Tu, Zi-Ao Ma, Tian Lan, Yuehao Zhao, Heyan Huang, Xian-Ling Mao. ArXiv preprint (2024) [Paper]
Multi-modal Retrieval Augmented Multi-modal Generation: A Benchmark, Evaluate Metrics and Strong Baselines Zi-Ao Ma, Tian Lan, Rong-Cheng Tu, Yong Hu, Heyan Huang, Xian-Ling Mao. ArXiv preprint (2024) [Paper]
LLM-AS-AN-INTERVIEWER: Beyond Static Testing Through Dynamic LLM Evaluation Eunsu Kim, Juyoung Suk, Seungone Kim, Niklas Muennighoff, Dongkwan Kim, Alice Oh. ArXiv preprint (2024) [Paper]
MAG-V: A Multi-Agent Framework for Synthetic Data Generation and Verification Saptarshi Sengupta, Kristal Curtis, Akshay Mallipeddi, Abhinav Mathur, Joseph Ross, Liang Gou. ArXiv preprint (2024) [Paper]
Law of the Weakest Link: Cross Capabilities of Large Language Models Ming Zhong, Aston Zhang, Xuewei Wang, Rui Hou, Wenhan Xiong, Chenguang Zhu, Zhengxing Chen, Liang Tan, Chloe Bi, Mike Lewis, Sravya Popuri, Sharan Narang, Melanie Kambadur, Dhruv Mahajan, Sergey Edunov, Jiawei Han, Laurens van der Maaten. ArXiv preprint (2024) [Paper]
MCQG-SRefine: Multiple Choice Question Generation and Evaluation with Iterative Self-Critique, Correction, and Comparison Feedback Zonghai Yao, Aditya Parashar, Huixue Zhou, Won Seok Jang, Feiyun Ouyang, Zhichao Yang, Hong Yu. ArXiv preprint (2024) [Paper]
ALMA: Alignment with Minimal Annotation Michihiro Yasunaga, Leonid Shamis, Chunting Zhou, Andrew Cohen, Jason Weston, Luke Zettlemoyer, Marjan Ghazvininejad. ArXiv preprint (2024) [Paper]
ConQRet: Benchmarking Fine-Grained Evaluation of Retrieval Augmented Argumentation with LLM Judges Kaustubh D. Dhole, Kai Shu, Eugene Agichtein. ArXiv preprint (2024) [Paper]
Evaluating and Aligning CodeLLMs on Human Preference Jian Yang, Jiaxi Yang, Ke Jin, Yibo Miao, Lei Zhang, Liqun Yang, Zeyu Cui, Yichang Zhang, Binyuan Hui, Junyang Lin. ArXiv preprint (2024) [Paper]
Benchmarking LLMs' Judgments with No Gold Standard Shengwei Xu, Yuxuan Lu, Grant Schoenebeck, Yuqing Kong. ArXiv preprint (2024) [Paper]
The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance? Sourav Banerjee, Ayushi Agarwal, Eishkaran Singh. ArXiv preprint (2024) [Paper]
Breaking Event Rumor Detection via Stance-Separated Multi-Agent Debate Mingqing Zhang, Haisong Gong, Qiang Liu, Shu Wu, Liang Wang. ArXiv preprint (2024) [Paper]
Can Large Language Models Serve as Evaluators for Code Summarization? Yang Wu, Yao Wan, Zhaoyang Chu, Wenting Zhao, Ye Liu, Hongyu Zhang, Xuanhua Shi, Philip S. Yu. ArXiv preprint (2024) [Paper]
FutureAGI agent-opt - Open-source library for automated optimization of AI agent workflows with evaluation-driven prompt and config tuning, including LLM-as-judge support.
Oceangpt: A large language model for ocean science tasks. Bi, Zhen, Zhang, Ningyu, Xue, Yida, Ou, Yixin, Ji, Daxiong, Zheng, Guozhou, and Chen, Huajun. ArXiv preprint (2023) [Paper]
Lawbench: Benchmarking legal knowledge of large language models. Fei, Zhiwei, Shen, Xiaoyu, Zhu, Dawei, Zhou, Fengzhe, Han, Zhuo, Zhang, Songyang, Chen, Kai, Shen, Zongwen, and Ge, Jidong. ArXiv preprint (2023) [Paper]
Sotopia: Interactive evaluation for social intelligence in language agents. Zhou, Xuhui, Zhu, Hao, Mathur, Leena, Zhang, Ruohong, Yu, Haofei, Qi, Zhengyang, Morency, Louis-Philippe, Bisk, Yonatan, Fried, Daniel, Neubig, Graham, and others. ArXiv preprint (2023) [Paper]
Can {C}hat{GPT} Defend its Belief in Truth? Evaluating {LLM} Reasoning via Debate. Wang, Boshi , Yue, Xiang , and Sun, Huan. Findings of the Association for Computational Linguistics: EMNLP 2023 (2023) [Paper]
On Evaluating the Integration of Reasoning and Action in {LLM} Agents with Database Question Answering. Nan, Linyong , Zhang, Ellen , Zou, Weijin , Zhao, Yilun , Zhou, Wenfei , and Cohan, Arman. Findings of the Association for Computational Linguistics: NAACL 2024 (2024) [Paper]
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Lianmin Zheng, Wei{-}Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 (2023) [Paper]
Human-like summarization evaluation with chatgpt. Gao, Mingqi, Ruan, Jie, Sun, Renliang, Yin, Xunjian, Yang, Shiping, and Wan, Xiaojun. ArXiv preprint (2023) [Paper]
Large language models are diverse role-players for summarization evaluation. Wu, Ning, Gong, Ming, Shou, Linjun, Liang, Shining, and Jiang, Daxin. CCF International Conference on Natural Language Processing and Chinese Computing (2023) [Paper]
Evaluating hallucinations in chinese large language models. Cheng, Qinyuan, Sun, Tianxiang, Zhang, Wenwei, Wang, Siyin, Liu, Xiangyang, Zhang, Mozhi, He, Junliang, Huang, Mianqiu, Yin, Zhangyue, Chen, Kai, and others. ArXiv preprint (2023) [Paper]
LLM-Eval: Unified Multi-Dimensional Automatic Evaluation for Open-Domain Conversations with Large Language Models. Lin, Yen-Ting , and Chen, Yun-Nung. Proceedings of the 5th Workshop on NLP for Conversational AI (NLP4ConvAI 2023) (2023) [Paper]
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models--A Survey. Mondorf, Philipp, and Plank, Barbara. ArXiv preprint (2024) [Paper]
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form Text. Badshah, Sher, and Sajjad, Hassan. ArXiv preprint (2024) [Paper]
Benchmarking Foundation Models with Language-Model-as-an-Examiner. Yushi Bai, Jiahao Ying, Yixin Cao, Xin Lv, Yuze He, Xiaozhi Wang, Jifan Yu, Kaisheng Zeng, Yijia Xiao, Haozhe Lyu, Jiayin Zhang, Juanzi Li, and Lei Hou. Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023 (2023) [Paper]
Decoding biases: Automated methods and llm judges for gender bias detection in language models. Kumar, Shachi H, Sahay, Saurav, Mazumder, Sahisnu, Okur, Eda, Manuvinakurike, Ramesh, Beckage, Nicole, Su, Hsuan, Lee, Hung-yi, and Nachman, Lama. ArXiv preprint (2024) [Paper]
Halu-J: Critique-Based Hallucination Judge. Wang, Binjie, Chern, Steffi, Chern, Ethan, and Liu, Pengfei. ArXiv preprint (2024) [Paper]
Salad-bench: A hierarchical and comprehensive safety benchmark for large language models. Li, Lijun, Dong, Bowen, Wang, Ruohui, Hu, Xuhao, Zuo, Wangmeng, Lin, Dahua, Qiao, Yu, and Shao, Jing. ArXiv preprint (2024) [Paper]
Sorry-bench: Systematically evaluating large language model safety refusal behaviors. Xie, Tinghao, Qi, Xiangyu, Zeng, Yi, Huang, Yangsibo, Sehwag, Udari Madhushani, Huang, Kaixuan, He, Luxi, Wei, Boyi, Li, Dacheng, Sheng, Ying, and others. ArXiv preprint (2024) [Paper]
ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate. Chan, Chi-Min, Chen, Weize, Su, Yusheng, Yu, Jianxuan, Xue, Wei, Zhang, Shanghang, Fu, Jie, and Liu, Zhiyuan. The Twelfth International Conference on Learning Representations (2023) [Paper]
Evaluating the Performance of Large Language Models via Debates. Moniri, Behrad, Hassani, Hamed, and Dobriban, Edgar. ArXiv preprint (2024) [Paper]
Evaluating Mathematical Reasoning Beyond Accuracy. Xia, Shijie, Li, Xuefeng, Liu, Yixin, Wu, Tongshuang, and Liu, Pengfei. ArXiv preprint (2024) [Paper]
Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning. Fatemi, Bahare, Kazemi, Mehran, Tsitsulin, Anton, Malkan, Karishma, Yim, Jinyeong, Palowitch, John, Seo, Sungyong, Halcrow, Jonathan, and Perozzi, Bryan. ArXiv preprint (2024) [Paper]
LogicBench: Towards systematic evaluation of logical reasoning ability of large language models. Parmar, Mihir, Patel, Nisarg, Varshney, Neeraj, Nakamura, Mutsumi, Luo, Man, Mashetty, Santosh, Mitra, Arindam, and Baral, Chitta. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (2024) [Paper]
Academically intelligent LLMs are not necessarily socially intelligent. Xu, Ruoxi, Lin, Hongyu, Han, Xianpei, Sun, Le, and Sun, Yingfei. ArXiv preprint (2024) [Paper]
LLaVA-Critic: Learning to Evaluate Multimodal Models. Xiong, Tianyi, Wang, Xiyao, Guo, Dong, Ye, Qinghao, Fan, Haoqi, Gu, Quanquan, Huang, Heng, and Li, Chunyuan. ArXiv preprint (2024) [Paper]
Automated evaluation of large vision-language models on self-driving corner cases. Chen, Kai, Li, Yanze, Zhang, Wenhua, Liu, Yanxin, Li, Pengxiang, Gao, Ruiyuan, Hong, Lanqing, Tian, Meng, Zhao, Xinhai, Li, Zhenguo, and others. ArXiv preprint (2024) [Paper]
CodeJudge-Eval: A Benchmark for Evaluating Code Generation. Zhao, John, and others. ArXiv preprint (2024) [Paper]
Prompt-Gaming: A Pilot Study on LLM-Evaluating Agent in a Meaningful Energy Game. Isaza-Giraldo, Andr{'e}s, Bala, Paulo, Campos, Pedro F, and Pereira, Lucas. Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (2024) [Paper]
HealthQ: Unveiling Questioning Capabilities of LLM Chains in Healthcare Conversations. Wang, Ziyu, Li, Hao, Huang, Di, and Rahmani, Amir M. ArXiv preprint (2024) [Paper]