The Rules of the Game: A Survey of Rubrics for Large Language Models
WeChat@机器之心 |
WeChat |
Xiaohongshu |
X
If you like our project, please give us a star ⭐ on GitHub.
🚀 Update Log
- [May 31, 2026]: The brief introduction of our survey can be found on a series of platforms like WeChat, Xiaohongshu and X.
- [May 18, 2026]: We release the first version of our survey paper, Paper link.
📄 Citation
If you find this work helpful, please consider citing:
@misc{liu2026rules,
title={The Rules of the Game: A Survey of Rubrics for Large Language Models},
author={Liu, Wenhan and Jin, Jiajie and Huang, Zhaoheng and Wen, Tongyu and Dong, Guanting and Zhao, Ziliang and Zhu, Yutao and Dou, Zhicheng and Wen, Ji-Rong},
url={https://openreview.net/pdf?id=FnSimngGYk},
year={2026}
}
👋 Introduction
As large language models (LLMs) are increasingly used for reasoning, tool use, agentic interaction, and high-stakes decision-making, it becomes harder to define what makes a model response “good”. Rubrics provide a structured way to express multi-dimensional quality standards, such as factuality, completeness, safety, reasoning soundness, evidence grounding, and practical utility, making them useful for both model training and evaluation.
This repository maintains the paper list for our survey, The Rules of the Game: A Survey of Rubrics for Large Language Models. The survey first formalizes rubrics and compares them with reward models, verifiable rewards, and LLM-as-a-judge. It then organizes existing work into three directions: rubrics construction, rubrics for model training, and rubrics for evaluation, followed by discussions on open challenges such as reward hacking, evaluation bias, personalization, and rubric safety.
Feel free to contact us if you find a mistake, missing paper, or have any suggestions.
📋 Table of Content
📄 Paper List
Rubrics Construction
Background
- FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets, Ye et al., ICLR 2024. [Paper]
- Prometheus: Inducing Fine-grained Evaluation Capability in Language Models, Kim et al., ICLR 2024. [Paper]
- Rule Based Rewards for Language Model Safety, Mu et al., NeurIPS 2024. [Paper]
- Reinforcement Learning with Rubric Anchors, Huang et al., arXiv 2025. [Paper]
Direct Generation
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains, Gunjal et al., ICLR 2026. [Paper]
- Checklists Are Better Than Reward Models For Aligning Language Models, Viswanathan et al., NeurIPS 2025. [Paper]
- CARMO: Dynamic Criteria Generation for Context Aware Reward Modelling, Gupta et al., Findings of ACL 2025. [Paper]
- WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild, Lin et al., ICLR 2025. [Paper]
- SedarEval: Automated Evaluation using Self-Adaptive Rubrics, Fan et al., Findings of EMNLP 2024. [Paper]
- WritingBench: A Comprehensive Benchmark for Generative Writing, Wu et al., arXiv 2025. [Paper]
Contrastive Generation
- CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward Modeling, Liu et al., arXiv 2026. [Paper]
- Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training, Zhang et al., ICLR 2026. [Paper]
- OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment, Liu et al., arXiv 2025. [Paper]
- Rationale Matters: Learning Transferable Rubrics via Proxy-Guided Critique for VLM Reward Models, Qiu et al., arXiv 2026. [Paper]
- Learning Query-Specific Rubrics from Human Preferences for DeepResearch Report Generation, Lv et al., arXiv 2026. [Paper]
Iterative Refinement
Verification-Driven Refinement
- Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling, Xie et al., arXiv 2025. [Paper]
- OptimSyn: Influence-Guided Rubrics Optimization for Synthetic Data Generation, Fan et al., ICLR 2026. [Paper]
Structural Decomposition
- Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks, Shen et al., arXiv 2026. [Paper]
- Qworld: Question-Specific Evaluation Criteria for LLMs, Gao et al., arXiv 2026. [Paper]
- An Efficient Rubric-based Generative Verifier for Search-Augmented LLMs, Ma et al., arXiv 2025. [Paper]
- RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation, Li et al., arXiv 2026. [Paper]
- Rubric Is All You Need: Improving LLM-Based Code Evaluation With Question-Specific Rubrics, Pathak et al., ICER 2025. [Paper]
- DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Report, Li et al., arXiv 2026. [Paper]
- RefGrader: Automated Grading of Mathematical Competition Proofs using Agentic Workflows, Mahdavi et al., arXiv 2025. [Paper]
- RubricRAG: Towards Interpretable and Reliable LLM Evaluation via Domain Knowledge Retrieval for Rubric Generation, Dhole et al., arXiv 2026. [Paper]
- InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training, Wang et al., ICML 2026. [Paper]
De-duplication and Compression
- Generating Data-Driven Reasoning Rubrics for Domain-Adaptive Reward Modeling, Sanders et al., arXiv 2026. [Paper]
- Confusion-Aware Rubric Optimization for LLM-based Automated Grading, Chu et al., arXiv 2026. [Paper]
Online and Co-evolving Generation
Rollout-Based Evolving Rubrics
- DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research, Shao et al., ICML 2026. [Paper]
Online and Alternating Optimization of Rubric Generators
- Online Rubrics Elicitation from Pairwise Comparisons, Rezaei et al., ICML 2026. [Paper]
- Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-Training, Xu et al., ICML 2026. [Paper]
Self-Evolving, Adversarial, and Memory-Driven Rubrics
- SibylSense: Adaptive Rubric Learning via Memory Tuning and Adversarial Probing, Xu et al., arXiv 2026. [Paper]
- Open Rubric System: Scaling Reinforcement Learning with Pairwise Adaptive Rubric, Jia et al., arXiv 2026. [Paper]
- Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics, Sheng et al., arXiv 2026. [Paper]
- AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning, Jia et al., arXiv 2025. [Paper]
- RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks, Wu et al., ICLR 2026. [Paper]
Evaluation for Rubrics
- Rift: A rubric failure mode taxonomy and automated diagnostics, Qi et al., ICLR 2026 Workshop DATA-FM, [Paper]
- RubricRAG: Towards interpretable and reliable llm evaluation via domain knowledge retrieval for rubric generation, Dhole et al., SIGIR 2026, [Paper]
- Rubric-guided fine-tuning of speechllms for multi-aspect, multi-rater l2 reading-speech assessment, Parikh et al., LREC 2026, [Paper]
- Comparing developer and llm biases in code evaluation, Mittal et al., [Paper]
Rubrics for Model Training
Rubrics for Policy Model Training
Standard Rubric-based RL
- Checklists are better than reward models for aligning language models, Viswanathan et al., NeurIPS 2025. [Paper]
- Training AI Co-Scientists Using Rubric Rewards, Goel et al., arXiv 2025. [Paper]
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains, Gunjal et al., ICLR 2026. [Paper]
- Open Rubric System: Scaling Reinforcement Learning with Pairwise Adaptive Rubric, Jia et al., arXiv 2026. [Paper]
- Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training, Zhang et al., ICLR 2026. [Paper]
- RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks, Wu et al., arXiv 2025. [Paper]
- Visual Preference Optimization with Rubric Rewards, Yu et al., arXiv 2026. [Paper]
- AutoRubric-R1V: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning, Jia et al., arXiv 2025. [Paper]
- Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics, Sheng et al., arXiv 2026. [Paper]
- Dr Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research, Shao et al., ICML 2026. [Paper]
- OralGPT-Plus: Learning to Use Visual Tools via Reinforcement Learning for Panoramic X-ray Analysis, Fan et al., arXiv 2026. [Paper]
Advanced Reward Design
- Rule Based Rewards for Language Model Safety, Mu et al., NeurIPS 2024. [Paper]
- Reinforcement Learning with Rubric Anchors, Huang et al., arXiv 2025. [Paper]
- Mock Worlds, Real Skills: Building Small Agentic Language Models with Synthetic Tasks, Simulated Environments, and Rubric-Based Rewards, Lyu et al., arXiv 2026. [Paper]
- StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning, Li et al., arXiv 2026. [Paper]
- RuCL: Stratified Rubric-Based Curriculum Learning for Multimodal Large Language Model Reasoning, Chen et al., ICML 2026. [Paper]
- Alternating Reinforcement Learning with Contextual Rubric Rewards, Lan, arXiv 2026. [Paper]
- Stabilizing Rubric Integration Training via Decoupled Advantage Normalization, Tan et al., arXiv 2026. [Paper]
- Rubrics to Tokens: Bridging Response-level Rubrics and Token-level Rewards in Instruction Following Tasks, Xu et al., arXiv 2026. [Paper]
- Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks, _
Rubrics as Policy Guidance
- Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning, Zhou et al., arXiv 2025. [Paper]
- Experience is the Best Teacher: Motivating Effective Exploration in Reinforcement Learning for LLMs, Zhang et al., arXiv 2026. [Paper]
- Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance, Yu et al., arXiv 2026. [Paper]
Rubrics for Reward Model Training
Rubrics for Interpretability
- R3: Robust Rubric-Agnostic Reward Models, Anugraha et al., arXiv 2025. [Paper]
- Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards, Yuan et al., arXiv 2025. [Paper]
- mR3: Multilingual Rubric-Agnostic Reward Reasoning Models, Anugraha et al., arXiv 2025. [Paper]
- Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis, Kong et al., arXiv 2026. [Paper]
- CDRRM: Contrast-Driven Rubric Generation for Reliable and Interpretable Reward Modeling, Liu et al., arXiv 2026. [Paper]
- C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences, Kawabata et al., arXiv 2026. [Paper]
- DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification, Liu et al., arXiv 2026. [Paper]
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts, Wang et al., ENNLP (Findings) 2024. [Paper]
- A Rubric-Supervised Critic from Sparse Real-World Outcomes, Wang et al., arXiv 2026. [Paper]
- Multidimensional Rubric-oriented Reward Model Learning via Geometric Projection Reference Constraints, Jin et al., arXiv 2025. [Paper]
Rubrics for Reward Signals
- Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models, Wang et al., arXiv 2026. [Paper]
- Rationale Matters: Learning Transferable Rubrics via Proxy-Guided Critique for VLM Reward Models, Qiu et al., arXiv 2026. [Paper]
Rubrics for Data Construction
- Robust Reward Modeling via Causal Rubrics, Srivastava et al., arXiv 2025. [Paper]
Rubrics for Evaluation
Rubrics for General Task Evaluation
Reasoning Capability Evaluation
- Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist, Zhou et al., arXiv 2024. [Paper]
- SedarEval: Automated Evaluation using Self-Adaptive Rubrics, Fan et al., Findings of EMNLP 2024. [Paper]
- RefGrader: Automated Grading of Mathematical Competition Proofs using Agentic Workflows, Mahdavi et al., arXiv 2025. [Paper]
- Rubric Is All You Need: Improving LLM-Based Code Evaluation With Question-Specific Rubrics, Pathak et al., ICER 2025. [Paper]
- Comparing Developer and LLM Biases in Code Evaluation, Mittal et al., arXiv 2026. [Paper]
- ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge, Wang et al., arXiv 2025. [Paper]
- $OneMillion-Bench: How Far are Language Agents from Human Experts?, Yang et al., arXiv 2026. [Paper]
- MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes, Chiu et al., arXiv 2025. [Paper]
- Qworld: Question-Specific Evaluation Criteria for LLMs, Gao et al., arXiv 2026. [Paper]
- An Efficient Rubric-based Generative Verifier for Search-Augmented LLMs, Ma et al., arXiv 2025. [Paper]
Deep Research and Open-Ended Generation Evaluation
- HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models, Que et al., arXiv 2024. [Paper]
- WritingBench: A Comprehensive Benchmark for Generative Writing, Wu et al., arXiv 2025. [Paper]
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents, Du et al., arXiv 2025. [Paper]
- DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Report, Li et al., arXiv 2026. [Paper]
- DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation, Han et al., arXiv 2026. [Paper]
- ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents, Sharma et al., arXiv 2025. [Paper]
- MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome, Ye et al., arXiv 2026. [Paper]
- Pencils Down! Automatic Rubric-based Evaluation of Retrieve/Generate Systems, Farzi et al., ICTIR 2024. [Paper]
- Auto-Rubric: Learning to Extract Generalizable Criteria for Reward Modeling, Xie et al., arXiv 2025. [Paper]
- RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation, Li et al., arXiv 2026. [Paper]
General Agent Capability Evaluation
- AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents, Ma et al., arXiv 2024. [Paper]
- AdaRubric: Task-Adaptive Rubrics for LLM Agent Evaluation, Ding, arXiv 2026. [Paper]
- TRAJECT-Bench: A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use, He et al., arXiv 2025. [Paper]
- MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers, Luo et al., arXiv 2025. [Paper]
- MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs, Sirdeshmukh et al., arXiv 2025. [Paper]
- SCRIBE: Structured Mid-Level Supervision for Tool-Using Language Models, Jiang et al., arXiv 2026. [Paper]
- ASTRA-bench: Evaluating Tool-Use Agent Reasoning and Action Planning with Personal User Context, Xiu et al., arXiv 2026. [Paper]
- PaperBench: Evaluating AI's Ability to Replicate AI Research, Starace et al., ICML 2025. [Paper]
- Dr Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research, Shao et al., ICML 2026. [Paper]
Alignment Evaluation
- FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets, Ye et al., ICLR 2024. [Paper]
- InFoBench: Evaluating Instruction Following Ability in Large Language Models, Qin et al., Findings of ACL 2024. [Paper]
- AdvancedIF: Rubric-Based Benchmarking and Reinforcement Learning for Advancing LLM Instruction Following, He et al., arXiv 2025. [Paper]
- WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild, Lin et al., ICLR 2025. [Paper]
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment, Liu et al., EMNLP 2023. [Paper]
- Prometheus: Inducing Fine-grained Evaluation Capability in Language Models, Kim et al., ICLR 2024. [Paper]
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, Zheng et al., NeurIPS 2023. [Paper]
- RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following, Pan et al., arXiv 2026. [Paper]
- RubricBench: Aligning Model-Generated Rubrics with Human Standards, Zhang et al., arXiv 2026. [Paper]
- JudgeBench: A Benchmark for Evaluating LLM-based Judges, Tan et al., ICLR 2025. [Paper]
- A StrongREJECT for Empty Jailbreaks, Souly et al., arXiv 2024. [Paper]
Rubrics for Specific Task Evaluation
- PaperBench: Evaluating AI's Ability to Replicate AI Research, Starace et al., ICML 2025. [Paper]
- MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes, Chiu et al., arXiv 2025. [Paper]
- ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge, Wang et al., arXiv 2025. [Paper]
- Rubric Is All You Need: Improving LLM-Based Code Evaluation With Question-Specific Rubrics, Pathak et al., ICER 2025. [Paper]
- SCRIBE: Structured Mid-Level Supervision for Tool-Using Language Models, Jiang et al., arXiv 2026. [Paper]
Rubrics for Final Outputs
Content Factuality
- Pencils Down! Automatic Rubric-based Evaluation of Retrieve/Generate Systems, Farzi et al., ICTIR 2024. [Paper]
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents, Du et al., arXiv 2025. [Paper]
- DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Report, Li et al., arXiv 2026. [Paper]
- HealthBench: Evaluating Large Language Models Towards Improved Human Health, Arora et al., arXiv 2025. [Paper]
- PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning, Akyurek et al., arXiv 2025. [Paper]
- ASTRA-bench: Evaluating Tool-Use Agent Reasoning and Action Planning with Personal User Context, Xiu et al., arXiv 2026. [Paper]
- SCRIBE: Structured Mid-Level Supervision for Tool-Using Language Models, Jiang et al., arXiv 2026. [Paper]
- TechImage-Bench: Rubric-Based Evaluation for Technical Image Generation, Ni et al., arXiv 2026. [Paper]
- A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos, Kurpath et al., arXiv 2025. [Paper]
Safety Auditing
- HealthBench: Evaluating Large Language Models Towards Improved Human Health, Arora et al., arXiv 2025. [Paper]
- RubricRAG: Towards Interpretable and Reliable LLM Evaluation via Domain Knowledge Retrieval for Rubric Generation, Dhole et al., arXiv 2026. [Paper]
- MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes, Chiu et al., arXiv 2025. [Paper]
- Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability, Winata et al., arXiv 2025. [Paper]
Professional Presentation and Structural Coherence
- DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation, Han et al., arXiv 2026. [Paper]
- ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge, Wang et al., arXiv 2025. [Paper]
- $OneMillion-Bench: How Far are Language Agents from Human Experts?, Yang et al., arXiv 2026. [Paper]
- From Intuition to Calibrated Judgment: A Rubric-Based Expert-Panel Study of Human Detection of LLM-Generated Korean Text, Park et al., arXiv 2026. [Paper]
- PresentBench: A Fine-Grained Rubric-Based Benchmark for Slide Generation, Chen et al., arXiv 2026. [Paper]
- Rubric-Guided Fine-tuning of SpeechLLMs for Multi-Aspect, Multi-Rater L2 Reading-Speech Assessment, Parikh et al., arXiv 2026. [Paper]
Practical Utility and Actionability
- DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation, Han et al., arXiv 2026. [Paper]
- MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes, Chiu et al., arXiv 2025. [Paper]
- PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning, Akyurek et al., arXiv 2025. [Paper]
- ASTRA-bench: Evaluating Tool-Use Agent Reasoning and Action Planning with Personal User Context, Xiu et al., arXiv 2026. [Paper]
- LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts, Hashemi et al., ACL 2024. [Paper]
Contributing
We welcome contributions to this repository.
You can contribute by:
- Adding missing papers.
- Fixing incorrect metadata.
- Updating paper links, code links, or project links.
- Suggesting better taxonomy or section organization.
- Opening issues for discussion.
For any questions or feedback, please reach out to us at lwh@ruc.edu.cn.