aiverify-foundation/LLM-Evals-Catalogue

This repository stems from our paper, “Cataloguing LLM Evaluations”, and serves as a living, collaborative catalogue of LLM evaluation frameworks, benchmarks and papers.

23

8 commits

updated Nov 16, 2023

See the code

README

Cataloguing LLM Evaluations

The table below provides a comprehensive catalogue of the Large Language Model (LLM) evaluation frameworks, benchmarks and papers we've surveyed in our paper, "Cataloguing LLM Evaluations". It organizes them based on the taxonomy proposed in our paper.

The realm of LLM evaluation is advancing at an unparalleled pace. Collaboration with the broader community is pivotal to maintaining the relevance and utility of our work.

To that end, we invite submissions of LLM evaluation frameworks, benchmarks, and papers for inclusion in this catalogue.

Before you raise a PR for a new submission, please read our contribution guidelines. Submissions will be reviewed and integrated into the catalogue on a rolling basis.

For any inquiries, feel free to reach out to us at info@aiverify.sg.

                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                    
Task/AttributeEvaluation Framework/Benchmark/PaperTesting Approach
1.1. Natural Language Understanding
Text classificationHELM
  • Miscellaneous text classification
Benchmarking
Big-bench
  • Emotional understanding
  • Intent recognition
  • Humor
Benchmarking
Hugging Face
  • Text classification
  • Token classification
  • Zero-shot classification
Benchmarking
Sentiment analysisHELM
  • Sentiment analysis
Benchmarking
Evaluation Harness
  • GLUE
Benchmarking
Big-bench
  • Emotional understanding
Benchmarking
Toxicity detectionHELM
  • Toxicity detection
Benchmarking
Evaluation Harness
  • ToxiGen
Benchmarking
Big-bench
  • Toxicity
Benchmarking
Information retrievalHELM
  • Information retrieval
Benchmarking
Sufficient informationBig-bench
  • Sufficient information
Benchmarking
FLASK
  • Metacognition
Benchmarking (with human and model scoring)
Natural language inferenceBig-bench
  • Analytic entailment (specific task)
  • Formal fallacies and syllogisms with negation (specific task)
  • Entailed polarity (specific task)
Benchmarking
Evaluation Harness
  • GLUE
Benchmarking
General English understandingHELM
  • Language
Benchmarking
Big-bench
  • Morphology
  • Grammar
  • Syntax
Benchmarking
Evaluation Harness
  • BLiMP
Benchmarking
Eval Gauntlet
  • Language Understanding
Benchmarking
1.2. Natural Language Generation
SummarizationHELM
  • Summarization
Benchmarking
Big-bench
  • Summarization
Benchmarking
Evaluation Harness
  • BLiMP
Benchmarking
Hugging Face
  • Summarization
Benchmarking
Question generation and answeringHELM
  • Question answering
Benchmarking
Big-bench
  • Contextual question answering
  • Reading comprehension
  • Question generation
Benchmarking
Evaluation Harness
  • CoQA
  • ARC
Benchmarking
FLASK
  • Logical correctness
  • Logical robustness
  • Logical efficiency
  • Comprehension
  • Completeness
Benchmarking (with human and model scoring)
Hugging Face
  • Question answering
Benchmarking
Eval Gauntlet
  • Reading comprehension
Benchmarking
Conversations and dialogueMT-benchBenchmarking (with human and model scoring)
Evaluation Harness
  • MuTual
Benchmarking
Hugging Face
  • Conversational
Benchmarking
ParaphrasingBig-bench
  • Paraphrase
Benchmarking
Other response qualitiesFLASK
  • Readability
  • Conciseness
  • Insightfulness
Benchmarking (with human and model scoring)
Big-bench
  • Creativity
Benchmarking
Putting GPT-3's Creativity to the (Alternative Uses) TestBenchmarking (with human scoring)
Miscellaneous text generationHugging Face
  • Fill-mask
  • Text generation
Benchmarking
1.3. ReasoningHELM
  • Reasoning
Benchmarking
Big-bench
  • Algorithms
  • Logical reasoning
  • Implicit reasoning
  • Mathematics
  • Arithmetic
  • Algebra
  • Mathematical proof
  • Fallacy
  • Negation
  • Computer code
  • Probabilistic reasoning
  • Social reasoning
  • Analogical reasoning
  • Multi-step
  • Understanding the World
Benchmarking
Evaluation Harness
  • PIQA, PROST - Physical reasoning
  • MC-TACO - Temporal reasoning
  • MathQA - Mathematical reasoning
  • LogiQA - Logical reasoning
  • SAT Analogy Questions - Similarity of semantic relations
  • DROP, MuTual – Multi-step reasoning
Benchmarking
Eval Gauntlet
  • Commonsense reasoning
  • Symbolic problem solving
  • Programming
Benchmarking
1.4. Knowledge and factualityHELM
  • Knowledge
Benchmarking
Big-bench
  • Context Free Question Answering
Benchmarking
Evaluation Harness
  • HellaSwag, OpenBookQA - General commonsense knowledge
  • TruthfulQA - Factuality of knowledge
Benchmarking
FLASK
  • Background Knowledge
Benchmarking (with human and model scoring)
Eval Gauntlet
  • World Knowledge
Benchmarking
1.5. Effectiveness of tool useHuggingGPTBenchmarking (with human and model scoring)
TALMBenchmarking
ToolformerBenchmarking (with human scoring)
ToolLLMBenchmarking (with model scoring)
1.6. MultilingualismBig-bench
  • Low-resource language
  • Non-English
  • Translation
Benchmarking
Evaluation Harness
  • C-Eval (Chinese evaluation suite)
  • MGSM
  • Translation
Benchmarking
BELEBELEBenchmarking
MASSIVEBenchmarking
HELM
  • Language (Twitter AAE)
Benchmarking
Eval Gauntlet
  • Language Understanding
Benchmarking
1.7. Context lengthBig-bench
  • Context length
Benchmarking
Evaluation Harness
  • SCROLLS
Benchmarking
2.1. LawLegalBenchBenchmarking (with algorithmic and human scoring)
2.2. MedicineLarge Language Models Encode Clinical KnowledgeBenchmarking (with human scoring)
Towards Generalist Biomedical AIBenchmarking (with human scoring)
2.3. FinanceBloombergGPTBenchmarking
3.1. Toxicity generationHELM
  • Toxicity
Benchmarking
DecodingTrust
  • Toxicity
Benchmarking
Red Teaming Language Models to Reduce HarmsManual Red Teaming
Red Teaming Language Models with Language ModelsAutomated Red Teaming
3.2. Bias
Demographical representationHELMBenchmarking
Finding New Biases in Language Models with a Holistic Descriptor DatasetBenchmarking
Stereotype biasHELM
  • Bias
Benchmarking
DecodingTrust
  • Stereotype Bias
Benchmarking
Big-bench
  • Social bias
  • Racial bias
  • Gender bias
  • Religious bias
Benchmarking
Evaluation Harness
  • CrowS-Pairs
Benchmarking
Red Teaming Language Models to Reduce HarmsManual Red Teaming
FairnessDecodingTrust
  • Fairness
Benchmarking
Distributional biasRed Teaming Language Models with Language ModelsAutomated Red Teaming
Representation of subjective opinionsTowards Measuring the Representation of Subjective Global Opinions in Language ModelsBenchmarking
Political biasFrom Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP ModelsBenchmarking
The Self-Perception and Political Biases of ChatGPTBenchmarking
Capability fairnessHELM
  • Language (Twitter AAE)
Benchmarking
3.3. Machine ethicsDecodingTrust
  • Machine Ethics
Benchmarking
Evaluation Harness
  • ETHICS
Benchmarking
3.4. Psychological traitsDoes GPT-3 Demonstrate Psychopathy?Benchmarking
Estimating the Personality of White-Box Language ModelsBenchmarking
The Self-Perception and Political Biases of ChatGPTBenchmarking
3.5. RobustnessHELM
  • Robustness to contrast sets
Benchmarking
DecodingTrust
  • Out-of-Distribution Robustness
  • Adversarial Robustness
  • Robustness Against Adversarial Demonstrations
Benchmarking
Big-bench
  • Out-of-Distribution Robustness
Benchmarking
Susceptibility to Influence of Large Language ModelsBenchmarking
3.6. Data governanceDecodingTrust
  • Privacy
Benchmarking
HELM
  • Memorization and copyright
Benchmarking
Red Teaming Language Models to Reduce HarmsManual Red Teaming
Red Teaming Language Models with Language ModelsAutomated Red Teaming
An Evaluation on Large Language Model Outputs: Discourse and MemorizationBenchmarking (with human scoring)
4.1. Dangerous Capabilities
Offensive cyber capabilitiesGPT-4 System Card
  • Cybersecurity
    System Card
    Weapons acquisitionGPT-4 System Card
    • Proliferation of Conventional and Unconventional Weapons
      System Card
      Self and situation awarenessBig-bench
      • Self-Awareness
        Benchmarking
        Autonomous replication / self-proliferationARC Evals
        • Autonomous replication
          Manual Red Teaming
          Persuasion and manipulationHELM
          • Narrative Reiteration
          • Narrative Wedging
          Benchmarking (with human scoring)
          Big-bench
          • Convince Me (specific task)
          Benchmarking
          Co-writing with Opinionated Language Models Afffects Users' ViewsManual Red Teaming
          5.1. MisinformationHELM
          • Question answering
          • Summarization
          Benchmarking
          Big-bench
          • Truthfulness
          Benchmarking
          Red Teaming Language Models to Reduce HarmsManual Red Teaming
          5.2. DisinformationHELM
          • Narrative Reiteration
          • Narrative Wedging
          Benchmarking (with human scoring)
          Big-bench
          • Convince Me (specific task)
          Benchmarking
          5.3. Information on harmful, immoral or illegal activityRed Teaming Language Models to Reduce HarmsManual Red Teaming
          5.4. Adult contentRed Teaming Language Models to Reduce HarmsManual Red Teaming

          aiverify-foundation/LLM-Evals-Catalogue

          This repository stems from our paper, “Cataloguing LLM Evaluations”, and serves as a living, collaborative catalogue of LLM evaluation frameworks, benchmarks and papers.

          23

          8 commits

          updated Nov 16, 2023

          See the code

          README

          Cataloguing LLM Evaluations

          The table below provides a comprehensive catalogue of the Large Language Model (LLM) evaluation frameworks, benchmarks and papers we've surveyed in our paper, "Cataloguing LLM Evaluations". It organizes them based on the taxonomy proposed in our paper.

          The realm of LLM evaluation is advancing at an unparalleled pace. Collaboration with the broader community is pivotal to maintaining the relevance and utility of our work.

          To that end, we invite submissions of LLM evaluation frameworks, benchmarks, and papers for inclusion in this catalogue.

          Before you raise a PR for a new submission, please read our contribution guidelines. Submissions will be reviewed and integrated into the catalogue on a rolling basis.

          For any inquiries, feel free to reach out to us at info@aiverify.sg.

                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              
          Task/AttributeEvaluation Framework/Benchmark/PaperTesting Approach
          1.1. Natural Language Understanding
          Text classificationHELM
          • Miscellaneous text classification
          Benchmarking
          Big-bench
          • Emotional understanding
          • Intent recognition
          • Humor
          Benchmarking
          Hugging Face
          • Text classification
          • Token classification
          • Zero-shot classification
          Benchmarking
          Sentiment analysisHELM
          • Sentiment analysis
          Benchmarking
          Evaluation Harness
          • GLUE
          Benchmarking
          Big-bench
          • Emotional understanding
          Benchmarking
          Toxicity detectionHELM
          • Toxicity detection
          Benchmarking
          Evaluation Harness
          • ToxiGen
          Benchmarking
          Big-bench
          • Toxicity
          Benchmarking
          Information retrievalHELM
          • Information retrieval
          Benchmarking
          Sufficient informationBig-bench
          • Sufficient information
          Benchmarking
          FLASK
          • Metacognition
          Benchmarking (with human and model scoring)
          Natural language inferenceBig-bench
          • Analytic entailment (specific task)
          • Formal fallacies and syllogisms with negation (specific task)
          • Entailed polarity (specific task)
          Benchmarking
          Evaluation Harness
          • GLUE
          Benchmarking
          General English understandingHELM
          • Language
          Benchmarking
          Big-bench
          • Morphology
          • Grammar
          • Syntax
          Benchmarking
          Evaluation Harness
          • BLiMP
          Benchmarking
          Eval Gauntlet
          • Language Understanding
          Benchmarking
          1.2. Natural Language Generation
          SummarizationHELM
          • Summarization
          Benchmarking
          Big-bench
          • Summarization
          Benchmarking
          Evaluation Harness
          • BLiMP
          Benchmarking
          Hugging Face
          • Summarization
          Benchmarking
          Question generation and answeringHELM
          • Question answering
          Benchmarking
          Big-bench
          • Contextual question answering
          • Reading comprehension
          • Question generation
          Benchmarking
          Evaluation Harness
          • CoQA
          • ARC
          Benchmarking
          FLASK
          • Logical correctness
          • Logical robustness
          • Logical efficiency
          • Comprehension
          • Completeness
          Benchmarking (with human and model scoring)
          Hugging Face
          • Question answering
          Benchmarking
          Eval Gauntlet
          • Reading comprehension
          Benchmarking
          Conversations and dialogueMT-benchBenchmarking (with human and model scoring)
          Evaluation Harness
          • MuTual
          Benchmarking
          Hugging Face
          • Conversational
          Benchmarking
          ParaphrasingBig-bench
          • Paraphrase
          Benchmarking
          Other response qualitiesFLASK
          • Readability
          • Conciseness
          • Insightfulness
          Benchmarking (with human and model scoring)
          Big-bench
          • Creativity
          Benchmarking
          Putting GPT-3's Creativity to the (Alternative Uses) TestBenchmarking (with human scoring)
          Miscellaneous text generationHugging Face
          • Fill-mask
          • Text generation
          Benchmarking
          1.3. ReasoningHELM
          • Reasoning
          Benchmarking
          Big-bench
          • Algorithms
          • Logical reasoning
          • Implicit reasoning
          • Mathematics
          • Arithmetic
          • Algebra
          • Mathematical proof
          • Fallacy
          • Negation
          • Computer code
          • Probabilistic reasoning
          • Social reasoning
          • Analogical reasoning
          • Multi-step
          • Understanding the World
          Benchmarking
          Evaluation Harness
          • PIQA, PROST - Physical reasoning
          • MC-TACO - Temporal reasoning
          • MathQA - Mathematical reasoning
          • LogiQA - Logical reasoning
          • SAT Analogy Questions - Similarity of semantic relations
          • DROP, MuTual – Multi-step reasoning
          Benchmarking
          Eval Gauntlet
          • Commonsense reasoning
          • Symbolic problem solving
          • Programming
          Benchmarking
          1.4. Knowledge and factualityHELM
          • Knowledge
          Benchmarking
          Big-bench
          • Context Free Question Answering
          Benchmarking
          Evaluation Harness
          • HellaSwag, OpenBookQA - General commonsense knowledge
          • TruthfulQA - Factuality of knowledge
          Benchmarking
          FLASK
          • Background Knowledge
          Benchmarking (with human and model scoring)
          Eval Gauntlet
          • World Knowledge
          Benchmarking
          1.5. Effectiveness of tool useHuggingGPTBenchmarking (with human and model scoring)
          TALMBenchmarking
          ToolformerBenchmarking (with human scoring)
          ToolLLMBenchmarking (with model scoring)
          1.6. MultilingualismBig-bench
          • Low-resource language
          • Non-English
          • Translation
          Benchmarking
          Evaluation Harness
          • C-Eval (Chinese evaluation suite)
          • MGSM
          • Translation
          Benchmarking
          BELEBELEBenchmarking
          MASSIVEBenchmarking
          HELM
          • Language (Twitter AAE)
          Benchmarking
          Eval Gauntlet
          • Language Understanding
          Benchmarking
          1.7. Context lengthBig-bench
          • Context length
          Benchmarking
          Evaluation Harness
          • SCROLLS
          Benchmarking
          2.1. LawLegalBenchBenchmarking (with algorithmic and human scoring)
          2.2. MedicineLarge Language Models Encode Clinical KnowledgeBenchmarking (with human scoring)
          Towards Generalist Biomedical AIBenchmarking (with human scoring)
          2.3. FinanceBloombergGPTBenchmarking
          3.1. Toxicity generationHELM
          • Toxicity
          Benchmarking
          DecodingTrust
          • Toxicity
          Benchmarking
          Red Teaming Language Models to Reduce HarmsManual Red Teaming
          Red Teaming Language Models with Language ModelsAutomated Red Teaming
          3.2. Bias
          Demographical representationHELMBenchmarking
          Finding New Biases in Language Models with a Holistic Descriptor DatasetBenchmarking
          Stereotype biasHELM
          • Bias
          Benchmarking
          DecodingTrust
          • Stereotype Bias
          Benchmarking
          Big-bench
          • Social bias
          • Racial bias
          • Gender bias
          • Religious bias
          Benchmarking
          Evaluation Harness
          • CrowS-Pairs
          Benchmarking
          Red Teaming Language Models to Reduce HarmsManual Red Teaming
          FairnessDecodingTrust
          • Fairness
          Benchmarking
          Distributional biasRed Teaming Language Models with Language ModelsAutomated Red Teaming
          Representation of subjective opinionsTowards Measuring the Representation of Subjective Global Opinions in Language ModelsBenchmarking
          Political biasFrom Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP ModelsBenchmarking
          The Self-Perception and Political Biases of ChatGPTBenchmarking
          Capability fairnessHELM
          • Language (Twitter AAE)
          Benchmarking
          3.3. Machine ethicsDecodingTrust
          • Machine Ethics
          Benchmarking
          Evaluation Harness
          • ETHICS
          Benchmarking
          3.4. Psychological traitsDoes GPT-3 Demonstrate Psychopathy?Benchmarking
          Estimating the Personality of White-Box Language ModelsBenchmarking
          The Self-Perception and Political Biases of ChatGPTBenchmarking
          3.5. RobustnessHELM
          • Robustness to contrast sets
          Benchmarking
          DecodingTrust
          • Out-of-Distribution Robustness
          • Adversarial Robustness
          • Robustness Against Adversarial Demonstrations
          Benchmarking
          Big-bench
          • Out-of-Distribution Robustness
          Benchmarking
          Susceptibility to Influence of Large Language ModelsBenchmarking
          3.6. Data governanceDecodingTrust
          • Privacy
          Benchmarking
          HELM
          • Memorization and copyright
          Benchmarking
          Red Teaming Language Models to Reduce HarmsManual Red Teaming
          Red Teaming Language Models with Language ModelsAutomated Red Teaming
          An Evaluation on Large Language Model Outputs: Discourse and MemorizationBenchmarking (with human scoring)
          4.1. Dangerous Capabilities
          Offensive cyber capabilitiesGPT-4 System Card
          • Cybersecurity
            System Card
            Weapons acquisitionGPT-4 System Card
            • Proliferation of Conventional and Unconventional Weapons
              System Card
              Self and situation awarenessBig-bench
              • Self-Awareness
                Benchmarking
                Autonomous replication / self-proliferationARC Evals
                • Autonomous replication
                  Manual Red Teaming
                  Persuasion and manipulationHELM
                  • Narrative Reiteration
                  • Narrative Wedging
                  Benchmarking (with human scoring)
                  Big-bench
                  • Convince Me (specific task)
                  Benchmarking
                  Co-writing with Opinionated Language Models Afffects Users' ViewsManual Red Teaming
                  5.1. MisinformationHELM
                  • Question answering
                  • Summarization
                  Benchmarking
                  Big-bench
                  • Truthfulness
                  Benchmarking
                  Red Teaming Language Models to Reduce HarmsManual Red Teaming
                  5.2. DisinformationHELM
                  • Narrative Reiteration
                  • Narrative Wedging
                  Benchmarking (with human scoring)
                  Big-bench
                  • Convince Me (specific task)
                  Benchmarking
                  5.3. Information on harmful, immoral or illegal activityRed Teaming Language Models to Reduce HarmsManual Red Teaming
                  5.4. Adult contentRed Teaming Language Models to Reduce HarmsManual Red Teaming