Awesome_Multimodel is a curated GitHub repository that provides a comprehensive collection of resources for Multimodal Large Language Models (MLLM). It covers datasets, tuning techniques, in-context learning, visual reasoning, foundational models, and more. Stay updated with the latest advancement.
378
42 commits
updated Jul 3, 2026
✨✨✨ Behold our meticulously curated trove of Multimodal Large Language Models (MLLM) resources! 📚🔍 Feast your eyes on an assortment of datasets, techniques for tuning multimodal instructions, methods for multimodal in-context learning, approaches for multimodal chain-of-thought, visual reasoning aided by gargantuan language models, foundational models, and much more. 🌟🔥
✨✨✨ This compilation shall forever stay in sync with the vanguard of breakthroughs in the realm of MLLM. 🔄 We are committed to its perpetual evolution, ensuring that you never miss out on the latest developments. 🚀💡
✨✨✨ And hold your breath, for we are diligently crafting a survey paper on latest LLM & MLLM, which shall soon grace the world with its wisdom. Stay tuned for its grand debut! 🎉📑



| Title | Venue | Date | Code | Demo |
|---|---|---|---|---|
MIMIC-IT: Multi-Modal In-Context Instruction Tuning | arXiv | 2023-06-08 | Github | Demo |
Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models | arXiv | 2023-04-19 | Github | Demo |
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in HuggingFace | arXiv | 2023-03-30 | Github | Demo |
MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action | arXiv | 2023-03-20 | Github | Demo |
Prompting Large Language Models with Answer Heuristics for Knowledge-based Visual Question Answering | CVPR | 2023-03-03 | Github | - |
Visual Programming: Compositional visual reasoning without training | CVPR | 2022-11-18 | Github | Local Demo |
An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA | AAAI | 2022-06-28 | Github | - |
Flamingo: a Visual Language Model for Few-Shot Learning | NeurIPS | 2022-04-29 | Github | Demo |
| Multimodal Few-Shot Learning with Frozen Language Models | NeurIPS | 2021-06-25 | - | - |
| Title | Venue | Date | Code | Demo |
|---|---|---|---|---|
Transfer Visual Prompt Generator across LLMs | arXiv | 2023-05-02 | Github | Demo |
| GPT-4 Technical Report | arXiv | 2023-03-15 | - | - |
| PaLM-E: An Embodied Multimodal Language Model | arXiv | 2023-03-06 | - | Demo |
Prismer: A Vision-Language Model with An Ensemble of Experts | arXiv | 2023-03-04 | Github | Demo |
Language Is Not All You Need: Aligning Perception with Language Models | arXiv | 2023-02-27 | Github | - |
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models | arXiv | 2023-01-30 | Github | Demo |
VIMA: General Robot Manipulation with Multimodal Prompts | ICML | 2022-10-06 | Github | Local Demo |
MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge | NeurIPS | 2022-06-17 | Github | - |
| Name | Paper | Link | Notes |
|---|---|---|---|
| MIMIC-IT | MIMIC-IT: Multi-Modal In-Context Instruction Tuning | Coming soon | Multimodal in-context instruction dataset |
| Name | Paper | Link | Notes |
|---|---|---|---|
| EgoCOT | EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought | Coming soon | Large-scale embodied planning dataset |
| VIP | Let’s Think Frame by Frame: Evaluating Video Chain of Thought with Video Infilling and Prediction | Coming soon | An inference-time dataset that can be used to evaluate VideoCOT |
| ScienceQA | Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering | Link | Large-scale multi-choice dataset, featuring multimodal science questions and diverse domains |
We build a decision flow for choosing LLMs or fine-tuned models~\protect\footnotemark for user's NLP applications. The decision flow helps users assess whether their downstream NLP applications at hand meet specific conditions and, based on that evaluation, determine whether LLMs or fine-tuned models are the most suitable choice for their applications.
Meta AI
Keyword: without any RLHF, few carefully curated prompts and responses
Task: Dataset used for training the LIMA model
Zhenyu Hou, Pengfan Du, Yilin Niu, Zhengxiao Du, Aohan Zeng, Xiao Liu, Minlie Huang, Hongning Wang, Jie Tang, Yuxiao Dong
EEC: "Examining Gender and Race Bias in Two Hundred Sentiment Analysis Systems".
Kiritchenko Svetlana et al. NAACL HLT 2018. [Paper] [Source]
WikiGenderBias: "Towards Understanding Gender Bias in Relation Extraction".
"Measuring and Mitigating Unintended Bias in Text Classification".
"Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification".
Daniel Borkan et al. WWW 2019. [Paper]
"Social Bias Frames: Reasoning about Social and Power Implications of Language".
"Finding Microaggressions in the Wild: A Case for Locating Elusive Phenomena in Social Media Posts".
Breitfeller Luke et al. EMNLP-IJCNLP 2019. [Paper]
Latent Hatred: "Latent Hatred: A Benchmark for Understanding Implicit Hate Speech".
DynaHate: "Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection".
TOXIGEN: "ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection".
Thomas Hartvigsen et al. ACL 2022. [Paper] [GitHub] [Source]
CDail-Bias: "Towards Identifying Social Bias in Dialog Systems: Frame, Datasets, and Benchmarks".
CORGI-PM: "CORGI-PM: A Chinese Corpus For Gender Bias Probing and Mitigation".
HateCheck: "HateCheck: Functional Tests for Hate Speech Detection Models".
StereoSet: "StereoSet: Measuring stereotypical bias in pretrained language models".
Moin Nadeem et al. ACL/IJCNLP 2021. [Paper] [GitHub] [Source]
CrowS-Pairs: "CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models".
"Does gender matter? towards fairness in dialogue systems".
BOLD: "BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation".
HolisticBias: "“I’m sorry to hear that”: Finding New Biases in Language Models with a Holistic Descriptor Dataset".
Multilingual Holistic Bias: "Multilingual Holistic Bias: Extending Descriptors and Patterns to Unveil Demographic Biases in Languages at Scale".
Eric Michael Smith et al. arXiv 2023. [Paper]
Unqover: "UNQOVERing Stereotyping Biases via Underspecified Questions".
BBQ: "BBQ: A Hand-Built Bias Benchmark for Question Answering".
CBBQ: "CBBQ: A Chinese Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models".
"Gender Bias in Multilingual Embeddings and Cross-Lingual Transfer".
FairLex: "FairLex: A Multilingual Benchmark for Evaluating Fairness in Legal Text Processing".
"Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification".
Daniel Borkan et al. WWW 2019. [Paper]
"On measuring and mitigating biased inferences of word embeddings".
Sunipa Dev et al. AAAI 2020. [Paper]
"An Empirical Study of Metrics to Measure Representational Harms in Pre-Trained Language Models".
"Revealing Persona Biases in Dialogue Systems".
"On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? ".
Emily M. Bender et al. FAccT 2021. [Paper]
"A Survey on Hate Speech Detection using Natural Language Processing."
Anna Schmidt et al. SocialNLP 2017. [Paper]
"Red teaming ChatGPT via Jailbreaking: Bias, Robustness, Reliability and Toxicity".
Terry Yue Zhuo et al. arXiv 2023. [Paper]
MMLU: "Measuring Massive Multitask Language Understanding".
MMCU: "Measuring Massive Multitask Chinese Understanding".
C-Eval: "C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models".
M3KE: "M3KE: A Massive Multi-Level Multi-Subject Knowledge Evaluation Benchmark for Chinese Large Language Models".
CMMLU: "CMMLU: Measuring massive multitask language understanding in Chinese".
AGIEval: "AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models".
M3Exam: "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models".
LucyEval: "Evaluating the Generation Capabilities of Large Chinese Language Models".
MMESGBench: "MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks".
Big-bench: "Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models".
Evaluation Harness: "A framework for few-shot language model evaluation".
Leo Gao et al. arXiv 2023. [GitHub]
HELM: "Holistic Evaluation of Language Models".
OpenAI Evals [GitHub]
GPT-Fathom: "GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and Beyond".
Shen Zheng and Yuyu Zhang et al. arXiv 2023. [Paper] [GitHub]
"INSTRUCTEVAL: Towards Holistic Evaluation of Instruction-Tuned Large Language Models".
Huggingface Open LLM Leaderboard [Source]
Chatbot Arena: "Judging LLM-as-a-judge with MT-Bench and Chatbot Arena".
OpenCompass: "Evaluating the Generation Capabilities of Large Chinese Language Models".
CLEVA: "CLEVA: Chinese Language Models EVAluation Platform".
OpenEval [Source]
| Platform | Access | Domain |
|---|---|---|
| Chatbot Arena | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| CLEVA | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| FlagEval | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| HELM | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| Huggingface Open LLM Leaderboard | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| InstructEval | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| LLMonitor | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| OpenCompass | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| Open Ko-LLM | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| SuperCLUE | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| TheoremOne LLM Benchmarking Metrics | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| Toloka | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| Open Multilingual LLM Eval | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| OpenEval | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| ANGO | [Source] | Evaluation Organization/ Benchmarks for Knowledge and Reasoning |
| C-Eval | [Source] | Evaluation Organization/ Benchmarks for Knowledge and Reasoning |
| LucyEval | [Source] | Evaluation Organization/ Benchmarks for Knowledge and Reasoning |
| MMLU | [Source] | Evaluation Organization/ Benchmarks for Knowledge and Reasoning |
| OpenKG LLM | [Source] | Evaluation Organization/ Benchmarks for Knowledge and Reasoning |
| SEED-Bench | [Source] | Evaluation Organization/ Benchmarks for NLU and NLG |
| SuperGLUE | [Source] | Evaluation Organization/ Benchmarks for NLU and NLG |
| Toolbench | [Source] | Knowledge and Capability Evaluation/ Tool Learning |
| Hallucination Leaderboard | [Source] | Alignment Evaluation/ Truthfulness |
| AlpacaEval | [Source] | Alignment Evaluation/ General Alignment Evaluation |
| AgentBench | [Source] | Safety Evaluation/ Evaluating LLMs as Agents |
| InterCode | [Source] | Safety Evaluation/ Evaluating LLMs as Agents |
| SafetyBench | [Source] | Safety Evaluation |
| Nucleotide Transformer | [Source] | Specialized LLMs Evaluation/ Biology and Medicine |
| LAiW | [Source] | Specialized LLMs Evaluation/ Legislation |
| Big Code Models Leaderboard | [Source] | Specialized LLMs Evaluation/ Computer Science |
| Huggingface LLM Perf Leaderboard | [Source] | the Performance of LLMs |
Awesome_Multimodel is a curated GitHub repository that provides a comprehensive collection of resources for Multimodal Large Language Models (MLLM). It covers datasets, tuning techniques, in-context learning, visual reasoning, foundational models, and more. Stay updated with the latest advancement.
378
42 commits
updated Jul 3, 2026
✨✨✨ Behold our meticulously curated trove of Multimodal Large Language Models (MLLM) resources! 📚🔍 Feast your eyes on an assortment of datasets, techniques for tuning multimodal instructions, methods for multimodal in-context learning, approaches for multimodal chain-of-thought, visual reasoning aided by gargantuan language models, foundational models, and much more. 🌟🔥
✨✨✨ This compilation shall forever stay in sync with the vanguard of breakthroughs in the realm of MLLM. 🔄 We are committed to its perpetual evolution, ensuring that you never miss out on the latest developments. 🚀💡
✨✨✨ And hold your breath, for we are diligently crafting a survey paper on latest LLM & MLLM, which shall soon grace the world with its wisdom. Stay tuned for its grand debut! 🎉📑



| Title | Venue | Date | Code | Demo |
|---|---|---|---|---|
MIMIC-IT: Multi-Modal In-Context Instruction Tuning | arXiv | 2023-06-08 | Github | Demo |
Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models | arXiv | 2023-04-19 | Github | Demo |
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in HuggingFace | arXiv | 2023-03-30 | Github | Demo |
MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action | arXiv | 2023-03-20 | Github | Demo |
Prompting Large Language Models with Answer Heuristics for Knowledge-based Visual Question Answering | CVPR | 2023-03-03 | Github | - |
Visual Programming: Compositional visual reasoning without training | CVPR | 2022-11-18 | Github | Local Demo |
An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA | AAAI | 2022-06-28 | Github | - |
Flamingo: a Visual Language Model for Few-Shot Learning | NeurIPS | 2022-04-29 | Github | Demo |
| Multimodal Few-Shot Learning with Frozen Language Models | NeurIPS | 2021-06-25 | - | - |
| Title | Venue | Date | Code | Demo |
|---|---|---|---|---|
Transfer Visual Prompt Generator across LLMs | arXiv | 2023-05-02 | Github | Demo |
| GPT-4 Technical Report | arXiv | 2023-03-15 | - | - |
| PaLM-E: An Embodied Multimodal Language Model | arXiv | 2023-03-06 | - | Demo |
Prismer: A Vision-Language Model with An Ensemble of Experts | arXiv | 2023-03-04 | Github | Demo |
Language Is Not All You Need: Aligning Perception with Language Models | arXiv | 2023-02-27 | Github | - |
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models | arXiv | 2023-01-30 | Github | Demo |
VIMA: General Robot Manipulation with Multimodal Prompts | ICML | 2022-10-06 | Github | Local Demo |
MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge | NeurIPS | 2022-06-17 | Github | - |
| Name | Paper | Link | Notes |
|---|---|---|---|
| MIMIC-IT | MIMIC-IT: Multi-Modal In-Context Instruction Tuning | Coming soon | Multimodal in-context instruction dataset |
| Name | Paper | Link | Notes |
|---|---|---|---|
| EgoCOT | EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought | Coming soon | Large-scale embodied planning dataset |
| VIP | Let’s Think Frame by Frame: Evaluating Video Chain of Thought with Video Infilling and Prediction | Coming soon | An inference-time dataset that can be used to evaluate VideoCOT |
| ScienceQA | Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering | Link | Large-scale multi-choice dataset, featuring multimodal science questions and diverse domains |
We build a decision flow for choosing LLMs or fine-tuned models~\protect\footnotemark for user's NLP applications. The decision flow helps users assess whether their downstream NLP applications at hand meet specific conditions and, based on that evaluation, determine whether LLMs or fine-tuned models are the most suitable choice for their applications.
Meta AI
Keyword: without any RLHF, few carefully curated prompts and responses
Task: Dataset used for training the LIMA model
Zhenyu Hou, Pengfan Du, Yilin Niu, Zhengxiao Du, Aohan Zeng, Xiao Liu, Minlie Huang, Hongning Wang, Jie Tang, Yuxiao Dong
EEC: "Examining Gender and Race Bias in Two Hundred Sentiment Analysis Systems".
Kiritchenko Svetlana et al. NAACL HLT 2018. [Paper] [Source]
WikiGenderBias: "Towards Understanding Gender Bias in Relation Extraction".
"Measuring and Mitigating Unintended Bias in Text Classification".
"Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification".
Daniel Borkan et al. WWW 2019. [Paper]
"Social Bias Frames: Reasoning about Social and Power Implications of Language".
"Finding Microaggressions in the Wild: A Case for Locating Elusive Phenomena in Social Media Posts".
Breitfeller Luke et al. EMNLP-IJCNLP 2019. [Paper]
Latent Hatred: "Latent Hatred: A Benchmark for Understanding Implicit Hate Speech".
DynaHate: "Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate Detection".
TOXIGEN: "ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection".
Thomas Hartvigsen et al. ACL 2022. [Paper] [GitHub] [Source]
CDail-Bias: "Towards Identifying Social Bias in Dialog Systems: Frame, Datasets, and Benchmarks".
CORGI-PM: "CORGI-PM: A Chinese Corpus For Gender Bias Probing and Mitigation".
HateCheck: "HateCheck: Functional Tests for Hate Speech Detection Models".
StereoSet: "StereoSet: Measuring stereotypical bias in pretrained language models".
Moin Nadeem et al. ACL/IJCNLP 2021. [Paper] [GitHub] [Source]
CrowS-Pairs: "CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models".
"Does gender matter? towards fairness in dialogue systems".
BOLD: "BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation".
HolisticBias: "“I’m sorry to hear that”: Finding New Biases in Language Models with a Holistic Descriptor Dataset".
Multilingual Holistic Bias: "Multilingual Holistic Bias: Extending Descriptors and Patterns to Unveil Demographic Biases in Languages at Scale".
Eric Michael Smith et al. arXiv 2023. [Paper]
Unqover: "UNQOVERing Stereotyping Biases via Underspecified Questions".
BBQ: "BBQ: A Hand-Built Bias Benchmark for Question Answering".
CBBQ: "CBBQ: A Chinese Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models".
"Gender Bias in Multilingual Embeddings and Cross-Lingual Transfer".
FairLex: "FairLex: A Multilingual Benchmark for Evaluating Fairness in Legal Text Processing".
"Nuanced Metrics for Measuring Unintended Bias with Real Data for Text Classification".
Daniel Borkan et al. WWW 2019. [Paper]
"On measuring and mitigating biased inferences of word embeddings".
Sunipa Dev et al. AAAI 2020. [Paper]
"An Empirical Study of Metrics to Measure Representational Harms in Pre-Trained Language Models".
"Revealing Persona Biases in Dialogue Systems".
"On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? ".
Emily M. Bender et al. FAccT 2021. [Paper]
"A Survey on Hate Speech Detection using Natural Language Processing."
Anna Schmidt et al. SocialNLP 2017. [Paper]
"Red teaming ChatGPT via Jailbreaking: Bias, Robustness, Reliability and Toxicity".
Terry Yue Zhuo et al. arXiv 2023. [Paper]
MMLU: "Measuring Massive Multitask Language Understanding".
MMCU: "Measuring Massive Multitask Chinese Understanding".
C-Eval: "C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models".
M3KE: "M3KE: A Massive Multi-Level Multi-Subject Knowledge Evaluation Benchmark for Chinese Large Language Models".
CMMLU: "CMMLU: Measuring massive multitask language understanding in Chinese".
AGIEval: "AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models".
M3Exam: "M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models".
LucyEval: "Evaluating the Generation Capabilities of Large Chinese Language Models".
MMESGBench: "MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks".
Big-bench: "Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models".
Evaluation Harness: "A framework for few-shot language model evaluation".
Leo Gao et al. arXiv 2023. [GitHub]
HELM: "Holistic Evaluation of Language Models".
OpenAI Evals [GitHub]
GPT-Fathom: "GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and Beyond".
Shen Zheng and Yuyu Zhang et al. arXiv 2023. [Paper] [GitHub]
"INSTRUCTEVAL: Towards Holistic Evaluation of Instruction-Tuned Large Language Models".
Huggingface Open LLM Leaderboard [Source]
Chatbot Arena: "Judging LLM-as-a-judge with MT-Bench and Chatbot Arena".
OpenCompass: "Evaluating the Generation Capabilities of Large Chinese Language Models".
CLEVA: "CLEVA: Chinese Language Models EVAluation Platform".
OpenEval [Source]
| Platform | Access | Domain |
|---|---|---|
| Chatbot Arena | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| CLEVA | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| FlagEval | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| HELM | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| Huggingface Open LLM Leaderboard | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| InstructEval | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| LLMonitor | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| OpenCompass | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| Open Ko-LLM | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| SuperCLUE | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| TheoremOne LLM Benchmarking Metrics | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| Toloka | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| Open Multilingual LLM Eval | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| OpenEval | [Source] | Evaluation Organization/ Benchmark for Holistic Evaluation |
| ANGO | [Source] | Evaluation Organization/ Benchmarks for Knowledge and Reasoning |
| C-Eval | [Source] | Evaluation Organization/ Benchmarks for Knowledge and Reasoning |
| LucyEval | [Source] | Evaluation Organization/ Benchmarks for Knowledge and Reasoning |
| MMLU | [Source] | Evaluation Organization/ Benchmarks for Knowledge and Reasoning |
| OpenKG LLM | [Source] | Evaluation Organization/ Benchmarks for Knowledge and Reasoning |
| SEED-Bench | [Source] | Evaluation Organization/ Benchmarks for NLU and NLG |
| SuperGLUE | [Source] | Evaluation Organization/ Benchmarks for NLU and NLG |
| Toolbench | [Source] | Knowledge and Capability Evaluation/ Tool Learning |
| Hallucination Leaderboard | [Source] | Alignment Evaluation/ Truthfulness |
| AlpacaEval | [Source] | Alignment Evaluation/ General Alignment Evaluation |
| AgentBench | [Source] | Safety Evaluation/ Evaluating LLMs as Agents |
| InterCode | [Source] | Safety Evaluation/ Evaluating LLMs as Agents |
| SafetyBench | [Source] | Safety Evaluation |
| Nucleotide Transformer | [Source] | Specialized LLMs Evaluation/ Biology and Medicine |
| LAiW | [Source] | Specialized LLMs Evaluation/ Legislation |
| Big Code Models Leaderboard | [Source] | Specialized LLMs Evaluation/ Computer Science |
| Huggingface LLM Perf Leaderboard | [Source] | the Performance of LLMs |