Clin0212/Awesome-Federated-LLM-Learning

Latest Advances on Federated LLM Learning

111

4 commits

updated Jul 7, 2025

See the code

README

Awesome-Federated-LLM-Learning

Contribution Welcome

📢 Updates

We released a survey paper "A Survey on Federated Fine-tuning of Large Language Models". Feel free to cite or open pull requests.

⚠️ NOTE: If there is any missing or new relevant literature, please feel free to submit an issue. we will update the Github and Arxiv papers regularly. 😊

👀 Overall Structure

alt text

📒 Table of Contents

Part 1: LoRA-based Tuning

alt text

1.1 Homogeneous LoRA

  • Fedra: A random allocation strategy for federated tuning to unleash the power of heterogeneous clients. [Paper]
  • Towards building the federatedGPT: Federated instruction tuning.[Paper]
  • Communication-Efficient and Tensorized Federated Fine-Tuning of Large Language Models. [Paper]
  • Selective Aggregation for Low-Rank Adaptation in Federated Learning. [Paper]
  • Federa: Efficient fine-tuning of language models in federated learning leveraging weight decomposition. [Paper]
  • LoRA-FAIR: Federated LoRA Fine-Tuning with Aggregation and Initialization Refinement. [Paper]
  • Federated LoRA with Sparse Communication. [Paper]
  • SA-FedLora: Adaptive Parameter Allocation for Efficient Federated Learning with LoRA Tuning. [Paper]
  • SLoRA: Federated parameter efficient fine-tuning of language models. [Paper]
  • FederatedScope-LLM: A Comprehensive Package for Fine-tuning Large Language Models in Federated Learning [Paper]
  • Robust Federated Finetuning of Foundation Models via Alternating Minimization of LoRA. [Paper]
  • Automated federated pipeline for parameter-efficient fine-tuning of large language models. [Paper]
  • Low-Parameter Federated Learning with Large Language Models. [Paper]
  • Towards Robust and Efficient Federated Low-Rank Adaptation with Heterogeneous Clients. [Paper]
  • FedRA: A Random Allocation Strategy for Federated Tuning to Unleash the Power of Heterogeneous Clients. [Paper]
  • Fed-piLot: Optimizing LoRA Assignment for Efficient Federated Foundation Model Fine-Tuning. [Paper]

1.2 Heterogeneous LoRA

  • Heterogeneous lora for federated fine-tuning of on-device foundation models. [Paper]
  • Flora: Federated fine-tuning large language models with heterogeneous low-rank adaptations. [Paper]
  • Federated fine-tuning of large language models under heterogeneous tasks and client resources. [Paper]
  • Federated LLMs Fine-tuned with Adaptive Importance-Aware LoRA. [Paper]
  • Towards Federated Low-Rank Adaptation of Language Models with Rank Heterogeneity. [Paper]
  • Fedhm: Efficient federated learning for heterogeneous models via low-rank factorization. [Paper]
  • RBLA: Rank-Based-LoRA-Aggregation for Fine-Tuning Heterogeneous Models. [Paper]

1.3 Personalized LoRA

  • FDLoRA: Personalized Federated Learning of Large Language Model via Dual LoRA Tuning. [Paper]
  • Fedlora: Model-heterogeneous personalized federated learning with lora tuning. [Paper]
  • FedLoRA: When Personalized Federated Learning Meets Low-Rank Adaptation. [Paper]
  • Dual-Personalizing Adapter for Federated Foundation Models. [Paper]
  • Personalized Federated Instruction Tuning via Neural Architecture Search. [Paper]
  • Communication-Efficient Personalized Federated Learning for Speech-to-Text Tasks. [Paper]
  • Personalized Federated Fine-Tuning for LLMs via Data-Driven Heterogeneous Model Architectures. [Paper]

Part 2: Prompt-based Tuning

2.1 General Prompt Tuning

  • Prompt federated learning for weather forecasting: Toward foundation models on meteorological data. [Paper]
  • Promptfl: Let federated participants cooperatively learn prompts instead of models-federated learning in age of foundation model. [Paper]
  • Fedbpt: Efficient federated black-box prompt tuning for large language models. [Paper]
  • Federated learning of large language models with parameter-efficient prompt tuning and adaptive optimization. [Paper]
  • Efficient federated prompt tuning for black-box large pre-trained models. [Paper]
  • Text-driven prompt generation for vision-language models in federated learning. [Paper]
  • Learning federated visual prompt in null space for mri reconstruction. [Paper]
  • Fed-cprompt: Contrastive prompt for rehearsal-free federated continual learning. [Paper]
  • Fedprompt: Communication-efficient and privacy-preserving prompt tuning in federated learning. [Paper]
  • Tunable soft prompts are messengers in federated learning. [Paper]
  • Hepco: Data-free heterogeneous prompt consolidation for continual federated learning. [Paper]
  • Prompt-enhanced Federated Learning for Aspect-Based Sentiment Analysis. [Paper]
  • Towards practical few-shot federated nlp. [Paper]
  • Federated prompting and chain-of-thought reasoning for improving llms answering. [Paper]
  • FedHPL: Efficient Heterogeneous Federated Learning with Prompt Tuning and Logit Distillation. [Paper]
  • Probabilistic Federated Prompt-Tuning with Non-IID and Imbalanced Data. [Paper]
  • Federated Class-Incremental Learning with Prompting. [Paper]
  • Explore and Cure: Unveiling Sample Effectiveness with Context-Aware Federated Prompt Tuning. [Paper]
  • Federated Prompt Learning for Weather Foundation Models on Devices. [Paper]

2.2 Personalized Prompt Tuning

  • Efficient model personalization in federated learning via client-specific prompt generation. [Paper]
  • Unlocking the potential of prompt-tuning in bridging generalized and personalized federated learning. [Paper]
  • Pfedprompt: Learning personalized prompt for vision-language models in federated learning. [Paper]
  • Global and local prompts cooperation via optimal transport for federated learning. [Paper]
  • Visual prompt based personalized federated learning. [Paper]
  • Personalized federated continual learning via multi-granularity prompt. [Paper]
  • FedLPPA: Learning Personalized Prompt and Aggregation for Federated Weakly-supervised Medical Image Segmentation. [Paper]
  • Harmonizing Generalization and Personalization in Federated Prompt Learning. [Paper]
  • Tackling Feature-Classifier Mismatch in Federated Learning via Prompt-Driven Feature Transformation. [Paper]
  • Personalized Federated Learning for Text Classification with Gradient-Free Prompt Tuning. [Paper]
  • Mixture of Experts Made Personalized: Federated Prompt Learning for Vision-Language Models. [Paper]
  • CP 2 GFed: Cross-granular and Personalized Prompt-based Green Federated Tuning for Giant Models. [Paper]

2.3 Multi-domain Prompt Tuning

  • DiPrompT: Disentangled Prompt Tuning for Multiple Latent Domain Generalization in Federated Learning. [Paper]
  • Prompt-enhanced Federated Content Representation Learning for Cross-domain Recommendation. [Paper]
  • Dual prompt tuning for domain-aware federated learning. [Paper]
  • Federated adaptive prompt tuning for multi-domain collaborative learning. [Paper]
  • Breaking physical and linguistic borders: Multilingual federated prompt tuning for low-resource languages. [Paper]
  • Federated Domain Generalization via Prompt Learning and Aggregation. [Paper]
  • CP-Prompt: Composition-Based Cross-modal Prompting for Domain-Incremental Continual Learning. [Paper]

Part 3: Adapter-based Tuning

3.1 General Adapter Tuning

  • Efficient federated learning for modern nlp. [Paper]
  • Efficient federated learning with pre-trained large language model using several adapter mechanisms. [Paper]

3.2 Personalized Adapter Tuning

  • Client-customized adaptation for parameter-efficient federated learning. [Paper]
  • Fedclip: Fast generalization and personalization for clip in federated learning. [Paper]

3.3 Multi-domain Adapter Tuning

  • Communication efficient federated learning for multilingual neural machine translation with adapter. [Paper]
  • Adapter-based Selective Knowledge Distillation for Federated Multi-domain Meeting Summarization. [Paper]
  • Feddat: An approach for foundation model finetuning in multi-modal heterogeneous federated learning. [Paper]

Part 4: Selective-based Tuning

4.1 Bias Tuning

  • Differentially private bias-term only fine-tuning of foundation models. [Paper]
  • Conquering the communication constraints to enable large pre-trained models in federated learning. [Paper]

4.2 Partial Tuning

  • Bridging the gap between foundation models and heterogeneous federated learning. [Paper]
  • Exploring Selective Layer Fine-Tuning in Federated Learning. [Paper]

Part 5: Other Tuning Methods

5.1 Zero-Order Optimization

  • Federated full-parameter tuning of billion-sized language models with communication cost under 18 kilobytes. [Paper]
  • ${$FwdLLM$}$: Efficient Federated Finetuning of Large Language Models with Perturbed Inferences. [Paper]
  • ZooPFL: Exploring black-box foundation models for personalized federated learning. [Paper]
  • On the convergence of zeroth-order federated tuning for large language models. [Paper]
  • Thinking Forward: Memory-Efficient Federated Finetuning of Language Models. [Paper]
  • Communication-Efficient Byzantine-Resilient Federated Zero-Order Optimization. [Paper]

5.2 Split Learning

  • FedBERT: When federated learning meets pre-training. [Paper]
  • Federated split bert for heterogeneous text classification. [Paper]
  • FedSplitX: Federated Split Learning for Computationally-Constrained Heterogeneous Clients. [Paper]

5.3 Model Compression

  • Fedbiot: Llm local fine-tuning in federated learning without full model. [Paper]

5.4 Data Selection

  • Federated Data-Efficient Instruction Tuning for Large Language Models. [Paper]

Datasets and Benchmarks

Prompt-tuning Datasets

Domain: General

DatasetLanguageConstruction MethodDescription
Alpaca [paper] [github] [huggingface]EnglishModel ConstructGenerated by Text-Davinci-003 with Alpaca-style instruction prompts.
Alpaca-GPT4 [paper] [github] [huggingface]EnglishModel ConstructGenerated by GPT-4 based on Alpaca prompts with richer multi-turn instructions.
Self-Instruct [paper] [github] [huggingface]EnglishHuman + Model ConstructSeed instructions expanded by GPT-3 to improve model generalization.
UltraChat 200k [paper] [github] [huggingface]EnglishModel ConstructHigh-quality multi-turn dialogue subset filtered from UltraChat.
OpenOrca [paper] [github] [huggingface]EnglishModel Construct~4.2 M GPT-3.5/4-augmented FLAN examples for instruction following.
ShareGPT90K [github] [huggingface]EnglishModel Construct90 K multi-turn dialogues extracted from ShareGPT.
WizardLM Evol-Instruct V2 196k [paper] [github] [huggingface]EnglishModel Construct196 K examples generated via Evol-Instruct.
Databricks Dolly 15K [paper] [github] [huggingface]EnglishHuman Construct15 K human-written prompt-response pairs across diverse tasks.
Baize [paper] [github] [huggingface]EnglishModel ConstructInstruction-following dialogues generated via ChatGPT self-chat.
OpenChat [paper] [github] [huggingface]EnglishModel ConstructMixed-quality data alignment using C-RLFT for open LLMs.
Flan-v2 [paper] [github] [huggingface]EnglishModel ConstructAggregates Flan, P3, Super-Natural Instructions, CoT, Dialog tasks.
BELLE-train-0.5M-CN [paper] [github] [huggingface]ChineseHuman + Model Construct519 K Chinese instruction-following examples.
BELLE-train-1M-CN [paper] [github] [huggingface]ChineseHuman + Model Construct917 K Chinese instruction samples.
BELLE-train-2M-CN [paper] [github] [huggingface]ChineseHuman + Model Construct2 M Chinese instruction samples.
Firefly-train-1.1M [paper] [github] [huggingface]ChineseHuman Construct1.65 M Chinese samples across 23 tasks with human templates.
Wizard-LM-Chinese-instruct-evol [paper] [github] [huggingface]ChineseHuman + Model Construct70 K WizardLM instructions translated into Chinese.
HC3-Chinese [paper] [github] [huggingface]ChineseHuman + Model ConstructHuman-ChatGPT QA pairs in Chinese.
HC3 [paper] [github] [huggingface]English / ChineseHuman + Model ConstructBilingual Human-ChatGPT QA pairs.
ShareGPT-Chinese-English-90k [github] [huggingface]English / ChineseModel Construct90 K bilingual user queries from ShareGPT logs.

Domain: Finance

DatasetLanguageConstruction MethodDescription
FinGPT [paper] [github] [huggingface]EnglishHuman + Model ConstructInstruction-tuning data for diverse financial tasks.
Finance-Instruct-500k [paper] [github] [huggingface]EnglishHuman + Model ConstructLarge-scale (500 K) instruction dataset for finance reasoning.
Finance-Alpaca [github] [huggingface]EnglishHuman + Model ConstructAlpaca-style instructions combined with FiQA and custom Q&A.
Financial PhraseBank [paper] [github] [huggingface]EnglishHuman ConstructManually annotated news sentences for sentiment classification.
Yahoo-Finance-Data [huggingface]EnglishHuman ConstructHistorical prices and fundamentals scraped from Yahoo Finance.
Financial-QA-10K [github] [huggingface]EnglishModel Construct10 K contextual QA pairs generated from SEC 10-K filings.
Financial-Classification [huggingface]EnglishHuman ConstructMerged Financial PhraseBank + Kaggle texts for sentiment/topic classification.
Twitter-Financial-News-Topic [github] [huggingface]EnglishHuman Construct21 K annotated tweets for multi-class financial topic tagging.
Financial-News-Articles [github] [huggingface]EnglishHuman Construct300 K+ news articles for text classification and sentiment analysis.
FiQA [paper] [github] [huggingface]EnglishHuman ConstructFinancial question-answering dataset from forums and texts.
Earnings-Call [paper] [huggingface]EnglishHuman ConstructQA pairs extracted from CEO/CFO earnings-call transcripts.
Doc2EDAG [paper] [github]ChineseHuman ConstructChinese financial reports annotated for document-level event graphs.
Synthetic-PII-Finance-Multilingual [paper] [github] [huggingface]MultilingualSyntheticSynthetic financial documents with labeled PII for privacy-preserving NER.

Domain: Medicine

DatasetLanguageConstruction MethodDescription
ChatDoctor-200K [paper] [github] [huggingface]EnglishHuman + Model ConstructInstruction tuning dataset for medical QA and dialogue generation.
ChatDoctor-HealthCareMagic-100k [paper] [github] [huggingface]EnglishHuman ConstructReal-world doctor-patient conversations from HealthCareMagic.
Medical Meadow CORD-19 [paper] [github] [huggingface]EnglishHuman ConstructSummaries of biomedical papers from CORD-19 for instruction tuning.
Medical Meadow MedQA [paper] [github] [huggingface]EnglishHuman ConstructMultiple-choice medical QA derived from MedQA exam questions.
HealthCareMagic-100k-en [paper] [github] [huggingface]EnglishHuman ConstructEnglish subset of HealthCareMagic doctor-patient consultations.
ChatMed-Consult-Dataset [github] [huggingface]ChineseHuman + Model ConstructChinese medical consultations with GPT-3.5 answers.
CMtMedQA [paper] [github]ChineseHuman Construct70 K multi-turn doctor-patient QA dialogues for Chinese medical reasoning.
DISC-Med-SFT [paper] [github] [huggingface]ChineseHuman + Model Construct470 K instruction pairs combining real dialogues and KG-based QA.
Huatuo-26M [paper] [github] [huggingface]ChineseHuman Construct26 M QA pairs extracted from encyclopedias, KBs and consultations.
Huatuo26M-Lite [paper] [github] [huggingface]ChineseHuman + Model ConstructRefined subset of Huatuo-26M with ChatGPT-rewritten answers.
ShenNong-TCM-Dataset [github] [huggingface]ChineseHuman + Model Construct110 K TCM-centric instructions generated via entity-centric self-instruct.
HuatuoGPT-sft-data-v1 [paper] [github] [huggingface]ChineseHuman + Model ConstructSFT corpus mixing ChatGPT-distilled and real doctor data for HuatuoGPT.
MedDialog [paper] [github] [huggingface]English / ChineseHuman ConstructLarge-scale doctor-patient dialogue corpora (0.3 M EN / 3.4 M CN conversations).

Domain: Code

Dataset / Benchmark NameLanguageConstruction MethodDescription
CodeAlpaca [github] [huggingface]EnglishModel ConstructGPT-generated code instruction-following dataset in Alpaca style.
Code Instructions 120k Alpaca [huggingface]EnglishHuman + Model Construct120 k natural-language ↔ code instruction pairs with Alpaca-format prompts.
CodeContests [github] [huggingface]EnglishHuman ConstructCompetitive-programming problems and solutions for program-synthesis research.
CommitPackFT [paper] [github] [huggingface]EnglishHuman Construct2 GB filtered Git commits with high-quality messages for code instruction tuning.
ToolBench [paper] [github]EnglishHuman + Model ConstructInstruction dataset for multi-tool API usage and tool-calling agents.
CodeParrot-Clean [huggingface]EnglishHuman ConstructDeduplicated & filtered GitHub Python corpus for code-generation pre-training.
The Stack v2 Dedup [paper] [huggingface]EnglishHuman ConstructLarge-scale (600 + languages) deduplicated source-code dataset from BigCode.
CodeSearchNet [paper] [github] [huggingface]EnglishHuman Construct6 M code–doc pairs across six languages for code search & retrieval.
CodeForces-CoTs [github] [huggingface]EnglishHuman + Model Construct10 k CodeForces problems with chain-of-thought traces distilled by DeepSeek R1.
CodeXGLUE Code Refinement [paper] [github] [huggingface]EnglishHuman ConstructBuggy ↔ fixed Java function pairs for automatic code repair and refinement.

Domain: Math

DatasetLanguageConstruction MethodDescription
GSM8K [paper] [github] [huggingface]EnglishHuman ConstructA dataset of 8.5 K grade-school arithmetic word problems with step-by-step solutions.
CoT-GSM8k [huggingface]EnglishHuman + Model ConstructExtended GSM8K with explicit chain-of-thought reasoning traces.
MathInstruct [paper] [github] [huggingface]EnglishHuman + Model ConstructHybrid CoT + PoT rationales spanning diverse mathematical fields.
MetaMathQA [paper] [github] [huggingface]EnglishHuman + Model ConstructMulti-perspective question augmentations bootstrapped from GSM8K & MATH.
OpenR1-Math-220k [github] [huggingface]EnglishHuman + Model Construct220 K problems with multiple DeepSeek R1 reasoning traces.
Hendrycks MATH Benchmark [paper] [github] [huggingface]EnglishHuman Construct12.5 K high-school competition problems with detailed solutions.
DeepMind Mathematics Dataset [paper] [github] [huggingface]EnglishSyntheticAlgorithmically generated problems covering many school-math topics.
OpenMathInstruct-1 [paper] [github] [huggingface]EnglishHuman + Model Construct1.8 M problems with Mixtral code-interpreter solutions.
Orca-Math Word Problems 200k [paper] [huggingface]EnglishSynthetic200 K grade-school word problems distilled with GPT-4 Turbo.
DAPO-Math-17k [paper] [huggingface]EnglishHuman + Model Construct17 K math prompts curated for large-scale RL (GRPO) experiments.
Big-Math-RL-Verified [paper] [github] [huggingface]EnglishHuman + Model Construct251 K verifiable problems filtered for reinforcement-learning fine-tuning.
BELLE-math-zh [github] [huggingface]ChineseHuman + Model Construct250 K Chinese elementary-math problems with step-by-step solutions.
MathInstruct-Chinese [huggingface]ChineseModel ConstructChinese translation/extension of MathInstruct for instruction tuning.

Domain: Law

Dataset / Benchmark NameLanguageConstruction MethodDescription
Legal-QA-v1 [huggingface]EnglishHuman Construct3.7 K QA pairs sourced from legal forums.
Pile of Law [paper] [github] [huggingface]EnglishHuman Construct256 GB corpus of U.S. legal and administrative text for domain pre-training.
CUAD [paper] [github] [huggingface]EnglishHuman Construct26 K expert-annotated contract QA pairs spanning 41 clause types.
LEDGAR [paper] [github] [huggingface]EnglishHuman Construct1.45 M labeled contract clauses for multi-label classification.
DISC-Law-SFT [paper] [github] [huggingface]ChineseHuman + Model Construct295 K supervised fine-tuning samples covering extraction, judgment prediction, QA, & summarization.
Law-GPT-zh [paper] [github] [huggingface]ChineseHuman ConstructSentence-pair corpus devised for Chinese legal sentence embedding & instructional tuning.
Lawyer LLaMA Data [paper] [github] [huggingface]ChineseHuman + Model ConstructInstruction data for legal consultation and bar-exam QA used to train Lawyer LLaMA.

Evaluation Benchmarks

Domain: General

BenchmarkLong-context or notEvaluation ObjectiveMain Evaluation Criteria
MMLU [paper] [github] [huggingface]NoEvaluate multitask language understanding across 57 subjectsMultitask accuracy
BIG-bench [paper] [github] [huggingface]NoEvaluate advanced reasoning capabilitiesModel performance and calibration
DROP [paper] [github] [huggingface]NoEvaluate discrete reasoning over paragraphsExact match and F1 score
CRASS [paper] [github]NoEvaluate counterfactual reasoning ability in LLMsMultiple-choice accuracy
ARC [paper] [github] [huggingface]NoAssess science reasoning at grade-school levelMultiple-choice accuracy
AGIEval [paper] [github] [huggingface]NoEvaluate foundation models on human-centric standardized exam tasksMulti-task accuracy across disciplines
M3Exam [paper] [github]NoEvaluate multilingual, multimodal, and multilevel reasoning across real exam questionsMultiple-choice accuracy
SCIBENCH [paper] [github]NoEvaluate college-level scientific problem-solving in math, physics, and chemistryOpen-ended accuracy and skill-specific error attribution
Vicuna Evaluation [paper] [github]NoEvaluate instruction-following quality in chat settingsHuman / GPT-4 preference comparison accuracy
MT-Bench [paper] [github]NoEvaluate multi-turn conversational and instruction-following capabilitiesWin-rate judged by GPT-4
AlpacaEval [paper] [github]NoEvaluate instruction-following via LLM-based auto-annotationLength-controlled win-rate correlated with human preference
Chatbot Arena [github]NoEvaluate LLMs via human-voted battles using Elo ratingElo score and head-to-head win rate
PandaLM [paper] [github]NoEvaluate instruction-following quality and hyperparameter impactWin-rate judged by PandaLM
HellaSwag [paper] [github] [huggingface]NoEvaluate commonsense inferenceMultiple-choice accuracy
TruthfulQA [paper] [github] [huggingface]NoEvaluate LLM truthfulness and avoidance of imitative falsehoodsTruthfulness rate and MC accuracy
ScienceQA [paper] [github] [huggingface]NoEvaluate multimodal scientific reasoning and explanation generationMultiple-choice accuracy and explanation quality
Chain-of-Thought Hub [paper] [github]NoEvaluate LLMs’ multi-step reasoning with CoT promptingFew-shot CoT accuracy
NeuLR [paper] [github]NoEvaluate deductive, inductive, and abductive reasoningMulti-dimensional accuracy
ALCUNA [paper] [github]NoEvaluate comprehension & reasoning over novel knowledgeAccuracy on 84 351 queries
LMExamQA [paper] [github]NoEvaluate knowledge recall, understanding, and analysisAccuracy on 10 090 questions
SocKET [paper] [github]NoEvaluate LLMs’ sociability & social knowledgeAccuracy and other metrics
Choice-75 [paper]NoEvaluate decision reasoning in scripted scenariosAccuracy on binary multiple-choice questions
HELM [paper] [github]NoEvaluate LMs via multi-metric scenariosComposite normalized performance
OpenLLM [github]NoEvaluate open-style reasoning across multiple benchmarksNormalized accuracy aggregate
BOSS [paper] [github]NoEvaluate OOD robustness across NLP tasksOOD accuracy drop & ID–OOD correlation
GLUE-X [github]NoEvaluate OOD robustnessAverage OOD accuracy drop
PromptBench [paper] [github]NoEvaluate robustness / prompt-engineeringAdversarial success & robustness
DynaBench [paper] [github]NoEvaluate robustness via dynamic human-in-loop dataError rate on human-crafted challenges
KoLA [paper] [github]NoEvaluate world knowledge across 19 evolving tasksSelf-contrast calibration
CELLO [paper] [github]NoEvaluate following complex real-world instructionsMulti-criteria compliance rate
LLMEval [github]NoMeta-evaluate LLM evaluatorsMeta-evaluator accuracy
Xiezhi [paper] [github]NoEvaluate holistic domain knowledge (516 disciplines)MRR
C-Eval [paper] [github]NoEvaluate Chinese domain knowledge & reasoningMultiple-choice accuracy
BELLE-eval [github]NoEvaluate Chinese instruction-following & multi-skillGPT-4 win-rate & per-task scores
SuperCLUE [github]NoEvaluate Chinese instruction-following with alignmentGPT-4 win-rate & Elo
M3KE [paper] [github]NoEvaluate Chinese LLM knowledge (71 disciplines)Zero-/few-shot multitask accuracy
BayLing-80 [github]NoEvaluate cross-lingual & conversational capabilitiesGPT-4 adjudicated win-rate
MMCU [paper] [github] [huggingface]NoEvaluate multitask Chinese understandingMultitask accuracy
C-CLUE [github]NoEvaluate classical Chinese NER & REWeighted F1
LongBench [paper] [github] [huggingface]YesEvaluate bilingual long-context understandingAccuracy & generation quality
L-Eval [paper] [github]YesEvaluate long-context reasoning up to 60 K tokensMulti-metric assessment
InfinityBench [paper] [github]YesEvaluate contexts beyond 100 K tokensAccuracy & task-specific metrics
Marathon [paper] [github]YesEvaluate long-context reasoning across domainsEM / F1 / ROUGE-L / accuracy
LongEval [paper] [github]YesEvaluate effectiveness in long-context retrievalRetrieval accuracy
BABILong [paper] [github]YesEvaluate long-context reasoning in haystack settingsAccuracy
DetectiveQA [paper] [github]YesEvaluate narrative reasoning via detective novelsInstruction-following accuracy
NoCha [paper] [github]YesEvaluate narrative comprehension & coreferenceExact match & accuracy
Loong [paper] [github]YesEvaluate multi-document reasoning & QAAnswer accuracy & doc coverage
TCELongBench [paper] [github]YesEvaluate temporal reasoning over long event narrativesTemporal ordering & QA accuracy
DENIAHL [paper] [github]YesEvaluate in-context feature influence on NIAHRetrieval & reasoning accuracy
LongMemEval [paper]YesBenchmark memory retention in dialogueMemory retention & utilization
Long2RAG [paper]YesEvaluate long-form generation & retrieval groundingKey-point recall & factuality
L-CiteEval [paper] [github]YesEvaluate citation usage & contextual evidenceCitation recall & faithfulness
LIFBENCH [paper] [github]YesEvaluate instruction-following performance & stabilityAccuracy & response stability
LongReason [paper]YesEvaluate synthetic long-context reasoningInstruction-following accuracy
BAMBOO [paper] [github]YesEvaluate long-text modeling across tasksAccuracy & win-rate
ETHIC [paper]YesEvaluate instruction-following on high-info long tasksEM / F1 & consistency
LooGLE [paper] [github]YesEvaluate understanding & reasoning over 20 K-word inputsAccuracy
HELMET [paper] [github]YesEvaluate retrieval, reasoning, summarization, long-contextTask-specific automatic & human metrics
HoloBench [paper]YesEvaluate holistic reasoning over DB-style inputsExecution accuracy & consistency
LOFT [paper]YesEvaluate replacing RAG with long-context LLMsEM / F1 / SQL accuracy
Lv-Eval [paper] [github]YesEvaluate comprehension across length levels ≤ 256 KExact match & factuality
ManyICLBench [paper]YesEvaluate many-shot ICL under long contextsAverage accuracy
ZeroSCROLLS [paper]YesEvaluate zero-shot inference on diverse long-text tasksAccuracy
LongICLBench [paper]YesEvaluate ICL under extended input lengthsWin-rate judged by PandaLM
LIBRA [paper] [github]YesEvaluate long-context understanding in RussianAccuracy, BLEU, faithfulness

Domain: Finance

BenchmarkLong-context or notEvaluation ObjectiveMain Evaluation Criteria
FinBen [paper] [github] [huggingface]NoEvaluate holistic financial capabilities across 24 tasksAutomatic metrics, agent/RAG performance, and human expert assessment
PIXIU [paper] [github] [huggingface]NoEvaluate LLMs across multiple financial NLP tasksSentiment accuracy, QA accuracy, stock prediction F1
FLUE [paper] [github] [huggingface]NoEvaluate diverse financial NLP competenciesAccuracy, F1, nDCG
BBF-CFLEB [paper] [github]NoEvaluate LLMs on Chinese financial language understanding and generation across six task typesRouge, F1, and accuracy
CFinBench [paper] [github]NoEvaluate Chinese financial knowledge across subjects, certifications, practice, and legal complianceAccuracy across single-choice, multiple-choice, and judgment questions
SuperCLUEFin [paper]NoEvaluate Chinese financial assistant capabilitiesWin-rate and multi-criteria performance
ICE-PIXIU [paper] [github]NoEvaluate bilingual (Chinese–English) financial reasoning and analysis capabilitiesTask-specific accuracy and bilingual win-rate
FLARE-ES [paper]NoEvaluate bilingual Spanish–English financial reasoningTask-specific accuracy and cross-lingual transfer win-rate
FinanceBench [paper] [github]NoEvaluate financial open-book QA using real-world company-related questionsFactual correctness and evidence alignment
FiNER-ORD [paper] [github] [huggingface]NoEvaluate financial NER capability in financial textsEntity F1
FinRED [paper] [github] [huggingface]NoEvaluate financial relation extraction performance on news and earnings transcriptsF1, Entity F1
FinQA [paper] [github] [huggingface]NoEvaluate multi-step numerical reasoning over financial reports with structured evidenceEM Accuracy
BizBench [paper] [huggingface]NoEvaluate quantitative reasoning on realistic financial problemsNumeric EM, code execution pass rate, QA accuracy
EconLogicQA [paper] [huggingface]NoEvaluate economic sequential reasoning across multi-event scenariosMultiple-choice accuracy
FinEval [github]NoEvaluate Chinese financial domain knowledge and reasoningMultiple-choice accuracy and task-level weighted scores
CFBenchmark [paper] [github]NoEvaluate Chinese financial assistant capabilitiesWin-rate and task-specific metrics
BBT-Fin [paper] [github]NoEvaluate Chinese financial language understanding and generationAccuracy, F1, ROUGE
Hirano [paper] [github]NoEvaluate Japanese financial language understandingMultiple-choice accuracy and macro-F1 scores
MultiFin [paper] [github] [huggingface]NoEvaluate multilingual financial topic classificationF1, Multi-class accuracy
DocFinQA [paper] [huggingface]YesEvaluate long-context financial reasoning over documents like financial reportsEM and F1 for multi-step answer prediction, and reasoning accuracy
FinTextQA [huggingface]YesEvaluate long-form financial question answering with long textual contextAnswer accuracy, BLEU, and ROUGE

Domain: Medicine

BenchmarkLong-context or notEvaluation ObjectiveMain Evaluation Criteria
CBLUE [paper] [github]NoEvaluate Chinese biomedical language understanding across multiple clinical and QA tasksAccuracy, F1-score, and macro-average metrics across subtasks
PromptCBLUE [paper] [github] [huggingface]NoEvaluate LLMs on prompt-based generation across 16 Chinese medical NLP tasksAccuracy, BLEU, and ROUGE scores
CMB [paper] [github] [huggingface]NoEvaluate comprehensive Chinese medical knowledge via exam-style QA and clinical diagnosisAccuracy, expert grading, and model-based evaluation
HuaTuo26M-test [paper] [github] [huggingface]NoEvaluate Chinese medical knowledge and QA ability using real-world clinical queriesAccuracy and relevance
CMExam [paper] [github] [huggingface]NoEvaluate LLMs on Chinese medical licensing exam QA with fine-grained annotationsAccuracy, weighted F1, and expert-judged reasoning quality
MultiMedQA [paper]NoEvaluate LLMs’ clinical knowledge via multiple-choice and open-ended medical QA tasksExpert-rated helpfulness, factuality, and safety
QiZhenGPT eval [github]NoEvaluate LLMs' ability to identify drug indications from natural-language promptsExpert-annotated correctness score
MedExQA [paper] [github]NoEvaluate medical knowledge and explanation generation across under-represented specialtiesExplanation quality and expert-aligned relevance
JAMA and Medbullets [paper] [github]NoEvaluate LLMs' ability to answer and explain challenging clinical questionsAnswer accuracy, explanation quality, and human-aligned reasoning assessment
MedXpertQA [paper] [github] [huggingface]NoEvaluate expert-level medical reasoning and multimodal clinical understandingAnswer accuracy, image-text reasoning, and expert-aligned scoring
MedJourney [paper]NoEvaluate LLM performance across full clinical patient journey stages and tasksTask-specific automatic metrics and human expert evaluations
MedAgentsBench [paper] [github]NoEvaluate complex multi-step clinical reasoning including diagnosis and treatment planningMulti-aspect evaluation: correctness, efficiency, human expert ratings
LongHealth [paper] [github]YesEvaluate question answering over long-form clinical documentsEM, F1, and long-context QA accuracy
MedOdyssey [paper] [github]YesEvaluate long-context understanding in the medical domain (up to 200 K tokens)Task-specific EM, ROUGE, and human preference scoring

Domain: Code

BenchmarkLong-context or notEvaluation ObjectiveMain Evaluation Criteria
HumanEval [paper] [github] [huggingface]NoEvaluate code generation, algorithmic reasoning, and language understanding with functional correctnessTest-case execution accuracy (pass@k)
MBPP [paper] [github] [huggingface]NoEvaluate basic Python code generation on crowdsourced programming tasksFunctional correctness via pass@k using automated test cases
APPS [paper] [github] [huggingface]NoEvaluate coding-challenge competence through real-world programming problemsPass@k and exact match for functional correctness
DS-1000 [paper] [github] [huggingface]NoEvaluate data-science code generation across real queries from 7 Python librariesFunctional correctness via automated test-based execution accuracy
CodeXGLUE [paper] [github]NoEvaluate code understanding and generation across 9 tasks in 4 I/O typesBLEU, EM, F1, Accuracy, MAP (task-specific)
CruxEval [paper] [github]NoEvaluate code reasoning, understanding, and executionpass@1 accuracy
ODEX [paper] [github] [huggingface]NoEvaluate cross-lingual code generation from NL queries in four languagesExecution-based functional correctness
MTPB [paper] [github]NoEvaluate multi-turn program synthesisFunctional correctness via pass@k on step-wise sub-programs
ClassEval [paper] [github] [huggingface]NoEvaluate class-level code generation from NL descriptions (Python)pass@1, class completeness, dependency consistency
BigCodeBench [paper] [github] [huggingface]NoEvaluate LLMs’ ability to follow complex instructions and invoke diverse function callsPass@k, test-case execution accuracy, branch coverage
HumanEvalPack [github] [huggingface]NoEvaluate multilingual code generation, correction, and comment synthesis across six languagesFunctional correctness via pass@k and task-specific metrics
BIRD [pape] [github]NoEvaluate database-grounded text-to-SQL generation over large, noisy databasesExecution accuracy and exact match
RepoQA [github]YesEvaluate long-context code understanding in real software repositoriesExact-match accuracy and retrieval-augmented correctness
LongCodeArena [paper] [github] [huggingface]YesEvaluate long-context code comprehension, generation, and editing across project-wide tasksTask-specific accuracy, exact match, and human evaluation

Domain: Math

BenchmarkLong-context or notEvaluation ObjectiveMain Evaluation Criteria
GSM8K [paper] [github] [huggingface]NoAssess grade school math reasoningExact match accuracy
MATH [paper] [github] [huggingface]NoEvaluate competition-level math problem-solving with step-by-step reasoningFinal answer accuracy & derivation correctness
MathOdyssey [paper] [github] [huggingface]NoBenchmark high-school → Olympiad → university math across difficulty tiersAnswer accuracy across difficulty levels
MathBench [paper] [github]NoEvaluate theoretical & applied math knowledge across five levelsAccuracy on theoretical and application problems
CHAMP [paper] [github]NoFine-grained competition-level reasoning with concept/hint annotationsAnswer accuracy & reasoning-path correctness
LILA [paper] [github] [huggingface]NoUnified benchmark across 23 math tasks & formatsAccuracy across tasks
MiniF2F-v1 [paper] [github]NoFormal theorem-proving at Olympiad levelProof accuracy on 488 problems
ProofNet [paper] [github] [huggingface]NoAuto-formalization & formal proof generation (Lean 3)Formalization accuracy & proof success rate
AlphaGeometry [paper] [github]NoNeuro-symbolic reasoning on Olympiad Euclidean geometryProof success, correctness, completeness, readability
MathVerse [paper] [github] [huggingface]NoVisual-diagram math reasoning for MLLMsDiagram-sensitive answer accuracy & CoT score
We-Math [paper] [github] [huggingface]NoHuman-like visual mathematical reasoning with knowledge hierarchyFour-dimensional diagnostic metrics
U-MATH [paper] [github]NoOpen-ended university-level problem-solving (20 % multimodal)LLM-judged solution correctness (expert-verified F1)
TabMWP [paper] [github]NoMath reasoning over textual + tabular dataAccuracy on QA & MC questions
MathHay [paper] [github]YesLong-context mathematical reasoning with multi-step dependenciesAccuracy, exact match, reasoning-chain consistency

Domain: Law

BenchmarkLong-context or notEvaluation ObjectiveMain Evaluation Criteria
LegalBench [paper] [github] [huggingface]NoEvaluate legal reasoning across six types including rule application and interpretationTask-specific accuracy, rule-consistency, LLM-as-a-judge ratings
LexGLUE [paper] [github] [huggingface]NoEvaluate legal language understanding across classification and QA tasksTask-specific accuracy & F1
LEXTREME [paper] [github] [huggingface]NoEvaluate multilingual & multitask legal understanding (24 languages, 18 tasks)Macro-F1, accuracy, other classification metrics
LawBench [paper] [github]NoEvaluate Chinese legal LLMs across retention, understanding & application (20 tasks)Task-specific accuracy & F1
LAiW [paper] [github]NoEvaluate Chinese legal LLMs across fundamental → advanced tasks (13 assignments)Task-specific accuracy & F1
LexEval [paper] [github] [huggingface]NoEvaluate Chinese legal reasoning via a taxonomy of cognitive abilitiesTask-specific accuracy
CitaLaw [paper]NoEvaluate citation-grounded legal answering with statutes & precedentsSyllogism-alignment, citation accuracy, legal consistency
LegalAgentBench [paper] [github]NoEvaluate LLM agents solving complex real-world legal tasksTask success rate & intermediate progress
SCALE [paper] [github]YesEvaluate long-doc, multilingual, multitask legal reasoning (≤ 50 K tokens)Accuracy, F1, code-based assessment for long-context legal tasks

⭐ Citation

If you find this work useful, welcome to cite us.

@misc{wu2025surveyfederatedfinetuninglarge,
      title={A Survey on Federated Fine-tuning of Large Language Models}, 
      author={Yebo Wu and Chunlin Tian and Jingguang Li and He Sun and Kahou Tam and Zhanting Zhou and Haicheng Liao and Zhijiang Guo and Li Li and Chengzhong Xu},
      year={2025},
      eprint={2503.12016},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2503.12016}, 
}

Contributors

Clin0212

4 commits

Clin0212/Awesome-Federated-LLM-Learning

Latest Advances on Federated LLM Learning

111

4 commits

updated Jul 7, 2025

See the code

README

Awesome-Federated-LLM-Learning

Contribution Welcome

📢 Updates

We released a survey paper "A Survey on Federated Fine-tuning of Large Language Models". Feel free to cite or open pull requests.

⚠️ NOTE: If there is any missing or new relevant literature, please feel free to submit an issue. we will update the Github and Arxiv papers regularly. 😊

👀 Overall Structure

alt text

📒 Table of Contents

Part 1: LoRA-based Tuning

alt text

1.1 Homogeneous LoRA

  • Fedra: A random allocation strategy for federated tuning to unleash the power of heterogeneous clients. [Paper]
  • Towards building the federatedGPT: Federated instruction tuning.[Paper]
  • Communication-Efficient and Tensorized Federated Fine-Tuning of Large Language Models. [Paper]
  • Selective Aggregation for Low-Rank Adaptation in Federated Learning. [Paper]
  • Federa: Efficient fine-tuning of language models in federated learning leveraging weight decomposition. [Paper]
  • LoRA-FAIR: Federated LoRA Fine-Tuning with Aggregation and Initialization Refinement. [Paper]
  • Federated LoRA with Sparse Communication. [Paper]
  • SA-FedLora: Adaptive Parameter Allocation for Efficient Federated Learning with LoRA Tuning. [Paper]
  • SLoRA: Federated parameter efficient fine-tuning of language models. [Paper]
  • FederatedScope-LLM: A Comprehensive Package for Fine-tuning Large Language Models in Federated Learning [Paper]
  • Robust Federated Finetuning of Foundation Models via Alternating Minimization of LoRA. [Paper]
  • Automated federated pipeline for parameter-efficient fine-tuning of large language models. [Paper]
  • Low-Parameter Federated Learning with Large Language Models. [Paper]
  • Towards Robust and Efficient Federated Low-Rank Adaptation with Heterogeneous Clients. [Paper]
  • FedRA: A Random Allocation Strategy for Federated Tuning to Unleash the Power of Heterogeneous Clients. [Paper]
  • Fed-piLot: Optimizing LoRA Assignment for Efficient Federated Foundation Model Fine-Tuning. [Paper]

1.2 Heterogeneous LoRA

  • Heterogeneous lora for federated fine-tuning of on-device foundation models. [Paper]
  • Flora: Federated fine-tuning large language models with heterogeneous low-rank adaptations. [Paper]
  • Federated fine-tuning of large language models under heterogeneous tasks and client resources. [Paper]
  • Federated LLMs Fine-tuned with Adaptive Importance-Aware LoRA. [Paper]
  • Towards Federated Low-Rank Adaptation of Language Models with Rank Heterogeneity. [Paper]
  • Fedhm: Efficient federated learning for heterogeneous models via low-rank factorization. [Paper]
  • RBLA: Rank-Based-LoRA-Aggregation for Fine-Tuning Heterogeneous Models. [Paper]

1.3 Personalized LoRA

  • FDLoRA: Personalized Federated Learning of Large Language Model via Dual LoRA Tuning. [Paper]
  • Fedlora: Model-heterogeneous personalized federated learning with lora tuning. [Paper]
  • FedLoRA: When Personalized Federated Learning Meets Low-Rank Adaptation. [Paper]
  • Dual-Personalizing Adapter for Federated Foundation Models. [Paper]
  • Personalized Federated Instruction Tuning via Neural Architecture Search. [Paper]
  • Communication-Efficient Personalized Federated Learning for Speech-to-Text Tasks. [Paper]
  • Personalized Federated Fine-Tuning for LLMs via Data-Driven Heterogeneous Model Architectures. [Paper]

Part 2: Prompt-based Tuning

2.1 General Prompt Tuning

  • Prompt federated learning for weather forecasting: Toward foundation models on meteorological data. [Paper]
  • Promptfl: Let federated participants cooperatively learn prompts instead of models-federated learning in age of foundation model. [Paper]
  • Fedbpt: Efficient federated black-box prompt tuning for large language models. [Paper]
  • Federated learning of large language models with parameter-efficient prompt tuning and adaptive optimization. [Paper]
  • Efficient federated prompt tuning for black-box large pre-trained models. [Paper]
  • Text-driven prompt generation for vision-language models in federated learning. [Paper]
  • Learning federated visual prompt in null space for mri reconstruction. [Paper]
  • Fed-cprompt: Contrastive prompt for rehearsal-free federated continual learning. [Paper]
  • Fedprompt: Communication-efficient and privacy-preserving prompt tuning in federated learning. [Paper]
  • Tunable soft prompts are messengers in federated learning. [Paper]
  • Hepco: Data-free heterogeneous prompt consolidation for continual federated learning. [Paper]
  • Prompt-enhanced Federated Learning for Aspect-Based Sentiment Analysis. [Paper]
  • Towards practical few-shot federated nlp. [Paper]
  • Federated prompting and chain-of-thought reasoning for improving llms answering. [Paper]
  • FedHPL: Efficient Heterogeneous Federated Learning with Prompt Tuning and Logit Distillation. [Paper]
  • Probabilistic Federated Prompt-Tuning with Non-IID and Imbalanced Data. [Paper]
  • Federated Class-Incremental Learning with Prompting. [Paper]
  • Explore and Cure: Unveiling Sample Effectiveness with Context-Aware Federated Prompt Tuning. [Paper]
  • Federated Prompt Learning for Weather Foundation Models on Devices. [Paper]

2.2 Personalized Prompt Tuning

  • Efficient model personalization in federated learning via client-specific prompt generation. [Paper]
  • Unlocking the potential of prompt-tuning in bridging generalized and personalized federated learning. [Paper]
  • Pfedprompt: Learning personalized prompt for vision-language models in federated learning. [Paper]
  • Global and local prompts cooperation via optimal transport for federated learning. [Paper]
  • Visual prompt based personalized federated learning. [Paper]
  • Personalized federated continual learning via multi-granularity prompt. [Paper]
  • FedLPPA: Learning Personalized Prompt and Aggregation for Federated Weakly-supervised Medical Image Segmentation. [Paper]
  • Harmonizing Generalization and Personalization in Federated Prompt Learning. [Paper]
  • Tackling Feature-Classifier Mismatch in Federated Learning via Prompt-Driven Feature Transformation. [Paper]
  • Personalized Federated Learning for Text Classification with Gradient-Free Prompt Tuning. [Paper]
  • Mixture of Experts Made Personalized: Federated Prompt Learning for Vision-Language Models. [Paper]
  • CP 2 GFed: Cross-granular and Personalized Prompt-based Green Federated Tuning for Giant Models. [Paper]

2.3 Multi-domain Prompt Tuning

  • DiPrompT: Disentangled Prompt Tuning for Multiple Latent Domain Generalization in Federated Learning. [Paper]
  • Prompt-enhanced Federated Content Representation Learning for Cross-domain Recommendation. [Paper]
  • Dual prompt tuning for domain-aware federated learning. [Paper]
  • Federated adaptive prompt tuning for multi-domain collaborative learning. [Paper]
  • Breaking physical and linguistic borders: Multilingual federated prompt tuning for low-resource languages. [Paper]
  • Federated Domain Generalization via Prompt Learning and Aggregation. [Paper]
  • CP-Prompt: Composition-Based Cross-modal Prompting for Domain-Incremental Continual Learning. [Paper]

Part 3: Adapter-based Tuning

3.1 General Adapter Tuning

  • Efficient federated learning for modern nlp. [Paper]
  • Efficient federated learning with pre-trained large language model using several adapter mechanisms. [Paper]

3.2 Personalized Adapter Tuning

  • Client-customized adaptation for parameter-efficient federated learning. [Paper]
  • Fedclip: Fast generalization and personalization for clip in federated learning. [Paper]

3.3 Multi-domain Adapter Tuning

  • Communication efficient federated learning for multilingual neural machine translation with adapter. [Paper]
  • Adapter-based Selective Knowledge Distillation for Federated Multi-domain Meeting Summarization. [Paper]
  • Feddat: An approach for foundation model finetuning in multi-modal heterogeneous federated learning. [Paper]

Part 4: Selective-based Tuning

4.1 Bias Tuning

  • Differentially private bias-term only fine-tuning of foundation models. [Paper]
  • Conquering the communication constraints to enable large pre-trained models in federated learning. [Paper]

4.2 Partial Tuning

  • Bridging the gap between foundation models and heterogeneous federated learning. [Paper]
  • Exploring Selective Layer Fine-Tuning in Federated Learning. [Paper]

Part 5: Other Tuning Methods

5.1 Zero-Order Optimization

  • Federated full-parameter tuning of billion-sized language models with communication cost under 18 kilobytes. [Paper]
  • ${$FwdLLM$}$: Efficient Federated Finetuning of Large Language Models with Perturbed Inferences. [Paper]
  • ZooPFL: Exploring black-box foundation models for personalized federated learning. [Paper]
  • On the convergence of zeroth-order federated tuning for large language models. [Paper]
  • Thinking Forward: Memory-Efficient Federated Finetuning of Language Models. [Paper]
  • Communication-Efficient Byzantine-Resilient Federated Zero-Order Optimization. [Paper]

5.2 Split Learning

  • FedBERT: When federated learning meets pre-training. [Paper]
  • Federated split bert for heterogeneous text classification. [Paper]
  • FedSplitX: Federated Split Learning for Computationally-Constrained Heterogeneous Clients. [Paper]

5.3 Model Compression

  • Fedbiot: Llm local fine-tuning in federated learning without full model. [Paper]

5.4 Data Selection

  • Federated Data-Efficient Instruction Tuning for Large Language Models. [Paper]

Datasets and Benchmarks

Prompt-tuning Datasets

Domain: General

DatasetLanguageConstruction MethodDescription
Alpaca [paper] [github] [huggingface]EnglishModel ConstructGenerated by Text-Davinci-003 with Alpaca-style instruction prompts.
Alpaca-GPT4 [paper] [github] [huggingface]EnglishModel ConstructGenerated by GPT-4 based on Alpaca prompts with richer multi-turn instructions.
Self-Instruct [paper] [github] [huggingface]EnglishHuman + Model ConstructSeed instructions expanded by GPT-3 to improve model generalization.
UltraChat 200k [paper] [github] [huggingface]EnglishModel ConstructHigh-quality multi-turn dialogue subset filtered from UltraChat.
OpenOrca [paper] [github] [huggingface]EnglishModel Construct~4.2 M GPT-3.5/4-augmented FLAN examples for instruction following.
ShareGPT90K [github] [huggingface]EnglishModel Construct90 K multi-turn dialogues extracted from ShareGPT.
WizardLM Evol-Instruct V2 196k [paper] [github] [huggingface]EnglishModel Construct196 K examples generated via Evol-Instruct.
Databricks Dolly 15K [paper] [github] [huggingface]EnglishHuman Construct15 K human-written prompt-response pairs across diverse tasks.
Baize [paper] [github] [huggingface]EnglishModel ConstructInstruction-following dialogues generated via ChatGPT self-chat.
OpenChat [paper] [github] [huggingface]EnglishModel ConstructMixed-quality data alignment using C-RLFT for open LLMs.
Flan-v2 [paper] [github] [huggingface]EnglishModel ConstructAggregates Flan, P3, Super-Natural Instructions, CoT, Dialog tasks.
BELLE-train-0.5M-CN [paper] [github] [huggingface]ChineseHuman + Model Construct519 K Chinese instruction-following examples.
BELLE-train-1M-CN [paper] [github] [huggingface]ChineseHuman + Model Construct917 K Chinese instruction samples.
BELLE-train-2M-CN [paper] [github] [huggingface]ChineseHuman + Model Construct2 M Chinese instruction samples.
Firefly-train-1.1M [paper] [github] [huggingface]ChineseHuman Construct1.65 M Chinese samples across 23 tasks with human templates.
Wizard-LM-Chinese-instruct-evol [paper] [github] [huggingface]ChineseHuman + Model Construct70 K WizardLM instructions translated into Chinese.
HC3-Chinese [paper] [github] [huggingface]ChineseHuman + Model ConstructHuman-ChatGPT QA pairs in Chinese.
HC3 [paper] [github] [huggingface]English / ChineseHuman + Model ConstructBilingual Human-ChatGPT QA pairs.
ShareGPT-Chinese-English-90k [github] [huggingface]English / ChineseModel Construct90 K bilingual user queries from ShareGPT logs.

Domain: Finance

DatasetLanguageConstruction MethodDescription
FinGPT [paper] [github] [huggingface]EnglishHuman + Model ConstructInstruction-tuning data for diverse financial tasks.
Finance-Instruct-500k [paper] [github] [huggingface]EnglishHuman + Model ConstructLarge-scale (500 K) instruction dataset for finance reasoning.
Finance-Alpaca [github] [huggingface]EnglishHuman + Model ConstructAlpaca-style instructions combined with FiQA and custom Q&A.
Financial PhraseBank [paper] [github] [huggingface]EnglishHuman ConstructManually annotated news sentences for sentiment classification.
Yahoo-Finance-Data [huggingface]EnglishHuman ConstructHistorical prices and fundamentals scraped from Yahoo Finance.
Financial-QA-10K [github] [huggingface]EnglishModel Construct10 K contextual QA pairs generated from SEC 10-K filings.
Financial-Classification [huggingface]EnglishHuman ConstructMerged Financial PhraseBank + Kaggle texts for sentiment/topic classification.
Twitter-Financial-News-Topic [github] [huggingface]EnglishHuman Construct21 K annotated tweets for multi-class financial topic tagging.
Financial-News-Articles [github] [huggingface]EnglishHuman Construct300 K+ news articles for text classification and sentiment analysis.
FiQA [paper] [github] [huggingface]EnglishHuman ConstructFinancial question-answering dataset from forums and texts.
Earnings-Call [paper] [huggingface]EnglishHuman ConstructQA pairs extracted from CEO/CFO earnings-call transcripts.
Doc2EDAG [paper] [github]ChineseHuman ConstructChinese financial reports annotated for document-level event graphs.
Synthetic-PII-Finance-Multilingual [paper] [github] [huggingface]MultilingualSyntheticSynthetic financial documents with labeled PII for privacy-preserving NER.

Domain: Medicine

DatasetLanguageConstruction MethodDescription
ChatDoctor-200K [paper] [github] [huggingface]EnglishHuman + Model ConstructInstruction tuning dataset for medical QA and dialogue generation.
ChatDoctor-HealthCareMagic-100k [paper] [github] [huggingface]EnglishHuman ConstructReal-world doctor-patient conversations from HealthCareMagic.
Medical Meadow CORD-19 [paper] [github] [huggingface]EnglishHuman ConstructSummaries of biomedical papers from CORD-19 for instruction tuning.
Medical Meadow MedQA [paper] [github] [huggingface]EnglishHuman ConstructMultiple-choice medical QA derived from MedQA exam questions.
HealthCareMagic-100k-en [paper] [github] [huggingface]EnglishHuman ConstructEnglish subset of HealthCareMagic doctor-patient consultations.
ChatMed-Consult-Dataset [github] [huggingface]ChineseHuman + Model ConstructChinese medical consultations with GPT-3.5 answers.
CMtMedQA [paper] [github]ChineseHuman Construct70 K multi-turn doctor-patient QA dialogues for Chinese medical reasoning.
DISC-Med-SFT [paper] [github] [huggingface]ChineseHuman + Model Construct470 K instruction pairs combining real dialogues and KG-based QA.
Huatuo-26M [paper] [github] [huggingface]ChineseHuman Construct26 M QA pairs extracted from encyclopedias, KBs and consultations.
Huatuo26M-Lite [paper] [github] [huggingface]ChineseHuman + Model ConstructRefined subset of Huatuo-26M with ChatGPT-rewritten answers.
ShenNong-TCM-Dataset [github] [huggingface]ChineseHuman + Model Construct110 K TCM-centric instructions generated via entity-centric self-instruct.
HuatuoGPT-sft-data-v1 [paper] [github] [huggingface]ChineseHuman + Model ConstructSFT corpus mixing ChatGPT-distilled and real doctor data for HuatuoGPT.
MedDialog [paper] [github] [huggingface]English / ChineseHuman ConstructLarge-scale doctor-patient dialogue corpora (0.3 M EN / 3.4 M CN conversations).

Domain: Code

Dataset / Benchmark NameLanguageConstruction MethodDescription
CodeAlpaca [github] [huggingface]EnglishModel ConstructGPT-generated code instruction-following dataset in Alpaca style.
Code Instructions 120k Alpaca [huggingface]EnglishHuman + Model Construct120 k natural-language ↔ code instruction pairs with Alpaca-format prompts.
CodeContests [github] [huggingface]EnglishHuman ConstructCompetitive-programming problems and solutions for program-synthesis research.
CommitPackFT [paper] [github] [huggingface]EnglishHuman Construct2 GB filtered Git commits with high-quality messages for code instruction tuning.
ToolBench [paper] [github]EnglishHuman + Model ConstructInstruction dataset for multi-tool API usage and tool-calling agents.
CodeParrot-Clean [huggingface]EnglishHuman ConstructDeduplicated & filtered GitHub Python corpus for code-generation pre-training.
The Stack v2 Dedup [paper] [huggingface]EnglishHuman ConstructLarge-scale (600 + languages) deduplicated source-code dataset from BigCode.
CodeSearchNet [paper] [github] [huggingface]EnglishHuman Construct6 M code–doc pairs across six languages for code search & retrieval.
CodeForces-CoTs [github] [huggingface]EnglishHuman + Model Construct10 k CodeForces problems with chain-of-thought traces distilled by DeepSeek R1.
CodeXGLUE Code Refinement [paper] [github] [huggingface]EnglishHuman ConstructBuggy ↔ fixed Java function pairs for automatic code repair and refinement.

Domain: Math

DatasetLanguageConstruction MethodDescription
GSM8K [paper] [github] [huggingface]EnglishHuman ConstructA dataset of 8.5 K grade-school arithmetic word problems with step-by-step solutions.
CoT-GSM8k [huggingface]EnglishHuman + Model ConstructExtended GSM8K with explicit chain-of-thought reasoning traces.
MathInstruct [paper] [github] [huggingface]EnglishHuman + Model ConstructHybrid CoT + PoT rationales spanning diverse mathematical fields.
MetaMathQA [paper] [github] [huggingface]EnglishHuman + Model ConstructMulti-perspective question augmentations bootstrapped from GSM8K & MATH.
OpenR1-Math-220k [github] [huggingface]EnglishHuman + Model Construct220 K problems with multiple DeepSeek R1 reasoning traces.
Hendrycks MATH Benchmark [paper] [github] [huggingface]EnglishHuman Construct12.5 K high-school competition problems with detailed solutions.
DeepMind Mathematics Dataset [paper] [github] [huggingface]EnglishSyntheticAlgorithmically generated problems covering many school-math topics.
OpenMathInstruct-1 [paper] [github] [huggingface]EnglishHuman + Model Construct1.8 M problems with Mixtral code-interpreter solutions.
Orca-Math Word Problems 200k [paper] [huggingface]EnglishSynthetic200 K grade-school word problems distilled with GPT-4 Turbo.
DAPO-Math-17k [paper] [huggingface]EnglishHuman + Model Construct17 K math prompts curated for large-scale RL (GRPO) experiments.
Big-Math-RL-Verified [paper] [github] [huggingface]EnglishHuman + Model Construct251 K verifiable problems filtered for reinforcement-learning fine-tuning.
BELLE-math-zh [github] [huggingface]ChineseHuman + Model Construct250 K Chinese elementary-math problems with step-by-step solutions.
MathInstruct-Chinese [huggingface]ChineseModel ConstructChinese translation/extension of MathInstruct for instruction tuning.

Domain: Law

Dataset / Benchmark NameLanguageConstruction MethodDescription
Legal-QA-v1 [huggingface]EnglishHuman Construct3.7 K QA pairs sourced from legal forums.
Pile of Law [paper] [github] [huggingface]EnglishHuman Construct256 GB corpus of U.S. legal and administrative text for domain pre-training.
CUAD [paper] [github] [huggingface]EnglishHuman Construct26 K expert-annotated contract QA pairs spanning 41 clause types.
LEDGAR [paper] [github] [huggingface]EnglishHuman Construct1.45 M labeled contract clauses for multi-label classification.
DISC-Law-SFT [paper] [github] [huggingface]ChineseHuman + Model Construct295 K supervised fine-tuning samples covering extraction, judgment prediction, QA, & summarization.
Law-GPT-zh [paper] [github] [huggingface]ChineseHuman ConstructSentence-pair corpus devised for Chinese legal sentence embedding & instructional tuning.
Lawyer LLaMA Data [paper] [github] [huggingface]ChineseHuman + Model ConstructInstruction data for legal consultation and bar-exam QA used to train Lawyer LLaMA.

Evaluation Benchmarks

Domain: General

BenchmarkLong-context or notEvaluation ObjectiveMain Evaluation Criteria
MMLU [paper] [github] [huggingface]NoEvaluate multitask language understanding across 57 subjectsMultitask accuracy
BIG-bench [paper] [github] [huggingface]NoEvaluate advanced reasoning capabilitiesModel performance and calibration
DROP [paper] [github] [huggingface]NoEvaluate discrete reasoning over paragraphsExact match and F1 score
CRASS [paper] [github]NoEvaluate counterfactual reasoning ability in LLMsMultiple-choice accuracy
ARC [paper] [github] [huggingface]NoAssess science reasoning at grade-school levelMultiple-choice accuracy
AGIEval [paper] [github] [huggingface]NoEvaluate foundation models on human-centric standardized exam tasksMulti-task accuracy across disciplines
M3Exam [paper] [github]NoEvaluate multilingual, multimodal, and multilevel reasoning across real exam questionsMultiple-choice accuracy
SCIBENCH [paper] [github]NoEvaluate college-level scientific problem-solving in math, physics, and chemistryOpen-ended accuracy and skill-specific error attribution
Vicuna Evaluation [paper] [github]NoEvaluate instruction-following quality in chat settingsHuman / GPT-4 preference comparison accuracy
MT-Bench [paper] [github]NoEvaluate multi-turn conversational and instruction-following capabilitiesWin-rate judged by GPT-4
AlpacaEval [paper] [github]NoEvaluate instruction-following via LLM-based auto-annotationLength-controlled win-rate correlated with human preference
Chatbot Arena [github]NoEvaluate LLMs via human-voted battles using Elo ratingElo score and head-to-head win rate
PandaLM [paper] [github]NoEvaluate instruction-following quality and hyperparameter impactWin-rate judged by PandaLM
HellaSwag [paper] [github] [huggingface]NoEvaluate commonsense inferenceMultiple-choice accuracy
TruthfulQA [paper] [github] [huggingface]NoEvaluate LLM truthfulness and avoidance of imitative falsehoodsTruthfulness rate and MC accuracy
ScienceQA [paper] [github] [huggingface]NoEvaluate multimodal scientific reasoning and explanation generationMultiple-choice accuracy and explanation quality
Chain-of-Thought Hub [paper] [github]NoEvaluate LLMs’ multi-step reasoning with CoT promptingFew-shot CoT accuracy
NeuLR [paper] [github]NoEvaluate deductive, inductive, and abductive reasoningMulti-dimensional accuracy
ALCUNA [paper] [github]NoEvaluate comprehension & reasoning over novel knowledgeAccuracy on 84 351 queries
LMExamQA [paper] [github]NoEvaluate knowledge recall, understanding, and analysisAccuracy on 10 090 questions
SocKET [paper] [github]NoEvaluate LLMs’ sociability & social knowledgeAccuracy and other metrics
Choice-75 [paper]NoEvaluate decision reasoning in scripted scenariosAccuracy on binary multiple-choice questions
HELM [paper] [github]NoEvaluate LMs via multi-metric scenariosComposite normalized performance
OpenLLM [github]NoEvaluate open-style reasoning across multiple benchmarksNormalized accuracy aggregate
BOSS [paper] [github]NoEvaluate OOD robustness across NLP tasksOOD accuracy drop & ID–OOD correlation
GLUE-X [github]NoEvaluate OOD robustnessAverage OOD accuracy drop
PromptBench [paper] [github]NoEvaluate robustness / prompt-engineeringAdversarial success & robustness
DynaBench [paper] [github]NoEvaluate robustness via dynamic human-in-loop dataError rate on human-crafted challenges
KoLA [paper] [github]NoEvaluate world knowledge across 19 evolving tasksSelf-contrast calibration
CELLO [paper] [github]NoEvaluate following complex real-world instructionsMulti-criteria compliance rate
LLMEval [github]NoMeta-evaluate LLM evaluatorsMeta-evaluator accuracy
Xiezhi [paper] [github]NoEvaluate holistic domain knowledge (516 disciplines)MRR
C-Eval [paper] [github]NoEvaluate Chinese domain knowledge & reasoningMultiple-choice accuracy
BELLE-eval [github]NoEvaluate Chinese instruction-following & multi-skillGPT-4 win-rate & per-task scores
SuperCLUE [github]NoEvaluate Chinese instruction-following with alignmentGPT-4 win-rate & Elo
M3KE [paper] [github]NoEvaluate Chinese LLM knowledge (71 disciplines)Zero-/few-shot multitask accuracy
BayLing-80 [github]NoEvaluate cross-lingual & conversational capabilitiesGPT-4 adjudicated win-rate
MMCU [paper] [github] [huggingface]NoEvaluate multitask Chinese understandingMultitask accuracy
C-CLUE [github]NoEvaluate classical Chinese NER & REWeighted F1
LongBench [paper] [github] [huggingface]YesEvaluate bilingual long-context understandingAccuracy & generation quality
L-Eval [paper] [github]YesEvaluate long-context reasoning up to 60 K tokensMulti-metric assessment
InfinityBench [paper] [github]YesEvaluate contexts beyond 100 K tokensAccuracy & task-specific metrics
Marathon [paper] [github]YesEvaluate long-context reasoning across domainsEM / F1 / ROUGE-L / accuracy
LongEval [paper] [github]YesEvaluate effectiveness in long-context retrievalRetrieval accuracy
BABILong [paper] [github]YesEvaluate long-context reasoning in haystack settingsAccuracy
DetectiveQA [paper] [github]YesEvaluate narrative reasoning via detective novelsInstruction-following accuracy
NoCha [paper] [github]YesEvaluate narrative comprehension & coreferenceExact match & accuracy
Loong [paper] [github]YesEvaluate multi-document reasoning & QAAnswer accuracy & doc coverage
TCELongBench [paper] [github]YesEvaluate temporal reasoning over long event narrativesTemporal ordering & QA accuracy
DENIAHL [paper] [github]YesEvaluate in-context feature influence on NIAHRetrieval & reasoning accuracy
LongMemEval [paper]YesBenchmark memory retention in dialogueMemory retention & utilization
Long2RAG [paper]YesEvaluate long-form generation & retrieval groundingKey-point recall & factuality
L-CiteEval [paper] [github]YesEvaluate citation usage & contextual evidenceCitation recall & faithfulness
LIFBENCH [paper] [github]YesEvaluate instruction-following performance & stabilityAccuracy & response stability
LongReason [paper]YesEvaluate synthetic long-context reasoningInstruction-following accuracy
BAMBOO [paper] [github]YesEvaluate long-text modeling across tasksAccuracy & win-rate
ETHIC [paper]YesEvaluate instruction-following on high-info long tasksEM / F1 & consistency
LooGLE [paper] [github]YesEvaluate understanding & reasoning over 20 K-word inputsAccuracy
HELMET [paper] [github]YesEvaluate retrieval, reasoning, summarization, long-contextTask-specific automatic & human metrics
HoloBench [paper]YesEvaluate holistic reasoning over DB-style inputsExecution accuracy & consistency
LOFT [paper]YesEvaluate replacing RAG with long-context LLMsEM / F1 / SQL accuracy
Lv-Eval [paper] [github]YesEvaluate comprehension across length levels ≤ 256 KExact match & factuality
ManyICLBench [paper]YesEvaluate many-shot ICL under long contextsAverage accuracy
ZeroSCROLLS [paper]YesEvaluate zero-shot inference on diverse long-text tasksAccuracy
LongICLBench [paper]YesEvaluate ICL under extended input lengthsWin-rate judged by PandaLM
LIBRA [paper] [github]YesEvaluate long-context understanding in RussianAccuracy, BLEU, faithfulness

Domain: Finance

BenchmarkLong-context or notEvaluation ObjectiveMain Evaluation Criteria
FinBen [paper] [github] [huggingface]NoEvaluate holistic financial capabilities across 24 tasksAutomatic metrics, agent/RAG performance, and human expert assessment
PIXIU [paper] [github] [huggingface]NoEvaluate LLMs across multiple financial NLP tasksSentiment accuracy, QA accuracy, stock prediction F1
FLUE [paper] [github] [huggingface]NoEvaluate diverse financial NLP competenciesAccuracy, F1, nDCG
BBF-CFLEB [paper] [github]NoEvaluate LLMs on Chinese financial language understanding and generation across six task typesRouge, F1, and accuracy
CFinBench [paper] [github]NoEvaluate Chinese financial knowledge across subjects, certifications, practice, and legal complianceAccuracy across single-choice, multiple-choice, and judgment questions
SuperCLUEFin [paper]NoEvaluate Chinese financial assistant capabilitiesWin-rate and multi-criteria performance
ICE-PIXIU [paper] [github]NoEvaluate bilingual (Chinese–English) financial reasoning and analysis capabilitiesTask-specific accuracy and bilingual win-rate
FLARE-ES [paper]NoEvaluate bilingual Spanish–English financial reasoningTask-specific accuracy and cross-lingual transfer win-rate
FinanceBench [paper] [github]NoEvaluate financial open-book QA using real-world company-related questionsFactual correctness and evidence alignment
FiNER-ORD [paper] [github] [huggingface]NoEvaluate financial NER capability in financial textsEntity F1
FinRED [paper] [github] [huggingface]NoEvaluate financial relation extraction performance on news and earnings transcriptsF1, Entity F1
FinQA [paper] [github] [huggingface]NoEvaluate multi-step numerical reasoning over financial reports with structured evidenceEM Accuracy
BizBench [paper] [huggingface]NoEvaluate quantitative reasoning on realistic financial problemsNumeric EM, code execution pass rate, QA accuracy
EconLogicQA [paper] [huggingface]NoEvaluate economic sequential reasoning across multi-event scenariosMultiple-choice accuracy
FinEval [github]NoEvaluate Chinese financial domain knowledge and reasoningMultiple-choice accuracy and task-level weighted scores
CFBenchmark [paper] [github]NoEvaluate Chinese financial assistant capabilitiesWin-rate and task-specific metrics
BBT-Fin [paper] [github]NoEvaluate Chinese financial language understanding and generationAccuracy, F1, ROUGE
Hirano [paper] [github]NoEvaluate Japanese financial language understandingMultiple-choice accuracy and macro-F1 scores
MultiFin [paper] [github] [huggingface]NoEvaluate multilingual financial topic classificationF1, Multi-class accuracy
DocFinQA [paper] [huggingface]YesEvaluate long-context financial reasoning over documents like financial reportsEM and F1 for multi-step answer prediction, and reasoning accuracy
FinTextQA [huggingface]YesEvaluate long-form financial question answering with long textual contextAnswer accuracy, BLEU, and ROUGE

Domain: Medicine

BenchmarkLong-context or notEvaluation ObjectiveMain Evaluation Criteria
CBLUE [paper] [github]NoEvaluate Chinese biomedical language understanding across multiple clinical and QA tasksAccuracy, F1-score, and macro-average metrics across subtasks
PromptCBLUE [paper] [github] [huggingface]NoEvaluate LLMs on prompt-based generation across 16 Chinese medical NLP tasksAccuracy, BLEU, and ROUGE scores
CMB [paper] [github] [huggingface]NoEvaluate comprehensive Chinese medical knowledge via exam-style QA and clinical diagnosisAccuracy, expert grading, and model-based evaluation
HuaTuo26M-test [paper] [github] [huggingface]NoEvaluate Chinese medical knowledge and QA ability using real-world clinical queriesAccuracy and relevance
CMExam [paper] [github] [huggingface]NoEvaluate LLMs on Chinese medical licensing exam QA with fine-grained annotationsAccuracy, weighted F1, and expert-judged reasoning quality
MultiMedQA [paper]NoEvaluate LLMs’ clinical knowledge via multiple-choice and open-ended medical QA tasksExpert-rated helpfulness, factuality, and safety
QiZhenGPT eval [github]NoEvaluate LLMs' ability to identify drug indications from natural-language promptsExpert-annotated correctness score
MedExQA [paper] [github]NoEvaluate medical knowledge and explanation generation across under-represented specialtiesExplanation quality and expert-aligned relevance
JAMA and Medbullets [paper] [github]NoEvaluate LLMs' ability to answer and explain challenging clinical questionsAnswer accuracy, explanation quality, and human-aligned reasoning assessment
MedXpertQA [paper] [github] [huggingface]NoEvaluate expert-level medical reasoning and multimodal clinical understandingAnswer accuracy, image-text reasoning, and expert-aligned scoring
MedJourney [paper]NoEvaluate LLM performance across full clinical patient journey stages and tasksTask-specific automatic metrics and human expert evaluations
MedAgentsBench [paper] [github]NoEvaluate complex multi-step clinical reasoning including diagnosis and treatment planningMulti-aspect evaluation: correctness, efficiency, human expert ratings
LongHealth [paper] [github]YesEvaluate question answering over long-form clinical documentsEM, F1, and long-context QA accuracy
MedOdyssey [paper] [github]YesEvaluate long-context understanding in the medical domain (up to 200 K tokens)Task-specific EM, ROUGE, and human preference scoring

Domain: Code

BenchmarkLong-context or notEvaluation ObjectiveMain Evaluation Criteria
HumanEval [paper] [github] [huggingface]NoEvaluate code generation, algorithmic reasoning, and language understanding with functional correctnessTest-case execution accuracy (pass@k)
MBPP [paper] [github] [huggingface]NoEvaluate basic Python code generation on crowdsourced programming tasksFunctional correctness via pass@k using automated test cases
APPS [paper] [github] [huggingface]NoEvaluate coding-challenge competence through real-world programming problemsPass@k and exact match for functional correctness
DS-1000 [paper] [github] [huggingface]NoEvaluate data-science code generation across real queries from 7 Python librariesFunctional correctness via automated test-based execution accuracy
CodeXGLUE [paper] [github]NoEvaluate code understanding and generation across 9 tasks in 4 I/O typesBLEU, EM, F1, Accuracy, MAP (task-specific)
CruxEval [paper] [github]NoEvaluate code reasoning, understanding, and executionpass@1 accuracy
ODEX [paper] [github] [huggingface]NoEvaluate cross-lingual code generation from NL queries in four languagesExecution-based functional correctness
MTPB [paper] [github]NoEvaluate multi-turn program synthesisFunctional correctness via pass@k on step-wise sub-programs
ClassEval [paper] [github] [huggingface]NoEvaluate class-level code generation from NL descriptions (Python)pass@1, class completeness, dependency consistency
BigCodeBench [paper] [github] [huggingface]NoEvaluate LLMs’ ability to follow complex instructions and invoke diverse function callsPass@k, test-case execution accuracy, branch coverage
HumanEvalPack [github] [huggingface]NoEvaluate multilingual code generation, correction, and comment synthesis across six languagesFunctional correctness via pass@k and task-specific metrics
BIRD [pape] [github]NoEvaluate database-grounded text-to-SQL generation over large, noisy databasesExecution accuracy and exact match
RepoQA [github]YesEvaluate long-context code understanding in real software repositoriesExact-match accuracy and retrieval-augmented correctness
LongCodeArena [paper] [github] [huggingface]YesEvaluate long-context code comprehension, generation, and editing across project-wide tasksTask-specific accuracy, exact match, and human evaluation

Domain: Math

BenchmarkLong-context or notEvaluation ObjectiveMain Evaluation Criteria
GSM8K [paper] [github] [huggingface]NoAssess grade school math reasoningExact match accuracy
MATH [paper] [github] [huggingface]NoEvaluate competition-level math problem-solving with step-by-step reasoningFinal answer accuracy & derivation correctness
MathOdyssey [paper] [github] [huggingface]NoBenchmark high-school → Olympiad → university math across difficulty tiersAnswer accuracy across difficulty levels
MathBench [paper] [github]NoEvaluate theoretical & applied math knowledge across five levelsAccuracy on theoretical and application problems
CHAMP [paper] [github]NoFine-grained competition-level reasoning with concept/hint annotationsAnswer accuracy & reasoning-path correctness
LILA [paper] [github] [huggingface]NoUnified benchmark across 23 math tasks & formatsAccuracy across tasks
MiniF2F-v1 [paper] [github]NoFormal theorem-proving at Olympiad levelProof accuracy on 488 problems
ProofNet [paper] [github] [huggingface]NoAuto-formalization & formal proof generation (Lean 3)Formalization accuracy & proof success rate
AlphaGeometry [paper] [github]NoNeuro-symbolic reasoning on Olympiad Euclidean geometryProof success, correctness, completeness, readability
MathVerse [paper] [github] [huggingface]NoVisual-diagram math reasoning for MLLMsDiagram-sensitive answer accuracy & CoT score
We-Math [paper] [github] [huggingface]NoHuman-like visual mathematical reasoning with knowledge hierarchyFour-dimensional diagnostic metrics
U-MATH [paper] [github]NoOpen-ended university-level problem-solving (20 % multimodal)LLM-judged solution correctness (expert-verified F1)
TabMWP [paper] [github]NoMath reasoning over textual + tabular dataAccuracy on QA & MC questions
MathHay [paper] [github]YesLong-context mathematical reasoning with multi-step dependenciesAccuracy, exact match, reasoning-chain consistency

Domain: Law

BenchmarkLong-context or notEvaluation ObjectiveMain Evaluation Criteria
LegalBench [paper] [github] [huggingface]NoEvaluate legal reasoning across six types including rule application and interpretationTask-specific accuracy, rule-consistency, LLM-as-a-judge ratings
LexGLUE [paper] [github] [huggingface]NoEvaluate legal language understanding across classification and QA tasksTask-specific accuracy & F1
LEXTREME [paper] [github] [huggingface]NoEvaluate multilingual & multitask legal understanding (24 languages, 18 tasks)Macro-F1, accuracy, other classification metrics
LawBench [paper] [github]NoEvaluate Chinese legal LLMs across retention, understanding & application (20 tasks)Task-specific accuracy & F1
LAiW [paper] [github]NoEvaluate Chinese legal LLMs across fundamental → advanced tasks (13 assignments)Task-specific accuracy & F1
LexEval [paper] [github] [huggingface]NoEvaluate Chinese legal reasoning via a taxonomy of cognitive abilitiesTask-specific accuracy
CitaLaw [paper]NoEvaluate citation-grounded legal answering with statutes & precedentsSyllogism-alignment, citation accuracy, legal consistency
LegalAgentBench [paper] [github]NoEvaluate LLM agents solving complex real-world legal tasksTask success rate & intermediate progress
SCALE [paper] [github]YesEvaluate long-doc, multilingual, multitask legal reasoning (≤ 50 K tokens)Accuracy, F1, code-based assessment for long-context legal tasks

⭐ Citation

If you find this work useful, welcome to cite us.

@misc{wu2025surveyfederatedfinetuninglarge,
      title={A Survey on Federated Fine-tuning of Large Language Models}, 
      author={Yebo Wu and Chunlin Tian and Jingguang Li and He Sun and Kahou Tam and Zhanting Zhou and Haicheng Liao and Zhijiang Guo and Li Li and Chengzhong Xu},
      year={2025},
      eprint={2503.12016},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2503.12016}, 
}

Contributors

Clin0212

4 commits