arpita8/Awesome-Mixture-of-Experts-Papers

Survey: A collection of AWESOME papers and resources on the latest research in Mixture of Experts.

146

18 commits

updated Aug 21, 2024

See the code

README

Mixture-of-Experts-Papers Awesome

A curated list of exceptional papers and resources on Mixture of Experts and related topics.

News: Our Mixture of Experts survey has been released. The Evolution of Mixture of Experts: A Survey from Basics to Breakthroughs

Editor

Mendeley | ResearchGate | PDF If our work has been of assistance to you, please feel free to cite our survey. Thank you.

@article{article,
author = {Vats, Arpita and Raja, Rahul and Jain, Vinija and Chadha, Aman},
year = {2024},
month = {08},
pages = {12},
title = {THE EVOLUTION OF MIXTURE OF EXPERTS: A SURVEY FROM BASICS TO BREAKTHROUGHS}
}

Table of Contents

Evolution in Sparse Mixture of Experts

Editor
NamePaperVenueYear
The Sparsely-Gated Mixture-of-Experts LayerOutrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts LayerarXiv2017

Collection of Recent MoE Papers

MoE in Visual Domain

MoE in LLMs

MoE for Scaling LLMs

NamePaperVenueYear
u-LLaVAu-LLaVA: Unifying Multi-Modal Tasks via Large Language ModelarXiv2024
MoLEQMoE: Practical Sub-1-Bit Compression of Trillion-Parameter ModelsarXiv2024
LoryLory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-trainingarXiv2024
Uni-MoEUni-MoE: Scaling Unified Multimodal LLMs with Mixture of ExpertsarXiv2024
MH-MoEMulti-Head Mixture-of-ExpertsarXiv2024
DeepSeekMoEDeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsarXiv2024
Mini-GeminiMini-Gemini: Mining the Potential of Multi-modality Vision Language ModelsarXiv2024
OpenMoEOpenMoE: An Early Effort on Open Mixture-of-Experts Language ModelsarXiv2024
TUTELTutel: Adaptive Mixture-of-Experts at ScalearXiv2023
QMoEQMoE: Practical Sub-1-Bit Compression of Trillion-Parameter ModelsarXiv2023
Switch-NeRFSwitch-NeRF: Learning Scene Decomposition with Mixture of Experts for Large-scale Neural Radiance FieldsICLR2023
SaMoESaMoE: Parameter Efficient MoE Language Models via Self-Adaptive Expert Combination ICLR2023
JetMoEJetMoE: Reaching Llama2 Performance with 0.1M DollarsarXiv2023
MegaBlocksMegaBlocks: Efficient Sparse Training with Mixture-of-ExpertsarXiv2022
ST-MoEST-MoE: Designing Stable and Transferable Sparse Expert ModelsarXiv2022
Uni-Perceiver-MoEUni-Perceiver-MoE: Learning Sparse Generalist Models with Conditional MoEs NeurIPS2022
SpeechMoESpeechMoE: Scaling to Large Acoustic Models with Dynamic Routing Mixture of ExpertsarXiv2021
Fully-Differential Sparse TransformerSparse is Enough in Scaling TransformersarXiv2021

MoE: Enhancing System Performance and Efficiency

NamePaperVenueYear
pMoEPMoE: Progressive Mixture of Experts with Asymmetric Transformer for Continual LearningarXiv2024
HyperMoEHyperMoE: Towards Better Mixture of Experts via Transferring Among ExpertsarXiv2024
BlackMambaBlackMamba: Mixture of Experts for State-Space ModelsarXiv2024
ScheMoEScheMoE: An Extensible Mixture-of-Experts Distributed Training System with Tasks SchedulingarXiv2024
Pre-Gates MoEPre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert InferencearXiv2024
MoE-MambaMoE-Mamba: Efficient Selective State Space Models with Mixture of ExpertsarXiv2024
Parameter-efficient MoEsPushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction TuningarXiv2023
SMoE-DropoutSparse MoE as the New Dropout: Scaling Dense and Self-Slimmable TransformersarXiv2023
StableMoEStableMoE: Stable Routing Strategy for Mixture of ExpertsarXiv2022
AlpaAlpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep LearningarXiv2022
BaGuaLuBaGuaLu: targeting brain scale pretrained models with over 37 million coresACM2022
MEFTMEFT: Memory-Efficient Fine-Tuning through Sparse AdapterarXiv2024
EdgeMoEEdgeMoE: Fast On-Device Inference of MoE-based Large Language ModelsarXiv2023
SE-MoESE-MoE: A Scalable and Efficient Mixture-of-Experts Distributed Training and Inference SystemarXiv2022
NLLBNo Language Left Behind: Scaling Human-Centered Machine TranslationarXiv2022
EvoMoEEvoMoE: An Evolutional Mixture-of-Experts Training Framework via Dense-To-Sparse GatearXiv2022
FastMoEFastMoE: A Fast Mixture-of-Expert Training SystemarXiv2021
ACEACE: Ally Complementary Experts for Solving Long-Tailed Recognition in One-ShotICCV2021
M6-10TM6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter PretrainingarXiv2021
GShardGShard: Scaling Giant Models with Conditional Computation and Automatic ShardingarXiv2020
PAD-NetPAD-Net: Multi-Tasks Guided Prediction-and-Distillation Network for Simultaneous Depth Estimation and Scene ParsingarXiv2018

Integrating Mixture of Experts into Recommendation Algorithms

Python Libraries for MoE


Hope our survey with collection of all the recent MoE can help your work.
awesome
computer-vision
deep-learning
large-language-models
llm
machine-learning
mixture-of-experts
rec
recsys

Contributors

arpita8

18 commits

arpita8/Awesome-Mixture-of-Experts-Papers

Survey: A collection of AWESOME papers and resources on the latest research in Mixture of Experts.

146

18 commits

updated Aug 21, 2024

See the code

README

Mixture-of-Experts-Papers Awesome

A curated list of exceptional papers and resources on Mixture of Experts and related topics.

News: Our Mixture of Experts survey has been released. The Evolution of Mixture of Experts: A Survey from Basics to Breakthroughs

Editor

Mendeley | ResearchGate | PDF If our work has been of assistance to you, please feel free to cite our survey. Thank you.

@article{article,
author = {Vats, Arpita and Raja, Rahul and Jain, Vinija and Chadha, Aman},
year = {2024},
month = {08},
pages = {12},
title = {THE EVOLUTION OF MIXTURE OF EXPERTS: A SURVEY FROM BASICS TO BREAKTHROUGHS}
}

Table of Contents

Evolution in Sparse Mixture of Experts

Editor
NamePaperVenueYear
The Sparsely-Gated Mixture-of-Experts LayerOutrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts LayerarXiv2017

Collection of Recent MoE Papers

MoE in Visual Domain

MoE in LLMs

MoE for Scaling LLMs

NamePaperVenueYear
u-LLaVAu-LLaVA: Unifying Multi-Modal Tasks via Large Language ModelarXiv2024
MoLEQMoE: Practical Sub-1-Bit Compression of Trillion-Parameter ModelsarXiv2024
LoryLory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-trainingarXiv2024
Uni-MoEUni-MoE: Scaling Unified Multimodal LLMs with Mixture of ExpertsarXiv2024
MH-MoEMulti-Head Mixture-of-ExpertsarXiv2024
DeepSeekMoEDeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsarXiv2024
Mini-GeminiMini-Gemini: Mining the Potential of Multi-modality Vision Language ModelsarXiv2024
OpenMoEOpenMoE: An Early Effort on Open Mixture-of-Experts Language ModelsarXiv2024
TUTELTutel: Adaptive Mixture-of-Experts at ScalearXiv2023
QMoEQMoE: Practical Sub-1-Bit Compression of Trillion-Parameter ModelsarXiv2023
Switch-NeRFSwitch-NeRF: Learning Scene Decomposition with Mixture of Experts for Large-scale Neural Radiance FieldsICLR2023
SaMoESaMoE: Parameter Efficient MoE Language Models via Self-Adaptive Expert Combination ICLR2023
JetMoEJetMoE: Reaching Llama2 Performance with 0.1M DollarsarXiv2023
MegaBlocksMegaBlocks: Efficient Sparse Training with Mixture-of-ExpertsarXiv2022
ST-MoEST-MoE: Designing Stable and Transferable Sparse Expert ModelsarXiv2022
Uni-Perceiver-MoEUni-Perceiver-MoE: Learning Sparse Generalist Models with Conditional MoEs NeurIPS2022
SpeechMoESpeechMoE: Scaling to Large Acoustic Models with Dynamic Routing Mixture of ExpertsarXiv2021
Fully-Differential Sparse TransformerSparse is Enough in Scaling TransformersarXiv2021

MoE: Enhancing System Performance and Efficiency

NamePaperVenueYear
pMoEPMoE: Progressive Mixture of Experts with Asymmetric Transformer for Continual LearningarXiv2024
HyperMoEHyperMoE: Towards Better Mixture of Experts via Transferring Among ExpertsarXiv2024
BlackMambaBlackMamba: Mixture of Experts for State-Space ModelsarXiv2024
ScheMoEScheMoE: An Extensible Mixture-of-Experts Distributed Training System with Tasks SchedulingarXiv2024
Pre-Gates MoEPre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert InferencearXiv2024
MoE-MambaMoE-Mamba: Efficient Selective State Space Models with Mixture of ExpertsarXiv2024
Parameter-efficient MoEsPushing Mixture of Experts to the Limit: Extremely Parameter Efficient MoE for Instruction TuningarXiv2023
SMoE-DropoutSparse MoE as the New Dropout: Scaling Dense and Self-Slimmable TransformersarXiv2023
StableMoEStableMoE: Stable Routing Strategy for Mixture of ExpertsarXiv2022
AlpaAlpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep LearningarXiv2022
BaGuaLuBaGuaLu: targeting brain scale pretrained models with over 37 million coresACM2022
MEFTMEFT: Memory-Efficient Fine-Tuning through Sparse AdapterarXiv2024
EdgeMoEEdgeMoE: Fast On-Device Inference of MoE-based Large Language ModelsarXiv2023
SE-MoESE-MoE: A Scalable and Efficient Mixture-of-Experts Distributed Training and Inference SystemarXiv2022
NLLBNo Language Left Behind: Scaling Human-Centered Machine TranslationarXiv2022
EvoMoEEvoMoE: An Evolutional Mixture-of-Experts Training Framework via Dense-To-Sparse GatearXiv2022
FastMoEFastMoE: A Fast Mixture-of-Expert Training SystemarXiv2021
ACEACE: Ally Complementary Experts for Solving Long-Tailed Recognition in One-ShotICCV2021
M6-10TM6-10T: A Sharing-Delinking Paradigm for Efficient Multi-Trillion Parameter PretrainingarXiv2021
GShardGShard: Scaling Giant Models with Conditional Computation and Automatic ShardingarXiv2020
PAD-NetPAD-Net: Multi-Tasks Guided Prediction-and-Distillation Network for Simultaneous Depth Estimation and Scene ParsingarXiv2018

Integrating Mixture of Experts into Recommendation Algorithms

Python Libraries for MoE


Hope our survey with collection of all the recent MoE can help your work.
awesome
computer-vision
deep-learning
large-language-models
llm
machine-learning
mixture-of-experts
rec
recsys

Contributors

arpita8

18 commits