Curated collection of papers in machine learning systems
656
52 commits
updated Sep 15, 2026
A curated list of machine learning systems papers published in major CS conferences (plus some workshops and journals).
Survey papers are annotated with
[Survey 🔍].
For arXiv preprints, please see README_arxiv.md.
Data pipeline optimization
Preprocessing stalls
Fetch stalls (I/O)
Specific workloads (GNN, DLRM)
Caching and distributed storage for ML training
LLM data plane
Data formats
Data pipeline fairness and correctness
Data labeling automation
PAI)Philly)[EuroSys'26] Suika: Efficient and High-quality Re-scheduling of 3D-parallelized LLM Training Jobs in Shared Clusters
[EuroSys'26] Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design
[EuroSys'26] Bridging the GPU Utilization Gap: Predictive Multi-Dimensional Resource Scheduling for AI Workloads
[EuroSys'26] AdaGen: Workload-Adaptive Cluster Scheduler for Latency-Optimal LLM Inference Serving
[OSDI'25] Decouple and Decompose: Scaling Resource Allocation with DeDe
[SoCC'25] Cuckoo: Deadline-Aware Job Packing on Heterogeneous GPUs for DL Model Training
[EuroSys'25] Eva: Cost-Efficient Cloud-Based Cluster Scheduling
[TACO'24] Taming Flexible Job Packing in Deep Learning Training Clusters
[SoCC'24] Kale: Elastic GPU Scheduling for Online DL Model Training
[SC'24] PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU Clusters
[OSDI'24] MAST: Global Scheduling of ML Training across Geo-Distributed Datacenters at Hyperscale
[ASPLOS'24] Heet: Accelerating Elastic Training in Heterogeneous Deep Learning Clusters
[Middleware'24] Optimal Resource Efficiency with Fairness in Heterogeneous GPU Clusters
[IPDPS'24] Hadar: Heterogeneity-Aware Optimization-Based Online Scheduling for Deep Learning Cluster
[EuroSys'24] Blox: A Modular Toolkit for Deep Learning Schedulers
[NSDI'24] Swing: Short-cutting Rings for Higher Bandwidth Allreduce
[NSDI'24] Towards Domain-Specific Network Transport for Distributed DNN Training
[NSDI'24] Vulcan: Automatic Query Planning for Live ML Analytics
[NSDI'24] CASSINI: Network-Aware Job Scheduling in Machine Learning Clusters
[Survey :mag:] [ACM CSUR'23] Deep Learning Workload Scheduling in GPU Datacenters: A Survey
[SC'23] EasyScale: Accuracy-consistent Elastic Training for Deep Learning
[ICPP'23] CoTrain: Efficient Scheduling for Large-Model Training upon GPU and CPU in Parallel
[ICPP'23] Embracing Uncertainty for Equity in Resource Allocation in ML Training
[SOSP'23] Sia: Heterogeneity-aware, goodput-optimized ML-cluster scheduling
[NSDI'23] Shockwave: Proactive, Fair, and Efficient Cluster Scheduling for Dynamic Adaptation in Machine Learning
[EuroSys'23] SiloD: A Co-design of Caching and Scheduling for Deep Learning Clusters
[EuroSys'23] Lyra: Elastic Scheduling for Deep Learning Clusters
[EuroSys'23] ElasticFlow: An Elastic Serverless Training Platform for Distributed Deep Learning
[ASPLOS'23] Lucid: A Non-intrusive, Scalable and Interpretable Scheduler for Deep Learning Training Jobs
[SoCC'22] ESCHER: Expressive Scheduling with Ephemeral Resources
[NSDI'22] MLaaS in the wild: workload analysis and scheduling in large-scale heterogeneous GPU clusters (PAI)
[OSDI'22] Looking Beyond GPUs for DNN Scheduling on Multi-Tenant Clusters (Synergy)
[SIGCOMM'22] Multi-resource interleaving for deep learning training (Muri)
[MLSys'21] Wavelet: Efficient DNN Training with Tick-Tock Scheduling
[SoCC'21] Chronus: A Novel Deadline-aware Scheduler for Deep Learning Training Jobs
[SC'21] Characterization and Prediction of Deep Learning Workloads in Large-Scale GPU Datacenters (Helios)
[OSDI'21] Privacy Budget Scheduling (DPF)
[NSDI'21] Elastic Resource Sharing for Distributed Deep Learning (AFS)
[OSDI'21] Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning
[EuroSys'20] Balancing efficiency and fairness in heterogeneous GPU clusters for deep learning (GandivaFair)
[NSDI'20] Themis: Fair and Efficient GPU Cluster Scheduling
[OSDI'20] HiveD: Sharing a GPU Cluster for Deep Learning with Guarantees
[OSDI'20] Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads (Gavel)
[EuroSys'20] AlloX: Compute Allocation in Hybrid Clusters
[MLSys'20] Resource Elasticity in Distributed Deep Learning
[NSDI'19] Tiresias: A GPU Cluster Manager for Distributed Deep Learning
[ATC'19] Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads (Philly)
[EuroSys'18] Optimus: an efficient dynamic resource scheduler for deep learning clusters
[OSDI'18] Gandiva: Introspective Cluster Scheduling for Deep Learning
[MLSys'26] HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
[ICML'26] When Data Is Scarce: Scaling Sparse Language Models with Repeated Training
[ICML'26] AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training
[EuroSys'26] HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
[HPCA'26] Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
[HPCA'26] AutoHAAP: Automated Heterogeneity-Aware Asymmetric Partitioning for LLM Training
[HPCA'26] WATOS: Efficient LLM Training Strategies and Architecture Co-exploration for Wafer-scale Chip
[ASPLOS'26] SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
[NeurIPS'25] Synergistic Tensor and Pipeline Parallelism
[NeurIPS'25] First Attentions Last: Better Exploiting First Attentions for Efficient Transformer Training
[SC'25] Hypertron: Efficiently Scaling Large Models by Exploring High-Dimensional Parallelization Space
[CLUSTER'25] BMPipe: Bubble-Memory Co-Optimization Strategy Planner for Very-Large DNN Training
[OSDI'25] WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training
[ISCA'25] FRED: A Wafer-scale Fabric for 3D Parallel DNN Training
[ISCA'25] MeshSlice: Efficient 2D Tensor Parallelism for Distributed DNN Training
[ISCA'25] Scaling Llama 3 Training with Efficient Parallelism Strategies
[MLSys'25] Radius: Range-based Gradient Sparsity for Large Foundation Model Pre-training
[ICLR'25] TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training
[INFOCOM'25] Espresso: Cost-Efficient Large Model Training by Exploiting GPU Heterogeneity in the Cloud
[ASPLOS'25] GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism
[ASPLOS'25] FlexSP: Accelerating Large Language Model Training via Flexible Sequence Parallelism
[ASPLOS'25] Spindle: Efficient Distributed Training of Multi-Task Large Models via Wavefront Scheduling
[EuroSys'25] JABAS: Joint Adaptive Batching and Automatic Scaling for DNN Training on Heterogeneous GPUs
[TPDS'24] UMPIPE: Unequal Microbatches-Based Pipeline Parallelism for Deep Neural Network Training
[Survey :mag:] [ACM CSUR'24] Resource-efficient Algorithms and Systems of Foundation Models: A Survey
[SOSP'24] Uncovering Nested Data Parallelism and Data Reuse in DNN Computation with FractalTensor
[SOSP'24] Enabling Parallelism Hot Switching for Efficient Training of Large Language Models
[TACO'24] ATP: Achieving Throughput Peak for DNN Training via Smart GPU Memory Management
[NeurIPS'24] Rethinking Memory and Communication Costs for Efficient Data Parallel Training of Large Language Models
[NeurIPS'24] SpeedLoader: An I/O efficient scheme for heterogeneous and distributed LLM operation
[SC'24] Accelerating Distributed DLRM Training with Optimized TT Decomposition and Micro-Batching
[SC'24] Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
[SoCC'24] Distributed training of large language models on AWS Trainium
[TPDS'24] AutoDDL: Automatic Distributed Deep Learning With Near-Optimal Bandwidth Cost
[SOSP'24] Enabling Parallelism Hot Switching for Efficient Training of Large Language Models
[SOSP'24] TENPLEX: Changing Resources of Deep Learning Jobs using Parallelizable Tensor Collections
[ICPP'24] AutoPipe: Automatic Configuration of Pipeline Parallelism in Shared GPU Cluster
[COLM'24] LightSeq: Sequence Level Parallelism for Distributed Training of Long Context Transformers
[OSDI'24] nnScaler: Constraint-Guided Parallelization Plan Generation for Deep Learning Training
[ATC'24] Metis: Fast Automatic Distributed Training on Heterogeneous GPUs
[ATC'24] FwdLLM: Efficient Federated Finetuning of Large Language Models with Perturbed Inferences
[ATC'24] OPER: Optimality-Guided Embedding Table Parallelization for Large-scale Recommendation Model
[HPDC'24] DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models
[ICML'24] Scaling Beyond the GPU Memory Limit for Large Mixture-of-Experts Model Training
[ICML'24] Integrated Hardware Architecture and Device Placement Search
[MLSys'24] DiffusionPipe: Training Large Diffusion Models with Efficient Pipelines
[MobiCom'24] Asteroid: Resource-Efficient Hybrid Pipeline Parallelism for Collaborative DNN Training on Heterogeneous Edge Devices
[EuroSys'24] DynaPipe: Optimizing Multi-task Training through Dynamic Pipelines
[EuroSys'24] ScheMoE: An Extensible Mixture-of-Experts Distributed Training System with Tasks Scheduling
[EuroMLSys@EuroSys'24] ML Training with Cloud GPU Shortages: Is Cross-Region the Answer?
[ASPLOS'24] AdaPipe: Optimizing Pipeline Parallelism with Adaptive Recomputation and Partitioning
[ASPLOS'24] PrimePar: Efficient Spatial-temporal Tensor Partitioning for Large Transformer Model Training
[EuroSys'24] Aceso: Efficient Parallel DNN Training through Iterative Bottleneck Alleviation
[NSDI'24] MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs
[NSDI'24] DISTMM: Accelerating Distributed Multi-modal Model Training
[NSDI'24] Accelerating Neural Recommendation Training with Embedding Scheduling
[NSDI'24] Resiliency at Scale: Managing Google’s TPUv4 Machine Learning Supercomputer
[NSDI'24] QuickUpdate: a Real-Time Personalization System for Large-Scale Recommendation Models
[NSDI'24] Scaling Large Language Model Training to More Than 10,000 GPUs
[TKDE'24] Improving Automatic Parallel Training via Balanced Memory Workload Optimization
[ICLR'24] CO2: Efficient Distributed Training with Full Communication-Computation Overlap
[AAMAS'24] Holonic Learning: A Flexible Agent-based Distributed Machine Learning Framework
[VLDB'24] Saturn: An Optimized Data System for Multi-Large-Model Deep Learning Workloads
[HPCA'24] Tessel: Boosting Distributed Execution of Large DNN Models via Flexible Schedule Search
[NSDI'24] Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible Instances
[EuroSys'24] HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program Synthesis
[ICPP'23] Mercury: Fast and Optimal Device Placement for Large Deep Learning Models
[IPDPS'23] MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism
[CLUSTER'23] Prophet: Fine-grained Load Balancing for Parallel Training of Large-scale MoE Models
[NeurIPS'23] ASPEN: Breaking Operator Barriers for Efficient Parallelization of Deep Neural Networks
[NeurIPS'23] DeepPCR: Parallelizing Sequential Operations in Neural Networks
[DAC'23] MixPipe: Efficient Bidirectional Pipeline Parallelism for Training Large-Scale Models
[SC'23] Hanayo: Harnessing Wave-like Pipeline Parallelism for Enhanced Large Model Training Efficiency
[SOSP'23] PIT: Optimization of Dynamic Sparse Deep Learning Models via Permutation Invariant Transformation
[SOSP'23] Oobleck: Resilient Distributed Training of Large Models Using Pipeline Templates
[MICRO'23] Grape: Practical and Efficient Graphed Execution for Dynamic Deep Neural Networks on GPUs
[HPCA'23] Phloem: Automatic Acceleration of Irregular Applications with Fine-Grain Pipeline Parallelism
[ACL'23] Sequence Parallelism: Long Sequence Training from System Perspective
[CCGrid'23] A Deep Learning Pipeline Parallel Optimization Method
[OSDI'23] MGG: Accelerating Graph Neural Networks with Fine-Grained Intra-Kernel Communication-Computation Pipelining on Multi-GPU Platforms
[ATC'23] Accelerating Distributed MoE Training and Inference with Lina
[ATC'23] SmartMoE: Efficiently Training Sparsely-Activated Models through Combining Offline and Online Parallelization
[ATC'23] MSRL: Distributed Reinforcement Learning with Dataflow Fragments
[Survey :mag:] [TPDS'23] A Survey on Auto-Parallelism of Large-Scale Deep Learning Training
[ICML'23] SWARM Parallelism: Training Large Models Can Be Surprisingly Communication-Efficient
[ICML'23] BPipe: Memory-Balanced Pipeline Parallelism for Training Large Language Models
[ICS'23] A Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training
[NSDI'23] TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training Jobs
[NSDI'23] Bamboo: Making Preemptible Instances Resilient for Affordable Training of Large DNNs
[NSDI'23] ARK: GPU-driven Code Execution for Distributed Deep Learning
[MLSys'23] On Optimizing the Communication of Model Parallelism
[MLSys'23] MegaBlocks: Efficient Sparse Training with Mixture-of-Experts
[MLSys'23] Tutel: Adaptive Mixture-of-Experts at Scale
[TPDS'23] Merak: An Efficient Distributed DNN Training Framework with Automated 3D Parallelism for Giant Foundation Models
[PPoPP'23] Elastic Averaging for Efficient Pipelined DNN Training
[PPoPP'23] Efficient All-Reduce for Distributed DNN Training in Optical Interconnect Systems
[VLDB'23] MiCS: Near-linear Scaling for Training Gigantic Model on Public Cloud
[VLDB'23] Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism
[ASPLOS'23] Mobius: Fine Tuning Large-Scale Models on Commodity GPU Servers
[ASPLOS'23] Optimus-CC: Efficient Large NLP Model Training with 3D Parallelism Aware Communication Compression
[ICPP'22] Tesseract: Parallelize the Tensor Parallelism Efficiently
[NeurIPS'22] Fine-tuning Language Models over Slow Networks using Activation Quantization with Guarantees
[SoCC'22] Accelerating Large-Scale Distributed Neural Network Training with SPMD Parallelism
[MLSys'22] Pathways: Asynchronous distributed dataflow for ML
[MLSys'22] SRIFTY: Swift and Thrifty Distributed Neural Network Training on the Cloud
[MLSys'22] Efficient Strong Scaling Through Burst Parallel Training
[EuroSys'22] Varuna: scalable, low-cost training of massive deep learning models
[ATC'22] Whale: Efficient Giant Model Training over Heterogeneous GPUs
[NeurIPS'22] AMP: Automatically Finding Model Parallel Strategies with Heterogeneity Awareness
[PPoPP'22] FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained models
[ICML'22] DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale
[ICML'22] GLaM: Efficient Scaling of Language Models with Mixture-of-Experts
[HPDC'22] Hare: Exploiting Inter-job and Intra-job Parallelism of Distributed Machine Learning on Heterogeneous GPUs
[OSDI'22] Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep Learning
[NSDI'22] Accelerating Collective Communication in Data Parallel Training across Deep Learning Frameworks
[JMLR'21] Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
[TPDS'21] TensorOpt: Exploring the Tradeoffs in Distributed DNN Training With Auto-Parallelism
[ATC'21] Fine-tuning giant neural networks on commodity hardware with automatic pipeline model parallelism
[SIGMOD'21] Heterogeneity-Aware Distributed Machine Learning Training via Partial Reduce(#210-communication-optimization)]
[MLSys'21] PipeMare: Asynchronous Pipeline Parallel DNN Training
[ICLR'21] GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
[NeurIPS'21] Piper: Multidimensional Planner for DNN Parallelization
[ICML'21] Memory-Efficient Pipeline-Parallel DNN Training
[ICML'21] TeraPipe: Token-Level Pipeline Parallelism for Training Large-Scale Language Models
[ICML'21] PipeTransformer: Automated Elastic Pipelining for Distributed Training of Large-scale Models
[SC'21] Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
[SC'21] Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM (PTD-P or Megatron-LM v2)
[FAST'21] Behemoth: A Flash-centric Training Accelerator for Extreme-scale DNNs
[PPoPP'21] DAPPLE: a pipelined data parallel approach for training large models
[VLDB'21] Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches
[HPCA'20] AccPar: Tensor Partitioning for Heterogeneous Deep Learning Accelerators
[NeurIPS'20] Efficient Algorithms for Device Placement of DNN Graph Operators
[KDD'20 Tutorial] DeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters
[VLDB'20] PyTorch Distributed: Experiences on Accelerating Data Parallel Training
[OSDI'20] A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU Clusters (BytePS)
[SOSP'19] PipeDream: Generalized Pipeline Parallelism for DNN Training
[NeurIPS'20] Language Models are Few-Shot Learners
[HPCA'19] HyPar: Towards Hybrid Parallelism for Deep Learning Accelerator Array
[IEEE MICRO'19] Optimizing Multi-GPU Parallelization Strategies for Deep Learning Training
[MLSys'19] Beyond data and model parallelism for deep neural networks (FlexFlow)
[MLSys'19] TicTac: Accelerating Distributed Deep Learning with Communication Scheduling
[EuroSys'19] Parallax: Sparsity-aware Data Parallel Training of Deep Neural Networks
[EuroSys'19] Supporting Very Large Models using Automatic Dataflow Graph Partitioning (Tofu)
[SOSP'19] A Generic Communication Scheduler for Distributed DNN Training Acceleration
[NeurIPS'19] Mesh-TensorFlow: Deep Learning for Supercomputers
[NeurIPS'19] GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
[ICML'18] Exploring Hidden Dimensions in Parallelizing Convolutional Neural Networks
[Survey :mag:] [IJCAI'22] Survey on Effcient Training of Large Neural Networks
[Survey :mag:] [ACM CSUR'19] Demystifying Parallel and Distributed Deep Learning
[Survey :mag:] [ACM CSUR'19] Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques, and Tools
For comprehensive list of GNN systems papers, refer to https://github.com/chwan1016/awesome-gnn-systems.
[KDD'26] OrionInfer: Low-Overhead Parallelism Switching and Live Migration for Efficient LLM Serving
[ISCA'26] CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIM
[ISCA'26] DynoPipe: Heterogeneous Edge-Cloud LLM Serving with Dynamically Orchestrated Pipeline Boundaries
[ISCA'26] Tetris: Efficient Long-context LLM Serving with Chunkwise Dynamic Sequence Parallelism
[ISCA'26] ConServe: Contiguity-Preserving Memory Management for Multi-Turn LLM Serving
[SIGOPS OSR'26] Rethinking LLM Deployment for Intent-Based Serving
[SIGOPS OSR'26] Elastic Memory Remapping for Multi-tenant LLM Serving
[ICML'26] Beyond Prediction: Tail-Aware Scheduling for LLM Inference
[MLSys'26] SHIP: SRAM-Based Huge Inference Pipelines for Fast LLM Serving
[MLSys'26] Dataflow Is All You Need
[MLSys'26] Optimizing Deployment Configurations for LLM Inference
[MobiSys'26] TimelyLLM: Time-sensitive LLM Serving System for Physical-I/O Limited Agents
[ISCA'26] Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
[SIGCOMM'26] KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
[ICML'26] PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
[EuroSys'26] Efficient Data Passing for Serverless Inference Workflows: A GPU-Centric Approach
[EuroSys'26] High Throughput and Low Latency LLM Serving via Adaptive KV Caching
[EuroSys'26] Automated End-to-End Model Serving with Cooperative Compilation and Scheduling
[SIGMOD'26] Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
[ASPLOS'26] BlendServe: Optimizing Offline Inference with Resource-Aware Batching
[ASPLOS'26] DFVG: A Heterogeneous Architecture for Speculative Decoding with Draft-on-FPGA and Verify-on-GPU
[ASPLOS'26] QoServe: Breaking the Silos of LLM Inference Serving
[ASPLOS'26] Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
[ASPLOS'26] SwiftSpec: Disaggregated Speculative Decoding and Fused Kernels for Low-Latency LLM Inference
[ASPLOS'26] ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
[HPCA'26] ELORA: Efficient LoRA and KV Cache Management for Multi-LoRA LLM Serving
[FAST'26] CacheSlide: Unlocking Cross Position-Aware KV Cache Reuse for Accelerating LLM Serving
[FAST'26] SolidAttention: Low-Latency SSD-based Serving on Memory-Constrained PCs
[HPCA'26] PASCAL: A Phase-Aware Scheduling Algorithm for Serving Reasoning-based Large Language Models
[MLSys'26] Meeting SLOs, Slashing Hours: Automated Enterprise LLM Optimization with OptiKIT
[VLDB'26] ORBITFLOW: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
[IEEE Computer'26] Challenges and Research Directions for Large Language Model Inference Hardware
[NSDI'26] FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
[NSDI'26] FastServe: Iteration-Level Preemptive Scheduling for Large Language Model Inference
[NSDI'26] HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
[FPGA'26] CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving
[ASPLOS'26] XY-Serve: End-to-End Versatile Production Serving for Dynamic LLM Workloads
[AAAI'26] Lethe: Layer- and Time-Adaptive KV Cache Pruning for Reasoning-Intensive LLM Serving
[EuroSys'26] FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
[EuroSys'26] KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
[EuroSys'26] TokenFlow: Responsive LLM Text Streaming Serving under Request Burst via Preemptive Scheduling
[SoCC'25] Multiplexed Heterogeneous LLM Serving via Stage-Aligned Parallelism
[Middleware'25] Argus: Quality-Aware High-Throughput Text-to-Image Inference Serving System
[NeurIPS'25] SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
[EMNLP'25] Distributed LLM Serving on Consumer-Grade GPUs by Reconciling Computation and Communication
[MICRO'25] MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
[MICRO'25] Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
[CLUSTER'25] Scalable and Fast Inference Serving via Hybrid Communication Scheduling on Heterogeneous Networks
[Survey :mag:] [ACM CSUR'25] Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
[SOSP'25] Aegaeon: Effective GPU Pooling for Concurrent LLM Serving on the Market
[SOSP'25] IC-Cache: Efficient Large Language Model Serving via In-context Caching
[SOSP'25] DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction
[COLM'25] OverFill: Two-Stage Models for Efficient Language Model Decoding
[ACM MM'25] TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
[SC'25] Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
[SIGCOMM'25] SCX: Stateless KV-Cache Encoding for Cloud-Scale Confidential Transformer Serving
[OSDI'25] BlitzScale: Fast and Live Large Model Autoscaling with O(1) Host Caching
[OSDI'25] WaferLLM: Large Language Model Inference at Wafer Scale
[OSDI'25] NanoFlow: Towards Optimal Large Language Model Serving Throughput
[ICML'25] Packrat: Automatic Reconfiguration for Latency Minimization in CPU-based DNN Serving
[ACL'25] SPECTRA: Faster Large Language Model Inference with Optimized Internal and External Speculation
[CODEML @ ICML'25] TorchAO: PyTorch-Native Training-to-Serving Model Optimization
[ICML'25] EPIC: Efficient Position-Independent Caching for Serving Large Language Models
[ATC'25] DEEPSERVE: Serverless Large Language Model Serving at Scale
[ISCA'25] WindServe: Efficient Phase-Disaggregated LLM Serving with Stream-based Dynamic Scheduling
[ISCA'25] Hybe: GPU-NPU Hybrid System for Efficient LLM Inference with Million-Token Context Window
[ICLR'25] TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
[OSDI'25] Clover: Exploiting Intra-device Parallelism for High Throughput Large Language Model Serving
[MLSys'25] SOLA: Optimizing SLO Attainment for Large Language Model Serving with State-Aware Scheduling
[MLSys'25] Marconi: Prefix Caching for the Era of Hybrid LLMs
[Mobicom'25] D2MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
[ISPASS'25] Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
[SIGMOD'25] Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
[EuroMLSys'25] Performance Aware LLM Load Balancer for Mixed Workloads
[MLSys'25] Rethinking Key-Value Cache Compression Techniques for Large Language Model Serving
[HPCA'25] PAISE: PIM-Accelerated Inference Scheduling Engine for Transformer-based LLM
[HPCA'25] throttLL'eM: Predictive GPU Throttling for Energy Efficient LLM Inference Serving
[ASPLOS'25] Aqua: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
[ASPLOS'25] Past-Future Scheduler for LLM Serving under SLA Guarantees
[ASPLOS'25] Accelerating LLM Serving for Multi-turn Dialogues with Efficient Resource Management
[EuroSys'25] SpInfer: Leveraging Low-Level Sparsity for Efficient Large Language Model Inference on GPUs
[EuroSys'25] Multiplexing Dynamic Deep Learning Workloads with SLO-awareness in GPU Clusters
[EuroSys'25] NeuStream: Bridging Deep Learning Serving and Stream Processing
[SoCC'25] ModServe: Scalable and Resource-Efficient Large Multimodal Model Serving
[ISCA'25] Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
[NSDI'25] SuperServe: Fine-Grained Inference Serving for Unpredictable Workloads
[MLSys'25] ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
[ICLR'25] HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
[EuroSys'25] SkyServe: Serving AI Models across Regions and Clouds with Spot Instances
[ASPLOS'25] Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow
[ASPLOS'25] Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective Elasticity
[MLSys'25] FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
[EuroSys'25] A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
[Survey :mag:] [ACM CSUR'24] Resource-efficient Algorithms and Systems of Foundation Models: A Survey
[ICML'25] SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization [Code]
[ICLR'25] SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration [Code]
[ICML'25] SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference [Code]
[ACL'24] LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
[ACL'24] SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget
[IPDPS'24] Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
[NeurIPS'24] Kangaroo: Lossless Self-Speculative Decoding for Accelerating LLMs via Double Early Exiting
[NeurIPS'24] Toward Efficient Inference for Mixture of Experts
[NeurIPS'24] Sequoia: Scalable and Robust Speculative Decoding
[SC'24] PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation
[SC'24] SMIless: Serving DAG-based Inference with Dynamic Invocations under Serverless Computing
[SenSys'24] LiteMoE: Customizing On-device LLM Serving via Proxy Submodel Tuning
[MICRO'24] Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUs
[PML4LRS @ ICLR2024] Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
[EuroSys'25] Fast State Restoration in LLM Serving with HCache
[HPCA'24] KRISP: Enabling Kernel-wise RIght-sizing for Spatial Partitioned GPU Inference Servers
[NeurIPS'24] Efficient LLM Scheduling by Learning to Rank
[SOSP'24] PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
[SOSP'24] LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
[SOSP'24] Improving DNN Inference Throughput Using Practical, Per-Input Compute Adaptation
[SOSP'24] Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving
[ICPP'24] GMM: An Efficient GPU Memory Management-based Model Serving System for Multiple DNN Inference Models
[SIGCOMM'24] CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
[ES-FoMO @ ICML'24] CO2: Precise Attention Score Observation for improving KV Cache Replacement in Large Language Models
[OSDI'24] dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM Serving
[OSDI'24] Parrot: Efficient Serving of LLM-based Applications with Semantic Variable
[OSDI'24] USHER: Holistic Interference Avoidance for Resource Optimized ML Inference
[OSDI'24] Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
[OSDI'24] ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
[OSDI'24] InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management
[OSDI'24] Llumnix: Dynamic Scheduling for Large Language Model Serving
[OSDI'24] DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
[ATC'24] Power-aware Deep Learning Model Serving with μ-Serve
[ATC'24] Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
[ATC'24] PUZZLE: Efficiently Aligning Large Language Models through Light-Weight Context Switch
[TPDS'24] ElasticBatch: A Learning-Augmented Elastic Scheduling System for Batch Inference on MIG
[OSDI'24] Parrot: Efficient Serving of LLM-based Applications with Semantic Variable
[ISCA'24] Splitwise: Efficient generative LLM inference using phase splitting
[ICML'24] Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
[ICML'24] Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
[ICML'24] HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
[ICML'24] EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
[ICML'24] MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
[MobiSys'24] ARISE: High-Capacity AR Offloading Inference Serving via Proactive Scheduling
[MobiSys'24] Pantheon: Preemptible Multi-DNN Inference on Mobile Edge GPUs
[MLSys'24] HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
[MLSys'24] S-LoRA: Serving Thousands of Concurrent LoRA Adapters
[MLSys'24] Vidur: A Large-Scale Simulation Framework For LLM Inference
[WWW'24] λGrapher: A Resource-Efficient Serverless System for GNN Serving through Graph Sharing
[ICML'24] CLLMs: Consistency Large Language Models
[EuroSys'24] Model Selection for Latency-Critical Inference Serving
[ISCA'24] Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
[ASPLOS'24] ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference
[ASPLOS'24] NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
[ICML'24] DéjàVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving
[ICLR'24] Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
[NSDI'24] Approximate Caching for Efficiently Serving Diffusion Models
[ASPLOS'24] SpotServe: Serving Generative Large Language Models on Preemptible Instances
[NeurIPS'23] SpecTr: Fast Speculative Decoding via Optimal Transport
[HPDC'23] Kairos: Building Cost-Efficient Machine Learning Inference Systems with Heterogeneous Cloud Resources
[SOSP'23] Paella: Low-latency Model Serving with Virtualized GPU Scheduling
[SOSP'23] Efficient Memory Management for Large Language Model Serving with PagedAttention
[MLSys'23] Efficiently Scaling Transformer Inference
[EuroSys'23] Fast and Efficient Model Serving Using Multi-GPUs with Direct-Host-Access
[EuroSys'23] Tabi: An Efficient Multi-Level Inference System for Large Language Models
[EuroSys'23] Pocket: ML Serving from the Edge
[OSDI'23] AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving
[NSDI'23] SHEPHERD: Serving DNNs in the Wild
[VLDB'23] Serving and Optimizing Machine Learning Workflows on Heterogeneous Infrastructures
[ICML'23] Fast Inference from Transformers via Speculative Decoding
[SIGMOD'22] Serverless Data Science - Are We There Yet? A Case Study of Model Serving
[OSDI'22] Orca: A Distributed Serving System for Transformer-Based Generative Models
[OSDI'22] Microsecond-scale Preemption for Concurrent GPU-accelerated DNN Inferences
[ATC'22] SOTER: Guarding Black-box Inference for General Neural Networks at the Edge
[ATC'22] Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal Sharing
[ATC'22] Tetris: Memory-efficient Serverless Inference through Tensor Sharing
[ATC'22] PetS: A Unified Framework for Parameter-Efficient Transformers Serving
[ATC'21] INFaaS: Automated Model-less Inference Serving
[SoCC'21] Morphling: Fast, Near-Optimal Auto-Configuration for Cloud-Native Model Serving
[MobiCom'20] SPINN: Synergistic Progressive Inference of Neural Networks over Device and Cloud
[CAL'26] HyGIN: Hybrid CPU/GPU-Initiated Communication for Mixture-of-Experts Training
[SIGCOMM'26] Balancing and Beyond: Communication-Centric Optimizations in Expert Parallelism
[SIGCOMM'26] UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods
[OSDI'26] Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design
[MLSys'26] From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
[OCML'26] EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference
[ICML'26] ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference
[ISCA'26] Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
[ISCA'26] Orders in Chaos: Enhancing Large-Scale MoE LLM Serving with Data Movement Forecasting
[ASPLOS'26] EARTH: An Efficient MoE Accelerator with Entropy-Aware Speculative Prefetch and Result Reuse
[ASPLOS'26] MoE-APEX: An Efficient MoE Inference System with Adaptive Precision Expert Offloading
[NSDI'26] SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
[ASPLOS'26] LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
[EuroSys'26] Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert Offloading
[EuroSys'26] MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
[SC'25] Diff-MoE: Efficient Batched MoE Inference with Priority-Driven Differential Expert Caching
[SC workshop'25] Compression Error Sensitivity Analysis for Different Experts in MoE Model Inference
[SC workshop'25] Batch Tiling on Attention: Efficient Mixture of Experts Training on Wafer-Scale Processors
[MICRO'25] Optimizing All-to-All Collective Communication with Fault Tolerance on Torus Networks
[SOSP'25] KTransformers: Unleashing the Full Potential of CPU/GPU Hybrid Inference for MoE Models
[ICML'25] Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
[NeurIPS'25] BrainMoE: Cognition Joint Embedding via Mixture-of-Expert Towards Robust Brain Foundation Model
[NeurIPS'25] S’MoRE: Structural Mixture of Residual Experts for Parameter-Efficient LLM Fine-tuning
[NeurIPS'25] The Omni-Expert: A Computationally Efficient Approach to Achieve a Mixture of Experts in a Single Expert Model
[NeurIPS'25] MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
[NeurIPS'25] FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-Experts
[NeurIPS'25] FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training
[NeurIPS'25] FlashMoE: Fast Distributed MoE in a Single Kernel [Code]
[SC'25] MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
[SIGCOMM'25] MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
[ICLR'25] Ada-K Routing: Boosting the Efficiency of MoE-based LLMs
[ICML'25] I2MoE: Interpretable Multimodal Interaction-aware Mixture-of-Experts
[SC'25] X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
[SIGCOMM'25] MixNet: A Runtime Reconfigurable Optical-Electrical Fabric for Distributed Mixture-of-Experts Training
[ACL'25] EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
[ACL'25] FOLDMOE: Efficient Long Sequence MoE Training via Attention-MoE Pipelining
[ICML'25] FloE: On-the-Fly MoE Inference on Memory-constrained GPU
[NAACL'25] Marrying LLMs with Dynamic Forecasting: A Graph Mixture-of-expert Perspective
[NAACL'25] Sparser Mixture-of-Adapters with Cross-Layer Generalization
[NAACL'25] SimSMoE: Toward Efficient Training Mixture of Experts via Solving Representational Collapse
[Mobicom'25] D2MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
[DAC'25] HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
[TKDE'25] A Survey on Mixture of Experts
[ICLR'25] NetMoE: Accelerating MoE Training through Dynamic Sample Placement
[EuroSys'25] Samoyeds: Accelerating MoE Models with Structured Sparsity Leveraging Sparse Tensor Cores
[EuroMLSys'25] Priority-Aware Preemptive Scheduling for Mixed-Priority Workloads in MoE Inference
[EuroMLSys'25] Accelerating MoE Model Inference with Expert Sharding
[KDD'25] ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration
[MLSys'25] Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
[CVPR'25] DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models
[ASPLOS'25] CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
[TPDS'25] EfficientMoE: Optimizing Mixture-of-Experts Model Training with Adaptive Load Balance
[NAACL'25] MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEs
[ASPLOS'25] FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models
[MICRO'24] SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts
[TPDS'24] MPMoE: Memory Efficient MoE for Pre-Trained Models With Adaptive Pipeline Parallelism
[MLArchSys'24 @ ISCA'24] MoE-ERAS: Expert Residency Aware Selection
[COLM'24] Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training
[ME-FoMo @ ICLR'24] Scaling Laws for Fine-Grained Mixture of Experts
[ML for Sys workshop @ NeurIPS'24] IFMoE: An Inference Framework Design for Fine-grained MoE
[ML for Sys workshop @ NeurIPS'24] TurboMoE: Enhancing MoE Model Training with Smart Kernel-Fusion and Data Transformation
[EMNLP'24] MiLoRA: Efficient Mixture of Low-Rank Adaptation for Large Language Models Fine-tuning
[EMNLP'24] Mixture of Diverse Size Experts
[EMNLP'24] AdaMOE: Token-Adaptive Routing with Null Experts for Mixture-of-Experts Language Models
[ACL'24] SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget
[SoCC'24] MoEsaic: Shared Mixture of Experts
[KDD'24] Efficient Mixture of Experts based on Large Language Models for Low-Resource Data Preprocessing
[IPDPS'24] Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
[NeurIPS'24] Toward Efficient Inference for Mixture of Experts
[SC'24] APTMoE: Affinity-Aware Pipeline Tuning for MoE Models on Bandwidth-Constrained GPU Nodes
[NeurIPS'24] GraphMETRO: Mitigating Complex Graph Distribution Shifts via Mixture of Aligned Experts
[NeurIPS'24] LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
[NeurIPS'24] Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
[PML4LRS @ ICLR'24] Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
[NeurIPS'24 (Splotlight)] Flex-MoE: Modeling Arbitrary Modality Combination via the Flexible Mixture-of-Experts
[SRW @ ACL'24] MoExtend: Tuning New Experts for Modality and Task Extension
[ICML'24] Scaling Beyond the GPU Memory Limit for Large Mixture-of-Experts Model Training
[MLSys'24] QMoE: Sub-1-Bit Compression of Trillion-Parameter Models
[SIGIR'24] M3oE: Multi-Domain Multi-Task Mixture-of Experts Recommendation Framework
[EuroSys'24] ScheMoE: An Extensible Mixture-of-Experts Distributed Training System with Tasks Scheduling
[ICLR'24] Mixture of LoRA Experts
[IJCAI'24] LocMoE: A Low-overhead MoE for Large Language Model Training
[ISCA'24] Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
[IPDPS'23] MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism
[EMNLP'23] Adaptive Gating in Mixture-of-Experts based Language Models
[ICLR'23] Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints
[ATC'23] Accelerating Distributed MoE Training and Inference with Lina
[OSDI'23] Optimizing Dynamic Neural Networks with Brainstorm
[SIGMOD'23] FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement
[ICS'23] A Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training
[MLSys'23] MegaBlocks: Efficient Sparse Training with Mixture-of-Experts
[MLSys'23] Tutel: Adaptive Mixture-of-Experts at Scale
[PPoPP'22] FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained models
[SustaiNLP @ EMNLP'22] Who Says Elephants Can't Run: Bringing Large Scale MoE Models into Cloud Scale Production
[NeurIPS'22] Mixture-of-Experts with Expert Choice Routing
[ICML'22] DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale
[ICML'22] GLaM: Efficient Scaling of Language Models with Mixture-of-Experts
[JMLR'22] Switch transformers: scaling to trillion parameter models with simple and efficient sparsity
[EMNLP'21] Beyond Distillation: Task-level Mixture-of-Experts for Efficient Inference
[ICLR'17] Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
P^2)CoCoNET)SCCL)BytePS)P3)ByteScheduler)For comprehensive list of quantization papers, refer to https://github.com/Efficient-ML/Awesome-Model-Quantization.
FrugalMCT)https://github.com/friedrichor/Awesome-Multimodal-Papers
Curated collection of papers in machine learning systems
656
52 commits
updated Sep 15, 2026
A curated list of machine learning systems papers published in major CS conferences (plus some workshops and journals).
Survey papers are annotated with
[Survey 🔍].
For arXiv preprints, please see README_arxiv.md.
Data pipeline optimization
Preprocessing stalls
Fetch stalls (I/O)
Specific workloads (GNN, DLRM)
Caching and distributed storage for ML training
LLM data plane
Data formats
Data pipeline fairness and correctness
Data labeling automation
PAI)Philly)[EuroSys'26] Suika: Efficient and High-quality Re-scheduling of 3D-parallelized LLM Training Jobs in Shared Clusters
[EuroSys'26] Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design
[EuroSys'26] Bridging the GPU Utilization Gap: Predictive Multi-Dimensional Resource Scheduling for AI Workloads
[EuroSys'26] AdaGen: Workload-Adaptive Cluster Scheduler for Latency-Optimal LLM Inference Serving
[OSDI'25] Decouple and Decompose: Scaling Resource Allocation with DeDe
[SoCC'25] Cuckoo: Deadline-Aware Job Packing on Heterogeneous GPUs for DL Model Training
[EuroSys'25] Eva: Cost-Efficient Cloud-Based Cluster Scheduling
[TACO'24] Taming Flexible Job Packing in Deep Learning Training Clusters
[SoCC'24] Kale: Elastic GPU Scheduling for Online DL Model Training
[SC'24] PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU Clusters
[OSDI'24] MAST: Global Scheduling of ML Training across Geo-Distributed Datacenters at Hyperscale
[ASPLOS'24] Heet: Accelerating Elastic Training in Heterogeneous Deep Learning Clusters
[Middleware'24] Optimal Resource Efficiency with Fairness in Heterogeneous GPU Clusters
[IPDPS'24] Hadar: Heterogeneity-Aware Optimization-Based Online Scheduling for Deep Learning Cluster
[EuroSys'24] Blox: A Modular Toolkit for Deep Learning Schedulers
[NSDI'24] Swing: Short-cutting Rings for Higher Bandwidth Allreduce
[NSDI'24] Towards Domain-Specific Network Transport for Distributed DNN Training
[NSDI'24] Vulcan: Automatic Query Planning for Live ML Analytics
[NSDI'24] CASSINI: Network-Aware Job Scheduling in Machine Learning Clusters
[Survey :mag:] [ACM CSUR'23] Deep Learning Workload Scheduling in GPU Datacenters: A Survey
[SC'23] EasyScale: Accuracy-consistent Elastic Training for Deep Learning
[ICPP'23] CoTrain: Efficient Scheduling for Large-Model Training upon GPU and CPU in Parallel
[ICPP'23] Embracing Uncertainty for Equity in Resource Allocation in ML Training
[SOSP'23] Sia: Heterogeneity-aware, goodput-optimized ML-cluster scheduling
[NSDI'23] Shockwave: Proactive, Fair, and Efficient Cluster Scheduling for Dynamic Adaptation in Machine Learning
[EuroSys'23] SiloD: A Co-design of Caching and Scheduling for Deep Learning Clusters
[EuroSys'23] Lyra: Elastic Scheduling for Deep Learning Clusters
[EuroSys'23] ElasticFlow: An Elastic Serverless Training Platform for Distributed Deep Learning
[ASPLOS'23] Lucid: A Non-intrusive, Scalable and Interpretable Scheduler for Deep Learning Training Jobs
[SoCC'22] ESCHER: Expressive Scheduling with Ephemeral Resources
[NSDI'22] MLaaS in the wild: workload analysis and scheduling in large-scale heterogeneous GPU clusters (PAI)
[OSDI'22] Looking Beyond GPUs for DNN Scheduling on Multi-Tenant Clusters (Synergy)
[SIGCOMM'22] Multi-resource interleaving for deep learning training (Muri)
[MLSys'21] Wavelet: Efficient DNN Training with Tick-Tock Scheduling
[SoCC'21] Chronus: A Novel Deadline-aware Scheduler for Deep Learning Training Jobs
[SC'21] Characterization and Prediction of Deep Learning Workloads in Large-Scale GPU Datacenters (Helios)
[OSDI'21] Privacy Budget Scheduling (DPF)
[NSDI'21] Elastic Resource Sharing for Distributed Deep Learning (AFS)
[OSDI'21] Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning
[EuroSys'20] Balancing efficiency and fairness in heterogeneous GPU clusters for deep learning (GandivaFair)
[NSDI'20] Themis: Fair and Efficient GPU Cluster Scheduling
[OSDI'20] HiveD: Sharing a GPU Cluster for Deep Learning with Guarantees
[OSDI'20] Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads (Gavel)
[EuroSys'20] AlloX: Compute Allocation in Hybrid Clusters
[MLSys'20] Resource Elasticity in Distributed Deep Learning
[NSDI'19] Tiresias: A GPU Cluster Manager for Distributed Deep Learning
[ATC'19] Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads (Philly)
[EuroSys'18] Optimus: an efficient dynamic resource scheduler for deep learning clusters
[OSDI'18] Gandiva: Introspective Cluster Scheduling for Deep Learning
[MLSys'26] HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
[ICML'26] When Data Is Scarce: Scaling Sparse Language Models with Repeated Training
[ICML'26] AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training
[EuroSys'26] HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters
[HPCA'26] Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
[HPCA'26] AutoHAAP: Automated Heterogeneity-Aware Asymmetric Partitioning for LLM Training
[HPCA'26] WATOS: Efficient LLM Training Strategies and Architecture Co-exploration for Wafer-scale Chip
[ASPLOS'26] SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
[NeurIPS'25] Synergistic Tensor and Pipeline Parallelism
[NeurIPS'25] First Attentions Last: Better Exploiting First Attentions for Efficient Transformer Training
[SC'25] Hypertron: Efficiently Scaling Large Models by Exploring High-Dimensional Parallelization Space
[CLUSTER'25] BMPipe: Bubble-Memory Co-Optimization Strategy Planner for Very-Large DNN Training
[OSDI'25] WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training
[ISCA'25] FRED: A Wafer-scale Fabric for 3D Parallel DNN Training
[ISCA'25] MeshSlice: Efficient 2D Tensor Parallelism for Distributed DNN Training
[ISCA'25] Scaling Llama 3 Training with Efficient Parallelism Strategies
[MLSys'25] Radius: Range-based Gradient Sparsity for Large Foundation Model Pre-training
[ICLR'25] TorchTitan: One-stop PyTorch native solution for production ready LLM pre-training
[INFOCOM'25] Espresso: Cost-Efficient Large Model Training by Exploiting GPU Heterogeneity in the Cloud
[ASPLOS'25] GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism
[ASPLOS'25] FlexSP: Accelerating Large Language Model Training via Flexible Sequence Parallelism
[ASPLOS'25] Spindle: Efficient Distributed Training of Multi-Task Large Models via Wavefront Scheduling
[EuroSys'25] JABAS: Joint Adaptive Batching and Automatic Scaling for DNN Training on Heterogeneous GPUs
[TPDS'24] UMPIPE: Unequal Microbatches-Based Pipeline Parallelism for Deep Neural Network Training
[Survey :mag:] [ACM CSUR'24] Resource-efficient Algorithms and Systems of Foundation Models: A Survey
[SOSP'24] Uncovering Nested Data Parallelism and Data Reuse in DNN Computation with FractalTensor
[SOSP'24] Enabling Parallelism Hot Switching for Efficient Training of Large Language Models
[TACO'24] ATP: Achieving Throughput Peak for DNN Training via Smart GPU Memory Management
[NeurIPS'24] Rethinking Memory and Communication Costs for Efficient Data Parallel Training of Large Language Models
[NeurIPS'24] SpeedLoader: An I/O efficient scheme for heterogeneous and distributed LLM operation
[SC'24] Accelerating Distributed DLRM Training with Optimized TT Decomposition and Micro-Batching
[SC'24] Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
[SoCC'24] Distributed training of large language models on AWS Trainium
[TPDS'24] AutoDDL: Automatic Distributed Deep Learning With Near-Optimal Bandwidth Cost
[SOSP'24] Enabling Parallelism Hot Switching for Efficient Training of Large Language Models
[SOSP'24] TENPLEX: Changing Resources of Deep Learning Jobs using Parallelizable Tensor Collections
[ICPP'24] AutoPipe: Automatic Configuration of Pipeline Parallelism in Shared GPU Cluster
[COLM'24] LightSeq: Sequence Level Parallelism for Distributed Training of Long Context Transformers
[OSDI'24] nnScaler: Constraint-Guided Parallelization Plan Generation for Deep Learning Training
[ATC'24] Metis: Fast Automatic Distributed Training on Heterogeneous GPUs
[ATC'24] FwdLLM: Efficient Federated Finetuning of Large Language Models with Perturbed Inferences
[ATC'24] OPER: Optimality-Guided Embedding Table Parallelization for Large-scale Recommendation Model
[HPDC'24] DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models
[ICML'24] Scaling Beyond the GPU Memory Limit for Large Mixture-of-Experts Model Training
[ICML'24] Integrated Hardware Architecture and Device Placement Search
[MLSys'24] DiffusionPipe: Training Large Diffusion Models with Efficient Pipelines
[MobiCom'24] Asteroid: Resource-Efficient Hybrid Pipeline Parallelism for Collaborative DNN Training on Heterogeneous Edge Devices
[EuroSys'24] DynaPipe: Optimizing Multi-task Training through Dynamic Pipelines
[EuroSys'24] ScheMoE: An Extensible Mixture-of-Experts Distributed Training System with Tasks Scheduling
[EuroMLSys@EuroSys'24] ML Training with Cloud GPU Shortages: Is Cross-Region the Answer?
[ASPLOS'24] AdaPipe: Optimizing Pipeline Parallelism with Adaptive Recomputation and Partitioning
[ASPLOS'24] PrimePar: Efficient Spatial-temporal Tensor Partitioning for Large Transformer Model Training
[EuroSys'24] Aceso: Efficient Parallel DNN Training through Iterative Bottleneck Alleviation
[NSDI'24] MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs
[NSDI'24] DISTMM: Accelerating Distributed Multi-modal Model Training
[NSDI'24] Accelerating Neural Recommendation Training with Embedding Scheduling
[NSDI'24] Resiliency at Scale: Managing Google’s TPUv4 Machine Learning Supercomputer
[NSDI'24] QuickUpdate: a Real-Time Personalization System for Large-Scale Recommendation Models
[NSDI'24] Scaling Large Language Model Training to More Than 10,000 GPUs
[TKDE'24] Improving Automatic Parallel Training via Balanced Memory Workload Optimization
[ICLR'24] CO2: Efficient Distributed Training with Full Communication-Computation Overlap
[AAMAS'24] Holonic Learning: A Flexible Agent-based Distributed Machine Learning Framework
[VLDB'24] Saturn: An Optimized Data System for Multi-Large-Model Deep Learning Workloads
[HPCA'24] Tessel: Boosting Distributed Execution of Large DNN Models via Flexible Schedule Search
[NSDI'24] Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible Instances
[EuroSys'24] HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program Synthesis
[ICPP'23] Mercury: Fast and Optimal Device Placement for Large Deep Learning Models
[IPDPS'23] MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism
[CLUSTER'23] Prophet: Fine-grained Load Balancing for Parallel Training of Large-scale MoE Models
[NeurIPS'23] ASPEN: Breaking Operator Barriers for Efficient Parallelization of Deep Neural Networks
[NeurIPS'23] DeepPCR: Parallelizing Sequential Operations in Neural Networks
[DAC'23] MixPipe: Efficient Bidirectional Pipeline Parallelism for Training Large-Scale Models
[SC'23] Hanayo: Harnessing Wave-like Pipeline Parallelism for Enhanced Large Model Training Efficiency
[SOSP'23] PIT: Optimization of Dynamic Sparse Deep Learning Models via Permutation Invariant Transformation
[SOSP'23] Oobleck: Resilient Distributed Training of Large Models Using Pipeline Templates
[MICRO'23] Grape: Practical and Efficient Graphed Execution for Dynamic Deep Neural Networks on GPUs
[HPCA'23] Phloem: Automatic Acceleration of Irregular Applications with Fine-Grain Pipeline Parallelism
[ACL'23] Sequence Parallelism: Long Sequence Training from System Perspective
[CCGrid'23] A Deep Learning Pipeline Parallel Optimization Method
[OSDI'23] MGG: Accelerating Graph Neural Networks with Fine-Grained Intra-Kernel Communication-Computation Pipelining on Multi-GPU Platforms
[ATC'23] Accelerating Distributed MoE Training and Inference with Lina
[ATC'23] SmartMoE: Efficiently Training Sparsely-Activated Models through Combining Offline and Online Parallelization
[ATC'23] MSRL: Distributed Reinforcement Learning with Dataflow Fragments
[Survey :mag:] [TPDS'23] A Survey on Auto-Parallelism of Large-Scale Deep Learning Training
[ICML'23] SWARM Parallelism: Training Large Models Can Be Surprisingly Communication-Efficient
[ICML'23] BPipe: Memory-Balanced Pipeline Parallelism for Training Large Language Models
[ICS'23] A Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training
[NSDI'23] TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training Jobs
[NSDI'23] Bamboo: Making Preemptible Instances Resilient for Affordable Training of Large DNNs
[NSDI'23] ARK: GPU-driven Code Execution for Distributed Deep Learning
[MLSys'23] On Optimizing the Communication of Model Parallelism
[MLSys'23] MegaBlocks: Efficient Sparse Training with Mixture-of-Experts
[MLSys'23] Tutel: Adaptive Mixture-of-Experts at Scale
[TPDS'23] Merak: An Efficient Distributed DNN Training Framework with Automated 3D Parallelism for Giant Foundation Models
[PPoPP'23] Elastic Averaging for Efficient Pipelined DNN Training
[PPoPP'23] Efficient All-Reduce for Distributed DNN Training in Optical Interconnect Systems
[VLDB'23] MiCS: Near-linear Scaling for Training Gigantic Model on Public Cloud
[VLDB'23] Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism
[ASPLOS'23] Mobius: Fine Tuning Large-Scale Models on Commodity GPU Servers
[ASPLOS'23] Optimus-CC: Efficient Large NLP Model Training with 3D Parallelism Aware Communication Compression
[ICPP'22] Tesseract: Parallelize the Tensor Parallelism Efficiently
[NeurIPS'22] Fine-tuning Language Models over Slow Networks using Activation Quantization with Guarantees
[SoCC'22] Accelerating Large-Scale Distributed Neural Network Training with SPMD Parallelism
[MLSys'22] Pathways: Asynchronous distributed dataflow for ML
[MLSys'22] SRIFTY: Swift and Thrifty Distributed Neural Network Training on the Cloud
[MLSys'22] Efficient Strong Scaling Through Burst Parallel Training
[EuroSys'22] Varuna: scalable, low-cost training of massive deep learning models
[ATC'22] Whale: Efficient Giant Model Training over Heterogeneous GPUs
[NeurIPS'22] AMP: Automatically Finding Model Parallel Strategies with Heterogeneity Awareness
[PPoPP'22] FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained models
[ICML'22] DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale
[ICML'22] GLaM: Efficient Scaling of Language Models with Mixture-of-Experts
[HPDC'22] Hare: Exploiting Inter-job and Intra-job Parallelism of Distributed Machine Learning on Heterogeneous GPUs
[OSDI'22] Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep Learning
[NSDI'22] Accelerating Collective Communication in Data Parallel Training across Deep Learning Frameworks
[JMLR'21] Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
[TPDS'21] TensorOpt: Exploring the Tradeoffs in Distributed DNN Training With Auto-Parallelism
[ATC'21] Fine-tuning giant neural networks on commodity hardware with automatic pipeline model parallelism
[SIGMOD'21] Heterogeneity-Aware Distributed Machine Learning Training via Partial Reduce(#210-communication-optimization)]
[MLSys'21] PipeMare: Asynchronous Pipeline Parallel DNN Training
[ICLR'21] GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
[NeurIPS'21] Piper: Multidimensional Planner for DNN Parallelization
[ICML'21] Memory-Efficient Pipeline-Parallel DNN Training
[ICML'21] TeraPipe: Token-Level Pipeline Parallelism for Training Large-Scale Language Models
[ICML'21] PipeTransformer: Automated Elastic Pipelining for Distributed Training of Large-scale Models
[SC'21] Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
[SC'21] Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM (PTD-P or Megatron-LM v2)
[FAST'21] Behemoth: A Flash-centric Training Accelerator for Extreme-scale DNNs
[PPoPP'21] DAPPLE: a pipelined data parallel approach for training large models
[VLDB'21] Distributed Deep Learning on Data Systems: A Comparative Analysis of Approaches
[HPCA'20] AccPar: Tensor Partitioning for Heterogeneous Deep Learning Accelerators
[NeurIPS'20] Efficient Algorithms for Device Placement of DNN Graph Operators
[KDD'20 Tutorial] DeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters
[VLDB'20] PyTorch Distributed: Experiences on Accelerating Data Parallel Training
[OSDI'20] A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU Clusters (BytePS)
[SOSP'19] PipeDream: Generalized Pipeline Parallelism for DNN Training
[NeurIPS'20] Language Models are Few-Shot Learners
[HPCA'19] HyPar: Towards Hybrid Parallelism for Deep Learning Accelerator Array
[IEEE MICRO'19] Optimizing Multi-GPU Parallelization Strategies for Deep Learning Training
[MLSys'19] Beyond data and model parallelism for deep neural networks (FlexFlow)
[MLSys'19] TicTac: Accelerating Distributed Deep Learning with Communication Scheduling
[EuroSys'19] Parallax: Sparsity-aware Data Parallel Training of Deep Neural Networks
[EuroSys'19] Supporting Very Large Models using Automatic Dataflow Graph Partitioning (Tofu)
[SOSP'19] A Generic Communication Scheduler for Distributed DNN Training Acceleration
[NeurIPS'19] Mesh-TensorFlow: Deep Learning for Supercomputers
[NeurIPS'19] GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
[ICML'18] Exploring Hidden Dimensions in Parallelizing Convolutional Neural Networks
[Survey :mag:] [IJCAI'22] Survey on Effcient Training of Large Neural Networks
[Survey :mag:] [ACM CSUR'19] Demystifying Parallel and Distributed Deep Learning
[Survey :mag:] [ACM CSUR'19] Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques, and Tools
For comprehensive list of GNN systems papers, refer to https://github.com/chwan1016/awesome-gnn-systems.
[KDD'26] OrionInfer: Low-Overhead Parallelism Switching and Live Migration for Efficient LLM Serving
[ISCA'26] CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIM
[ISCA'26] DynoPipe: Heterogeneous Edge-Cloud LLM Serving with Dynamically Orchestrated Pipeline Boundaries
[ISCA'26] Tetris: Efficient Long-context LLM Serving with Chunkwise Dynamic Sequence Parallelism
[ISCA'26] ConServe: Contiguity-Preserving Memory Management for Multi-Turn LLM Serving
[SIGOPS OSR'26] Rethinking LLM Deployment for Intent-Based Serving
[SIGOPS OSR'26] Elastic Memory Remapping for Multi-tenant LLM Serving
[ICML'26] Beyond Prediction: Tail-Aware Scheduling for LLM Inference
[MLSys'26] SHIP: SRAM-Based Huge Inference Pipelines for Fast LLM Serving
[MLSys'26] Dataflow Is All You Need
[MLSys'26] Optimizing Deployment Configurations for LLM Inference
[MobiSys'26] TimelyLLM: Time-sensitive LLM Serving System for Physical-I/O Limited Agents
[ISCA'26] Understanding Inference Scaling for LLMs: Bottlenecks, Trade-offs, and Performance Principles
[SIGCOMM'26] KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
[ICML'26] PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative Decoding
[EuroSys'26] Efficient Data Passing for Serverless Inference Workflows: A GPU-Centric Approach
[EuroSys'26] High Throughput and Low Latency LLM Serving via Adaptive KV Caching
[EuroSys'26] Automated End-to-End Model Serving with Cooperative Compilation and Scheduling
[SIGMOD'26] Serving Hybrid LLM Loads with SLO Guarantees Using CPU-GPU Attention Piggybacking
[ASPLOS'26] BlendServe: Optimizing Offline Inference with Resource-Aware Batching
[ASPLOS'26] DFVG: A Heterogeneous Architecture for Speculative Decoding with Draft-on-FPGA and Verify-on-GPU
[ASPLOS'26] QoServe: Breaking the Silos of LLM Inference Serving
[ASPLOS'26] Towards High-Goodput LLM Serving with Prefill-decode Multiplexing
[ASPLOS'26] SwiftSpec: Disaggregated Speculative Decoding and Fused Kernels for Low-Latency LLM Inference
[ASPLOS'26] ZipServ: Fast and Memory-Efficient LLM Inference with Hardware-Aware Lossless Compression
[HPCA'26] ELORA: Efficient LoRA and KV Cache Management for Multi-LoRA LLM Serving
[FAST'26] CacheSlide: Unlocking Cross Position-Aware KV Cache Reuse for Accelerating LLM Serving
[FAST'26] SolidAttention: Low-Latency SSD-based Serving on Memory-Constrained PCs
[HPCA'26] PASCAL: A Phase-Aware Scheduling Algorithm for Serving Reasoning-based Large Language Models
[MLSys'26] Meeting SLOs, Slashing Hours: Automated Enterprise LLM Optimization with OptiKIT
[VLDB'26] ORBITFLOW: SLO-Aware Long-Context LLM Serving with Fine-Grained KV Cache Reconfiguration
[IEEE Computer'26] Challenges and Research Directions for Large Language Model Inference Hardware
[NSDI'26] FlexLLM: Token-Level Co-Serving of LLM Inference and Finetuning with SLO Guarantees
[NSDI'26] FastServe: Iteration-Level Preemptive Scheduling for Large Language Model Inference
[NSDI'26] HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
[FPGA'26] CXL-SpecKV: A Disaggregated FPGA Speculative KV-Cache for Datacenter LLM Serving
[ASPLOS'26] XY-Serve: End-to-End Versatile Production Serving for Dynamic LLM Workloads
[AAAI'26] Lethe: Layer- and Time-Adaptive KV Cache Pruning for Reasoning-Intensive LLM Serving
[EuroSys'26] FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
[EuroSys'26] KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
[EuroSys'26] TokenFlow: Responsive LLM Text Streaming Serving under Request Burst via Preemptive Scheduling
[SoCC'25] Multiplexed Heterogeneous LLM Serving via Stage-Aligned Parallelism
[Middleware'25] Argus: Quality-Aware High-Throughput Text-to-Image Inference Serving System
[NeurIPS'25] SuffixDecoding: Extreme Speculative Decoding for Emerging AI Applications
[EMNLP'25] Distributed LLM Serving on Consumer-Grade GPUs by Reconciling Computation and Communication
[MICRO'25] MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
[MICRO'25] Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
[CLUSTER'25] Scalable and Fast Inference Serving via Hybrid Communication Scheduling on Heterogeneous Networks
[Survey :mag:] [ACM CSUR'25] Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
[SOSP'25] Aegaeon: Effective GPU Pooling for Concurrent LLM Serving on the Market
[SOSP'25] IC-Cache: Efficient Large Language Model Serving via In-context Caching
[SOSP'25] DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction
[COLM'25] OverFill: Two-Stage Models for Efficient Language Model Decoding
[ACM MM'25] TinyServe: Query-Aware Cache Selection for Efficient LLM Serving
[SC'25] Hetis: Serving LLMs in Heterogeneous GPU Clusters with Fine-grained and Dynamic Parallelism
[SIGCOMM'25] SCX: Stateless KV-Cache Encoding for Cloud-Scale Confidential Transformer Serving
[OSDI'25] BlitzScale: Fast and Live Large Model Autoscaling with O(1) Host Caching
[OSDI'25] WaferLLM: Large Language Model Inference at Wafer Scale
[OSDI'25] NanoFlow: Towards Optimal Large Language Model Serving Throughput
[ICML'25] Packrat: Automatic Reconfiguration for Latency Minimization in CPU-based DNN Serving
[ACL'25] SPECTRA: Faster Large Language Model Inference with Optimized Internal and External Speculation
[CODEML @ ICML'25] TorchAO: PyTorch-Native Training-to-Serving Model Optimization
[ICML'25] EPIC: Efficient Position-Independent Caching for Serving Large Language Models
[ATC'25] DEEPSERVE: Serverless Large Language Model Serving at Scale
[ISCA'25] WindServe: Efficient Phase-Disaggregated LLM Serving with Stream-based Dynamic Scheduling
[ISCA'25] Hybe: GPU-NPU Hybrid System for Efficient LLM Inference with Million-Token Context Window
[ICLR'25] TidalDecode: Fast and Accurate LLM Decoding with Position Persistent Sparse Attention
[OSDI'25] Clover: Exploiting Intra-device Parallelism for High Throughput Large Language Model Serving
[MLSys'25] SOLA: Optimizing SLO Attainment for Large Language Model Serving with State-Aware Scheduling
[MLSys'25] Marconi: Prefix Caching for the Era of Hybrid LLMs
[Mobicom'25] D2MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
[ISPASS'25] Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
[SIGMOD'25] Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
[EuroMLSys'25] Performance Aware LLM Load Balancer for Mixed Workloads
[MLSys'25] Rethinking Key-Value Cache Compression Techniques for Large Language Model Serving
[HPCA'25] PAISE: PIM-Accelerated Inference Scheduling Engine for Transformer-based LLM
[HPCA'25] throttLL'eM: Predictive GPU Throttling for Energy Efficient LLM Inference Serving
[ASPLOS'25] Aqua: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains
[ASPLOS'25] Past-Future Scheduler for LLM Serving under SLA Guarantees
[ASPLOS'25] Accelerating LLM Serving for Multi-turn Dialogues with Efficient Resource Management
[EuroSys'25] SpInfer: Leveraging Low-Level Sparsity for Efficient Large Language Model Inference on GPUs
[EuroSys'25] Multiplexing Dynamic Deep Learning Workloads with SLO-awareness in GPU Clusters
[EuroSys'25] NeuStream: Bridging Deep Learning Serving and Stream Processing
[SoCC'25] ModServe: Scalable and Resource-Efficient Large Multimodal Model Serving
[ISCA'25] Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
[NSDI'25] SuperServe: Fine-Grained Inference Serving for Unpredictable Workloads
[MLSys'25] ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
[ICLR'25] HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
[EuroSys'25] SkyServe: Serving AI Models across Regions and Clouds with Spot Instances
[ASPLOS'25] Helix: Serving Large Language Models over Heterogeneous GPUs and Network via Max-Flow
[ASPLOS'25] Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective Elasticity
[MLSys'25] FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
[EuroSys'25] A House United Within Itself: SLO-Awareness for On-Premises Containerized ML Inference Clusters via Faro
[Survey :mag:] [ACM CSUR'24] Resource-efficient Algorithms and Systems of Foundation Models: A Survey
[ICML'25] SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization [Code]
[ICLR'25] SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration [Code]
[ICML'25] SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference [Code]
[ACL'24] LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
[ACL'24] SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget
[IPDPS'24] Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
[NeurIPS'24] Kangaroo: Lossless Self-Speculative Decoding for Accelerating LLMs via Double Early Exiting
[NeurIPS'24] Toward Efficient Inference for Mixture of Experts
[NeurIPS'24] Sequoia: Scalable and Robust Speculative Decoding
[SC'24] PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation
[SC'24] SMIless: Serving DAG-based Inference with Dynamic Invocations under Serverless Computing
[SenSys'24] LiteMoE: Customizing On-device LLM Serving via Proxy Submodel Tuning
[MICRO'24] Pushing the Performance Envelope of DNN-based Recommendation Systems Inference on GPUs
[PML4LRS @ ICLR2024] Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
[EuroSys'25] Fast State Restoration in LLM Serving with HCache
[HPCA'24] KRISP: Enabling Kernel-wise RIght-sizing for Spatial Partitioned GPU Inference Servers
[NeurIPS'24] Efficient LLM Scheduling by Learning to Rank
[SOSP'24] PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
[SOSP'24] LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
[SOSP'24] Improving DNN Inference Throughput Using Practical, Per-Input Compute Adaptation
[SOSP'24] Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving
[ICPP'24] GMM: An Efficient GPU Memory Management-based Model Serving System for Multiple DNN Inference Models
[SIGCOMM'24] CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving
[ES-FoMO @ ICML'24] CO2: Precise Attention Score Observation for improving KV Cache Replacement in Large Language Models
[OSDI'24] dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM Serving
[OSDI'24] Parrot: Efficient Serving of LLM-based Applications with Semantic Variable
[OSDI'24] USHER: Holistic Interference Avoidance for Resource Optimized ML Inference
[OSDI'24] Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
[OSDI'24] ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
[OSDI'24] InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management
[OSDI'24] Llumnix: Dynamic Scheduling for Large Language Model Serving
[OSDI'24] DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
[ATC'24] Power-aware Deep Learning Model Serving with μ-Serve
[ATC'24] Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
[ATC'24] PUZZLE: Efficiently Aligning Large Language Models through Light-Weight Context Switch
[TPDS'24] ElasticBatch: A Learning-Augmented Elastic Scheduling System for Batch Inference on MIG
[OSDI'24] Parrot: Efficient Serving of LLM-based Applications with Semantic Variable
[ISCA'24] Splitwise: Efficient generative LLM inference using phase splitting
[ICML'24] Break the Sequential Dependency of LLM Inference Using Lookahead Decoding
[ICML'24] Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
[ICML'24] HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
[ICML'24] EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
[ICML'24] MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving
[MobiSys'24] ARISE: High-Capacity AR Offloading Inference Serving via Proactive Scheduling
[MobiSys'24] Pantheon: Preemptible Multi-DNN Inference on Mobile Edge GPUs
[MLSys'24] HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
[MLSys'24] S-LoRA: Serving Thousands of Concurrent LoRA Adapters
[MLSys'24] Vidur: A Large-Scale Simulation Framework For LLM Inference
[WWW'24] λGrapher: A Resource-Efficient Serverless System for GNN Serving through Graph Sharing
[ICML'24] CLLMs: Consistency Large Language Models
[EuroSys'24] Model Selection for Latency-Critical Inference Serving
[ISCA'24] Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
[ASPLOS'24] ExeGPT: Constraint-Aware Resource Scheduling for LLM Inference
[ASPLOS'24] NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
[ICML'24] DéjàVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving
[ICLR'24] Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
[NSDI'24] Approximate Caching for Efficiently Serving Diffusion Models
[ASPLOS'24] SpotServe: Serving Generative Large Language Models on Preemptible Instances
[NeurIPS'23] SpecTr: Fast Speculative Decoding via Optimal Transport
[HPDC'23] Kairos: Building Cost-Efficient Machine Learning Inference Systems with Heterogeneous Cloud Resources
[SOSP'23] Paella: Low-latency Model Serving with Virtualized GPU Scheduling
[SOSP'23] Efficient Memory Management for Large Language Model Serving with PagedAttention
[MLSys'23] Efficiently Scaling Transformer Inference
[EuroSys'23] Fast and Efficient Model Serving Using Multi-GPUs with Direct-Host-Access
[EuroSys'23] Tabi: An Efficient Multi-Level Inference System for Large Language Models
[EuroSys'23] Pocket: ML Serving from the Edge
[OSDI'23] AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving
[NSDI'23] SHEPHERD: Serving DNNs in the Wild
[VLDB'23] Serving and Optimizing Machine Learning Workflows on Heterogeneous Infrastructures
[ICML'23] Fast Inference from Transformers via Speculative Decoding
[SIGMOD'22] Serverless Data Science - Are We There Yet? A Case Study of Model Serving
[OSDI'22] Orca: A Distributed Serving System for Transformer-Based Generative Models
[OSDI'22] Microsecond-scale Preemption for Concurrent GPU-accelerated DNN Inferences
[ATC'22] SOTER: Guarding Black-box Inference for General Neural Networks at the Edge
[ATC'22] Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal Sharing
[ATC'22] Tetris: Memory-efficient Serverless Inference through Tensor Sharing
[ATC'22] PetS: A Unified Framework for Parameter-Efficient Transformers Serving
[ATC'21] INFaaS: Automated Model-less Inference Serving
[SoCC'21] Morphling: Fast, Near-Optimal Auto-Configuration for Cloud-Native Model Serving
[MobiCom'20] SPINN: Synergistic Progressive Inference of Neural Networks over Device and Cloud
[CAL'26] HyGIN: Hybrid CPU/GPU-Initiated Communication for Mixture-of-Experts Training
[SIGCOMM'26] Balancing and Beyond: Communication-Centric Optimizations in Expert Parallelism
[SIGCOMM'26] UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods
[OSDI'26] Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design
[MLSys'26] From Tokens to Layers: Redefining Stall-Free Scheduling for MoE Serving with Layered Prefill
[OCML'26] EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference
[ICML'26] ReMoE: Boosting Expert Reuse through Router Fine-Tuning in Memory-Constrained MoE LLM Inference
[ISCA'26] Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs
[ISCA'26] Orders in Chaos: Enhancing Large-Scale MoE LLM Serving with Data Movement Forecasting
[ASPLOS'26] EARTH: An Efficient MoE Accelerator with Entropy-Aware Speculative Prefetch and Result Reuse
[ASPLOS'26] MoE-APEX: An Efficient MoE Inference System with Adaptive Precision Expert Offloading
[NSDI'26] SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State Decoupling
[ASPLOS'26] LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training
[EuroSys'26] Taming Latency-Memory Trade-Off in MoE-Based LLM Serving via Fine-Grained Expert Offloading
[EuroSys'26] MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
[SC'25] Diff-MoE: Efficient Batched MoE Inference with Priority-Driven Differential Expert Caching
[SC workshop'25] Compression Error Sensitivity Analysis for Different Experts in MoE Model Inference
[SC workshop'25] Batch Tiling on Attention: Efficient Mixture of Experts Training on Wafer-Scale Processors
[MICRO'25] Optimizing All-to-All Collective Communication with Fault Tolerance on Torus Networks
[SOSP'25] KTransformers: Unleashing the Full Potential of CPU/GPU Hybrid Inference for MoE Models
[ICML'25] Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
[NeurIPS'25] BrainMoE: Cognition Joint Embedding via Mixture-of-Expert Towards Robust Brain Foundation Model
[NeurIPS'25] S’MoRE: Structural Mixture of Residual Experts for Parameter-Efficient LLM Fine-tuning
[NeurIPS'25] The Omni-Expert: A Computationally Efficient Approach to Achieve a Mixture of Experts in a Single Expert Model
[NeurIPS'25] MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE
[NeurIPS'25] FlyLoRA: Boosting Task Decoupling and Parameter Efficiency via Implicit Rank-Wise Mixture-of-Experts
[NeurIPS'25] FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training
[NeurIPS'25] FlashMoE: Fast Distributed MoE in a Single Kernel [Code]
[SC'25] MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
[SIGCOMM'25] MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism
[ICLR'25] Ada-K Routing: Boosting the Efficiency of MoE-based LLMs
[ICML'25] I2MoE: Interpretable Multimodal Interaction-aware Mixture-of-Experts
[SC'25] X-MoE: Enabling Scalable Training for Emerging Mixture-of-Experts Architectures on HPC Platforms
[SIGCOMM'25] MixNet: A Runtime Reconfigurable Optical-Electrical Fabric for Distributed Mixture-of-Experts Training
[ACL'25] EAC-MoE: Expert-Selection Aware Compressor for Mixture-of-Experts Large Language Models
[ACL'25] FOLDMOE: Efficient Long Sequence MoE Training via Attention-MoE Pipelining
[ICML'25] FloE: On-the-Fly MoE Inference on Memory-constrained GPU
[NAACL'25] Marrying LLMs with Dynamic Forecasting: A Graph Mixture-of-expert Perspective
[NAACL'25] Sparser Mixture-of-Adapters with Cross-Layer Generalization
[NAACL'25] SimSMoE: Toward Efficient Training Mixture of Experts via Solving Representational Collapse
[Mobicom'25] D2MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
[DAC'25] HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference
[TKDE'25] A Survey on Mixture of Experts
[ICLR'25] NetMoE: Accelerating MoE Training through Dynamic Sample Placement
[EuroSys'25] Samoyeds: Accelerating MoE Models with Structured Sparsity Leveraging Sparse Tensor Cores
[EuroMLSys'25] Priority-Aware Preemptive Scheduling for Mixed-Priority Workloads in MoE Inference
[EuroMLSys'25] Accelerating MoE Model Inference with Expert Sharding
[KDD'25] ResMoE: Space-efficient Compression of Mixture of Experts LLMs via Residual Restoration
[MLSys'25] Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
[CVPR'25] DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models
[ASPLOS'25] CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory
[TPDS'25] EfficientMoE: Optimizing Mixture-of-Experts Model Training with Adaptive Load Balance
[NAACL'25] MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEs
[ASPLOS'25] FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models
[MICRO'24] SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts
[TPDS'24] MPMoE: Memory Efficient MoE for Pre-Trained Models With Adaptive Pipeline Parallelism
[MLArchSys'24 @ ISCA'24] MoE-ERAS: Expert Residency Aware Selection
[COLM'24] Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training
[ME-FoMo @ ICLR'24] Scaling Laws for Fine-Grained Mixture of Experts
[ML for Sys workshop @ NeurIPS'24] IFMoE: An Inference Framework Design for Fine-grained MoE
[ML for Sys workshop @ NeurIPS'24] TurboMoE: Enhancing MoE Model Training with Smart Kernel-Fusion and Data Transformation
[EMNLP'24] MiLoRA: Efficient Mixture of Low-Rank Adaptation for Large Language Models Fine-tuning
[EMNLP'24] Mixture of Diverse Size Experts
[EMNLP'24] AdaMOE: Token-Adaptive Routing with Null Experts for Mixture-of-Experts Language Models
[ACL'24] SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget
[SoCC'24] MoEsaic: Shared Mixture of Experts
[KDD'24] Efficient Mixture of Experts based on Large Language Models for Low-Resource Data Preprocessing
[IPDPS'24] Exploiting Inter-Layer Expert Affinity for Accelerating Mixture-of-Experts Model Inference
[NeurIPS'24] Toward Efficient Inference for Mixture of Experts
[SC'24] APTMoE: Affinity-Aware Pipeline Tuning for MoE Models on Bandwidth-Constrained GPU Nodes
[NeurIPS'24] GraphMETRO: Mitigating Complex Graph Distribution Shifts via Mixture of Aligned Experts
[NeurIPS'24] LSH-MoE: Communication-efficient MoE Training via Locality-Sensitive Hashing
[NeurIPS'24] Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design
[PML4LRS @ ICLR'24] Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
[NeurIPS'24 (Splotlight)] Flex-MoE: Modeling Arbitrary Modality Combination via the Flexible Mixture-of-Experts
[SRW @ ACL'24] MoExtend: Tuning New Experts for Modality and Task Extension
[ICML'24] Scaling Beyond the GPU Memory Limit for Large Mixture-of-Experts Model Training
[MLSys'24] QMoE: Sub-1-Bit Compression of Trillion-Parameter Models
[SIGIR'24] M3oE: Multi-Domain Multi-Task Mixture-of Experts Recommendation Framework
[EuroSys'24] ScheMoE: An Extensible Mixture-of-Experts Distributed Training System with Tasks Scheduling
[ICLR'24] Mixture of LoRA Experts
[IJCAI'24] LocMoE: A Low-overhead MoE for Large Language Model Training
[ISCA'24] Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
[IPDPS'23] MPipeMoE: Memory Efficient MoE for Pre-trained Models with Adaptive Pipeline Parallelism
[EMNLP'23] Adaptive Gating in Mixture-of-Experts based Language Models
[ICLR'23] Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints
[ATC'23] Accelerating Distributed MoE Training and Inference with Lina
[OSDI'23] Optimizing Dynamic Neural Networks with Brainstorm
[SIGMOD'23] FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement
[ICS'23] A Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training
[MLSys'23] MegaBlocks: Efficient Sparse Training with Mixture-of-Experts
[MLSys'23] Tutel: Adaptive Mixture-of-Experts at Scale
[PPoPP'22] FasterMoE: modeling and optimizing training of large-scale dynamic pre-trained models
[SustaiNLP @ EMNLP'22] Who Says Elephants Can't Run: Bringing Large Scale MoE Models into Cloud Scale Production
[NeurIPS'22] Mixture-of-Experts with Expert Choice Routing
[ICML'22] DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale
[ICML'22] GLaM: Efficient Scaling of Language Models with Mixture-of-Experts
[JMLR'22] Switch transformers: scaling to trillion parameter models with simple and efficient sparsity
[EMNLP'21] Beyond Distillation: Task-level Mixture-of-Experts for Efficient Inference
[ICLR'17] Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
P^2)CoCoNET)SCCL)BytePS)P3)ByteScheduler)For comprehensive list of quantization papers, refer to https://github.com/Efficient-ML/Awesome-Model-Quantization.
FrugalMCT)https://github.com/friedrichor/Awesome-Multimodal-Papers