(MLSys'24) HeteGen: Efficient Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices paperInference | Parallelism | NUS]
(MLSys'24) FlashDecoding++: Faster Large Language Model Inference with Asynchronization, Flat GEMM Optimization, and Heuristics paperInference | Tsinghua | SJTU]
(MLSys'24) VIDUR: A LARGE-SCALE SIMULATION FRAMEWORK FOR LLM INFERENCE -Inference | Simulation Framework | Microsoft]
(MLSys'24) UniDM: A Unified Framework for Data Manipulation with Large Language Models paperInference | Memory | Long Context | Alibaba]
(MLSys'24) SiDA: Sparsity-Inspired Data-Aware Serving for Efficient and Scalable Large Mixture-of-Experts Models paperServing | MoE
(MLSys'24) Keyformer: KV Cache reduction through key tokens selection for Efficient Generative Inference paperInference | KV Cache
(MLSys'24) Q-Hitter: A Better Token Oracle for Efficient LLM Inference via Sparse-Quantized KV Cache paperInference | KV Cache
(MLSys'24) Lancet: Accelerating Mixture-of-Experts Training by Overlapping Weight Gradient Computation and All-to-All Communication -Training | MoE | HKU]
(MLSys'24) DiffusionPipe: Training Large Diffusion Models with Efficient Pipelines -Training | Diffusion | HKU]
ML Serving
(MLSys'24) FLASH: Fast Model Adaptation in ML-Centric Cloud Platforms papercode [MLsys | UIUC]
(MLSys'24) ACROBAT: Optimizing Auto-batching of Dynamic Deep Learning at Compile Time paperCompiling | Batching | CMU]
(MLSys'24) On Latency Predictors for Neural Architecture Search paper [Google]
(MLSys'24) vMCU: Coordinated Memory Management and Kernel Optimization for DNN Inference on MCUs paperDNN Inference | PKU]
Retrieval-Augmented Generation
(arxiv) RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation paper
(NSDI'24) Fast Vector Query Processing for Large Datasets Beyond GPU Memory with Reordered Pipelining paper
(Sigcomm'24) CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving paper
(EuroSys'25) CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion papercode
(EuroSys'25) Fast State Restoration in LLM Serving with HCache paper
(OSDI'24) Parrot: Efficient Serving of LLM-based Applications with Semantic Variable paper
(MLSys'24) HeteGen: Efficient Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices paperInference | Parallelism | NUS]
(MLSys'24) FlashDecoding++: Faster Large Language Model Inference with Asynchronization, Flat GEMM Optimization, and Heuristics paperInference | Tsinghua | SJTU]
(MLSys'24) VIDUR: A LARGE-SCALE SIMULATION FRAMEWORK FOR LLM INFERENCE -Inference | Simulation Framework | Microsoft]
(MLSys'24) UniDM: A Unified Framework for Data Manipulation with Large Language Models paperInference | Memory | Long Context | Alibaba]
(MLSys'24) SiDA: Sparsity-Inspired Data-Aware Serving for Efficient and Scalable Large Mixture-of-Experts Models paperServing | MoE
(MLSys'24) Keyformer: KV Cache reduction through key tokens selection for Efficient Generative Inference paperInference | KV Cache
(MLSys'24) Q-Hitter: A Better Token Oracle for Efficient LLM Inference via Sparse-Quantized KV Cache paperInference | KV Cache
(MLSys'24) Lancet: Accelerating Mixture-of-Experts Training by Overlapping Weight Gradient Computation and All-to-All Communication -Training | MoE | HKU]
(MLSys'24) DiffusionPipe: Training Large Diffusion Models with Efficient Pipelines -Training | Diffusion | HKU]
ML Serving
(MLSys'24) FLASH: Fast Model Adaptation in ML-Centric Cloud Platforms papercode [MLsys | UIUC]
(MLSys'24) ACROBAT: Optimizing Auto-batching of Dynamic Deep Learning at Compile Time paperCompiling | Batching | CMU]
(MLSys'24) On Latency Predictors for Neural Architecture Search paper [Google]
(MLSys'24) vMCU: Coordinated Memory Management and Kernel Optimization for DNN Inference on MCUs paperDNN Inference | PKU]
Retrieval-Augmented Generation
(arxiv) RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation paper
(NSDI'24) Fast Vector Query Processing for Large Datasets Beyond GPU Memory with Reordered Pipelining paper
(Sigcomm'24) CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving paper
(EuroSys'25) CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion papercode
(EuroSys'25) Fast State Restoration in LLM Serving with HCache paper
(OSDI'24) Parrot: Efficient Serving of LLM-based Applications with Semantic Variable paper