Must-read papers on improving efficiency for LLM serving clusters
34
26 commits
updated May 28, 2025
🔥 Must-read papers for managing LLM serving clusters.
This repository contains papers on optimizing the efficiency of large language model (LLM) serving clusters, including request scheduling, auto-scaling, and serverless computing-oriented storage optimization and communication optimization techniques.
Papers on large language model cluster management typically involve optimizations from multiple perspectives simultaneously. For example, they may jointly optimize request scheduling and resource allocation. We categorize these papers based on their primary focus.
We sincerely welcome everyone to collect papers presented at top-tier conferences in the fields of systems and artificial intelligence and submit pull requests.
You can directly click on the title to jump to the corresponding PDF link location
Survey of LLM Serving Systems
Inter-instance and intra-instance request/job scheduling
Autoscaling and placement of LLM serving instances
Accelerating LLM loading and service startup
Enhancing p/d disaggregation-based LLM serving
Optimizing transmission of KV Cache
Please contact me if I miss your name on the list, and I will add you back ASAP!
Must-read papers on improving efficiency for LLM serving clusters
34
26 commits
updated May 28, 2025
🔥 Must-read papers for managing LLM serving clusters.
This repository contains papers on optimizing the efficiency of large language model (LLM) serving clusters, including request scheduling, auto-scaling, and serverless computing-oriented storage optimization and communication optimization techniques.
Papers on large language model cluster management typically involve optimizations from multiple perspectives simultaneously. For example, they may jointly optimize request scheduling and resource allocation. We categorize these papers based on their primary focus.
We sincerely welcome everyone to collect papers presented at top-tier conferences in the fields of systems and artificial intelligence and submit pull requests.
You can directly click on the title to jump to the corresponding PDF link location
Survey of LLM Serving Systems
Inter-instance and intra-instance request/job scheduling
Autoscaling and placement of LLM serving instances
Accelerating LLM loading and service startup
Enhancing p/d disaggregation-based LLM serving
Optimizing transmission of KV Cache
Please contact me if I miss your name on the list, and I will add you back ASAP!