Ketonmi/Awesome-Large-Scale-LLM-Serving

Must-read papers on improving efficiency for LLM serving clusters

34

26 commits

updated May 28, 2025

See the code

README

Awesome-Large-Scale-LLM-Serving

Awesome LICENSE commit PR GitHub Repo stars

🔥 Must-read papers for managing LLM serving clusters.

This repository contains papers on optimizing the efficiency of large language model (LLM) serving clusters, including request scheduling, auto-scaling, and serverless computing-oriented storage optimization and communication optimization techniques.

Papers on large language model cluster management typically involve optimizations from multiple perspectives simultaneously. For example, they may jointly optimize request scheduling and resource allocation. We categorize these papers based on their primary focus.

We sincerely welcome everyone to collect papers presented at top-tier conferences in the fields of systems and artificial intelligence and submit pull requests.

Contents

  1. Survey
  2. Request & Job Scheduling
  3. Resource Management
  4. Serverless LLM Serving
  5. Prefill-Decoding Disaggregation
  6. Communication Optimization

📜 Papers

You can directly click on the title to jump to the corresponding PDF link location

0. Survey

Survey of LLM Serving Systems

1. Request & Job Scheduling

Inter-instance and intra-instance request/job scheduling

2. Resource Management

Autoscaling and placement of LLM serving instances

3. Serverless LLM Serving

Accelerating LLM loading and service startup

4. Prefill-Decoding Disaggregation

Enhancing p/d disaggregation-based LLM serving

5. Communication Optimization

Optimizing transmission of KV Cache

Acknowledegments

Please contact me if I miss your name on the list, and I will add you back ASAP!

Contributors

Star History

Star History Chart

Ketonmi/Awesome-Large-Scale-LLM-Serving

Must-read papers on improving efficiency for LLM serving clusters

34

26 commits

updated May 28, 2025

See the code

README

Awesome-Large-Scale-LLM-Serving

Awesome LICENSE commit PR GitHub Repo stars

🔥 Must-read papers for managing LLM serving clusters.

This repository contains papers on optimizing the efficiency of large language model (LLM) serving clusters, including request scheduling, auto-scaling, and serverless computing-oriented storage optimization and communication optimization techniques.

Papers on large language model cluster management typically involve optimizations from multiple perspectives simultaneously. For example, they may jointly optimize request scheduling and resource allocation. We categorize these papers based on their primary focus.

We sincerely welcome everyone to collect papers presented at top-tier conferences in the fields of systems and artificial intelligence and submit pull requests.

Contents

  1. Survey
  2. Request & Job Scheduling
  3. Resource Management
  4. Serverless LLM Serving
  5. Prefill-Decoding Disaggregation
  6. Communication Optimization

📜 Papers

You can directly click on the title to jump to the corresponding PDF link location

0. Survey

Survey of LLM Serving Systems

1. Request & Job Scheduling

Inter-instance and intra-instance request/job scheduling

2. Resource Management

Autoscaling and placement of LLM serving instances

3. Serverless LLM Serving

Accelerating LLM loading and service startup

4. Prefill-Decoding Disaggregation

Enhancing p/d disaggregation-based LLM serving

5. Communication Optimization

Optimizing transmission of KV Cache

Acknowledegments

Please contact me if I miss your name on the list, and I will add you back ASAP!

Contributors

Star History

Star History Chart