cr7258/ai-infra-learning

This repository organizes materials, recordings, and schedules related to AI-infra learning meetings.

596

57 commits

updated Mar 1, 2026

See the code

README

AI Infra 学习会议

主题时间预习资料录频文档问题反馈 & 课后思考题
vLLM Quickstart2025-05-11Doc: vLLMAI INFRA 学习 01 - LLM 全景图介绍/vLLM 快速入门01-vllm-quickstart
PagedAttention2025-05-25Blog: vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention

Paper: Efficient Memory Management for Large Language Model Serving with PagedAttention

Video: Fast LLM Serving with vLLM and PagedAttention
AI INFRA 学习 02 - vLLM PagedAttention 论文精读02-pagedattention02-PagedAttention 问题反馈
Prefix Caching2025-06-08Doc: Automatic Prefix Caching

Design Doc: Automatic Prefix Caching

Paper: SGLang: Efficient Execution of Structured Language Model Programs
AI INFRA 学习 03 - Prefix Caching 原理详解03-prefix-caching
Speculative Decoding2025-06-22Doc: Speculative Decoding

Blog: How Speculative Decoding Boosts vLLM Performance by up to 2.8x

Video: Hacker's Guide to Speculative Decoding in VLLM

Video: Speculative Decoding in vLLM

Paper: Accelerating Large Language Model Decoding with Speculative Sampling

Paper: Fast Inference from Transformers via Speculative Decoding
AI INFRA 学习 04 - Speculative Decoding 实现方案04-speculative-decoding
Chunked-Prefills2025-07-13Doc: vLLM Chunked Prefill

Paper: SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Paper: DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Paper: Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
AI INFRA 学习 05 - Chunked-Prefills 分块预填充05-chunked-prefills05-Chunked-Prefills 问题反馈 & 课后思考题
Disaggregating Prefill and Decoding2025-09-21Doc: Disaggregated Prefilling

Doc: vLLM Production Stack Disaggregated Prefill

Paper: DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving

Paper: Splitwise: Efficient generative LLM inference using phase splitting

Video: vLLM Office Hours - Disaggregated Prefill and KV Cache Storage in vLLM
AI INFRA 学习 06 - PD 分离推理架构详解06-disaggregating-prefill-and-decoding06-PD 分离问题反馈
Inference Platform2026-03-01llm-d

NVIDIA Dynamo

AIBrix

Kthena

RoleBasedGroup
AI INFRA 学习 07 - 推理平台全景07-inference-platform
LoRA AdaptersDoc: LoRA Adapters
Paper: LoRA: Low-Rank Adaptation of Large Language Models
Quantization

交流群(加群请备注来意)

<img src=https://github.com/user-attachments/assets/b0451ab2-b16e-4079-8b0a-b5893097572a width=60% />

微信公众号

<img src=https://github.com/user-attachments/assets/d2362785-c05a-4b5b-aaa7-49e939ccfc02 width=50% />

搜索框传播样式-白色版

Contributors

cr7258

57 commits

cr7258/ai-infra-learning

This repository organizes materials, recordings, and schedules related to AI-infra learning meetings.

596

57 commits

updated Mar 1, 2026

See the code

README

AI Infra 学习会议

主题时间预习资料录频文档问题反馈 & 课后思考题
vLLM Quickstart2025-05-11Doc: vLLMAI INFRA 学习 01 - LLM 全景图介绍/vLLM 快速入门01-vllm-quickstart
PagedAttention2025-05-25Blog: vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention

Paper: Efficient Memory Management for Large Language Model Serving with PagedAttention

Video: Fast LLM Serving with vLLM and PagedAttention
AI INFRA 学习 02 - vLLM PagedAttention 论文精读02-pagedattention02-PagedAttention 问题反馈
Prefix Caching2025-06-08Doc: Automatic Prefix Caching

Design Doc: Automatic Prefix Caching

Paper: SGLang: Efficient Execution of Structured Language Model Programs
AI INFRA 学习 03 - Prefix Caching 原理详解03-prefix-caching
Speculative Decoding2025-06-22Doc: Speculative Decoding

Blog: How Speculative Decoding Boosts vLLM Performance by up to 2.8x

Video: Hacker's Guide to Speculative Decoding in VLLM

Video: Speculative Decoding in vLLM

Paper: Accelerating Large Language Model Decoding with Speculative Sampling

Paper: Fast Inference from Transformers via Speculative Decoding
AI INFRA 学习 04 - Speculative Decoding 实现方案04-speculative-decoding
Chunked-Prefills2025-07-13Doc: vLLM Chunked Prefill

Paper: SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Paper: DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference

Paper: Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
AI INFRA 学习 05 - Chunked-Prefills 分块预填充05-chunked-prefills05-Chunked-Prefills 问题反馈 & 课后思考题
Disaggregating Prefill and Decoding2025-09-21Doc: Disaggregated Prefilling

Doc: vLLM Production Stack Disaggregated Prefill

Paper: DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving

Paper: Splitwise: Efficient generative LLM inference using phase splitting

Video: vLLM Office Hours - Disaggregated Prefill and KV Cache Storage in vLLM
AI INFRA 学习 06 - PD 分离推理架构详解06-disaggregating-prefill-and-decoding06-PD 分离问题反馈
Inference Platform2026-03-01llm-d

NVIDIA Dynamo

AIBrix

Kthena

RoleBasedGroup
AI INFRA 学习 07 - 推理平台全景07-inference-platform
LoRA AdaptersDoc: LoRA Adapters
Paper: LoRA: Low-Rank Adaptation of Large Language Models
Quantization

交流群(加群请备注来意)

<img src=https://github.com/user-attachments/assets/b0451ab2-b16e-4079-8b0a-b5893097572a width=60% />

微信公众号

<img src=https://github.com/user-attachments/assets/d2362785-c05a-4b5b-aaa7-49e939ccfc02 width=50% />

搜索框传播样式-白色版

Contributors

cr7258

57 commits