A curated list of Large Language Model systems related academic papers, articles, tutorials, slides and projects. Star this repository, and then you can keep abreast of the latest developments of this booming research field.
Three eras, one lens: what unit of work the system optimizes — a request (2024), a session or reasoning trace (2025), a whole agent trajectory (2026).

Serving is still the largest area in absolute terms, but its share of the list fell from 49% to 33% — the growth went to kernel/model co-design (3 → 29 papers), agentic systems (4 → 22), AI-for-systems (5 → 16) and edge (2 → 14).


The fastest-rising techniques by share of the year's papers: agentic / multi-agent (1.0% → 13.1%), compiler / kernel / megakernel (0.0% → 10.2%), speculative decoding (1.0% → 5.7%), sparse attention (1.0% → 4.5%) and energy / power (1.0% → 3.3%).
→ Full analysis, per-technique numbers and reproduction scripts: trends/
2026
2026
2026
2026
2026
2026
2026
A curated collection of NeurIPS 2025 papers focused on efficient systems for generative AI models. The collection includes papers on:
See the full NeurIPS 2025 collection for detailed categorization and paper summaries.
DeepSpeed: a deep learning optimization library that makes distributed training and inference easy, efficient, and effective | Microsoft
Accelerate | Hugging Face
Megatron | Nvidia
NeMo | Nvidia
torchtitan | PyTorch
torchtune: PyTorch-native fine-tuning library for LLMs with minimal dependencies | PyTorch
veScale | ByteDance
VeOmni: Scaling any Modality Model Training
Cornstarch: Distributed Multimodal Training Must Be Multimodality-Aware | UMich
GPT-NeoX: Model-parallel autoregressive LLM training combining Megatron and DeepSpeed | EleutherAI
nanotron: Minimalistic 3D-parallel (tensor/pipeline/data) LLM training framework | Hugging Face
litgpt: 20+ LLM implementations with pre-training and fine-tuning recipes | Lightning AI
LLaMA-Factory: Unified efficient fine-tuning of 100+ LLMs and VLMs via LoRA, full fine-tuning, and RL methods | ACL' 24
Unsloth: 2-5x faster LLM fine-tuning with ~80% less memory via custom Triton/CUDA kernels
Post-Training
Python
100.0%
A curated list of Large Language Model systems related academic papers, articles, tutorials, slides and projects. Star this repository, and then you can keep abreast of the latest developments of this booming research field.
Three eras, one lens: what unit of work the system optimizes — a request (2024), a session or reasoning trace (2025), a whole agent trajectory (2026).

Serving is still the largest area in absolute terms, but its share of the list fell from 49% to 33% — the growth went to kernel/model co-design (3 → 29 papers), agentic systems (4 → 22), AI-for-systems (5 → 16) and edge (2 → 14).


The fastest-rising techniques by share of the year's papers: agentic / multi-agent (1.0% → 13.1%), compiler / kernel / megakernel (0.0% → 10.2%), speculative decoding (1.0% → 5.7%), sparse attention (1.0% → 4.5%) and energy / power (1.0% → 3.3%).
→ Full analysis, per-technique numbers and reproduction scripts: trends/
2026
2026
2026
2026
2026
2026
2026
A curated collection of NeurIPS 2025 papers focused on efficient systems for generative AI models. The collection includes papers on:
See the full NeurIPS 2025 collection for detailed categorization and paper summaries.
DeepSpeed: a deep learning optimization library that makes distributed training and inference easy, efficient, and effective | Microsoft
Accelerate | Hugging Face
Megatron | Nvidia
NeMo | Nvidia
torchtitan | PyTorch
torchtune: PyTorch-native fine-tuning library for LLMs with minimal dependencies | PyTorch
veScale | ByteDance
VeOmni: Scaling any Modality Model Training
Cornstarch: Distributed Multimodal Training Must Be Multimodality-Aware | UMich
GPT-NeoX: Model-parallel autoregressive LLM training combining Megatron and DeepSpeed | EleutherAI
nanotron: Minimalistic 3D-parallel (tensor/pipeline/data) LLM training framework | Hugging Face
litgpt: 20+ LLM implementations with pre-training and fine-tuning recipes | Lightning AI
LLaMA-Factory: Unified efficient fine-tuning of 100+ LLMs and VLMs via LoRA, full fine-tuning, and RL methods | ACL' 24
Unsloth: 2-5x faster LLM fine-tuning with ~80% less memory via custom Triton/CUDA kernels
Post-Training
Python
100.0%