mosharaf/cse585

Advanced Scalable Systems for X

98

35 commits

updated Sep 28, 2026

See the code

README

CSE 585: Advanced Scalable Systems for Agentic AI (F'26)

Administrivia

  • Catalog Number: 29242
  • Lectures/Discussion: 1005 DOW, MoWe: 10:30 AM - 12:00 PM
  • Projects/Makeup: 1005 DOW, F 1:30 PM - 2:30 PM
  • Counts as: Software Breadth and Depth (PhD); Technical Elective and 500-Level (MS/E)

Important links:

Team

Member (uniqname)RoleOffice Hours
Mosharaf Chowdhury (mosharaf)Faculty4156 LEIN. By appointments only.
Kevin Xue (kaiwenx)GSI4828 BBB, F 11:30 AM -12:30 PM.

Communication

ALL communication regarding this course must be via Ed. This includes questions, discussions, announcements, as well as private messages.

Presentation slides and paper summaries should be emailed to cse585-staff@umich.edu.

Course Description

This iteration of CSE585 will introduce you to the key concepts and the state-of-the-art in practical, scalable, and fault-tolerant systems for Agentic and Generative AI and encourage you to think about either building new tools or how to apply the existing ones.

Since datacenters and cloud computing form the backbone of modern AI, we will start with an overview of the two. We will then take a deep dive into systems for the Agentic and Generative AI landscape, focusing on different types of problems. Our topics will include: basics on generative models and agentic AI from a systems perspective; systems for the AI lifecycle including pre-training, post-training, and inference serving; serving systems for text, multimodal, and agentic workloads; state management, system interfaces, and security for agents; and the operational realities of running AI at scale, including capacity and cost, reliability and fault tolerance, and power and energy. We will cover topics primarily from top conferences that take a systems view to the relevant challenges.

Note that this course is NOT focused on AI methods. Instead, we will focus on how one can build systems so that existing AI methods can be used in practice and new AI methods can emerge.

Prerequisites

Students are expected to have good programming skills and must have taken at least one undergraduate-level systems-related course (from operating systems/EECS482, databases/EECS484, distributed systems/EECS491, and networking/EECS489). Having an undergraduate ML/AI course may be helpful, but not required or necessary.

Textbook

This course has no textbooks. We will read recent papers from top venues to understand trends in scalable GenAI and agentic systems, and their applications.

Tentative Schedule and Reading List

This is an evolving list and subject to changes due to the breakneck pace of agentic and generative AI innovations.

DateReadingsPresenationSummaryReview
Aug 31IntroductionMosharaf
Hints and Principles for Computer System Design (Required)
Machine Learning Fleet Efficiency: Analyzing and Optimizing Large-Scale Google TPU Systems with ML Productivity Goodput (Required)
The Datacenter as a Computer (Chapters 1 and 2)
Heterogeneity at Hyperscale: Characterization and Scheduling of Large Production AI Clusters at Alibaba
Sep 2No Class: Find Project Groups
How to Read a Paper (Required)
How to Give a Bad Talk (Required)
Sep 7Labor Day
Sep 9Systems for AI (Agents) BasicsKevin
The Illustrated Transformer (Required)
Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems (Required)
Anthropic, When AI builds itself
Kimi K3: Open Frontier Intelligence
Sep 14Distributed Training BasicsKevin
Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM (Required)
The Ultra-Scale Playbook: Training LLMs on GPU Clusters
Sep 16Pre-Training at ScaleEllie, Michela, Marilyn, Ruiqi;
(Context Parallelism Supplementary)
Angela, Saitej, JackZiming, Boe, Nikolai, Rashon
Tessera: A Holistic Pipeline Parallelism Framework for Trillion-Parameter Heterogeneous MoE Training (Required)
Scaling Llama 3 Training with Efficient Parallelism Strategies (Required)
WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training
Sep 21Post-TrainingSri, Samir, Arshdeep, AidenRohit, Arihan, Kushagra, JatinGautham, Mythri, Alanna, Alexander
HybridFlow: A Flexible and Efficient RLHF Framework (Required)
RollArt: Disaggregated Multi-Task Agentic RL Training at Scale (Required)
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Weave: Efficient Co-Scheduling for Disaggregated RL Post-Training
Sep 23No Class: Work on Project Proposals
Writing Reviews for Systems Conferences (Required)
Worse is Better (Required)
Sep 28Inference BasicsKevin
Efficient Memory Management for Large Language Model Serving with PagedAttention (Required)
DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving (Required)
Orca: A Distributed Serving System for Transformer-Based Generative Models
On Evaluating Performance of LLM Inference Serving Systems
Sep 30Inference: Disaggregation and FusionAngela, Saitej, Jack, YashMax, Lars, Jagger, SrinitishWei-Chun, Yun-De, Abhi
MPK: A Compiler and Runtime for Mega-Kernelizing Tensor Programs (Required)
Mooncake: Trading More Storage for Less Computation — A KVCache-centric Architecture for Serving LLM Chatbot (Required)
Not All Prefills Are Equal: PPD Disaggregation for Multi-turn LLM Serving
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
Oct 5Inference: Beyond TextPreetom, Savini, Wenquan, KennedyVaelone, Ayan, Himanish, MohammedRohit, Arihan, Kushagra, Jatin
Cornfigurator: Automated Planning for Any-to-Any Multimodal Model Serving (Required)
DiFlow: A System for Micro-Serving Text-to-Image Diffusion Workflows (Required)
TetriServe: Efficiently Serving Mixed DiT Workloads
Oct 7Agents as a New System WorkloadRuijie, Xinyi, ChenglinHaripreeth, Shreya, Janani, BhargavSri, Samir, Arshdeep, Aiden
The Cost of Dynamic Reasoning: Demystifying AI Agents and Test-Time Scaling from an AI Infrastructure Perspective (Required)
Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective (Required)
What Limits Agentic Systems Efficiency?
Oct 12Serving Systems for AgentsHongkwon, Donna, Dhruv, TobyEllie, Michela, Marilyn, RuiqiRuijie, Xinyi, Chenglin
Pie: A Programmable Serving System for Emerging LLM Applications (Required)
Murakkab: Resource-Efficient Agentic Workflow Orchestration in Cloud Platforms (Required)
Towards End-to-End Optimization of LLM-based Applications with Ayo
Oct 14Agentic State ManagementVaelone, Ayan, Himanish, MohammedPreetom, Savini, Wenquan, KennedyHaripreeth, Shreya, Janani, Bhargav
Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live (Required)
Strata: Hierarchical Context Caching for Long-Context LLM Serving (Required)
Conveyor: Efficient Tool-aware LLM Serving with Tool Partial Execution
Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM Agents
Oct 19Fall Study Break
Oct 21No Lecture: Work on Presentations
Oct 26Test-Time Compute as Resource AllocationRohit, Arihan, Kushagra, JatinRuthesh, Eric, HangPreetom, Savini, Wenquan, Kennedy
Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration (Required)
Speculative Actions: A Lossless Framework for Faster Agentic Systems (Required)
Act While Thinking: Accelerating LLM Agents via Pattern-Aware Speculative Tool Execution
Oct 28System Interfaces for AgentsGautham, Mythri, Alanna, AlexanderZiming, Boe, Nikolai, RashonRuthesh, Eric, Hang
CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution (Required)
From Imperative to Declarative: Towards LLM-friendly OS Interfaces for Boosted Computer-Use Agents (Required)
Branch-and-Browse: Efficient and Controllable Web Exploration with Tree-Structured Reasoning and Action Memory
AgentCgroup: Understanding and Controlling OS Resources of AI Agents
Nov 2Mid-Semester Presentations
Nov 4Mid-Semester Presentations
Nov 9Agents in the Physical WorldWei-Chun, Yun-De, AbhiHongkwon, Donna, Dhruv, TobyEllie, Michela, Marilyn, Ruiqi
ASPIRE: Agentic /Skills Discovery for Robotics (Required)
TimelyLLM: Time-sensitive LLM Serving System for Physical-I/O Limited Agents (Required)
VLA-Perf: Demystifying VLA Inference Performance
Nov 11Multi-Agent ExecutionHaripreeth, Shreya, Janani, BhargavWei-Chun, Yun-De, AbhiHongkwon, Donna, Dhruv, Toby
Towards a Science of Scaling Agent Systems (Required)
FlashAgents: Accelerating Multi-Agent LLM Systems via Streaming Prefill Overlap (Required)
Orla: A Library for Serving LLM-Based Multi-Agent Systems
Nov 16Agent SecurityZiming, Boe, Nikolai, RashonSri, Samir, Arshdeep, AidenVaelone, Ayan, Himanish, Mohammed
FragFuse: Bypassing Access Control of Large Language Model Agents via Memory-Based Query Fragmentation and Fusion (Required)
Towards Automating Data Access Permissions in AI Agents (Required)
ParaCell: Paravirtualized Secure Containers with Lightweight Intra-Container Isolation
Nov 18Operations: Capacity and CostJonathan, Alex, AudreyRuijie, Xinyi, ChenglinMax, Lars, Jagger, Srinitish
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework (Required)
Quota Marketplace: Dynamic Pricing for Efficient Allocation of ML Training Resources (Required)
Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning
Nov 23Operations: Reliability and Fault ToleranceRuthesh, Eric, HangJonathan, Alex, AudreyAngela, Saitej, Jack, Yash
SDCs in the Wild: Characterizing and Diagnosing SDC-Defective GPUs in Production LLM Training (Required)
LogAct: Enabling Agentic Reliability via Shared Logs (Required)
RobustRL: Role-Based Fault Tolerance System for RL Post-Training
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving
Nov 25Thanksgiving
Nov 30Operations: Power and EnergyMax, Lars, Jagger, SrinitishGautham, Mythri, Alanna, AlexanderJonathan, Alex, Audrey
KAIROS: Stateful, Context-Aware, Power-Efficient Agentic Inference Serving (Required)
Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster (Required)
Where Do the Joules Go? Diagnosing Inference Energy Consumption
Energy Calculus: A Compositional Algebra of Energy in Computational Systems
Dec 2Wrap UpMosharaf
Barbarians at the Gate: How AI is Upending Systems Research (Required)
We Need a New Ethics for a World of AI Agents (Required)
ECO: An AI-Driven Code Efficiency Optimizer for Warehouse Scale Computers
Dec 7No Lecture: Work on Posters
Creating an Effective Poster (Required)
How to Write a Great Research Paper (Required)
Dec 9Final Poster Presentations (TBD)

Policies

Honor Code

The Engineering Honor Code applies to all activities related to this course.

Groups

All activities of this course will be performed in groups of 4 students.

Required Reading

Each lecture will have two required readings that everyone must read.
There will be one or more optional related reading(s) that only the presenter(s) should be familiar with. They are optional for the rest of the class.

Student Lectures

The course will be conducted as a seminar. Only one group will present in each class. Each group will be assigned at least one lecture over the course of the semester. Presentations should succinctly cover all required papers for that lecture. The duration of the presentation should be at most 35 minutes with short clarifying questions and interruptions. The rest of the lecture time will be dedicated toward discussion on the papers and the broader topic(s) covered by the papers.

In the presentation, you should:

  • Provide necessary background and motivate the problem.
  • Present the high level idea, approach, and/or insight (using examples, whenever appropriate) in the required reading.
  • Discuss technical details so that one can understand key details without carefully reading.
  • Explain the differences between related works.
  • Identify strengths and weaknesses of the required reading and propose directions of future research.

The instructor team will review and suggest improvements for the presentations before each lecture. Therefore, the slides for a presentation must be emailed to the instructor team at least 24 hours prior to the corresponding class. To enable suggestions, use Google Slides and allow the instructor team give in-line comments.

Lecture Summaries

Each group will also be assigned to write summaries for at least one lecture. The summary assigned to a group will not be the reading they gave the lecture on. The group will write a summary for all presented papers (required readings) for that lecture.

The requirement for writing the summary is available here. Summaries violating the requirements will not be graded.

The paper summary of a paper must be emailed to the instructor team within 24 hours after its presentation. Late summaries will not be graded.

Post-Presentation Panel Discussion

To foster a deeper understanding of the papers and encourage critical thinking, each lecture will be followed by a panel discussion. This discussion will involve three distinct roles played by different student groups, simulating an interactive and dynamic scholarly exchange.

Roles and Responsibilities

  1. The Authors
  • Group Assignment: The group that presents the paper and the group that writes the summary will play the role of the paper's authors.
  • Responsibility: As authors, you are expected to defend your paper against critiques, answer questions, and discuss how you might improve or extend your research in the future, akin to writing a rebuttal during the peer-review process.
  1. The Reviewers
  • Group Assignment: Each group will be assigned to one slot to play the role of reviewers for all presented papers (required readings) of that lecture.
  • Responsibility: Reviewers critically assess the paper, posing challenging questions and highlighting potential weaknesses or areas for further investigation. Your goal is to engage in a constructive critique of the paper, simulating a peer review scenario.
  1. Rest of the Class
  • Responsibility:
    • You are required to submit one insightful question for each presented paper before each class.
    • During the panel discussions, feel free to actively ask questions and engage in the dialogue.

Participation

Given the discussion-based nature of this course, participation is required both for your own understanding and to improve the overall quality of the course. You are expected to attend all lectures (you may skip up to 2 lectures due to legitimate reasons), and more importantly, participate in class discussions. There will be random events to gauge attendance.

Not everyone has to add something every day, but it is expected that everyone has something to say over the semester.

Project

You will have to complete substantive work on an instructor-approved problem and have original contribution. Surveys are not permitted as projects; instead, each project must contain a survey of background and related work.

You must meet the following milestones (unless otherwise specified in future announcements) to ensure a high-quality project at the end of the semester:

  • Form a group and declare your group's membership and paper preferences by September 14. After this date, we will form groups from the remaining students.
  • Email a 2-page draft proposal (including references) by September 30. Remember to include the names and Michigan email addresses of the group members.
  • Each group must present mid-semester progress during class hours on November 2 and November 4.
  • Each group must turn in an 8-page final report and your code via email on or before 1:00PM EST on December 17. The report must be submitted as a PDF file, with formatting similar to that of the papers you've read in the class. It should point to a git repository with all the code along with a README file with a step-by-step guide on how to compile and run the code.
  • You can find how to access GPU resources here.

Tentative Grading

Weight
Paper Presentation20%
Paper Summary10%
Participation10%
Project Report40%
Project Presentations20%

mosharaf/cse585

Advanced Scalable Systems for X

98

35 commits

updated Sep 28, 2026

See the code

README

CSE 585: Advanced Scalable Systems for Agentic AI (F'26)

Administrivia

  • Catalog Number: 29242
  • Lectures/Discussion: 1005 DOW, MoWe: 10:30 AM - 12:00 PM
  • Projects/Makeup: 1005 DOW, F 1:30 PM - 2:30 PM
  • Counts as: Software Breadth and Depth (PhD); Technical Elective and 500-Level (MS/E)

Important links:

Team

Member (uniqname)RoleOffice Hours
Mosharaf Chowdhury (mosharaf)Faculty4156 LEIN. By appointments only.
Kevin Xue (kaiwenx)GSI4828 BBB, F 11:30 AM -12:30 PM.

Communication

ALL communication regarding this course must be via Ed. This includes questions, discussions, announcements, as well as private messages.

Presentation slides and paper summaries should be emailed to cse585-staff@umich.edu.

Course Description

This iteration of CSE585 will introduce you to the key concepts and the state-of-the-art in practical, scalable, and fault-tolerant systems for Agentic and Generative AI and encourage you to think about either building new tools or how to apply the existing ones.

Since datacenters and cloud computing form the backbone of modern AI, we will start with an overview of the two. We will then take a deep dive into systems for the Agentic and Generative AI landscape, focusing on different types of problems. Our topics will include: basics on generative models and agentic AI from a systems perspective; systems for the AI lifecycle including pre-training, post-training, and inference serving; serving systems for text, multimodal, and agentic workloads; state management, system interfaces, and security for agents; and the operational realities of running AI at scale, including capacity and cost, reliability and fault tolerance, and power and energy. We will cover topics primarily from top conferences that take a systems view to the relevant challenges.

Note that this course is NOT focused on AI methods. Instead, we will focus on how one can build systems so that existing AI methods can be used in practice and new AI methods can emerge.

Prerequisites

Students are expected to have good programming skills and must have taken at least one undergraduate-level systems-related course (from operating systems/EECS482, databases/EECS484, distributed systems/EECS491, and networking/EECS489). Having an undergraduate ML/AI course may be helpful, but not required or necessary.

Textbook

This course has no textbooks. We will read recent papers from top venues to understand trends in scalable GenAI and agentic systems, and their applications.

Tentative Schedule and Reading List

This is an evolving list and subject to changes due to the breakneck pace of agentic and generative AI innovations.

DateReadingsPresenationSummaryReview
Aug 31IntroductionMosharaf
Hints and Principles for Computer System Design (Required)
Machine Learning Fleet Efficiency: Analyzing and Optimizing Large-Scale Google TPU Systems with ML Productivity Goodput (Required)
The Datacenter as a Computer (Chapters 1 and 2)
Heterogeneity at Hyperscale: Characterization and Scheduling of Large Production AI Clusters at Alibaba
Sep 2No Class: Find Project Groups
How to Read a Paper (Required)
How to Give a Bad Talk (Required)
Sep 7Labor Day
Sep 9Systems for AI (Agents) BasicsKevin
The Illustrated Transformer (Required)
Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems (Required)
Anthropic, When AI builds itself
Kimi K3: Open Frontier Intelligence
Sep 14Distributed Training BasicsKevin
Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM (Required)
The Ultra-Scale Playbook: Training LLMs on GPU Clusters
Sep 16Pre-Training at ScaleEllie, Michela, Marilyn, Ruiqi;
(Context Parallelism Supplementary)
Angela, Saitej, JackZiming, Boe, Nikolai, Rashon
Tessera: A Holistic Pipeline Parallelism Framework for Trillion-Parameter Heterogeneous MoE Training (Required)
Scaling Llama 3 Training with Efficient Parallelism Strategies (Required)
WLB-LLM: Workload-Balanced 4D Parallelism for Large Language Model Training
Sep 21Post-TrainingSri, Samir, Arshdeep, AidenRohit, Arihan, Kushagra, JatinGautham, Mythri, Alanna, Alexander
HybridFlow: A Flexible and Efficient RLHF Framework (Required)
RollArt: Disaggregated Multi-Task Agentic RL Training at Scale (Required)
AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Weave: Efficient Co-Scheduling for Disaggregated RL Post-Training
Sep 23No Class: Work on Project Proposals
Writing Reviews for Systems Conferences (Required)
Worse is Better (Required)
Sep 28Inference BasicsKevin
Efficient Memory Management for Large Language Model Serving with PagedAttention (Required)
DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving (Required)
Orca: A Distributed Serving System for Transformer-Based Generative Models
On Evaluating Performance of LLM Inference Serving Systems
Sep 30Inference: Disaggregation and FusionAngela, Saitej, Jack, YashMax, Lars, Jagger, SrinitishWei-Chun, Yun-De, Abhi
MPK: A Compiler and Runtime for Mega-Kernelizing Tensor Programs (Required)
Mooncake: Trading More Storage for Less Computation — A KVCache-centric Architecture for Serving LLM Chatbot (Required)
Not All Prefills Are Equal: PPD Disaggregation for Multi-turn LLM Serving
LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism
Oct 5Inference: Beyond TextPreetom, Savini, Wenquan, KennedyVaelone, Ayan, Himanish, MohammedRohit, Arihan, Kushagra, Jatin
Cornfigurator: Automated Planning for Any-to-Any Multimodal Model Serving (Required)
DiFlow: A System for Micro-Serving Text-to-Image Diffusion Workflows (Required)
TetriServe: Efficiently Serving Mixed DiT Workloads
Oct 7Agents as a New System WorkloadRuijie, Xinyi, ChenglinHaripreeth, Shreya, Janani, BhargavSri, Samir, Arshdeep, Aiden
The Cost of Dynamic Reasoning: Demystifying AI Agents and Test-Time Scaling from an AI Infrastructure Perspective (Required)
Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective (Required)
What Limits Agentic Systems Efficiency?
Oct 12Serving Systems for AgentsHongkwon, Donna, Dhruv, TobyEllie, Michela, Marilyn, RuiqiRuijie, Xinyi, Chenglin
Pie: A Programmable Serving System for Emerging LLM Applications (Required)
Murakkab: Resource-Efficient Agentic Workflow Orchestration in Cloud Platforms (Required)
Towards End-to-End Optimization of LLM-based Applications with Ayo
Oct 14Agentic State ManagementVaelone, Ayan, Himanish, MohammedPreetom, Savini, Wenquan, KennedyHaripreeth, Shreya, Janani, Bhargav
Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live (Required)
Strata: Hierarchical Context Caching for Long-Context LLM Serving (Required)
Conveyor: Efficient Tool-aware LLM Serving with Tool Partial Execution
Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM Agents
Oct 19Fall Study Break
Oct 21No Lecture: Work on Presentations
Oct 26Test-Time Compute as Resource AllocationRohit, Arihan, Kushagra, JatinRuthesh, Eric, HangPreetom, Savini, Wenquan, Kennedy
Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration (Required)
Speculative Actions: A Lossless Framework for Faster Agentic Systems (Required)
Act While Thinking: Accelerating LLM Agents via Pattern-Aware Speculative Tool Execution
Oct 28System Interfaces for AgentsGautham, Mythri, Alanna, AlexanderZiming, Boe, Nikolai, RashonRuthesh, Eric, Hang
CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution (Required)
From Imperative to Declarative: Towards LLM-friendly OS Interfaces for Boosted Computer-Use Agents (Required)
Branch-and-Browse: Efficient and Controllable Web Exploration with Tree-Structured Reasoning and Action Memory
AgentCgroup: Understanding and Controlling OS Resources of AI Agents
Nov 2Mid-Semester Presentations
Nov 4Mid-Semester Presentations
Nov 9Agents in the Physical WorldWei-Chun, Yun-De, AbhiHongkwon, Donna, Dhruv, TobyEllie, Michela, Marilyn, Ruiqi
ASPIRE: Agentic /Skills Discovery for Robotics (Required)
TimelyLLM: Time-sensitive LLM Serving System for Physical-I/O Limited Agents (Required)
VLA-Perf: Demystifying VLA Inference Performance
Nov 11Multi-Agent ExecutionHaripreeth, Shreya, Janani, BhargavWei-Chun, Yun-De, AbhiHongkwon, Donna, Dhruv, Toby
Towards a Science of Scaling Agent Systems (Required)
FlashAgents: Accelerating Multi-Agent LLM Systems via Streaming Prefill Overlap (Required)
Orla: A Library for Serving LLM-Based Multi-Agent Systems
Nov 16Agent SecurityZiming, Boe, Nikolai, RashonSri, Samir, Arshdeep, AidenVaelone, Ayan, Himanish, Mohammed
FragFuse: Bypassing Access Control of Large Language Model Agents via Memory-Based Query Fragmentation and Fusion (Required)
Towards Automating Data Access Permissions in AI Agents (Required)
ParaCell: Paravirtualized Secure Containers with Lightweight Intra-Container Isolation
Nov 18Operations: Capacity and CostJonathan, Alex, AudreyRuijie, Xinyi, ChenglinMax, Lars, Jagger, Srinitish
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework (Required)
Quota Marketplace: Dynamic Pricing for Efficient Allocation of ML Training Resources (Required)
Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning
Nov 23Operations: Reliability and Fault ToleranceRuthesh, Eric, HangJonathan, Alex, AudreyAngela, Saitej, Jack, Yash
SDCs in the Wild: Characterizing and Diagnosing SDC-Defective GPUs in Production LLM Training (Required)
LogAct: Enabling Agentic Reliability via Shared Logs (Required)
RobustRL: Role-Based Fault Tolerance System for RL Post-Training
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving
Nov 25Thanksgiving
Nov 30Operations: Power and EnergyMax, Lars, Jagger, SrinitishGautham, Mythri, Alanna, AlexanderJonathan, Alex, Audrey
KAIROS: Stateful, Context-Aware, Power-Efficient Agentic Inference Serving (Required)
Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster (Required)
Where Do the Joules Go? Diagnosing Inference Energy Consumption
Energy Calculus: A Compositional Algebra of Energy in Computational Systems
Dec 2Wrap UpMosharaf
Barbarians at the Gate: How AI is Upending Systems Research (Required)
We Need a New Ethics for a World of AI Agents (Required)
ECO: An AI-Driven Code Efficiency Optimizer for Warehouse Scale Computers
Dec 7No Lecture: Work on Posters
Creating an Effective Poster (Required)
How to Write a Great Research Paper (Required)
Dec 9Final Poster Presentations (TBD)

Policies

Honor Code

The Engineering Honor Code applies to all activities related to this course.

Groups

All activities of this course will be performed in groups of 4 students.

Required Reading

Each lecture will have two required readings that everyone must read.
There will be one or more optional related reading(s) that only the presenter(s) should be familiar with. They are optional for the rest of the class.

Student Lectures

The course will be conducted as a seminar. Only one group will present in each class. Each group will be assigned at least one lecture over the course of the semester. Presentations should succinctly cover all required papers for that lecture. The duration of the presentation should be at most 35 minutes with short clarifying questions and interruptions. The rest of the lecture time will be dedicated toward discussion on the papers and the broader topic(s) covered by the papers.

In the presentation, you should:

  • Provide necessary background and motivate the problem.
  • Present the high level idea, approach, and/or insight (using examples, whenever appropriate) in the required reading.
  • Discuss technical details so that one can understand key details without carefully reading.
  • Explain the differences between related works.
  • Identify strengths and weaknesses of the required reading and propose directions of future research.

The instructor team will review and suggest improvements for the presentations before each lecture. Therefore, the slides for a presentation must be emailed to the instructor team at least 24 hours prior to the corresponding class. To enable suggestions, use Google Slides and allow the instructor team give in-line comments.

Lecture Summaries

Each group will also be assigned to write summaries for at least one lecture. The summary assigned to a group will not be the reading they gave the lecture on. The group will write a summary for all presented papers (required readings) for that lecture.

The requirement for writing the summary is available here. Summaries violating the requirements will not be graded.

The paper summary of a paper must be emailed to the instructor team within 24 hours after its presentation. Late summaries will not be graded.

Post-Presentation Panel Discussion

To foster a deeper understanding of the papers and encourage critical thinking, each lecture will be followed by a panel discussion. This discussion will involve three distinct roles played by different student groups, simulating an interactive and dynamic scholarly exchange.

Roles and Responsibilities

  1. The Authors
  • Group Assignment: The group that presents the paper and the group that writes the summary will play the role of the paper's authors.
  • Responsibility: As authors, you are expected to defend your paper against critiques, answer questions, and discuss how you might improve or extend your research in the future, akin to writing a rebuttal during the peer-review process.
  1. The Reviewers
  • Group Assignment: Each group will be assigned to one slot to play the role of reviewers for all presented papers (required readings) of that lecture.
  • Responsibility: Reviewers critically assess the paper, posing challenging questions and highlighting potential weaknesses or areas for further investigation. Your goal is to engage in a constructive critique of the paper, simulating a peer review scenario.
  1. Rest of the Class
  • Responsibility:
    • You are required to submit one insightful question for each presented paper before each class.
    • During the panel discussions, feel free to actively ask questions and engage in the dialogue.

Participation

Given the discussion-based nature of this course, participation is required both for your own understanding and to improve the overall quality of the course. You are expected to attend all lectures (you may skip up to 2 lectures due to legitimate reasons), and more importantly, participate in class discussions. There will be random events to gauge attendance.

Not everyone has to add something every day, but it is expected that everyone has something to say over the semester.

Project

You will have to complete substantive work on an instructor-approved problem and have original contribution. Surveys are not permitted as projects; instead, each project must contain a survey of background and related work.

You must meet the following milestones (unless otherwise specified in future announcements) to ensure a high-quality project at the end of the semester:

  • Form a group and declare your group's membership and paper preferences by September 14. After this date, we will form groups from the remaining students.
  • Email a 2-page draft proposal (including references) by September 30. Remember to include the names and Michigan email addresses of the group members.
  • Each group must present mid-semester progress during class hours on November 2 and November 4.
  • Each group must turn in an 8-page final report and your code via email on or before 1:00PM EST on December 17. The report must be submitted as a PDF file, with formatting similar to that of the papers you've read in the class. It should point to a git repository with all the code along with a README file with a step-by-step guide on how to compile and run the code.
  • You can find how to access GPU resources here.

Tentative Grading

Weight
Paper Presentation20%
Paper Summary10%
Participation10%
Project Report40%
Project Presentations20%