SJTU-ReArch-Group/Paper-Reading-List

158

216 commits

updated Jul 24, 2026

See the code

README

ReArch Group Paper Reading List

 

Seminars

Spring 2021

Summer 2021

Fall 2021

Spring 2022

DatePaper TitlePresenterNotes
3.10Speculation Attack: Meltdown, Spectre, Pinned-LoadsZihan LiuSlides
3.24SparTA: Deep-Learning Model Sparsity via Tensor-with-Sparsity-AttributeYue Guan
3.31ROLLER: Fast and Efficient Tensor Compilation for Deep LearningYijia DiaoLink
4.07Adaptable Register File Organization for Vector ProcessorsZhihui Zhang
4.14CORTEX: A COMPILER FOR RECURSIVE DEEP LEARNING MODELSYangjie ZhouSlides
4.21Zero-Knowledge Succinct Non-Interactive Argument of KnowledgeShuwen LuSlides
5.05Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep LearningRunzhe ChenSlides

Fall 2022

DatePaper TitlePresenterNotes
9.20ANT: Exploiting Adaptive Numerical Data Type for Low-bit Deep Neural Network QuantizationCong GuoSlides
9.27X-cache: a modular architecture for domain-specific cachesZihan LiuSlides
10.18Automatically Discovering ML OptimizationsYangjie ZhouSlides
11.8Privacy Preserving Machine Learning--inferenceZhengyi LiSlides
11.15Dynamic Tensor CompilersYijia DiaoSlides

Spring 2023

Fall 2023

DatePaper TitlePresenterNotes
9.21GPU Warp Scheduling and Control CodeWeiming HuSlides
9.28Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured SparsityYue GuanSlides
10.12Shared SIMD unit: Occamy, Two Out-of-Order Commit CPU: NOREBA and OrinocoZihan LiuSlides
10.19Multitasking on GPU: PreemptionYijia DiaoSlides
10.26SecretFlow-SPU: A Performant and User-Friendly Framework for Privacy-Preserving Machine LearningZhengyi LiSlides
11.09Efficient large-scale language model training on GPU clusters using megatron-LM; ZeRO: Memory Optimizations Toward Training Trillion Parameter Models; ZeRO-Offload: Democratizing Billion-Scale Model Training; ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep LearningJiale XuSlides
11.16ATOM: LOW-BIT QUANTIZATION FOR EFFICIENT AND ACCURATE LLM SERVINGHaoyan ZhangSlides
11.23DFU: Dataflow Processing UnitRenyang GuanSlides
12.07WaveScalar;Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning WorkloadsGonglin XuSlides
12.14Fast Inference from Transformers via Speculative Decoding;SpecInfer: Accelerating Generative Large Language Model Serving with Speculative Inference and Token Tree Verification;LLMCad: Fast and Scalable On-device Large Language Model InferenceChangming YuSlides
12.28A Framework for Fine-Grained Synchronization of Dependent GPU Kernels;Fast Fine-Grained Global Synchronization on GPUs;AutoScratch: ML-Optimized Cache Management for Inference-Oriented GPUsZiyu HuangSlides

Spring 2024

DatePaper TitlePresenterNotes
03.14LLM Attack and DefenseZhengyi LiSlides
03.21Transparent GPU Sharing in Container Clouds for Deep Learning WorkloadsYijia DiaoLink
03.28DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingShuwen LuSlide
05.098-bit Transformer Inference and Fine-tuning for Edge AcceleratorsWeiming HuSlide

Fall 2024

DatePaper TitlePresenterNotes
07.26TEE-SGX IntroductionZhengyi LiSlides
08.15Accelerating mixture of experts model InferenceShuwen LuSlides
08.22Accelerating Stable Diffusion-based Video GenerationYuge ChengSlides
09.05TCP: A Tensor Contraction Processor for AI WorkloadsWeiming HuSlides
10.11dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM ServingHaoyan ZhangSlides
10.18Dataflow Chips and a CompilerRenyang GuanSlides
11.01ByteCheckpoint A Unified Checkpointing System for Large Foundation Model DevelopmentGonglin XuSlides
11.15Opensora architecture and its computational reuseHaosong LiuSlides
11.22LLM QuantizationWenxuan MiaoSlides
11.29Survey: Large-scale 3DGSZheng LiuSlides
12.05Enhance Efficiency: 3D Gaussian Splatting for Speed and Memory OptimizationXiaotong HuangSlides
12.13Stealing Part of a Production Language ModelZhengyi LiSlides
12.20Communication-Compute Co-Optimization in Distributed TrainingYijia DiaoSlides
12.27Byte Latent Transformer: Patches Scale Better Than TokensShuyong BaoSlides
01.03Gemini Mapping and Architecture Co-exploration for Large-scale DNN Chiplet AcceleratorsRenyang GuanSlides

Spring 2025

DatePaper TitlePresenterNotes
01.17HybridFlow: A Flexible and Efficient RLHF FrameworkGonglin XuSlides
02.28Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators FusionZiyu HuangSlides
03.14Auto-Vectorization in Compilers: Leveraging SIMD for HighPerformance ComputingShihan FangSlides
03.21Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionXing MaSlides
03.28Taming Load Balancing in Distributed LLM TrainingJiale XuSlides
04.11SparseAttn for Video GenerationYulin SunSlides
04.18Towards End-to-End Optimization of LLM-based Applications with AyoJiawei HuangSlides
04.25FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts ModelsXiaotong HuangSlides
05.23EXION: Exploiting Inter- and Intra-Iteration Output Sparsity for Diffusion ModelsYuge ChenSlides
05.30Modern Programming Model for Writing Kernels on GPUsXinhao LuoSlides
06.06Prefix Sharing LLM Inference and SGLangYitong DingSlides
06.13Ditto: Accelerating Diffusion Model via Temporal Value Similarity CMC: Video Transformer Acceleration via CODEC Assisted Matrix CondensingHaosong LiuSlides
06.21-25ISCA conference notesLink
06.27Speeding up LLM and GEMMWenxuan MiaoSlides
07.25Modeling and SimulationWeiming HuSlides

Fall 2025

DatePaper TitlePresenterNotes
09.18Fine Grained Comm Comp OverlapZiyu HuangSlides
09.26A Sample-Free Compilation Framework for Efficient Dynamic Tensor ComputationYangjie ZhouSlides
10.17Hydra: Harnessing Expert Popularity for Efficient Mixture-of-Expert Inference on Chiplet SystemGonglin XuSlides
10.24LLM for cuda codegenMa XingSlides
11.14A Survey of 3DGS SLAMXiaotong HuangSlides
11.21Survey of Vision-Language-Action (VLA) ModelsZheng LiuSlides
11.28Modern DSLs Compile WorkflowXinhao LuoSlides

Spring 2026

DatePaper TitlePresenterNotes
03.06Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPCZhengyi LiSlides
03.13Where LLMs Fit, and Where We Still MatterYijia DiaoSlides1, Paper2
03.27GPU Profiling for OptimizationWu SunSlides
05.08WarpDrive: GPU-Based Fully Homomorphic Encryption Acceleration Leveraging Tensor and CUDA CoresLiukun YuSlides1,Paper2
05.14New Paradigms for KV Cache ReuseXing MaSlides
05.29从 prompt 到 context 到 harnessJiawei HuangSlides
06.06From MHA to DSA: The Evolution of Attention Paradigms in Long-Context LLMsXiaotong HuangSlides
06.12Attention Quantization in Video Diffusion ModelsYuge ChengSlides
07.03Sort-free 3dgsHe ZhuSlides
07.17LLM for Kernel generationXinhao LuoSlides

DNN Architecture

Link

 

Deep Learning Compiler

List Contributed by Zihan Liu

 

Past Architecture Papers

List Contributed by Jingwen Leng

 

List Contributed by Shuwen Lu

Quantization, Data Type, Compression, Acceleration

List Contributed by Weiming Hu

LLM for Coding Papers

List Contributed by Ma Xing and Yangjie Zhou

Reading List From Other Groups

architecture
paperlist

Contributors

2436824987

58 commits

ziyuhuang123

51 commits

huweim

19 commits

SubjectNoi

13 commits

SJTU-ReArch-Group/Paper-Reading-List

158

216 commits

updated Jul 24, 2026

See the code

README

ReArch Group Paper Reading List

 

Seminars

Spring 2021

Summer 2021

Fall 2021

Spring 2022

DatePaper TitlePresenterNotes
3.10Speculation Attack: Meltdown, Spectre, Pinned-LoadsZihan LiuSlides
3.24SparTA: Deep-Learning Model Sparsity via Tensor-with-Sparsity-AttributeYue Guan
3.31ROLLER: Fast and Efficient Tensor Compilation for Deep LearningYijia DiaoLink
4.07Adaptable Register File Organization for Vector ProcessorsZhihui Zhang
4.14CORTEX: A COMPILER FOR RECURSIVE DEEP LEARNING MODELSYangjie ZhouSlides
4.21Zero-Knowledge Succinct Non-Interactive Argument of KnowledgeShuwen LuSlides
5.05Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep LearningRunzhe ChenSlides

Fall 2022

DatePaper TitlePresenterNotes
9.20ANT: Exploiting Adaptive Numerical Data Type for Low-bit Deep Neural Network QuantizationCong GuoSlides
9.27X-cache: a modular architecture for domain-specific cachesZihan LiuSlides
10.18Automatically Discovering ML OptimizationsYangjie ZhouSlides
11.8Privacy Preserving Machine Learning--inferenceZhengyi LiSlides
11.15Dynamic Tensor CompilersYijia DiaoSlides

Spring 2023

Fall 2023

DatePaper TitlePresenterNotes
9.21GPU Warp Scheduling and Control CodeWeiming HuSlides
9.28Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured SparsityYue GuanSlides
10.12Shared SIMD unit: Occamy, Two Out-of-Order Commit CPU: NOREBA and OrinocoZihan LiuSlides
10.19Multitasking on GPU: PreemptionYijia DiaoSlides
10.26SecretFlow-SPU: A Performant and User-Friendly Framework for Privacy-Preserving Machine LearningZhengyi LiSlides
11.09Efficient large-scale language model training on GPU clusters using megatron-LM; ZeRO: Memory Optimizations Toward Training Trillion Parameter Models; ZeRO-Offload: Democratizing Billion-Scale Model Training; ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep LearningJiale XuSlides
11.16ATOM: LOW-BIT QUANTIZATION FOR EFFICIENT AND ACCURATE LLM SERVINGHaoyan ZhangSlides
11.23DFU: Dataflow Processing UnitRenyang GuanSlides
12.07WaveScalar;Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning WorkloadsGonglin XuSlides
12.14Fast Inference from Transformers via Speculative Decoding;SpecInfer: Accelerating Generative Large Language Model Serving with Speculative Inference and Token Tree Verification;LLMCad: Fast and Scalable On-device Large Language Model InferenceChangming YuSlides
12.28A Framework for Fine-Grained Synchronization of Dependent GPU Kernels;Fast Fine-Grained Global Synchronization on GPUs;AutoScratch: ML-Optimized Cache Management for Inference-Oriented GPUsZiyu HuangSlides

Spring 2024

DatePaper TitlePresenterNotes
03.14LLM Attack and DefenseZhengyi LiSlides
03.21Transparent GPU Sharing in Container Clouds for Deep Learning WorkloadsYijia DiaoLink
03.28DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingShuwen LuSlide
05.098-bit Transformer Inference and Fine-tuning for Edge AcceleratorsWeiming HuSlide

Fall 2024

DatePaper TitlePresenterNotes
07.26TEE-SGX IntroductionZhengyi LiSlides
08.15Accelerating mixture of experts model InferenceShuwen LuSlides
08.22Accelerating Stable Diffusion-based Video GenerationYuge ChengSlides
09.05TCP: A Tensor Contraction Processor for AI WorkloadsWeiming HuSlides
10.11dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM ServingHaoyan ZhangSlides
10.18Dataflow Chips and a CompilerRenyang GuanSlides
11.01ByteCheckpoint A Unified Checkpointing System for Large Foundation Model DevelopmentGonglin XuSlides
11.15Opensora architecture and its computational reuseHaosong LiuSlides
11.22LLM QuantizationWenxuan MiaoSlides
11.29Survey: Large-scale 3DGSZheng LiuSlides
12.05Enhance Efficiency: 3D Gaussian Splatting for Speed and Memory OptimizationXiaotong HuangSlides
12.13Stealing Part of a Production Language ModelZhengyi LiSlides
12.20Communication-Compute Co-Optimization in Distributed TrainingYijia DiaoSlides
12.27Byte Latent Transformer: Patches Scale Better Than TokensShuyong BaoSlides
01.03Gemini Mapping and Architecture Co-exploration for Large-scale DNN Chiplet AcceleratorsRenyang GuanSlides

Spring 2025

DatePaper TitlePresenterNotes
01.17HybridFlow: A Flexible and Efficient RLHF FrameworkGonglin XuSlides
02.28Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators FusionZiyu HuangSlides
03.14Auto-Vectorization in Compilers: Leveraging SIMD for HighPerformance ComputingShihan FangSlides
03.21Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionXing MaSlides
03.28Taming Load Balancing in Distributed LLM TrainingJiale XuSlides
04.11SparseAttn for Video GenerationYulin SunSlides
04.18Towards End-to-End Optimization of LLM-based Applications with AyoJiawei HuangSlides
04.25FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts ModelsXiaotong HuangSlides
05.23EXION: Exploiting Inter- and Intra-Iteration Output Sparsity for Diffusion ModelsYuge ChenSlides
05.30Modern Programming Model for Writing Kernels on GPUsXinhao LuoSlides
06.06Prefix Sharing LLM Inference and SGLangYitong DingSlides
06.13Ditto: Accelerating Diffusion Model via Temporal Value Similarity CMC: Video Transformer Acceleration via CODEC Assisted Matrix CondensingHaosong LiuSlides
06.21-25ISCA conference notesLink
06.27Speeding up LLM and GEMMWenxuan MiaoSlides
07.25Modeling and SimulationWeiming HuSlides

Fall 2025

DatePaper TitlePresenterNotes
09.18Fine Grained Comm Comp OverlapZiyu HuangSlides
09.26A Sample-Free Compilation Framework for Efficient Dynamic Tensor ComputationYangjie ZhouSlides
10.17Hydra: Harnessing Expert Popularity for Efficient Mixture-of-Expert Inference on Chiplet SystemGonglin XuSlides
10.24LLM for cuda codegenMa XingSlides
11.14A Survey of 3DGS SLAMXiaotong HuangSlides
11.21Survey of Vision-Language-Action (VLA) ModelsZheng LiuSlides
11.28Modern DSLs Compile WorkflowXinhao LuoSlides

Spring 2026

DatePaper TitlePresenterNotes
03.06Breaking the Layer Barrier: Remodeling Private Transformer Inference with Hybrid CKKS and MPCZhengyi LiSlides
03.13Where LLMs Fit, and Where We Still MatterYijia DiaoSlides1, Paper2
03.27GPU Profiling for OptimizationWu SunSlides
05.08WarpDrive: GPU-Based Fully Homomorphic Encryption Acceleration Leveraging Tensor and CUDA CoresLiukun YuSlides1,Paper2
05.14New Paradigms for KV Cache ReuseXing MaSlides
05.29从 prompt 到 context 到 harnessJiawei HuangSlides
06.06From MHA to DSA: The Evolution of Attention Paradigms in Long-Context LLMsXiaotong HuangSlides
06.12Attention Quantization in Video Diffusion ModelsYuge ChengSlides
07.03Sort-free 3dgsHe ZhuSlides
07.17LLM for Kernel generationXinhao LuoSlides

DNN Architecture

Link

 

Deep Learning Compiler

List Contributed by Zihan Liu

 

Past Architecture Papers

List Contributed by Jingwen Leng

 

List Contributed by Shuwen Lu

Quantization, Data Type, Compression, Acceleration

List Contributed by Weiming Hu

LLM for Coding Papers

List Contributed by Ma Xing and Yangjie Zhou

Reading List From Other Groups

architecture
paperlist

Contributors

2436824987

58 commits

ziyuhuang123

51 commits

huweim

19 commits

SubjectNoi

13 commits