Jayanaka-98/cse598-3-compilers-for-ai

0

19 commits

updated Dec 15, 2025

See the code

README

CSE 598 (3) Compilers for AI

Project Report (with @savini @zichen)

PromptCompEval: Evaluation Methodology for Comparing Prompt Compilation Frameworks

Vibe_logs for papers.

#Paper TitlePaper LinkVibelog
1DSPy: Compiling Declarative Language Model Calls into Self-Improving PipelinesDSPyLink
2GEPA: Reflective Prompt Evolution Can Outperform Reinforcement LearningGEPALink
3MTP: A Meaning-Typed Language Abstraction for AI-Integrated ApplicationsMTPLink
4TVM: An Automated End-to-End Optimizing Compiler for Deep LearningTVMLink
5Relay: A High-Level Compiler for Deep LearningRelayLink
6Ansor: Generating High-Performance Tensor Programs for Deep LearningAnsorLink
7PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode and Graph Compilation for DNNsPyTorch 2Link
8TorchBench: Benchmarking PyTorch with High API Surface CoverageTorchBenchLink
9TorchTitan: One-stop PyTorch Native Solution for Production-Ready LLM PretrainingTorchTitanLink
10ECLIP: Energy-efficient and Practical Co-Location of ML Inference Pipelines on GPUsECLIPLink
11Triton: An Intermediate Language and Compiler for Tiled Neural Network ComputationsTritonLink
12Geak: Introducing Triton Kernel AI Agent & Evaluation BenchmarksGeakLink
13Operator Fusion in XLA: Analysis and EvaluationOpFusionLink
14Memory Safe Computations with XLA CompilerMemSafeXLALink
15MLIR: A Compiler Infrastructure for the End of Moore's LawMLIRLink
16Glow: Graph Lowering Compiler Techniques for Neural NetworksGlowLink
17Efficient Memory Management for Large Language Model Serving with PagedAttentionPagedAttentionLink
18Effective Memory Management for Serving LLMs with HeterogeneityEffLLMServLink
19Demystifying the NVIDIA Ampere Architecture through Microbenchmarking and Instruction-Level AnalysisNVIDIA AmpereLink
20Optimizing sDTW for AMD GPUsAMD sDTWLink
21TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for EmbeddingsTPU v4Link
22MTIA: First Generation Silicon Targeting Meta's Recommendation SystemsMTIALink
23Machine Learning Fleet Efficiency: Analyzing and Optimizing Large-Scale Google TPU Systems with ML Productivity GoodputML Fleet EfficiencyLink

Contributors

Jayanaka-98

19 commits

Jayanaka-98/cse598-3-compilers-for-ai

0

19 commits

updated Dec 15, 2025

See the code

README

CSE 598 (3) Compilers for AI

Project Report (with @savini @zichen)

PromptCompEval: Evaluation Methodology for Comparing Prompt Compilation Frameworks

Vibe_logs for papers.

#Paper TitlePaper LinkVibelog
1DSPy: Compiling Declarative Language Model Calls into Self-Improving PipelinesDSPyLink
2GEPA: Reflective Prompt Evolution Can Outperform Reinforcement LearningGEPALink
3MTP: A Meaning-Typed Language Abstraction for AI-Integrated ApplicationsMTPLink
4TVM: An Automated End-to-End Optimizing Compiler for Deep LearningTVMLink
5Relay: A High-Level Compiler for Deep LearningRelayLink
6Ansor: Generating High-Performance Tensor Programs for Deep LearningAnsorLink
7PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode and Graph Compilation for DNNsPyTorch 2Link
8TorchBench: Benchmarking PyTorch with High API Surface CoverageTorchBenchLink
9TorchTitan: One-stop PyTorch Native Solution for Production-Ready LLM PretrainingTorchTitanLink
10ECLIP: Energy-efficient and Practical Co-Location of ML Inference Pipelines on GPUsECLIPLink
11Triton: An Intermediate Language and Compiler for Tiled Neural Network ComputationsTritonLink
12Geak: Introducing Triton Kernel AI Agent & Evaluation BenchmarksGeakLink
13Operator Fusion in XLA: Analysis and EvaluationOpFusionLink
14Memory Safe Computations with XLA CompilerMemSafeXLALink
15MLIR: A Compiler Infrastructure for the End of Moore's LawMLIRLink
16Glow: Graph Lowering Compiler Techniques for Neural NetworksGlowLink
17Efficient Memory Management for Large Language Model Serving with PagedAttentionPagedAttentionLink
18Effective Memory Management for Serving LLMs with HeterogeneityEffLLMServLink
19Demystifying the NVIDIA Ampere Architecture through Microbenchmarking and Instruction-Level AnalysisNVIDIA AmpereLink
20Optimizing sDTW for AMD GPUsAMD sDTWLink
21TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for EmbeddingsTPU v4Link
22MTIA: First Generation Silicon Targeting Meta's Recommendation SystemsMTIALink
23Machine Learning Fleet Efficiency: Analyzing and Optimizing Large-Scale Google TPU Systems with ML Productivity GoodputML Fleet EfficiencyLink

Contributors

Jayanaka-98

19 commits