PromptCompEval: Evaluation Methodology for Comparing Prompt Compilation Frameworks
| # | Paper Title | Paper Link | Vibelog |
|---|---|---|---|
| 1 | DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines | DSPy | Link |
| 2 | GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning | GEPA | Link |
| 3 | MTP: A Meaning-Typed Language Abstraction for AI-Integrated Applications | MTP | Link |
| 4 | TVM: An Automated End-to-End Optimizing Compiler for Deep Learning | TVM | Link |
| 5 | Relay: A High-Level Compiler for Deep Learning | Relay | Link |
| 6 | Ansor: Generating High-Performance Tensor Programs for Deep Learning | Ansor | Link |
| 7 | PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode and Graph Compilation for DNNs | PyTorch 2 | Link |
| 8 | TorchBench: Benchmarking PyTorch with High API Surface Coverage | TorchBench | Link |
| 9 | TorchTitan: One-stop PyTorch Native Solution for Production-Ready LLM Pretraining | TorchTitan | Link |
| 10 | ECLIP: Energy-efficient and Practical Co-Location of ML Inference Pipelines on GPUs | ECLIP | Link |
| 11 | Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations | Triton | Link |
| 12 | Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks | Geak | Link |
| 13 | Operator Fusion in XLA: Analysis and Evaluation | OpFusion | Link |
| 14 | Memory Safe Computations with XLA Compiler | MemSafeXLA | Link |
| 15 | MLIR: A Compiler Infrastructure for the End of Moore's Law | MLIR | Link |
| 16 | Glow: Graph Lowering Compiler Techniques for Neural Networks | Glow | Link |
| 17 | Efficient Memory Management for Large Language Model Serving with PagedAttention | PagedAttention | Link |
| 18 | Effective Memory Management for Serving LLMs with Heterogeneity | EffLLMServ | Link |
| 19 | Demystifying the NVIDIA Ampere Architecture through Microbenchmarking and Instruction-Level Analysis | NVIDIA Ampere | Link |
| 20 | Optimizing sDTW for AMD GPUs | AMD sDTW | Link |
| 21 | TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for Embeddings | TPU v4 | Link |
| 22 | MTIA: First Generation Silicon Targeting Meta's Recommendation Systems | MTIA | Link |
| 23 | Machine Learning Fleet Efficiency: Analyzing and Optimizing Large-Scale Google TPU Systems with ML Productivity Goodput | ML Fleet Efficiency | Link |
19 commits
PromptCompEval: Evaluation Methodology for Comparing Prompt Compilation Frameworks
| # | Paper Title | Paper Link | Vibelog |
|---|---|---|---|
| 1 | DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines | DSPy | Link |
| 2 | GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning | GEPA | Link |
| 3 | MTP: A Meaning-Typed Language Abstraction for AI-Integrated Applications | MTP | Link |
| 4 | TVM: An Automated End-to-End Optimizing Compiler for Deep Learning | TVM | Link |
| 5 | Relay: A High-Level Compiler for Deep Learning | Relay | Link |
| 6 | Ansor: Generating High-Performance Tensor Programs for Deep Learning | Ansor | Link |
| 7 | PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode and Graph Compilation for DNNs | PyTorch 2 | Link |
| 8 | TorchBench: Benchmarking PyTorch with High API Surface Coverage | TorchBench | Link |
| 9 | TorchTitan: One-stop PyTorch Native Solution for Production-Ready LLM Pretraining | TorchTitan | Link |
| 10 | ECLIP: Energy-efficient and Practical Co-Location of ML Inference Pipelines on GPUs | ECLIP | Link |
| 11 | Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations | Triton | Link |
| 12 | Geak: Introducing Triton Kernel AI Agent & Evaluation Benchmarks | Geak | Link |
| 13 | Operator Fusion in XLA: Analysis and Evaluation | OpFusion | Link |
| 14 | Memory Safe Computations with XLA Compiler | MemSafeXLA | Link |
| 15 | MLIR: A Compiler Infrastructure for the End of Moore's Law | MLIR | Link |
| 16 | Glow: Graph Lowering Compiler Techniques for Neural Networks | Glow | Link |
| 17 | Efficient Memory Management for Large Language Model Serving with PagedAttention | PagedAttention | Link |
| 18 | Effective Memory Management for Serving LLMs with Heterogeneity | EffLLMServ | Link |
| 19 | Demystifying the NVIDIA Ampere Architecture through Microbenchmarking and Instruction-Level Analysis | NVIDIA Ampere | Link |
| 20 | Optimizing sDTW for AMD GPUs | AMD sDTW | Link |
| 21 | TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for Embeddings | TPU v4 | Link |
| 22 | MTIA: First Generation Silicon Targeting Meta's Recommendation Systems | MTIA | Link |
| 23 | Machine Learning Fleet Efficiency: Analyzing and Optimizing Large-Scale Google TPU Systems with ML Productivity Goodput | ML Fleet Efficiency | Link |
19 commits