Efficient Attention Mechanisms for LLMs

10 repos

Libraries and implementations of optimized attention algorithms for large language models, with a focus on reducing computational cost and memory usage during inference and training. The cluster centers on CUDA-accelerated approaches like Flash Attention, which achieve efficiency gains through algorithmic improvements and hardware-aware optimizations. Most repositories are Python-based implementations and research codebases exploring variants of fast attention, making this a practical resource for practitioners building or fine-tuning LLM inference systems.

Python · 7
C++ · 1
Cuda · 1
cuda ·3,855
efficient-attention ·3,716
attention ·3,716
quantization ·3,716
triton ·3,716
video-generate ·3,716
video-generation ·3,716
vit ·3,716
inference-acceleration ·3,716
llm ·3,716