16 repos
Fast inference techniques for large language models, particularly the EAGLE family of speculative decoding methods that accelerate text generation across various LLaMA model sizes and versions. Repositories in this cluster focus on optimizing inference speed and throughput for LLM endpoints through PyTorch-based implementations, with practical tools for deploying efficiently to compatible inference services. The central EAGLE variants demonstrate how speculative execution can reduce latency across model scales from 8B to 70B parameters.