Text Generation & Inference Optimization

49 repos across 3 sub-areas

Libraries, models, and frameworks for efficient large language model inference and text generation at scale. The cluster centers on techniques like speculative decoding and optimized tensor operations (via safetensors), with a focus on making transformer-based models faster and more practical to deploy. Primary repos include Qwen-series models optimized for inference performance, alongside text-generation-inference frameworks and related transformer utilities for production applications.