12 repos
Infrastructure and tooling for deploying and serving large language models at scale. The cluster centers on transformer-based text generation systems, with a focus on inference optimization, model hosting via compatible endpoints, and efficient serialization formats like SafeTensors. The most connected repositories are Gemma model variants (2B through 27B parameters) in both base and instruction-tuned versions, which serve as reference implementations and benchmarks for the broader text-generation-inference ecosystem.