8 repos
Quantized and specialized variants of large language models (13B and 7B parameter scales) optimized for efficient deployment and task-specific performance. The cluster centers on fine-tuned implementations of popular base models like Skywork, Llama 2, Mistral, and Zephyr, with emphasis on 8-bit quantization for reduced memory footprint and extended context handling. Repositories here focus on practical LLM engineering: model compression, supervised fine-tuning workflows, and inference optimization for production use cases including math reasoning and long-context tasks.