5 repos
Optimized deployment and inference of lightweight language models on edge devices using quantized formats like GGUF. The cluster centers on the GLM Edge model family and related edge-optimized architectures (Qwen, DeepSeek variants), with repositories providing both pre-quantized model weights and inference infrastructure. Resources here are oriented toward running smaller language models efficiently on resource-constrained hardware rather than cloud-based inference.