Edge AI Model Inference

5 repos

Optimized deployment and inference of lightweight language models on edge devices using quantized formats like GGUF. The cluster centers on the GLM Edge model family and related edge-optimized architectures (Qwen, DeepSeek variants), with repositories providing both pre-quantized model weights and inference infrastructure. Resources here are oriented toward running smaller language models efficiently on resource-constrained hardware rather than cloud-based inference.

conversational ·75
gguf ·75
en ·49
zh ·49
edge ·49
image-text-to-text ·31
endpoints_compatible ·26
text-generation ·18