39 repos
Optimized and quantized versions of large language models, primarily focused on reducing model size and computational requirements through techniques like 4-bit and mixed-precision quantization. The cluster centers on pre-quantized model repositories in formats like MLX and safetensors, enabling efficient inference on resource-constrained hardware. This represents the applied side of model compression and deployment optimization for making state-of-the-art language models accessible on edge devices and consumer hardware.