Quantized LLM Model Optimization

39 repos

Optimized and quantized versions of large language models, primarily focused on reducing model size and computational requirements through techniques like 4-bit and mixed-precision quantization. The cluster centers on pre-quantized model repositories in formats like MLX and safetensors, enabling efficient inference on resource-constrained hardware. This represents the applied side of model compression and deployment optimization for making state-of-the-art language models accessible on edge devices and consumer hardware.

safetensors ·17
text-generation ·17
development ·17
quantized ·17
4-bit ·17
mixed-precision ·17
conversational ·17
apple-silicon ·17
axquant ·17
vision ·17