Quantized LLM Model Deployment

126 repos across 8 sub-areas

Quantized large language model implementations optimized for Apple Silicon and resource-constrained environments. This cluster focuses on compressed model formats (4-bit, 6-bit quantization) and deployment strategies using frameworks like MLX, enabling efficient inference of conversational AI models on edge devices. The central repositories feature pre-quantized Qwen model variants with different precision levels and optimization techniques.

Quantized LLM Model Optimization

39 repos

Optimized and quantized versions of large language models, primarily focused on reducing model size and computational requirements through techniques like 4-bit and mixed-precision quantization. The cluster centers on pre-quantized model repositories in formats like MLX and safetensors, enabling efficient inference on resource-constrained hardware. This represents the applied side of model compression and deployment optimization for making state-of-the-art language models accessible on edge devices and consumer hardware.

Cluster 647073

30 repos

Qwen Model Quantization & Optimization

15 repos

Quantized and optimized variants of Alibaba's Qwen language models across multiple sizes and configurations, focusing on efficient deployment through quantization and safetensors format. The cluster includes specialized versions optimized for reasoning tasks and memory efficiency, with naming conventions indicating model size, quantization level, and architectural variations. These repositories represent production-ready implementations designed to reduce computational and memory requirements while maintaining model capability.

Quantized LLM Model Variants

13 repos

Quantized and optimized versions of large language models across multiple architectures (Gemma, Qwen, and related variants), primarily using 4-bit quantization formats like AWQ for efficient inference. The cluster focuses on model compression and deployment-ready formats compatible with standard transformer endpoints, enabling efficient serving of state-of-the-art language models with reduced memory and compute requirements.

Quantized LLM Model Optimization

12 repos

Low-bit quantization (2-bit, 4-bit, 6-bit) techniques for compressing and optimizing large language and vision models, particularly for deployment on resource-constrained hardware like Apple Silicon via MLX. The cluster focuses on making state-of-the-art models (Qwen, DeepSeek, MiniMax, Muse) inference-efficient through aggressive quantization while maintaining usable accuracy, with AXQ emerging as a common quantization standard across implementations.

Apple Silicon LLM Quantization

9 repos

Model quantization and optimization for large language models running on Apple Silicon hardware (MLX framework). These repositories contain pre-quantized versions of popular models like GPT-OSS and Gemma at 4-bit and 6-bit precision levels, enabling efficient inference on Apple devices. The cluster focuses on making state-of-the-art language models accessible and performant on consumer Mac hardware through aggressive quantization techniques.

Quantized MLX Model Deployments

6 repos

Optimized inference implementations of Qwen vision and speech models using Apple's MLX framework, with aggressive quantization (4-bit and 6-bit) for efficient deployment on Apple Silicon. The cluster focuses on making large multimodal and language models practical for edge inference through post-training quantization and hardware-specific optimization, with repositories representing different model variants (vision language models like Qwen3-VL and automatic speech recognition models like Qwen3-ASR) all compiled to run efficiently on Mac hardware.

Cluster 647070

2 repos