17 repos
Quantized and efficient variants of multimodal language models that process image and text inputs. The cluster centers on the MLX framework implementations of models like Inkling, focusing on 4-bit and 8-bit quantization techniques to reduce model size while maintaining conversational and image-text reasoning capabilities. Repos here showcase practical approaches to deploying vision-language models with reduced computational footprint, using the safetensors format for efficient model distribution.