Multimodal Language Models & Quantization

17 repos

Quantized and efficient variants of multimodal language models that process image and text inputs. The cluster centers on the MLX framework implementations of models like Inkling, focusing on 4-bit and 8-bit quantization techniques to reduce model size while maintaining conversational and image-text reasoning capabilities. Repos here showcase practical approaches to deploying vision-language models with reduced computational footprint, using the safetensors format for efficient model distribution.

conversational ·2,398
safetensors ·2,398
moe ·2,389
image-text-to-text ·2,273
inkling_mm_model ·2,273
audio-text-to-text ·2,268
endpoints_compatible ·2,266
transformers ·2,266
eval-results ·2,175
mlx ·132