Mobile Vision-Language Models

11 repos

Lightweight multimodal AI models optimized for mobile and edge devices, combining vision and language understanding in compact form factors. This cluster centers on efficient implementations of vision-language models like MobileVLM and related mobile-optimized architectures, built on PyTorch and compatible with standard transformer inference endpoints. Repositories here focus on making sophisticated visual reasoning and text generation accessible on resource-constrained devices.

Jupyter Notebook · 1
Python · 1
endpoints_compatible ·126
pytorch ·126
text-generation ·126
transformers ·126
mobilevlm ·74
text-generation-inference ·52
llama ·52
MobileVLM V2 ·45
MobileVLM ·29