11 repos
Lightweight multimodal AI models optimized for mobile and edge devices, combining vision and language understanding in compact form factors. This cluster centers on efficient implementations of vision-language models like MobileVLM and related mobile-optimized architectures, built on PyTorch and compatible with standard transformer inference endpoints. Repositories here focus on making sophisticated visual reasoning and text generation accessible on resource-constrained devices.