12 repos
Optimized model implementations and serving infrastructure for multimodal vision-language tasks, with a focus on efficient image-to-text and image-text-to-text generation. The cluster centers on PaliGemma and Gemma model variants adapted for vision capabilities, alongside tooling for loading, quantizing, and serving these models via standardized inference endpoints. Contributors here are working on making large multimodal models practical for deployment through optimizations in model architecture, weight formats, and inference serving.