Vision-Language Model Inference

12 repos

Optimized model implementations and serving infrastructure for multimodal vision-language tasks, with a focus on efficient image-to-text and image-text-to-text generation. The cluster centers on PaliGemma and Gemma model variants adapted for vision capabilities, alongside tooling for loading, quantizing, and serving these models via standardized inference endpoints. Contributors here are working on making large multimodal models practical for deployment through optimizations in model architecture, weight formats, and inference serving.

safetensors ·4,186
endpoints_compatible ·4,186
transformers ·4,186
image-text-to-text ·4,186
text-generation-inference ·3,814
gemma3 ·2,905
conversational ·2,821
eval-results ·1,480
paligemma ·909
t5gemma2 ·203