Vision-Language Models and Multimodal AI

5 repos

Large language models augmented with visual understanding capabilities, enabling systems to process and reason about both images and text. This cluster centers on instruction-tuned multimodal models like Qwen2-VL and specialized variants for document understanding (OCR), along with supporting infrastructure for model deployment via safetensors format and compatible endpoints. Researchers and practitioners exploring this cluster will find state-of-the-art model weights, reference implementations, and tools for building applications that combine vision and language understanding.

conversational ·2,029
en ·2,029
endpoints_compatible ·2,029
eval-results ·2,029
image-text-to-text ·2,029
text-generation-inference ·2,029
transformers ·2,029
qwen2_vl ·2,029
safetensors ·2,029
multimodal ·1,323