5 repos
Large language models augmented with visual understanding capabilities, enabling systems to process and reason about both images and text. This cluster centers on instruction-tuned multimodal models like Qwen2-VL and specialized variants for document understanding (OCR), along with supporting infrastructure for model deployment via safetensors format and compatible endpoints. Researchers and practitioners exploring this cluster will find state-of-the-art model weights, reference implementations, and tools for building applications that combine vision and language understanding.