Multimodal Vision-Language Models

6 repos

Feature extraction and transformer-based architectures for processing visual and textual information jointly, with emphasis on video understanding and instruction-following capabilities. The cluster centers on efficient implementations of vision-language models (including variants based on Qwen and MOSS architectures) that combine video/image encoders with language models, often using safetensors for model serialization and featuring custom inference optimizations for different resolutions and parameter scales.

video-text-to-text ·66
en ·66
model-index ·66
multimodal ·66
safetensors ·66
transformers ·66
custom_code ·66
feature-extraction ·66
videochat_flash_qwen ·58
internvl_chat ·6