6 repos
Feature extraction and transformer-based architectures for processing visual and textual information jointly, with emphasis on video understanding and instruction-following capabilities. The cluster centers on efficient implementations of vision-language models (including variants based on Qwen and MOSS architectures) that combine video/image encoders with language models, often using safetensors for model serialization and featuring custom inference optimizations for different resolutions and parameter scales.