Long-Context Vision-Language Models

10 repos

Extended-context multimodal AI models that process long sequences of visual and textual information together. This cluster focuses on vision-language model architectures and techniques designed to handle significantly longer input contexts than standard approaches, enabling processing of lengthy documents, videos, or image collections with associated text. The Long-VITA family of repositories represents the primary implementation focus, with variants optimized for different context lengths and frameworks.

Python · 1
long-context ·306
mllm ·306
vision-language-model ·306
custom_code ·3
long_vita ·3
safetensors ·3