10 repos
Extended-context multimodal AI models that process long sequences of visual and textual information together. This cluster focuses on vision-language model architectures and techniques designed to handle significantly longer input contexts than standard approaches, enabling processing of lengthy documents, videos, or image collections with associated text. The Long-VITA family of repositories represents the primary implementation focus, with variants optimized for different context lengths and frameworks.