16 repos
Tools and frameworks for building systems that combine large language models with visual understanding capabilities. This cluster spans implementations for processing images, diagrams, and visual content alongside text — from multimodal transformers to vision-language model applications. The central SVG-focused repositories represent a specialized sub-theme of visual asset generation and manipulation within this broader multimodal AI context.