10 repos
Multi-modal models that combine visual and textual understanding for tasks like image-text retrieval, visual question answering, and unified vision-language representation learning. The cluster centers on pre-trained foundation models (including variants optimized with LoRA fine-tuning) and their evaluation through vision-language benchmarks, with particular emphasis on Chinese language support and cross-modal alignment techniques.