12 repos
Visual question answering and multimodal AI systems designed for Japanese language processing. This cluster centers on adapting large vision-language models like LLaVA to Japanese contexts, including benchmarks for evaluating Japanese VQA performance, datasets for Japanese visual grounding and multi-image reasoning, and conversational AI systems that combine image understanding with Japanese natural language. The repositories span dataset creation, model fine-tuning, and evaluation frameworks for Japanese-specific multimodal tasks.