16 repos
Systems and datasets for combining language models with vision capabilities to perform complex reasoning tasks, including visual question answering, object detection, and scene understanding. The cluster centers on recent advances in reasoning-focused multimodal models (particularly DeepSeek-R1 and similar architectures) and associated evaluation datasets, with applications spanning reinforcement learning-enhanced reasoning and multi-object visual understanding.