Multimodal Reasoning with LLMs

16 repos

Systems and datasets for combining language models with vision capabilities to perform complex reasoning tasks, including visual question answering, object detection, and scene understanding. The cluster centers on recent advances in reasoning-focused multimodal models (particularly DeepSeek-R1 and similar architectures) and associated evaluation datasets, with applications spanning reinforcement learning-enhanced reasoning and multi-object visual understanding.

Python · 5
Jupyter Notebook · 1
reinforcement-learning ·4,189
reasoning ·4,189
llm ·3,476
skywork-r1v ·3,170
r1v ·3,170
multimodal-r1 ·3,170
multimodal-understanding ·3,170
deepseek-r1 ·3,170
grpo ·3,170
vlm ·3,170