A collection of multimodal reasoning papers, codes, datasets, benchmarks and resources.
43
115 commits
updated Jul 31, 2026
On 31 May 2026, this repository was targeted by an apparent coordinated fake-star attack that also affected many other open-source repositories. Its star count rose abnormally from approximately 673 to more than 14,000 within a single day, despite no promotion or involvement from the maintainers. We reported the incident to GitHub but have not received a substantive response. We remain sincerely grateful to the 670+ genuine supporters whose earlier stars may no longer be displayed—your support and trust have always been deeply appreciated. 💙
👏 Welcome to the Awesome-MLLM-Reasoning-Collections repository! This repository is a carefully curated collection of papers, code, datasets, benchmarks, and resources focused on reasoning within Multimodal Large Language Models (MLLMs).
Feel free to ⭐ star and fork this repository to keep up with the latest advancements and contribute to the community.

A conceptual trajectory of multimodal reasoning, evolving from static image-level understanding, through temporal video and audio reasoning, to holistic omni-level reasoning, and finally toward embodied embedding reasoning with perception–action interaction. This progression reflects increasing reasoning scope, compositionality, and interactivity.
If you find this repository or our survey useful for your research, please consider citing:
@article{hu2026static,
title = {From Static Perception to Interactive Decision: A Survey of Multimodal Reasoning},
author = {Hu, Jian and Cheng, Zixu and Ma, Yinghao and Dixit, Satvik and Pan, Bikang and Chen, Lei and Ma, Lin and Zeng, Zhixiong and Wang, Jiangya and Benetos, Emmanouil and others},
journal = {researchgate preprint},
year = {2026}
}

An evolutionary landscape of several representative multi-modal reasoning frameworks from 2022 to 2025.
26.02 VLANeXt: Recipes for Building Strong VLA Models | Paper📑 Code🖥️ Model🤗
26.02 SimVLA: A Simple VLA Baseline for Robotic Manipulation | Paper📑 Code🖥️ Model🤗
26.02 GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement Learning | Paper📑 Code🖥️ Project🌐
26.02 Recurrent-Depth VLA: Implicit Test-Time Compute Scaling of Vision-Language-Action Models via Latent Iterative Reasoning | Paper📑 Code🖥️ Project🌐
26.02 VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model | Paper📑 Code🖥️ Model🤗
26.02 DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos | Paper📑 Model🤗 Project🌐
26.02 ABot-N0: Technical Report on the VLA Foundation Model for Versatile Embodied Navigation | Paper📑 Code🖥️ Project🌐
26.02 TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic Environments | Paper📑 Code🖥️ Project🌐
26.02 QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models | Paper📑 Code🖥️
26.02 FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment | Paper📑 Code🖥️ Model🤗
26.02 TactAlign: Human-to-Robot Policy Transfer via Tactile Alignment | Paper📑 Project🌐
26.02 World Guidance: World Modeling in Condition Space for Action Generation | Paper📑 Project🌐
26.02 Green-VLA: Staged Vision-Language-Action Model for Generalist Robots | Paper📑 Code🖥️
26.02 Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs | Paper📑 Code🖥️
26.02 RISE: Self-Improving Robot Policy with Compositional World Model | Paper📑
26.02 chi_0: Resource-Aware Robust Manipulation via Taming Distributional Inconsistencies | Paper📑 Code🖥️ Model🤗
26.02 EgoHumanoid: Unlocking In-the-Wild Loco-Manipulation with Robot-Free Egocentric Demonstration | Paper📑
26.02 MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation | Paper📑 Code🖥️
26.02 ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning | Paper📑 Code🖥️ Model🤗
26.02 RLinf-Co: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models | Paper📑
26.02 Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution | Paper📑 Code🖥️
26.02 GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning | Paper📑 Code🖥️
26.02 RynnBrain: Open Embodied Foundation Models | Paper📑 Code🖥️ Model🤗 Dataset🤗
26.02 Learning Humanoid End-Effector Control for Open-Vocabulary Visual Loco-Manipulation | Paper📑
26.02 World Action Models are Zero-shot Policies | Paper📑 Code🖥️
26.02 Learning Native Continuation for Action Chunking Flow Policies | Paper📑
26.02 BiManiBench: A Hierarchical Benchmark for Evaluating Bimanual Coordination of Multimodal Large Language Models | Paper📑 Code🖥️
26.03 RoboPocket: Improve Robot Policies Instantly with Your Phone | Paper📑 Project🌐
26.03 UltraDexGrasp: Learning Universal Dexterous Grasping for Bimanual Robots with Synthetic Data | Paper📑 Project🌐
26.03 EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding | Paper📑 Project🌐
26.03 Lightweight Visual Reasoning for Socially-Aware Robots | Paper📑
26.01 ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models | Paper📑
26.01 Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning | Paper📑 Model🤗
26.01 DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation | Paper📑 Code🖥️ Dataset🤗
26.01 SOP: A Scalable Online Post-Training System for Vision-Language-Action Models | Paper📑 Project🌐
26.01 FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-Language Navigation | Paper📑 Code🖥️ Model🤗
26.01 RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation | Paper📑 Code🖥️
26.01 VLingNav: Embodied Navigation with Adaptive Reasoning and Visual-Assisted Linguistic Memory | Paper📑
25.12 DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action | Paper📑
25.12 HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for VLA Models | Paper📑
25.12 LEO-RobotAgent: A General-purpose Robotic Agent for Language-driven Embodied Operator | Paper📑
25.12 Steering VLA Models as Anti-Exploration: A Test-Time Scaling Approach | Paper📑
25.11 WMPO: World Model-based Policy Optimization for Vision-Language-Action Models | Paper📑
25.11 RynnVLA-002: A Unified Vision-Language-Action and World Model | Paper📑
25.11 Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight | Paper📑
25.11 MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots | Paper📑
25.10 VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards | Paper📑
25.10 InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy | Paper📑
25.10 X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model | Paper📑
25.10 GigaBrain-0: A World Model-Powered Vision-Language-Action Model | Paper📑
25.09 Robix: A Unified Model for Robot Interaction, Reasoning and Planning | Paper📑
25.09 FLOWER: Democratizing Generalist Robot Policies with Efficient VLA Flow Policies | Paper📑
25.08 RynnEC: Bringing MLLMs into Embodied World | Paper📑
25.08 Do What? Teaching Vision-Language-Action Models to Reject the Impossible | Paper📑
25.08 Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in VLA Policies | Paper📑
23.07 RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control | Paper📑 Project🌐
24.05 Octo: An Open-Source Generalist Robot Policy | Paper📑 Code🖥️ Project🌐 Model🤗
24.06 OpenVLA: An Open-Source Vision-Language-Action Model | Paper📑 Code🖥️ Project🌐 Model🤗
24.10 π₀: A Vision-Language-Action Flow Model for General Robot Control | Paper📑 Code🖥️
25.01 FAST: Efficient Action Tokenization for Vision-Language-Action Models | Paper📑 Code🖥️
25.02 Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models | Paper📑
25.03 Gemini Robotics: Bringing AI into the Physical World | Paper📑 Code🖥️ Project🌐 Dataset🤗
25.03 COT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models | Paper📑 Project🌐
25.03 GR00T: A Foundation Model for General-Purpose Robotics | Paper📑 Code🖥️ Model🤗 Dataset🤗
25.04 π0.5: a Vision-Language-Action Model with Open-World Generalization | Paper📑
25.06 Chain-of-Action: Faithful and Deterministic Robot Policy via Language-guided State-Action Augmentation | Paper📑 Code🖥️ Project🌐 Model🤗
25.07 Vision-Language-Action Instruction Tuning: From Understanding to Manipulation | Paper📑 Code🖥️ Project🌐 Model🤗
25.07 MinD: Learning A Dual-System World Model for Real-Time Planning and Implicit Risk Analysis | Paper📑 Code🖥️ Project🌐
| Date | Project | Task | Links |
|---|---|---|---|
| 26.03 | MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image Reasoning | Multi-image Reasoning | [📑 Paper] [🌐 Project] |
| 26.03 | RIVER: A Benchmark for Real-World Video Reasoning in Long and Short Contexts | Video Temporal Reasoning | [📑 Paper] |
| 26.03 | UniG2U-Bench: Comprehensive Benchmark for Unified Generation and Understanding MLLMs | Multimodal Evaluation | [📑 Paper] |
| 26.03 | AgentVista: Generalizable Multi-Task Agent with Diverse Visual Manipulation | Multi-Task Agent | [📑 Paper] |
| 26.02 | A Very Big Video Reasoning Suite (VBVR): 1M+ video clips across 200 reasoning tasks | Video Reasoning | [📑 Paper] [🤗 Model] [🤗 Data] |
| 26.02 | OmniGAIA: Omni-Modal AI Agent Benchmark with hindsight-guided exploration | Omni-Modal Agent Reasoning | [📑 Paper] [💻 Code] [🤗 Data] |
| 26.02 | SpatiaLab: Wild Spatial Reasoning benchmark across 6 VQA categories | Spatial Reasoning | [📑 Paper] [💻 Code] [🤗 Data] |
| 26.02 | MuRGAt: Multimodal Fact-Level Attribution benchmark for verifiable reasoning | Multimodal Attribution | [📑 Paper] [💻 Code] |
| 26.02 | DeepVision-103K: Verifiable multimodal math dataset for RLVR training | Math Reasoning | [📑 Paper] [💻 Code] [🤗 Data] |
| 26.02 | UniVBench: Unified evaluation for video foundation models across understanding, generation, editing | Video Foundation Model Evaluation | [📑 Paper] [💻 Code] |
| 26.02 | RISE-Video: Benchmark for video generators decoding implicit world rules | Video Generation Reasoning | [📑 Paper] [💻 Code] [🤗 Data] |
| 26.02 | SAW-Bench: Egocentric Situated Awareness evaluation with 786 smart-glass videos and 2,071+ QA pairs | Spatial Reasoning | [📑 Paper] |
| 26.02 | BrowseComp-V3: 300-question visual benchmark for complex multi-hop multimodal web search | Multimodal Browsing | [📑 Paper] |
| 26.02 | BiManiBench: Hierarchical benchmark for bimanual coordination evaluation in MLLMs | Bimanual Robotics | [📑 Paper] [💻 Code] |
| 26.01 | MMFineReason: Closing the Multimodal Reasoning Gap via Open Data-Centric Methods | Multimodal Reasoning | [📑 Paper] [🤗 Model] [🤗 Data] |
| 26.01 | ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis | Chart Reasoning | [📑 Paper] [💻 Code] [🤗 Model] [🤗 Data] |
| 26.01 | VideoLoom: Joint Spatial-Temporal Understanding with LoomBench | Spatial-Temporal Reasoning | [📑 Paper] [💻 Code] [🤗 Model] |
| 26.01 | PROGRESSLM: Towards Progress Reasoning in Vision-Language Models | Task Progress Reasoning | [📑 Paper] [💻 Code] [🤗 Data] |
| 26.01 | FutureOmni: Evaluating Future Forecasting from Omni-Modal Context | Omni-Modal Temporal Reasoning | [📑 Paper] |
| 26.01 | Afri-MCQA: Multimodal Cultural Question Answering for African Languages | Multilingual Multimodal Reasoning | [📑 Paper] |
| 26.01 | AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark | Cultural Multimodal Reasoning | [📑 Paper] |
| 25.12 | HERBench: Multi-Evidence Integration in Video Question Answering | Video Reasoning | [📑 Paper] |
| 25.12 | SVBench: Evaluation of Video Generation Models on Social Reasoning | Video Social Reasoning | [📑 Paper] |
| 25.12 | IF-Bench: Benchmarking MLLMs for Infrared Images | Infrared Image Understanding | [📑 Paper] |
| 25.12 | VABench: Comprehensive Benchmark for Audio-Video Generation | Audio-Video Generation | [📑 Paper] |
| 25.11 | MME-CC: Challenging Multi-Modal Evaluation Benchmark of Cognitive Capacity | Cognitive Capacity | [📑 Paper] |
| 25.11 | GGBench: Geometric Generative Reasoning Benchmark for Unified Multimodal Models | Geometric Reasoning | [📑 Paper] |
| 25.11 | WEAVE: Benchmarking In-context Interleaved Comprehension and Generation | Multimodal Comprehension & Generation | [📑 Paper] |
| 25.10 | Uni-MMMU: Massive Multi-discipline Multimodal Unified Benchmark | Multimodal Multi-discipline Reasoning | [📑 Paper] |
| 25.10 | PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs | Physical Tool Understanding | [📑 Paper] |
| 25.10 | BEAR: Benchmarking Multimodal Language Models for Atomic Embodied Capabilities | Embodied AI Capabilities | [📑 Paper] |
| 25.10 | OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs | Long-context, Video-Audio Unerstanding & Reasonin | [📑 Paper] [💻 Code] [🌐 Project] [🤗 Data] |
| 25.10 | XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models | Capability Balancing among Different Modalities | [📑 Paper] [💻 Code] [🌐 Project] |
| 25.10 | StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA | Termporal Reasoning | [📑 Paper] |
| 25.10 | Valor32k-AVQA v2.0: Open-Ended Audio-Visual Question Answering Dataset and Benchmark | Common Sense Omni Reasoning | [📑 Paper] |
| 25.09 | MARS2 2025 Challenge on Multimodal Reasoning | Multimodal Reasoning Challenge | [📑 Paper] |
| 25.09 | Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images | Table Reasoning | [📑 Paper] |
| 25.09 | AHELM: A Holistic Evaluation of Audio-Language Models | Audio-Language Understanding | [📑 Paper] |
| 25.09 | MDAR: A Multi-scene Dynamic Audio Reasoning Benchmark | Complex, Multi-scene, & Dynamically Evolving Speech & Audio Reasonin | [📑 Paper] [💻 Code] |
| 25.09 | MiMo-Audio-Eval Toolkit | Speech/Sound/Music Reasoning | [💻 Code] |
| 25.08 | SpeechR: A Benchmark for Speech Reasoning in Large Audio-Language Models | Speech Reasoning | [📑 Paper] [💻 Code] [Data] |
| 25.08 | MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence | Long-form, Spatial, and Multi-audio Reasoning on Speech/Music/Sound | [📑 Paper] [🤗 Data] |
| 25.08 | R²-AVSBench: Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation | Segmentation Reasoning | [📑 Paper] [🤗 Data] |
| 25.07 | Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding | Video Reasoning and Understanding | [📑 Paper]. [🌐 Project] [🤗 Data] |
| 25.06 | FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation | Financial Multi-Modal Reasoning Reasoning | [📑 Paper]. [💻 Code]. [🤗 Data] |
| 25.06 | MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos | Video Reasoning | [📑 Paper]. [💻 Code]. [🌐 Project] [🤗 Data] |
| 25.06 | OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models | Spatial Reasoning | [📑 Paper]. [💻 Code]. [🌐 Project] [🤗 Data] |
| 25.06 | MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark | Phonatics, Prosody, Rhetoric, Syntactics, Semantics, and Paralinguistics in Speech Understanding & Reasoning | [📑 Paper] [💻 Code] [🤗 Data] |
| 25.05 | Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities | Video&Audio Reasoning | [📑 Paper] [💻 Code] [🌐 Project] [🤗 Data] |
| 25.05 | MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix | Multi-step Audio Reasoning | [📑 Paper]. [💻 Code]. [🎥 demo] [🤗 Data] |
| 25.05 | On Path to Multimodal Generalist: General-Level and General-Bench | Multimodal Generation | [🌐 Project] [📑 Paper] [🤗 Data] |
| 25.04 | VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models | Visual Reasoning | [🌐 Project] [📑 Paper] [💻 Code] [🤗 Data] |
| 25.04 | IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs | Image-Grounded Video Perception and Reasoning | [📑 Paper] [💻 Code] |
| 25.04 | Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing | Reasoning-Informed viSual Editing | [📑 Paper] [💻 Code] |
| 25.04 | CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following | Music Information Retrieval & Knowledge | [📑 Paper] [💻 Code] |
| 25.03 | MAVERIX: Multimodal Audio-Visual Evaluation Reasoning IndeX | Common Sense Omni Reasoning | [📑 Paper] [🌐 Project] |
| 25.03 | V-STaR : Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning | Spatio-temporal Reasoning | [🌐 Project] [📑 Paper] [🤗 Data] |
| 25.03 | MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs | Spatio-temporal Understanding | [📑Paper] |
| 25.03 | Integrating Chain-of-Thought for Multimodal Alignment: A Study on 3D Vision-Language Learning | 3D-CoT | [📑 Paper] [🤗 Data] |
| 25.02 | MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models | MM-IQ | [📑 Paper] [💻 Code] |
| 25.02 | MM-RLHF: The Next Step Forward in Multimodal LLM Alignment | MM-RLHF-RewardBench, MM-RLHF-SafetyBench | [📑 Paper] |
| 25.02 | ZeroBench: An Impossible* Visual Benchmark for Contemporary Large Multimodal Models | ZeroBench | [🌐 Project] [🤗 Dataset] [💻 Code] |
| 25.02 | MME-CoT: Benchmarking Chain-of-Thought in LMMs for Reasoning Quality, Robustness, and Efficiency | MME-CoT | [📑 Paper] [💻 Code] |
| 25.02 | OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference | MM-AlignBench | [📑 Paper] [💻 Code] |
| 25.01 | AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs | Adversarial attack, Compositional reasoning, and Modality-specific dependency in Visual&Audio | [📑 Paper] |
| 25.01 | LlamaV-o1: Rethinking Step-By-Step Visual Reasoning in LLM |
Truncated — view the full README on GitHub.
A collection of multimodal reasoning papers, codes, datasets, benchmarks and resources.
43
115 commits
updated Jul 31, 2026
On 31 May 2026, this repository was targeted by an apparent coordinated fake-star attack that also affected many other open-source repositories. Its star count rose abnormally from approximately 673 to more than 14,000 within a single day, despite no promotion or involvement from the maintainers. We reported the incident to GitHub but have not received a substantive response. We remain sincerely grateful to the 670+ genuine supporters whose earlier stars may no longer be displayed—your support and trust have always been deeply appreciated. 💙
👏 Welcome to the Awesome-MLLM-Reasoning-Collections repository! This repository is a carefully curated collection of papers, code, datasets, benchmarks, and resources focused on reasoning within Multimodal Large Language Models (MLLMs).
Feel free to ⭐ star and fork this repository to keep up with the latest advancements and contribute to the community.

A conceptual trajectory of multimodal reasoning, evolving from static image-level understanding, through temporal video and audio reasoning, to holistic omni-level reasoning, and finally toward embodied embedding reasoning with perception–action interaction. This progression reflects increasing reasoning scope, compositionality, and interactivity.
If you find this repository or our survey useful for your research, please consider citing:
@article{hu2026static,
title = {From Static Perception to Interactive Decision: A Survey of Multimodal Reasoning},
author = {Hu, Jian and Cheng, Zixu and Ma, Yinghao and Dixit, Satvik and Pan, Bikang and Chen, Lei and Ma, Lin and Zeng, Zhixiong and Wang, Jiangya and Benetos, Emmanouil and others},
journal = {researchgate preprint},
year = {2026}
}

An evolutionary landscape of several representative multi-modal reasoning frameworks from 2022 to 2025.
26.02 VLANeXt: Recipes for Building Strong VLA Models | Paper📑 Code🖥️ Model🤗
26.02 SimVLA: A Simple VLA Baseline for Robotic Manipulation | Paper📑 Code🖥️ Model🤗
26.02 GigaBrain-0.5M*: a VLA That Learns From World Model-Based Reinforcement Learning | Paper📑 Code🖥️ Project🌐
26.02 Recurrent-Depth VLA: Implicit Test-Time Compute Scaling of Vision-Language-Action Models via Latent Iterative Reasoning | Paper📑 Code🖥️ Project🌐
26.02 VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model | Paper📑 Code🖥️ Model🤗
26.02 DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos | Paper📑 Model🤗 Project🌐
26.02 ABot-N0: Technical Report on the VLA Foundation Model for Versatile Embodied Navigation | Paper📑 Code🖥️ Project🌐
26.02 TIC-VLA: A Think-in-Control Vision-Language-Action Model for Robot Navigation in Dynamic Environments | Paper📑 Code🖥️ Project🌐
26.02 QuantVLA: Scale-Calibrated Post-Training Quantization for Vision-Language-Action Models | Paper📑 Code🖥️
26.02 FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment | Paper📑 Code🖥️ Model🤗
26.02 TactAlign: Human-to-Robot Policy Transfer via Tactile Alignment | Paper📑 Project🌐
26.02 World Guidance: World Modeling in Condition Space for Action Generation | Paper📑 Project🌐
26.02 Green-VLA: Staged Vision-Language-Action Model for Generalist Robots | Paper📑 Code🖥️
26.02 Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMs | Paper📑 Code🖥️
26.02 RISE: Self-Improving Robot Policy with Compositional World Model | Paper📑
26.02 chi_0: Resource-Aware Robust Manipulation via Taming Distributional Inconsistencies | Paper📑 Code🖥️ Model🤗
26.02 EgoHumanoid: Unlocking In-the-Wild Loco-Manipulation with Robot-Free Egocentric Demonstration | Paper📑
26.02 MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation | Paper📑 Code🖥️
26.02 ABot-M0: VLA Foundation Model for Robotic Manipulation with Action Manifold Learning | Paper📑 Code🖥️ Model🤗
26.02 RLinf-Co: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models | Paper📑
26.02 Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution | Paper📑 Code🖥️
26.02 GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning | Paper📑 Code🖥️
26.02 RynnBrain: Open Embodied Foundation Models | Paper📑 Code🖥️ Model🤗 Dataset🤗
26.02 Learning Humanoid End-Effector Control for Open-Vocabulary Visual Loco-Manipulation | Paper📑
26.02 World Action Models are Zero-shot Policies | Paper📑 Code🖥️
26.02 Learning Native Continuation for Action Chunking Flow Policies | Paper📑
26.02 BiManiBench: A Hierarchical Benchmark for Evaluating Bimanual Coordination of Multimodal Large Language Models | Paper📑 Code🖥️
26.03 RoboPocket: Improve Robot Policies Instantly with Your Phone | Paper📑 Project🌐
26.03 UltraDexGrasp: Learning Universal Dexterous Grasping for Bimanual Robots with Synthetic Data | Paper📑 Project🌐
26.03 EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding | Paper📑 Project🌐
26.03 Lightweight Visual Reasoning for Socially-Aware Robots | Paper📑
26.01 ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models | Paper📑
26.01 Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning | Paper📑 Model🤗
26.01 DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation | Paper📑 Code🖥️ Dataset🤗
26.01 SOP: A Scalable Online Post-Training System for Vision-Language-Action Models | Paper📑 Project🌐
26.01 FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-Language Navigation | Paper📑 Code🖥️ Model🤗
26.01 RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation | Paper📑 Code🖥️
26.01 VLingNav: Embodied Navigation with Adaptive Reasoning and Visual-Assisted Linguistic Memory | Paper📑
25.12 DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action | Paper📑
25.12 HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for VLA Models | Paper📑
25.12 LEO-RobotAgent: A General-purpose Robotic Agent for Language-driven Embodied Operator | Paper📑
25.12 Steering VLA Models as Anti-Exploration: A Test-Time Scaling Approach | Paper📑
25.11 WMPO: World Model-based Policy Optimization for Vision-Language-Action Models | Paper📑
25.11 RynnVLA-002: A Unified Vision-Language-Action and World Model | Paper📑
25.11 Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight | Paper📑
25.11 MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots | Paper📑
25.10 VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards | Paper📑
25.10 InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy | Paper📑
25.10 X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model | Paper📑
25.10 GigaBrain-0: A World Model-Powered Vision-Language-Action Model | Paper📑
25.09 Robix: A Unified Model for Robot Interaction, Reasoning and Planning | Paper📑
25.09 FLOWER: Democratizing Generalist Robot Policies with Efficient VLA Flow Policies | Paper📑
25.08 RynnEC: Bringing MLLMs into Embodied World | Paper📑
25.08 Do What? Teaching Vision-Language-Action Models to Reject the Impossible | Paper📑
25.08 Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in VLA Policies | Paper📑
23.07 RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control | Paper📑 Project🌐
24.05 Octo: An Open-Source Generalist Robot Policy | Paper📑 Code🖥️ Project🌐 Model🤗
24.06 OpenVLA: An Open-Source Vision-Language-Action Model | Paper📑 Code🖥️ Project🌐 Model🤗
24.10 π₀: A Vision-Language-Action Flow Model for General Robot Control | Paper📑 Code🖥️
25.01 FAST: Efficient Action Tokenization for Vision-Language-Action Models | Paper📑 Code🖥️
25.02 Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models | Paper📑
25.03 Gemini Robotics: Bringing AI into the Physical World | Paper📑 Code🖥️ Project🌐 Dataset🤗
25.03 COT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models | Paper📑 Project🌐
25.03 GR00T: A Foundation Model for General-Purpose Robotics | Paper📑 Code🖥️ Model🤗 Dataset🤗
25.04 π0.5: a Vision-Language-Action Model with Open-World Generalization | Paper📑
25.06 Chain-of-Action: Faithful and Deterministic Robot Policy via Language-guided State-Action Augmentation | Paper📑 Code🖥️ Project🌐 Model🤗
25.07 Vision-Language-Action Instruction Tuning: From Understanding to Manipulation | Paper📑 Code🖥️ Project🌐 Model🤗
25.07 MinD: Learning A Dual-System World Model for Real-Time Planning and Implicit Risk Analysis | Paper📑 Code🖥️ Project🌐
| Date | Project | Task | Links |
|---|---|---|---|
| 26.03 | MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image Reasoning | Multi-image Reasoning | [📑 Paper] [🌐 Project] |
| 26.03 | RIVER: A Benchmark for Real-World Video Reasoning in Long and Short Contexts | Video Temporal Reasoning | [📑 Paper] |
| 26.03 | UniG2U-Bench: Comprehensive Benchmark for Unified Generation and Understanding MLLMs | Multimodal Evaluation | [📑 Paper] |
| 26.03 | AgentVista: Generalizable Multi-Task Agent with Diverse Visual Manipulation | Multi-Task Agent | [📑 Paper] |
| 26.02 | A Very Big Video Reasoning Suite (VBVR): 1M+ video clips across 200 reasoning tasks | Video Reasoning | [📑 Paper] [🤗 Model] [🤗 Data] |
| 26.02 | OmniGAIA: Omni-Modal AI Agent Benchmark with hindsight-guided exploration | Omni-Modal Agent Reasoning | [📑 Paper] [💻 Code] [🤗 Data] |
| 26.02 | SpatiaLab: Wild Spatial Reasoning benchmark across 6 VQA categories | Spatial Reasoning | [📑 Paper] [💻 Code] [🤗 Data] |
| 26.02 | MuRGAt: Multimodal Fact-Level Attribution benchmark for verifiable reasoning | Multimodal Attribution | [📑 Paper] [💻 Code] |
| 26.02 | DeepVision-103K: Verifiable multimodal math dataset for RLVR training | Math Reasoning | [📑 Paper] [💻 Code] [🤗 Data] |
| 26.02 | UniVBench: Unified evaluation for video foundation models across understanding, generation, editing | Video Foundation Model Evaluation | [📑 Paper] [💻 Code] |
| 26.02 | RISE-Video: Benchmark for video generators decoding implicit world rules | Video Generation Reasoning | [📑 Paper] [💻 Code] [🤗 Data] |
| 26.02 | SAW-Bench: Egocentric Situated Awareness evaluation with 786 smart-glass videos and 2,071+ QA pairs | Spatial Reasoning | [📑 Paper] |
| 26.02 | BrowseComp-V3: 300-question visual benchmark for complex multi-hop multimodal web search | Multimodal Browsing | [📑 Paper] |
| 26.02 | BiManiBench: Hierarchical benchmark for bimanual coordination evaluation in MLLMs | Bimanual Robotics | [📑 Paper] [💻 Code] |
| 26.01 | MMFineReason: Closing the Multimodal Reasoning Gap via Open Data-Centric Methods | Multimodal Reasoning | [📑 Paper] [🤗 Model] [🤗 Data] |
| 26.01 | ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis | Chart Reasoning | [📑 Paper] [💻 Code] [🤗 Model] [🤗 Data] |
| 26.01 | VideoLoom: Joint Spatial-Temporal Understanding with LoomBench | Spatial-Temporal Reasoning | [📑 Paper] [💻 Code] [🤗 Model] |
| 26.01 | PROGRESSLM: Towards Progress Reasoning in Vision-Language Models | Task Progress Reasoning | [📑 Paper] [💻 Code] [🤗 Data] |
| 26.01 | FutureOmni: Evaluating Future Forecasting from Omni-Modal Context | Omni-Modal Temporal Reasoning | [📑 Paper] |
| 26.01 | Afri-MCQA: Multimodal Cultural Question Answering for African Languages | Multilingual Multimodal Reasoning | [📑 Paper] |
| 26.01 | AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark | Cultural Multimodal Reasoning | [📑 Paper] |
| 25.12 | HERBench: Multi-Evidence Integration in Video Question Answering | Video Reasoning | [📑 Paper] |
| 25.12 | SVBench: Evaluation of Video Generation Models on Social Reasoning | Video Social Reasoning | [📑 Paper] |
| 25.12 | IF-Bench: Benchmarking MLLMs for Infrared Images | Infrared Image Understanding | [📑 Paper] |
| 25.12 | VABench: Comprehensive Benchmark for Audio-Video Generation | Audio-Video Generation | [📑 Paper] |
| 25.11 | MME-CC: Challenging Multi-Modal Evaluation Benchmark of Cognitive Capacity | Cognitive Capacity | [📑 Paper] |
| 25.11 | GGBench: Geometric Generative Reasoning Benchmark for Unified Multimodal Models | Geometric Reasoning | [📑 Paper] |
| 25.11 | WEAVE: Benchmarking In-context Interleaved Comprehension and Generation | Multimodal Comprehension & Generation | [📑 Paper] |
| 25.10 | Uni-MMMU: Massive Multi-discipline Multimodal Unified Benchmark | Multimodal Multi-discipline Reasoning | [📑 Paper] |
| 25.10 | PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs | Physical Tool Understanding | [📑 Paper] |
| 25.10 | BEAR: Benchmarking Multimodal Language Models for Atomic Embodied Capabilities | Embodied AI Capabilities | [📑 Paper] |
| 25.10 | OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs | Long-context, Video-Audio Unerstanding & Reasonin | [📑 Paper] [💻 Code] [🌐 Project] [🤗 Data] |
| 25.10 | XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models | Capability Balancing among Different Modalities | [📑 Paper] [💻 Code] [🌐 Project] |
| 25.10 | StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA | Termporal Reasoning | [📑 Paper] |
| 25.10 | Valor32k-AVQA v2.0: Open-Ended Audio-Visual Question Answering Dataset and Benchmark | Common Sense Omni Reasoning | [📑 Paper] |
| 25.09 | MARS2 2025 Challenge on Multimodal Reasoning | Multimodal Reasoning Challenge | [📑 Paper] |
| 25.09 | Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images | Table Reasoning | [📑 Paper] |
| 25.09 | AHELM: A Holistic Evaluation of Audio-Language Models | Audio-Language Understanding | [📑 Paper] |
| 25.09 | MDAR: A Multi-scene Dynamic Audio Reasoning Benchmark | Complex, Multi-scene, & Dynamically Evolving Speech & Audio Reasonin | [📑 Paper] [💻 Code] |
| 25.09 | MiMo-Audio-Eval Toolkit | Speech/Sound/Music Reasoning | [💻 Code] |
| 25.08 | SpeechR: A Benchmark for Speech Reasoning in Large Audio-Language Models | Speech Reasoning | [📑 Paper] [💻 Code] [Data] |
| 25.08 | MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence | Long-form, Spatial, and Multi-audio Reasoning on Speech/Music/Sound | [📑 Paper] [🤗 Data] |
| 25.08 | R²-AVSBench: Think Before You Segment: An Object-aware Reasoning Agent for Referring Audio-Visual Segmentation | Segmentation Reasoning | [📑 Paper] [🤗 Data] |
| 25.07 | Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding | Video Reasoning and Understanding | [📑 Paper]. [🌐 Project] [🤗 Data] |
| 25.06 | FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation | Financial Multi-Modal Reasoning Reasoning | [📑 Paper]. [💻 Code]. [🤗 Data] |
| 25.06 | MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos | Video Reasoning | [📑 Paper]. [💻 Code]. [🌐 Project] [🤗 Data] |
| 25.06 | OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models | Spatial Reasoning | [📑 Paper]. [💻 Code]. [🌐 Project] [🤗 Data] |
| 25.06 | MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark | Phonatics, Prosody, Rhetoric, Syntactics, Semantics, and Paralinguistics in Speech Understanding & Reasoning | [📑 Paper] [💻 Code] [🤗 Data] |
| 25.05 | Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities | Video&Audio Reasoning | [📑 Paper] [💻 Code] [🌐 Project] [🤗 Data] |
| 25.05 | MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix | Multi-step Audio Reasoning | [📑 Paper]. [💻 Code]. [🎥 demo] [🤗 Data] |
| 25.05 | On Path to Multimodal Generalist: General-Level and General-Bench | Multimodal Generation | [🌐 Project] [📑 Paper] [🤗 Data] |
| 25.04 | VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models | Visual Reasoning | [🌐 Project] [📑 Paper] [💻 Code] [🤗 Data] |
| 25.04 | IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs | Image-Grounded Video Perception and Reasoning | [📑 Paper] [💻 Code] |
| 25.04 | Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing | Reasoning-Informed viSual Editing | [📑 Paper] [💻 Code] |
| 25.04 | CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following | Music Information Retrieval & Knowledge | [📑 Paper] [💻 Code] |
| 25.03 | MAVERIX: Multimodal Audio-Visual Evaluation Reasoning IndeX | Common Sense Omni Reasoning | [📑 Paper] [🌐 Project] |
| 25.03 | V-STaR : Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning | Spatio-temporal Reasoning | [🌐 Project] [📑 Paper] [🤗 Data] |
| 25.03 | MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs | Spatio-temporal Understanding | [📑Paper] |
| 25.03 | Integrating Chain-of-Thought for Multimodal Alignment: A Study on 3D Vision-Language Learning | 3D-CoT | [📑 Paper] [🤗 Data] |
| 25.02 | MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models | MM-IQ | [📑 Paper] [💻 Code] |
| 25.02 | MM-RLHF: The Next Step Forward in Multimodal LLM Alignment | MM-RLHF-RewardBench, MM-RLHF-SafetyBench | [📑 Paper] |
| 25.02 | ZeroBench: An Impossible* Visual Benchmark for Contemporary Large Multimodal Models | ZeroBench | [🌐 Project] [🤗 Dataset] [💻 Code] |
| 25.02 | MME-CoT: Benchmarking Chain-of-Thought in LMMs for Reasoning Quality, Robustness, and Efficiency | MME-CoT | [📑 Paper] [💻 Code] |
| 25.02 | OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference | MM-AlignBench | [📑 Paper] [💻 Code] |
| 25.01 | AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs | Adversarial attack, Compositional reasoning, and Modality-specific dependency in Visual&Audio | [📑 Paper] |
| 25.01 | LlamaV-o1: Rethinking Step-By-Step Visual Reasoning in LLM |
Truncated — view the full README on GitHub.