mll-lab-nu/Awesome-Spatial-Intelligence-in-VLM

A paper list for spatial reasoning

786

83 commits

updated Aug 23, 2026

See the code

README

Awesome Spatial Intelligence in VLMs

This carefully curated list brings together key methods, datasets, and benchmarks in the field of spatial intelligence for VLMs.

With the development of multimodal models, evaluating and enhancing their spatial intelligence has become a key research frontier. This list aims to provide researchers and engineers with a quick index to track the latest advancements in the field.

We welcome contributions of excellent resources you find via Pull Request!

Table of Contents

Methods

Visual-based methods

TitleIntroductionDateCode

Planning with the Views
image2026-05Github

RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
image2025-12Github

N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models
image2025-12Github

SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery
image2025-12Github

Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
image2025-12Github

CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning
image2025-12-

S2-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
image2025-12-

EagleVision: A Dual-Stage Framework with BEV-grounding-based Chain-of-Thought for Spatial Intelligence
image2025-12-

Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation
image2025-12-

Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning
image2025-11Github

G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
image2025-11Github

Video Spatial Reasoning with Object-Centric 3D Rollout
image2025-11-

Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
image2025-11-

SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models
image2025-11-

Cambrian-S: Towards Spatial Supersensing in Video
image2025-11Github

Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
image2025-11Github

SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards
image2025-11Github

TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
image2025-10Github

Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
image2025-10Github

Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
image2025-10Github

Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
image2025-10Github

Euclid’s Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks
image2025-10Github

SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
image2025-10Github

SpaceVista: All-Scale Visual Spatial Reasoning from mm to km
image2025-10Github
Publish
SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
image2025-09Github
Publish
See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
image2025-09-

3D Aware Region Prompted Vision Language Model
image2025-09Github

UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
image2025-08Github

SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
image2025-08Github

Enhancing Spatial Reasoning through Visual and Textual Thinking
image2025-07-
Publish
MindJourney: Test-Time Scaling with World Models for Spatial Reasoning
image2025-07Github
Publish
Struct2D: A Perception-Guided Framework for Spatial Reasoning in Large Multimodal Models
image2025-06Github
Publish
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
image2025-06Github
Publish
SpatialLM: Training Large Language Models for Structured Indoor Modeling
image2025-06Github

Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
image2025-06Github

Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
image2025-06Github

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
image2025-05Github
Publish
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
image2025-05Github
Publish
Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
image2025-05Github
Publish
SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
image2025-05Github
Publish
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
image2025-05Github

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs
image2025-05Github

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
image2025-05Github

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
image2025-05-
Publish
SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
image2025-05Github
Publish
Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation
image2025-04Github

Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipe
image2025-04-

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
image2025-04Github

Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning
image2025-04Github
Publish
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
image2025-04Github
Publish
ROSS3D: Reconstructive Visual Instruction Tuning with 3D-Awareness
image2025-04Github
Publish
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data
image2025-04-

Visual Agentic AI for Spatial Reasoning with a Dynamic API
image2025-02Github
Publish
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding
image2025-01Github

SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning
image2025-01Github
Publish
LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Model
image2024-12Github

COARSE CORRESPONDENCES Boost Spatial-Temporal Reasoning in Multimodal Language Model
image2024-08Github
Star Publish
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
image2024-06GitHub
Star Publish
SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
image2024-06Github
Star Publish
SpatialBot: Precise Spatial Understanding with Vision Language Models
image2024-06Github
Publish
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs
image2024-04Github

SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
image2024-03-
Publish
Can Transformers Capture Spatial Relations between Objects?
image2024-03Github
Star Publish
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
image2024-01Github

Proximity QA: Unleashing the Power of Multi-Modal Large Language Models for Spatial Proximity Analysis
image2024-01Github

3DAxiesPrompts: Unleashing the 3D Spatial Task Capabilities of GPT-4V
image2023-12-

Text-based methods

Datasets & Benchmarks

Visual-based data

TitleIntroductionDateCode

Planning with the Views
image2026-05Github
Publish
Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
image2026-02Github

EscherVerse: An Open World Benchmark and Dataset for Teleo-Spatial Intelligence with Physical-Dynamic and Intent-Driven Understanding
image2026-1Github

RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
image2025-12Github

Towards Cross-View Point Correspondence in Vision-Language Models
image2025-12Github
Publish
ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints
image2025-11-

Scaling Spatial Intelligence with Multimodal Foundation Models
image2025-11Github

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
image2025-11Github

Visual Spatial Tuning
image2025-11Github

Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
image2025-10Github

DSI-Bench: A Benchmark for Dynamic Spatial Intelligence
image2025-10Github

Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
image2025-10-

NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions
image2025-10Github

SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
image2025-09Github

Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
image2025-09Github

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture
image2025-09Github

VisualTrans: A Benchmark for Real-World Visual Transformation Reasoning
image2025-08Github

SpatialVID: A Large-Scale Video Dataset with Spatial Annotations
image2025-09Github
Publish
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
image2025-08Github

11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis
image2025-08-
Publish
Towards Scalable Spatial Intelligence via 2D-to-3D Data Lifting
image2025-07Github

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models
image2025-07-

SpatialViz-Bench: An MLLM Benchmark for Spatial Visualization
image2025-07Github

Spatial Mental Modeling from Limited Views
image2025-06Github

SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
image2025-06-
Publish
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
image2025-06Github
Publish
From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes
image2025-06Github
Publish
PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly
image2025-06Github

Can Vision Language Models Infer Human Gaze Direction? A Controlled Study
image2025-06Github

SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence
image2025-06Github

Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
image2025-06Github

OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
image2025-06Github

InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models
image2025-06-

MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
image2025-05Github
Publish
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
image2025-05Github

Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
image2025-05Github

SpatialScore: Towards Unified Evaluation for Multimodal Spatial Understanding
image2025-05Github

MIRAGE:A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
image2025-05Github

Can Multimodal Large Language Models Understand Spatial Relations
image2025-05Github

Visuospatial Cognitive Assistant
image2025-05Github

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?
image2025-05Github

Vision language models have difficulty recognizing virtual objects
image2025-05-

ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
image2025-05Github

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames
image2025-05-
Publish
SITE: towards Spatial Intelligence Thorough Evaluation
image2025-05Github

CameraBench: Towards Understanding Camera Motions in Any Video
image2025-04Github

Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
image2025-04Github

From Flatland to Space:Teaching Vision-Language Models to Perceive and Reason in 3D
image2025-03Github

MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLM
image2025-03-

Open3DVQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
image2025-03Github
Publish
STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
image2025-03Github
Publish
CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language Models
image2025-03Github

Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
image2025-03Github

LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?
image2025-03Github
Publish
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
image2025-02Github

FoREST: Frame of Reference Evaluation in Spatial Reasoning Tasks
image2025-02-

iVISPAR — An Interactive Visual-Spatial Reasoning Benchmark for VLMs
image2025-02Github

Defining and Evaluating Visual Language Models' Basic Spatial Abilities: A Perspective from Psychometrics
image2025-02-
Publish
SAT: Spatial Aptitude Training for Multimodal Language Models
image2024-12Github
Publish
SPHERE: A Hierarchical Evaluation on Spatial Perception and Reasoning for Vision-Language Models
image2024-12Github
Publish
3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
image2024-12Github
StarPublish
Thinking in Space: How Multimodal Large Language Models See, Remember and Recall Spaces
image2024-12Github
Publish
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
image2024-11Github
Publish
An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models
image2024-11-
Publish
IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos
image2024-11Github

Is ‘Right’ Right? Enhancing Object Orientation Understanding in Multimodal Language Models through Egocentric Instruction Tuning
image2024-10Github
Publish
ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models
image2024-10Github
Publish
DOES SPATIAL COGNITION EMERGE IN FRONTIER MODELS?
image2024-10-
Publish
Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities
image2024-10Github

R2D3: Imparting Spatial Reasoning by Reconstructing 3D Scenes from 2D Images
image2024-10Github
Publish
Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models
image2024-09Github
Publish
Can Vision Language Models Learn from Visual Demonstrations of Ambiguous Spatial Reasoning?
image2024-09-

VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs
image2024-07Github
Publish
EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models
image2024-06Github
Publish
TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
image2024-06Github
Publish
Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models
image2024-06Github
Publish
GSR-Bench: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs
image2024-06-
Publish
Reframing Spatial Reasoning Evaluation in Language Models: A Real-World Simulation Benchmark for Qualitative Reasoning
image2024-05Github
Publish
Visually Descriptive Language Model for Vector Graphics Reasoning
image2024-04-
PublishStar
SQA3D: Situated Question Answering in 3D Scenes
image2022-10Github
PublishStar
Things not Written in Text: Exploring Spatial Commonsense from Visual Signals
image2022-03Github
Publish
SPARE3D: A Dataset for SPAtial REasoning on Three-View Line Drawings
image2020-03Github

Text-based data

Findings

Applications

awesome-list
spatial-intelligence
spatial-reasoning
vision-language-models

Contributors

yyyybq

52 commits

Grady10086

5 commits

Zhoues

3 commits

jeasinema

2 commits

mll-lab-nu/Awesome-Spatial-Intelligence-in-VLM

A paper list for spatial reasoning

786

83 commits

updated Aug 23, 2026

See the code

README

Awesome Spatial Intelligence in VLMs

This carefully curated list brings together key methods, datasets, and benchmarks in the field of spatial intelligence for VLMs.

With the development of multimodal models, evaluating and enhancing their spatial intelligence has become a key research frontier. This list aims to provide researchers and engineers with a quick index to track the latest advancements in the field.

We welcome contributions of excellent resources you find via Pull Request!

Table of Contents

Methods

Visual-based methods

TitleIntroductionDateCode

Planning with the Views
image2026-05Github

RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
image2025-12Github

N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models
image2025-12Github

SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery
image2025-12Github

Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
image2025-12Github

CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning
image2025-12-

S2-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
image2025-12-

EagleVision: A Dual-Stage Framework with BEV-grounding-based Chain-of-Thought for Spatial Intelligence
image2025-12-

Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation
image2025-12-

Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning
image2025-11Github

G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
image2025-11Github

Video Spatial Reasoning with Object-Centric 3D Rollout
image2025-11-

Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
image2025-11-

SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models
image2025-11-

Cambrian-S: Towards Spatial Supersensing in Video
image2025-11Github

Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
image2025-11Github

SpatialThinker: Reinforcing 3D Reasoning in Multimodal LLMs via Spatial Rewards
image2025-11Github

TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
image2025-10Github

Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
image2025-10Github

Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
image2025-10Github

Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
image2025-10Github

Euclid’s Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks
image2025-10Github

SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
image2025-10Github

SpaceVista: All-Scale Visual Spatial Reasoning from mm to km
image2025-10Github
Publish
SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
image2025-09Github
Publish
See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
image2025-09-

3D Aware Region Prompted Vision Language Model
image2025-09Github

UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
image2025-08Github

SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
image2025-08Github

Enhancing Spatial Reasoning through Visual and Textual Thinking
image2025-07-
Publish
MindJourney: Test-Time Scaling with World Models for Spatial Reasoning
image2025-07Github
Publish
Struct2D: A Perception-Guided Framework for Spatial Reasoning in Large Multimodal Models
image2025-06Github
Publish
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
image2025-06Github
Publish
SpatialLM: Training Large Language Models for Structured Indoor Modeling
image2025-06Github

Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
image2025-06Github

Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
image2025-06Github

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
image2025-05Github
Publish
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
image2025-05Github
Publish
Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
image2025-05Github
Publish
SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
image2025-05Github
Publish
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
image2025-05Github

STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs
image2025-05Github

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
image2025-05Github

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
image2025-05-
Publish
SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
image2025-05Github
Publish
Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation
image2025-04Github

Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipe
image2025-04-

SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
image2025-04Github

Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning
image2025-04Github
Publish
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
image2025-04Github
Publish
ROSS3D: Reconstructive Visual Instruction Tuning with 3D-Awareness
image2025-04Github
Publish
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data
image2025-04-

Visual Agentic AI for Spatial Reasoning with a Dynamic API
image2025-02Github
Publish
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding
image2025-01Github

SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning
image2025-01Github
Publish
LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Model
image2024-12Github

COARSE CORRESPONDENCES Boost Spatial-Temporal Reasoning in Multimodal Language Model
image2024-08Github
Star Publish
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
image2024-06GitHub
Star Publish
SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
image2024-06Github
Star Publish
SpatialBot: Precise Spatial Understanding with Vision Language Models
image2024-06Github
Publish
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs
image2024-04Github

SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors
image2024-03-
Publish
Can Transformers Capture Spatial Relations between Objects?
image2024-03Github
Star Publish
SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
image2024-01Github

Proximity QA: Unleashing the Power of Multi-Modal Large Language Models for Spatial Proximity Analysis
image2024-01Github

3DAxiesPrompts: Unleashing the 3D Spatial Task Capabilities of GPT-4V
image2023-12-

Text-based methods

Datasets & Benchmarks

Visual-based data

TitleIntroductionDateCode

Planning with the Views
image2026-05Github
Publish
Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration?
image2026-02Github

EscherVerse: An Open World Benchmark and Dataset for Teleo-Spatial Intelligence with Physical-Dynamic and Intent-Driven Understanding
image2026-1Github

RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
image2025-12Github

Towards Cross-View Point Correspondence in Vision-Language Models
image2025-12Github
Publish
ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints
image2025-11-

Scaling Spatial Intelligence with Multimodal Foundation Models
image2025-11Github

SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
image2025-11Github

Visual Spatial Tuning
image2025-11Github

Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
image2025-10Github

DSI-Bench: A Benchmark for Dynamic Spatial Intelligence
image2025-10Github

Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
image2025-10-

NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions
image2025-10Github

SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
image2025-09Github

Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
image2025-09Github

Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture
image2025-09Github

VisualTrans: A Benchmark for Real-World Visual Transformation Reasoning
image2025-08Github

SpatialVID: A Large-Scale Video Dataset with Spatial Annotations
image2025-09Github
Publish
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
image2025-08Github

11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis
image2025-08-
Publish
Towards Scalable Spatial Intelligence via 2D-to-3D Data Lifting
image2025-07Github

Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models
image2025-07-

SpatialViz-Bench: An MLLM Benchmark for Spatial Visualization
image2025-07Github

Spatial Mental Modeling from Limited Views
image2025-06Github

SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
image2025-06-
Publish
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
image2025-06Github
Publish
From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes
image2025-06Github
Publish
PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly
image2025-06Github

Can Vision Language Models Infer Human Gaze Direction? A Controlled Study
image2025-06Github

SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence
image2025-06Github

Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
image2025-06Github

OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
image2025-06Github

InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models
image2025-06-

MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
image2025-05Github
Publish
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
image2025-05Github

Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
image2025-05Github

SpatialScore: Towards Unified Evaluation for Multimodal Spatial Understanding
image2025-05Github

MIRAGE:A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
image2025-05Github

Can Multimodal Large Language Models Understand Spatial Relations
image2025-05Github

Visuospatial Cognitive Assistant
image2025-05Github

Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?
image2025-05Github

Vision language models have difficulty recognizing virtual objects
image2025-05-

ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
image2025-05Github

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames
image2025-05-
Publish
SITE: towards Spatial Intelligence Thorough Evaluation
image2025-05Github

CameraBench: Towards Understanding Camera Motions in Any Video
image2025-04Github

Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
image2025-04Github

From Flatland to Space:Teaching Vision-Language Models to Perceive and Reason in 3D
image2025-03Github

MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLM
image2025-03-

Open3DVQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
image2025-03Github
Publish
STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
image2025-03Github
Publish
CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language Models
image2025-03Github

Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
image2025-03Github

LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?
image2025-03Github
Publish
Spatial457: A Diagnostic Benchmark for 6D Spatial Reasoning of Large Multimodal Models
image2025-02Github

FoREST: Frame of Reference Evaluation in Spatial Reasoning Tasks
image2025-02-

iVISPAR — An Interactive Visual-Spatial Reasoning Benchmark for VLMs
image2025-02Github

Defining and Evaluating Visual Language Models' Basic Spatial Abilities: A Perspective from Psychometrics
image2025-02-
Publish
SAT: Spatial Aptitude Training for Multimodal Language Models
image2024-12Github
Publish
SPHERE: A Hierarchical Evaluation on Spatial Perception and Reasoning for Vision-Language Models
image2024-12Github
Publish
3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
image2024-12Github
StarPublish
Thinking in Space: How Multimodal Large Language Models See, Remember and Recall Spaces
image2024-12Github
Publish
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
image2024-11Github
Publish
An Empirical Analysis on Spatial Reasoning Capabilities of Large Multimodal Models
image2024-11-
Publish
IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos
image2024-11Github

Is ‘Right’ Right? Enhancing Object Orientation Understanding in Multimodal Language Models through Egocentric Instruction Tuning
image2024-10Github
Publish
ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models
image2024-10Github
Publish
DOES SPATIAL COGNITION EMERGE IN FRONTIER MODELS?
image2024-10-
Publish
Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference Under Ambiguities
image2024-10Github

R2D3: Imparting Spatial Reasoning by Reconstructing 3D Scenes from 2D Images
image2024-10Github
Publish
Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models
image2024-09Github
Publish
Can Vision Language Models Learn from Visual Demonstrations of Ambiguous Spatial Reasoning?
image2024-09-

VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs
image2024-07Github
Publish
EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models
image2024-06Github
Publish
TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
image2024-06Github
Publish
Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models
image2024-06Github
Publish
GSR-Bench: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs
image2024-06-
Publish
Reframing Spatial Reasoning Evaluation in Language Models: A Real-World Simulation Benchmark for Qualitative Reasoning
image2024-05Github
Publish
Visually Descriptive Language Model for Vector Graphics Reasoning
image2024-04-
PublishStar
SQA3D: Situated Question Answering in 3D Scenes
image2022-10Github
PublishStar
Things not Written in Text: Exploring Spatial Commonsense from Visual Signals
image2022-03Github
Publish
SPARE3D: A Dataset for SPAtial REasoning on Three-View Line Drawings
image2020-03Github

Text-based data

Findings

Applications

awesome-list
spatial-intelligence
spatial-reasoning
vision-language-models

Contributors

yyyybq

52 commits

Grady10086

5 commits

Zhoues

3 commits

jeasinema

2 commits