Collection of the latest spatial, 3D, and video/temporal reasoning papers
37
15 commits
updated Sep 25, 2026
| Name | Modalities | Description | Train | Test | HF |
|---|---|---|---|---|---|
| SpatialGen-Bench (ProVisE) | Image, Text | Protocol-constrained visual-answer evaluation of spatial cognition for image-generation models and VLMs; 470 samples across 14 spatial subtasks | No | Yes | Code |
| Zebra-CoT | Multi-Image, Text | Interleaved Vision Language Reasoning | Yes | - | Link |
| Video-R1 | Video, Text | Video and text-based reasoning | Yes | - | Code |
| VSI-Bench | Video, Text | Video walk through of apartment, questions about spatial orientations and planning | No | Yes | No |
| CV-Bench | Image, Text | Spatial relationship QA on images | No | Yes | link |
| SAT | Multi-image, Text | Complex spatial QAs that require reasoning about action/motion causality on synthetic images | Yes | Yes | link |
| BLINK | Multi-image, Text | Complex visual QA with spatial perception splits | No | Yes | No |
| ProVision | Image, Text | Pseudo-annotated spatial relationship QAs on images | Yes | Yes | No |
| SpatialRGPT | Image, Text | Pseudo-annotated spatial relationship QAs on images | Yes | Yes | No |
| RoboPoint | Image, Text | spatial affordance prediction | Yes | Yes | link |
| RoboSpatial | Image, Text | spatial task QAs | Yes | Yes | No |
| PhysBench | Multi-image, Text | physical object properties, relationships, physics-driven dynamics | Yes | Yes | No |
| Cambrian-10M | Image, Text | MLM pretraining QA on images | Yes | No | link |
| PixMO | Image, Text | MLM pretraining QA on images | Yes | Yes | link |
| MegaBench | Image, Text | MLM benchmark with splits on spatial reasoning | No | Yes | No |
| MultiSpatialLLM | Image/multi-image, Text | Dynamic spatial QAs | Yes | Yes | - |
| PEVideo | Video, Text | dense action video annotations | Yes | Yes | - |
| SPaRC | Text | 2D pathfinding dataset, requiring multi-step spatial and rule-based reasoning | Yes | Yes | - |
| Paper | Venue/Date | Code |
|---|---|---|
| Perception Encoder | 2025 | |
| TIPS: Text-Image Pretraining with Spatial Awareness | ICLR 2025 |
Collection of the latest spatial, 3D, and video/temporal reasoning papers
37
15 commits
updated Sep 25, 2026
| Name | Modalities | Description | Train | Test | HF |
|---|---|---|---|---|---|
| SpatialGen-Bench (ProVisE) | Image, Text | Protocol-constrained visual-answer evaluation of spatial cognition for image-generation models and VLMs; 470 samples across 14 spatial subtasks | No | Yes | Code |
| Zebra-CoT | Multi-Image, Text | Interleaved Vision Language Reasoning | Yes | - | Link |
| Video-R1 | Video, Text | Video and text-based reasoning | Yes | - | Code |
| VSI-Bench | Video, Text | Video walk through of apartment, questions about spatial orientations and planning | No | Yes | No |
| CV-Bench | Image, Text | Spatial relationship QA on images | No | Yes | link |
| SAT | Multi-image, Text | Complex spatial QAs that require reasoning about action/motion causality on synthetic images | Yes | Yes | link |
| BLINK | Multi-image, Text | Complex visual QA with spatial perception splits | No | Yes | No |
| ProVision | Image, Text | Pseudo-annotated spatial relationship QAs on images | Yes | Yes | No |
| SpatialRGPT | Image, Text | Pseudo-annotated spatial relationship QAs on images | Yes | Yes | No |
| RoboPoint | Image, Text | spatial affordance prediction | Yes | Yes | link |
| RoboSpatial | Image, Text | spatial task QAs | Yes | Yes | No |
| PhysBench | Multi-image, Text | physical object properties, relationships, physics-driven dynamics | Yes | Yes | No |
| Cambrian-10M | Image, Text | MLM pretraining QA on images | Yes | No | link |
| PixMO | Image, Text | MLM pretraining QA on images | Yes | Yes | link |
| MegaBench | Image, Text | MLM benchmark with splits on spatial reasoning | No | Yes | No |
| MultiSpatialLLM | Image/multi-image, Text | Dynamic spatial QAs | Yes | Yes | - |
| PEVideo | Video, Text | dense action video annotations | Yes | Yes | - |
| SPaRC | Text | 2D pathfinding dataset, requiring multi-step spatial and rule-based reasoning | Yes | Yes | - |
| Paper | Venue/Date | Code |
|---|---|---|
| Perception Encoder | 2025 | |
| TIPS: Text-Image Pretraining with Spatial Awareness | ICLR 2025 |