arijitray1993/awesome-spatial-reasoning

Collection of the latest spatial, 3D, and video/temporal reasoning papers

37

15 commits

updated Sep 25, 2026

See the code

README

Collection of the latest spatial, 3D, and video reasoning papers

Datasets

NameModalitiesDescriptionTrainTestHF
SpatialGen-Bench (ProVisE)Image, TextProtocol-constrained visual-answer evaluation of spatial cognition for image-generation models and VLMs; 470 samples across 14 spatial subtasksNoYesCode
Zebra-CoTMulti-Image, TextInterleaved Vision Language ReasoningYes-Link
Video-R1Video, TextVideo and text-based reasoningYes-Code
VSI-BenchVideo, TextVideo walk through of apartment, questions about spatial orientations and planningNoYesNo
CV-BenchImage, TextSpatial relationship QA on imagesNoYeslink
SATMulti-image, TextComplex spatial QAs that require reasoning about action/motion causality on synthetic imagesYesYeslink
BLINKMulti-image, TextComplex visual QA with spatial perception splitsNoYesNo
ProVisionImage, TextPseudo-annotated spatial relationship QAs on imagesYesYesNo
SpatialRGPTImage, TextPseudo-annotated spatial relationship QAs on imagesYesYesNo
RoboPointImage, Textspatial affordance predictionYesYeslink
RoboSpatialImage, Textspatial task QAsYesYesNo
PhysBenchMulti-image, Textphysical object properties, relationships, physics-driven dynamicsYesYesNo
Cambrian-10MImage, TextMLM pretraining QA on imagesYesNolink
PixMOImage, TextMLM pretraining QA on imagesYesYeslink
MegaBenchImage, TextMLM benchmark with splits on spatial reasoningNoYesNo
MultiSpatialLLMImage/multi-image, TextDynamic spatial QAsYesYes-
PEVideoVideo, Textdense action video annotationsYesYes-
SPaRCText2D pathfinding dataset, requiring multi-step spatial and rule-based reasoningYesYes-

Topics

Vision-langauge Models

Image/video question-answering

Robotics and action

Reasoning, chain-of-thought, RL

Image-text representations


Multi-view 2D to 3D


Spatial Scene Generation


Contributors

arijitray1993

11 commits

lkaesberg

2 commits

zwq2018

2 commits

arijitray1993/awesome-spatial-reasoning

Collection of the latest spatial, 3D, and video/temporal reasoning papers

37

15 commits

updated Sep 25, 2026

See the code

README

Collection of the latest spatial, 3D, and video reasoning papers

Datasets

NameModalitiesDescriptionTrainTestHF
SpatialGen-Bench (ProVisE)Image, TextProtocol-constrained visual-answer evaluation of spatial cognition for image-generation models and VLMs; 470 samples across 14 spatial subtasksNoYesCode
Zebra-CoTMulti-Image, TextInterleaved Vision Language ReasoningYes-Link
Video-R1Video, TextVideo and text-based reasoningYes-Code
VSI-BenchVideo, TextVideo walk through of apartment, questions about spatial orientations and planningNoYesNo
CV-BenchImage, TextSpatial relationship QA on imagesNoYeslink
SATMulti-image, TextComplex spatial QAs that require reasoning about action/motion causality on synthetic imagesYesYeslink
BLINKMulti-image, TextComplex visual QA with spatial perception splitsNoYesNo
ProVisionImage, TextPseudo-annotated spatial relationship QAs on imagesYesYesNo
SpatialRGPTImage, TextPseudo-annotated spatial relationship QAs on imagesYesYesNo
RoboPointImage, Textspatial affordance predictionYesYeslink
RoboSpatialImage, Textspatial task QAsYesYesNo
PhysBenchMulti-image, Textphysical object properties, relationships, physics-driven dynamicsYesYesNo
Cambrian-10MImage, TextMLM pretraining QA on imagesYesNolink
PixMOImage, TextMLM pretraining QA on imagesYesYeslink
MegaBenchImage, TextMLM benchmark with splits on spatial reasoningNoYesNo
MultiSpatialLLMImage/multi-image, TextDynamic spatial QAsYesYes-
PEVideoVideo, Textdense action video annotationsYesYes-
SPaRCText2D pathfinding dataset, requiring multi-step spatial and rule-based reasoningYesYes-

Topics

Vision-langauge Models

Image/video question-answering

Robotics and action

Reasoning, chain-of-thought, RL

Image-text representations


Multi-view 2D to 3D


Spatial Scene Generation


Contributors

arijitray1993

11 commits

lkaesberg

2 commits

zwq2018

2 commits