Klingsor-tyx/Awesome-Spatial-Cognitive-Map

A curated list of papers analyzed in the survey: Spatial Intelligence from a Cognitive Map Perspective: A Survey

24

53 commits

updated Jun 1, 2026

See the code

README

Spatial Intelligence from a Cognitive Map Perspective:
A Survey

Yuxuan Tian* 1,2, Yuheng Ji*†1,2, Xiaolong Zheng*✉1,2, Ziheng Qin1,2, Yipu Wang1,2, Xinyi Zheng3,
Yuyang Liu1,2, Shuanghao Bai4, Zhe Li5, Liang Wang1,2, Daniel Dajun Zeng1,2

1CASIA 2UCAS 3BUAA 4XJTU 5NTU
* Equal Contribution    † Project Leader    ✉ Corresponding Author

🌐 Project Page    |    📄 Paper PDF    |    ⭐ Awesome List


📚 Overview

Overview of spatial intelligence from a cognitive map perspective.

This survey revisits Spatial Intelligence from the perspective of Cognitive Maps, treating cognitive maps as the representational blueprint that connects spatial perception, reasoning, and generation. Under this view, diverse research directions can be understood through a shared question: how an internal spatial representation is constructed, maintained, reasoned over, and realized.

We organize the literature into three cognitive-map-centric processes:

  • Perception: Construction of Cognitive Maps — how agents build internal spatial representations from local, partial, and multimodal observations.
  • Reasoning: Inference with Cognitive Maps — how cognitive maps are used as embeddings, prompts, or APIs to support spatial inference and decision-making.
  • Generation: Realization of Cognitive Maps — how internal maps are externalized into 3D scenes and dynamic world simulations.

This repository provides a curated list of papers analyzed in the survey, following the taxonomy proposed in Spatial Intelligence from a Cognitive Map Perspective: A Survey.

📖 Table of Contents


1. Perception: Construction of Cognitive Maps

1.1 Metric Representation

1.1.1 Explicit Geometry-based

2D Planar-based
TitleYearVenue
History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation2026AAAI
What You See is What You Reach: Towards Spatial Navigation with High-Level Human Instructions2026AAAI
One Map to Find Them All: Real-time Open-Vocabulary Mapping for Zero-shot Multi-Object Navigation2025ICRA
3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning2025CVPR
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models2025arXiv
MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation2025ACL
Ali-UI: Enhancing Complex Vision-Language Navigation with Alignment of Unified Map and Instruction Parsing2025MM
OVL-MAP: An Online Visual Language Map Approach for Vision-and-Language Navigation in Continuous Environments2025RAL
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning2025arXiv
Human-like Navigation in a World Built for Humans2025CoRL
Enhancing Embodied Object Detection with Spatial Feature Memory2025WACV
TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making2025NeurIPS
Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal Navigation2025NeurIPS
Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models2024NeurIPS
Explore until Confident: Efficient Exploration for Embodied Question Answering2024arXiv
GridMM: Grid Memory Map for Vision-and-Language Navigation2023ICCV
Visual Language Maps for Robot Navigation2023ICRA
Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied Agents2023ICCV
Open-vocabulary Queryable Scene Representations for Real World Planning2023ICRA
Cross-modal Map Learning for Vision and Language Navigation2022CVPR
Weakly-Supervised Multi-Granularity Map Learning for Vision-and-Language Navigation2022NeurIPS
Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric Views2021AAAI
Cognitive Mapping and Planning for Visual Navigation2017CVPR
3D Geometry-based

1.1.2 Parametric Coordinate-based

1.2 Relational Representation

1.2.1 Structured Graph-based

TitleYearVenue
Integrated Exploration and Sequential Manipulation on Scene Graph with LLM-based Situated Replanning2026ICRA
VPN: Visual Prompt Navigation2026AAAI
RAG-3DSG: Enhancing 3D Scene Graphs with Re-Shot Guided Retrieval-Augmented Generation2026arXiv
BrainyMP: Enhancing Motion Planning Using Graph Neural Network Inspired by Brain Spatial Relational Memory2025TITS
GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation2025CoRL
HiGS: Hierarchical Generative Scene Framework for Multi-Step Associative Semantic Spatial Composition2025arXiv
Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation2025NeurIPS
Planner3D: LLM-enhanced graph prior meets 3D indoor scene explicit regularization2025TPAMI
Visuomotor Navigation for Embodied Robots With Spatial Memory and Semantic Reasoning Cognition2025TNNLS
TB-HSU: Hierarchical 3D Scene Understanding with Contextual Affordances2025AAAI
InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior2024ICLR
MemoNav: Working Memory Model for Visual Navigation2024CVPR
Multiview Scene Graph2024NeurIPS
Scene-Driven Multimodal Knowledge Graph Construction for Embodied AI2024TKDE
3D Question Answering for City Scene Understanding2024MM
Bridging Visual and Textual Semantics: Towards Consistency for Unbiased Scene Graph Generation2024TPAMI
X-RefSeg3D: Enhancing Referring 3D Instance Segmentation via Structured Cross-Modal Graph Neural Networks2024AAAI
TopoNav: Topological Navigation for Efficient Exploration in Sparse Reward Environments2024IROS
CommonScenes: Generating Commonsense 3D Indoor Scenes with Scene Graph Diffusion2023NeurIPS
GTR: A Grafting-Then-Reassembling Framework for Dynamic Scene Graph Generation2023IJCAI
Topological Semantic Graph Memory for Image-Goal Navigation2023CoRL
Modeling Dynamic Environments with Scene Graph Memory2023ICML
Scene Graph Contrastive Learning for Embodied Navigation2023ICCV
Visual Graph Memory with Unsupervised Representation for Visual Navigation2021ICCV
Graph-to-3D: End-to-End Generation and Manipulation of 3D Scenes Using Scene Graphs2021ICCV
End-to-End Optimization of Scene Layout2020CVPR
PlanIT: Planning and Instantiating Indoor Scenes with Relation Graph and Spatial Prior Networks2019TOG
Adaptive Synthesis of Indoor Scenes via Activity-Associated Object Relation Graphs2017TOG

1.2.2 Serialized Graph-based

1.3 Hybrid Representation

1.3.1 Hierarchical Architecture

TitleYearVenue
GeoNav: Empowering MLLMs with Explicit Geospatial Reasoning Abilities for Language-Goal Aerial Navigation2026PR
Map2Thought: Explicit 3D Spatial Reasoning via Metric Cognitive Maps2026arXiv
OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation2026ICLR
SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical Planning2026AAAI
Stairway to Success: An Online Floor-Aware Zero-Shot Object-Goal Navigation Framework via LLM-Driven Coarse-to-Fine Exploration2026RAL
CAUSALNAV: A Long-term Embodied Navigation System for Autonomous Mobile Robots in Dynamic Outdoor Scenarios2026RAL
LOG-Nav: Efficient Layout-Aware Object-Goal Navigation with Hierarchical Planning2026AAAI
CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMs2025ICCV
FSR-VLN: Fast and Slow Reasoning for Vision-Language Navigation with Hierarchical Multi-modal Scene Graph2025arXiv
Agentic 3D Scene Generation with Spatially Contextualized VLMs2025arXiv
From reactive to cognitive: brain-inspired spatial intelligence for embodied agents2025arXiv
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning2025arXiv
GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering2025CoRL
Mem4Nav: Boosting Vision-and-Language Navigation in Urban Environments with a Hierarchical Spatial-Cognition Long-Short Memory System2025arXiv
MG-Nav: Dual-Scale Visual Navigation via Sparse Spatial Memory2025arXiv
ObjectReact: Learning Object-Relative Control for Visual Navigation2025CoRL
RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration2025arXiv
RoboOS-NeXT: A Unified Memory-based Framework for Lifelong, Scalable, and Robust Multi-Robot Collaboration2025arXiv
Controllable 3D Outdoor Scene Generation via Scene Graphs2025ICCV
Spatial Understanding from Videos: Structured Prompts Meet Simulation Data2025NeurIPS
Struct2D: A Perception-Guided Framework for Spatial Reasoning in Large Multimodal Models2025NeurIPS
Embodied-RAG: General Non-parametric Embodied Memory for Retrieval and Generation2024arXiv
BEVBert: Multimodal Map Pre-training for Language-guided Navigation2023ICCV
4D Panoptic Scene Graph Generation2023NeurIPS
Hydra: A Real-time Spatial Perception System for 3D Scene Graph Construction and Optimization2022RSS
No RL, No Simulation: Learning to Navigate without Navigating2021NeurIPS
SCENEHGN: Hierarchical Graph Networks for 3D Indoor Scene Generation with Fine-Grained Geometry2021TPAMI
3D Dynamic Scene Graphs: Actionable Spatial Perception with Places, Objects, and Humans2020RSS
GRAINS: Generative Recursive Autoencoders for INdoor Scenes2019TOG

1.3.2 Feature Fusion


2. Reasoning: Inference with Cognitive Maps

2.1 Map as Embedding

2.1.1 Structured State Propagation

2.1.2 Latent Feature Matching

TitleYearVenue
VPN: Visual Prompt Navigation2026AAAI
History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation2026AAAI
SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical Planning2026AAAI
Agent Journey Beyond RGB: Hierarchical Semantic-Spatial Representation Enrichment for Vision-and-Language Navigation2026AAAI
MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation2025ACL
Ali-UI: Enhancing Complex Vision-Language Navigation with Alignment of Unified Map and Instruction Parsing2025MM
Mem4Nav: Boosting Vision-and-Language Navigation in Urban Environments with a Hierarchical Spatial-Cognition Long-Short Memory System2025arXiv
Learning Bird’s Eye View scene graph and knowledge-inspired policy for embodied visual navigation2025KBS
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation2025ICCV
OVL-MAP: An Online Visual Language Map Approach for Vision-and-Language Navigation in Continuous Environments2025RAL
Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal Navigation2025NeurIPS
Scene-Driven Multimodal Knowledge Graph Construction for Embodied AI2024TKDE
3D Question Answering for City Scene Understanding2024MM
MemoNav: Working Memory Model for Visual Navigation2024CVPR
BEVBert: Multimodal Map Pre-training for Language-guided Navigation2023ICCV
Bird’s-Eye-View Scene Graph for Vision-Language Navigation2023ICCV
Learning Navigational Visual Representations with Semantic Map Supervision2023ICCV
GridMM: Grid Memory Map for Vision-and-Language Navigation2023ICCV
Scene Graph Contrastive Learning for Embodied Navigation2023ICCV
Topological Semantic Graph Memory for Image-Goal Navigation2023CoRL
Cross-modal Map Learning for Vision and Language Navigation2022CVPR
Weakly-Supervised Multi-Granularity Map Learning for Vision-and-Language Navigation2022NeurIPS
Visual Graph Memory with Unsupervised Representation for Visual Navigation2021ICCV

2.2 Map as Prompt

2.2.1 Textual Prompting

TitleYearVenue
Thinking with Blueprints: Assisting Vision-Language Models in Spatial Reasoning via Structured Object Representation2026arXiv
LOG-Nav: Efficient Layout-Aware Object-Goal Navigation with Hierarchical Planning2026AAAI
Map2Thought: Explicit 3D Spatial Reasoning via Metric Cognitive Maps2026arXiv
Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning2025arXiv
EfficientNav: Towards On-Device Object-Goal Navigation with Navigation Map Caching and Retrieval2025NeurIPS
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning2025arXiv
MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning2025arXiv
TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making2025NeurIPS
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes2025arXiv
Embodied VSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks2025arXiv
Hi-Dyna Graph: Hierarchical Dynamic Scene Graph for Robotic Autonomy in Human-Centric Environments2025arXiv
KARMA: Augmenting Embodied AI Agents with Long-and-short Term Memory Systems2025ICRA
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning2025arXiv
CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMs2025ICCV
Scene Map-based Prompt Tuning for Navigation Instruction Generation2025CVPR
Cog-GA: A Large Language Models-based Generative Agent for Vision-Language Navigation in Continuous Environments2024arXiv
Embodied-RAG: General Non-parametric Embodied Memory for Retrieval and Generation2024arXiv
MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation2024ACL
Open-vocabulary Queryable Scene Representations for Real World Planning2023ICRA
SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning2023CoRL

2.2.2 Visual Prompting

2.2.3 Multimodal Prompting

2.3 Map as API

2.3.1 Real-time State Snapshot

2.3.2 Persistent Spatial Memory


3. Generation: Realization of Cognitive Maps

3.1 Static Scene Synthesis

3.1.1 Map-based Retrieval

3.1.2 Map-to-Scene Generation

3.2 Dynamic World Simulation


4. Application Domains

4.1 Open-Loop Spatial Cognition

4.1.1 Spatial Question Answering

4.1.2 Indoor Scene Synthesis

TitleYearVenue
SPATIALGEN: Layout-guided 3D Indoor Scene Generation20263DV
LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models2025CVPR
MMGDreamer: Mixed-Modality Graph for Geometry-Controllable 3D Indoor Scene Generation2025AAAI
HiGS: Hierarchical Generative Scene Framework for Multi-Step Associative Semantic Spatial Composition2025arXiv
LayoutAgent: A Vision-Language Agent Guided Compositional Diffusion for Spatial Layout Planning2025arXiv
Planner3D: LLM-enhanced graph prior meets 3D indoor scene explicit regularization2025TPAMI
Ctrl-Room: Controllable Text-to-3D Room Meshes Generation with Layout Constraints20253DV
Learning 3D Persistent Embodied World Models2025NeurIPS
FOREST2SEQ: Revitalizing Order Prior for Sequential Indoor Scene Synthesis2024ECCV
HOLODECK: Language Guided Generation of 3D Embodied AI Environments2024CVPR
INSTRUCTSCENE: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior2024ICLR
External Knowledge Enhanced 3D Scene Generation from Sketch2024ECCV
GraphDreamer: Compositional 3D Scene Synthesis from Scene Graphs2024CVPR
CC3D: Layout-Conditioned Generation of Compositional 3D Scenes2023ICCV
CommonScenes: Generating Commonsense 3D Indoor Scenes with Scene Graph Diffusion2023NeurIPS
Graph-to-3D: End-to-End Generation and Manipulation of 3D Scenes Using Scene Graphs2021ICCV
SCENEHGN: Hierarchical Graph Networks for 3D Indoor Scene Generation with Fine-Grained Geometry2021TPAMI
End-to-End Optimization of Scene Layout2020CVPR
GRAINS: Generative Recursive Autoencoders for INdoor Scenes2019TOG
PlanIT: Planning and Instantiating Indoor Scenes with Relation Graph and Spatial Prior Networks2019TOG
Language-Driven Synthesis of 3D Scenes from Scene Databases2018TOG
Adaptive Synthesis of Indoor Scenes via Activity-Associated Object Relation Graphs2017TOG

4.1.3 Open-ended World Generation

4.2 Closed-Loop Spatial Interaction

4.2.1 Embodied Navigation

TitleYearVenue
GeoNav: Empowering MLLMs with Explicit Geospatial Reasoning Abilities for Language-Goal Aerial Navigation2026PR
OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation2026ICLR
What You See is What You Reach: Towards Spatial Navigation with High-Level Human Instructions2026AAAI
CAUSALNAV: A Long-term Embodied Navigation System for Autonomous Mobile Robots in Dynamic Outdoor Scenarios2026RAL
LOG-Nav: Efficient Layout-Aware Object-Goal Navigation with Hierarchical Planning2026AAAI
PanoNav: Mapless Zero-Shot Object Navigation with Panoramic Scene Parsing and Dynamic Memory2026AAAI
VPN: Visual Prompt Navigation2026AAAI
History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation2026AAAI
SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical Planning2026AAAI
Agent Journey Beyond RGB: Hierarchical Semantic-Spatial Representation Enrichment for Vision-and-Language Navigation2026AAAI
Stairway to Success: An Online Floor-Aware Zero-Shot Object-Goal Navigation Framework via LLM-Driven Coarse-to-Fine Exploration2026RAL
FSR-VLN: Fast and Slow Reasoning for Vision-Language Navigation with Hierarchical Multi-modal Scene Graph2025arXiv
GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering2025CoRL
Human-like Navigation in a World Built for Humans2025CoRL
3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning2025CVPR
Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation2025NeurIPS
EfficientNav: Towards On-Device Object-Goal Navigation with Navigation Map Caching and Retrieval2025NeurIPS
TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making2025NeurIPS
Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal Navigation2025NeurIPS
Embodied VSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks2025arXiv
CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMs2025ICCV
Scene Map-based Prompt Tuning for Navigation Instruction Generation2025CVPR
BeliefMapNav: 3D Voxel-Based Belief Map for Zero-Shot Object Navigation2025NeurIPS
From reactive to cognitive: brain-inspired spatial intelligence for embodied agents2025arXiv
GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation2025CoRL
MG-Nav: Dual-Scale Visual Navigation via Sparse Spatial Memory2025arXiv
3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model2025arXiv
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding2025ICCV
LagMemo: Language 3D Gaussian Splatting Memory for Multi-modal Open-vocabulary Multi-goal Visual Navigation2025arXiv
Meta-Memory: Retrieving and Integrating Semantic-Spatial Memories for Robot Spatial Reasoning2025arXiv
MrSteve: Instruction-Following Agents in Minecraft with What-Where-When Memory2025ICLR
ReMEmbR: Building and Reasoning Over Long-Horizon Spatio-Temporal Memory for Robot Navigation2025ICRA
RoboMemory: A Brain-inspired Multi-memory Agentic Framework for Interactive Environmental Learning in Physical Embodied Systems2025arXiv
Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation2025NeurIPS
Visuomotor Navigation for Embodied Robots With Spatial Memory and Semantic Reasoning Cognition2025TNNLS
ObjectReact: Learning Object-Relative Control for Visual Navigation2025CoRL
Ali-UI: Enhancing Complex Vision-Language Navigation with Alignment of Unified Map and Instruction Parsing2025MM
MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation2025ACL
Mem4Nav: Boosting Vision-and-Language Navigation in Urban Environments with a Hierarchical Spatial-Cognition Long-Short Memory System2025arXiv
Learning Bird’s Eye View scene graph and knowledge-inspired policy for embodied visual navigation2025KBS
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation2025ICCV
OVL-MAP: An Online Visual Language Map Approach for Vision-and-Language Navigation in Continuous Environments2025RAL
MemoNav: Working Memory Model for Visual Navigation2024CVPR
Embodied Contrastive Learning with Geometric Consistency and Behavioral Awareness for Object Navigation2024MM
Explore until Confident: Efficient Exploration for Embodied Question Answering2024arXiv
TopoNav: Topological Navigation for Efficient Exploration in Sparse Reward Environments2024IROS
Cog-GA: A Large Language Models-based Generative Agent for Vision-Language Navigation in Continuous Environments2024arXiv
Embodied-RAG: General Non-parametric Embodied Memory for Retrieval and Generation2024arXiv
MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation2024ACL
Visual Language Maps for Robot Navigation2023ICRA
Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied Agents2023ICCV
BEVBert: Multimodal Map Pre-training for Language-guided Navigation2023ICCV
Bird’s-Eye-View Scene Graph for Vision-Language Navigation2023ICCV
Learning Navigational Visual Representations with Semantic Map Supervision2023ICCV
GridMM: Grid Memory Map for Vision-and-Language Navigation2023ICCV
Scene Graph Contrastive Learning for Embodied Navigation2023ICCV
Topological Semantic Graph Memory for Image-Goal Navigation2023CoRL
Cross-modal Map Learning for Vision and Language Navigation2022CVPR
Weakly-Supervised Multi-Granularity Map Learning for Vision-and-Language Navigation2022NeurIPS
Visual Graph Memory with Unsupervised Representation for Visual Navigation2021ICCV
No RL, No Simulation: Learning to Navigate without Navigating2021NeurIPS
Cognitive Mapping and Planning for Visual Navigation2017CVPR

4.2.2 Embodied Manipulation


📌 Citation

If you find this survey or repository useful for your research, please consider citing:

@article{tian2026spatial,
  title={Spatial Intelligence from a Cognitive Map Perspective: A Survey},
  author={Tian, Yuxuan and Ji, Yuheng and Zheng, Xiaolong and Qin, Ziheng and Wang, Yipu and Zheng, Xinyi and Liu, Yuyang and Bai, Shuanghao and Li, Zhe and Wang, Liang and others},
  year={2026},
  publisher={Preprints}
}

Klingsor-tyx/Awesome-Spatial-Cognitive-Map

A curated list of papers analyzed in the survey: Spatial Intelligence from a Cognitive Map Perspective: A Survey

24

53 commits

updated Jun 1, 2026

See the code

README

Spatial Intelligence from a Cognitive Map Perspective:
A Survey

Yuxuan Tian* 1,2, Yuheng Ji*†1,2, Xiaolong Zheng*✉1,2, Ziheng Qin1,2, Yipu Wang1,2, Xinyi Zheng3,
Yuyang Liu1,2, Shuanghao Bai4, Zhe Li5, Liang Wang1,2, Daniel Dajun Zeng1,2

1CASIA 2UCAS 3BUAA 4XJTU 5NTU
* Equal Contribution    † Project Leader    ✉ Corresponding Author

🌐 Project Page    |    📄 Paper PDF    |    ⭐ Awesome List


📚 Overview

Overview of spatial intelligence from a cognitive map perspective.

This survey revisits Spatial Intelligence from the perspective of Cognitive Maps, treating cognitive maps as the representational blueprint that connects spatial perception, reasoning, and generation. Under this view, diverse research directions can be understood through a shared question: how an internal spatial representation is constructed, maintained, reasoned over, and realized.

We organize the literature into three cognitive-map-centric processes:

  • Perception: Construction of Cognitive Maps — how agents build internal spatial representations from local, partial, and multimodal observations.
  • Reasoning: Inference with Cognitive Maps — how cognitive maps are used as embeddings, prompts, or APIs to support spatial inference and decision-making.
  • Generation: Realization of Cognitive Maps — how internal maps are externalized into 3D scenes and dynamic world simulations.

This repository provides a curated list of papers analyzed in the survey, following the taxonomy proposed in Spatial Intelligence from a Cognitive Map Perspective: A Survey.

📖 Table of Contents


1. Perception: Construction of Cognitive Maps

1.1 Metric Representation

1.1.1 Explicit Geometry-based

2D Planar-based
TitleYearVenue
History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation2026AAAI
What You See is What You Reach: Towards Spatial Navigation with High-Level Human Instructions2026AAAI
One Map to Find Them All: Real-time Open-Vocabulary Mapping for Zero-shot Multi-Object Navigation2025ICRA
3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning2025CVPR
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models2025arXiv
MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation2025ACL
Ali-UI: Enhancing Complex Vision-Language Navigation with Alignment of Unified Map and Instruction Parsing2025MM
OVL-MAP: An Online Visual Language Map Approach for Vision-and-Language Navigation in Continuous Environments2025RAL
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning2025arXiv
Human-like Navigation in a World Built for Humans2025CoRL
Enhancing Embodied Object Detection with Spatial Feature Memory2025WACV
TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making2025NeurIPS
Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal Navigation2025NeurIPS
Mind's Eye of LLMs: Visualization-of-Thought Elicits Spatial Reasoning in Large Language Models2024NeurIPS
Explore until Confident: Efficient Exploration for Embodied Question Answering2024arXiv
GridMM: Grid Memory Map for Vision-and-Language Navigation2023ICCV
Visual Language Maps for Robot Navigation2023ICRA
Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied Agents2023ICCV
Open-vocabulary Queryable Scene Representations for Real World Planning2023ICRA
Cross-modal Map Learning for Vision and Language Navigation2022CVPR
Weakly-Supervised Multi-Granularity Map Learning for Vision-and-Language Navigation2022NeurIPS
Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric Views2021AAAI
Cognitive Mapping and Planning for Visual Navigation2017CVPR
3D Geometry-based

1.1.2 Parametric Coordinate-based

1.2 Relational Representation

1.2.1 Structured Graph-based

TitleYearVenue
Integrated Exploration and Sequential Manipulation on Scene Graph with LLM-based Situated Replanning2026ICRA
VPN: Visual Prompt Navigation2026AAAI
RAG-3DSG: Enhancing 3D Scene Graphs with Re-Shot Guided Retrieval-Augmented Generation2026arXiv
BrainyMP: Enhancing Motion Planning Using Graph Neural Network Inspired by Brain Spatial Relational Memory2025TITS
GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation2025CoRL
HiGS: Hierarchical Generative Scene Framework for Multi-Step Associative Semantic Spatial Composition2025arXiv
Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation2025NeurIPS
Planner3D: LLM-enhanced graph prior meets 3D indoor scene explicit regularization2025TPAMI
Visuomotor Navigation for Embodied Robots With Spatial Memory and Semantic Reasoning Cognition2025TNNLS
TB-HSU: Hierarchical 3D Scene Understanding with Contextual Affordances2025AAAI
InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior2024ICLR
MemoNav: Working Memory Model for Visual Navigation2024CVPR
Multiview Scene Graph2024NeurIPS
Scene-Driven Multimodal Knowledge Graph Construction for Embodied AI2024TKDE
3D Question Answering for City Scene Understanding2024MM
Bridging Visual and Textual Semantics: Towards Consistency for Unbiased Scene Graph Generation2024TPAMI
X-RefSeg3D: Enhancing Referring 3D Instance Segmentation via Structured Cross-Modal Graph Neural Networks2024AAAI
TopoNav: Topological Navigation for Efficient Exploration in Sparse Reward Environments2024IROS
CommonScenes: Generating Commonsense 3D Indoor Scenes with Scene Graph Diffusion2023NeurIPS
GTR: A Grafting-Then-Reassembling Framework for Dynamic Scene Graph Generation2023IJCAI
Topological Semantic Graph Memory for Image-Goal Navigation2023CoRL
Modeling Dynamic Environments with Scene Graph Memory2023ICML
Scene Graph Contrastive Learning for Embodied Navigation2023ICCV
Visual Graph Memory with Unsupervised Representation for Visual Navigation2021ICCV
Graph-to-3D: End-to-End Generation and Manipulation of 3D Scenes Using Scene Graphs2021ICCV
End-to-End Optimization of Scene Layout2020CVPR
PlanIT: Planning and Instantiating Indoor Scenes with Relation Graph and Spatial Prior Networks2019TOG
Adaptive Synthesis of Indoor Scenes via Activity-Associated Object Relation Graphs2017TOG

1.2.2 Serialized Graph-based

1.3 Hybrid Representation

1.3.1 Hierarchical Architecture

TitleYearVenue
GeoNav: Empowering MLLMs with Explicit Geospatial Reasoning Abilities for Language-Goal Aerial Navigation2026PR
Map2Thought: Explicit 3D Spatial Reasoning via Metric Cognitive Maps2026arXiv
OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation2026ICLR
SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical Planning2026AAAI
Stairway to Success: An Online Floor-Aware Zero-Shot Object-Goal Navigation Framework via LLM-Driven Coarse-to-Fine Exploration2026RAL
CAUSALNAV: A Long-term Embodied Navigation System for Autonomous Mobile Robots in Dynamic Outdoor Scenarios2026RAL
LOG-Nav: Efficient Layout-Aware Object-Goal Navigation with Hierarchical Planning2026AAAI
CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMs2025ICCV
FSR-VLN: Fast and Slow Reasoning for Vision-Language Navigation with Hierarchical Multi-modal Scene Graph2025arXiv
Agentic 3D Scene Generation with Spatially Contextualized VLMs2025arXiv
From reactive to cognitive: brain-inspired spatial intelligence for embodied agents2025arXiv
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning2025arXiv
GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering2025CoRL
Mem4Nav: Boosting Vision-and-Language Navigation in Urban Environments with a Hierarchical Spatial-Cognition Long-Short Memory System2025arXiv
MG-Nav: Dual-Scale Visual Navigation via Sparse Spatial Memory2025arXiv
ObjectReact: Learning Object-Relative Control for Visual Navigation2025CoRL
RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration2025arXiv
RoboOS-NeXT: A Unified Memory-based Framework for Lifelong, Scalable, and Robust Multi-Robot Collaboration2025arXiv
Controllable 3D Outdoor Scene Generation via Scene Graphs2025ICCV
Spatial Understanding from Videos: Structured Prompts Meet Simulation Data2025NeurIPS
Struct2D: A Perception-Guided Framework for Spatial Reasoning in Large Multimodal Models2025NeurIPS
Embodied-RAG: General Non-parametric Embodied Memory for Retrieval and Generation2024arXiv
BEVBert: Multimodal Map Pre-training for Language-guided Navigation2023ICCV
4D Panoptic Scene Graph Generation2023NeurIPS
Hydra: A Real-time Spatial Perception System for 3D Scene Graph Construction and Optimization2022RSS
No RL, No Simulation: Learning to Navigate without Navigating2021NeurIPS
SCENEHGN: Hierarchical Graph Networks for 3D Indoor Scene Generation with Fine-Grained Geometry2021TPAMI
3D Dynamic Scene Graphs: Actionable Spatial Perception with Places, Objects, and Humans2020RSS
GRAINS: Generative Recursive Autoencoders for INdoor Scenes2019TOG

1.3.2 Feature Fusion


2. Reasoning: Inference with Cognitive Maps

2.1 Map as Embedding

2.1.1 Structured State Propagation

2.1.2 Latent Feature Matching

TitleYearVenue
VPN: Visual Prompt Navigation2026AAAI
History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation2026AAAI
SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical Planning2026AAAI
Agent Journey Beyond RGB: Hierarchical Semantic-Spatial Representation Enrichment for Vision-and-Language Navigation2026AAAI
MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation2025ACL
Ali-UI: Enhancing Complex Vision-Language Navigation with Alignment of Unified Map and Instruction Parsing2025MM
Mem4Nav: Boosting Vision-and-Language Navigation in Urban Environments with a Hierarchical Spatial-Cognition Long-Short Memory System2025arXiv
Learning Bird’s Eye View scene graph and knowledge-inspired policy for embodied visual navigation2025KBS
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation2025ICCV
OVL-MAP: An Online Visual Language Map Approach for Vision-and-Language Navigation in Continuous Environments2025RAL
Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal Navigation2025NeurIPS
Scene-Driven Multimodal Knowledge Graph Construction for Embodied AI2024TKDE
3D Question Answering for City Scene Understanding2024MM
MemoNav: Working Memory Model for Visual Navigation2024CVPR
BEVBert: Multimodal Map Pre-training for Language-guided Navigation2023ICCV
Bird’s-Eye-View Scene Graph for Vision-Language Navigation2023ICCV
Learning Navigational Visual Representations with Semantic Map Supervision2023ICCV
GridMM: Grid Memory Map for Vision-and-Language Navigation2023ICCV
Scene Graph Contrastive Learning for Embodied Navigation2023ICCV
Topological Semantic Graph Memory for Image-Goal Navigation2023CoRL
Cross-modal Map Learning for Vision and Language Navigation2022CVPR
Weakly-Supervised Multi-Granularity Map Learning for Vision-and-Language Navigation2022NeurIPS
Visual Graph Memory with Unsupervised Representation for Visual Navigation2021ICCV

2.2 Map as Prompt

2.2.1 Textual Prompting

TitleYearVenue
Thinking with Blueprints: Assisting Vision-Language Models in Spatial Reasoning via Structured Object Representation2026arXiv
LOG-Nav: Efficient Layout-Aware Object-Goal Navigation with Hierarchical Planning2026AAAI
Map2Thought: Explicit 3D Spatial Reasoning via Metric Cognitive Maps2026arXiv
Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning2025arXiv
EfficientNav: Towards On-Device Object-Goal Navigation with Navigation Map Caching and Retrieval2025NeurIPS
SpaceR: Reinforcing MLLMs in Video Spatial Reasoning2025arXiv
MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Model for Embodied Task Planning2025arXiv
TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making2025NeurIPS
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes2025arXiv
Embodied VSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks2025arXiv
Hi-Dyna Graph: Hierarchical Dynamic Scene Graph for Robotic Autonomy in Human-Centric Environments2025arXiv
KARMA: Augmenting Embodied AI Agents with Long-and-short Term Memory Systems2025ICRA
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning2025arXiv
CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMs2025ICCV
Scene Map-based Prompt Tuning for Navigation Instruction Generation2025CVPR
Cog-GA: A Large Language Models-based Generative Agent for Vision-Language Navigation in Continuous Environments2024arXiv
Embodied-RAG: General Non-parametric Embodied Memory for Retrieval and Generation2024arXiv
MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation2024ACL
Open-vocabulary Queryable Scene Representations for Real World Planning2023ICRA
SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning2023CoRL

2.2.2 Visual Prompting

2.2.3 Multimodal Prompting

2.3 Map as API

2.3.1 Real-time State Snapshot

2.3.2 Persistent Spatial Memory


3. Generation: Realization of Cognitive Maps

3.1 Static Scene Synthesis

3.1.1 Map-based Retrieval

3.1.2 Map-to-Scene Generation

3.2 Dynamic World Simulation


4. Application Domains

4.1 Open-Loop Spatial Cognition

4.1.1 Spatial Question Answering

4.1.2 Indoor Scene Synthesis

TitleYearVenue
SPATIALGEN: Layout-guided 3D Indoor Scene Generation20263DV
LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models2025CVPR
MMGDreamer: Mixed-Modality Graph for Geometry-Controllable 3D Indoor Scene Generation2025AAAI
HiGS: Hierarchical Generative Scene Framework for Multi-Step Associative Semantic Spatial Composition2025arXiv
LayoutAgent: A Vision-Language Agent Guided Compositional Diffusion for Spatial Layout Planning2025arXiv
Planner3D: LLM-enhanced graph prior meets 3D indoor scene explicit regularization2025TPAMI
Ctrl-Room: Controllable Text-to-3D Room Meshes Generation with Layout Constraints20253DV
Learning 3D Persistent Embodied World Models2025NeurIPS
FOREST2SEQ: Revitalizing Order Prior for Sequential Indoor Scene Synthesis2024ECCV
HOLODECK: Language Guided Generation of 3D Embodied AI Environments2024CVPR
INSTRUCTSCENE: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior2024ICLR
External Knowledge Enhanced 3D Scene Generation from Sketch2024ECCV
GraphDreamer: Compositional 3D Scene Synthesis from Scene Graphs2024CVPR
CC3D: Layout-Conditioned Generation of Compositional 3D Scenes2023ICCV
CommonScenes: Generating Commonsense 3D Indoor Scenes with Scene Graph Diffusion2023NeurIPS
Graph-to-3D: End-to-End Generation and Manipulation of 3D Scenes Using Scene Graphs2021ICCV
SCENEHGN: Hierarchical Graph Networks for 3D Indoor Scene Generation with Fine-Grained Geometry2021TPAMI
End-to-End Optimization of Scene Layout2020CVPR
GRAINS: Generative Recursive Autoencoders for INdoor Scenes2019TOG
PlanIT: Planning and Instantiating Indoor Scenes with Relation Graph and Spatial Prior Networks2019TOG
Language-Driven Synthesis of 3D Scenes from Scene Databases2018TOG
Adaptive Synthesis of Indoor Scenes via Activity-Associated Object Relation Graphs2017TOG

4.1.3 Open-ended World Generation

4.2 Closed-Loop Spatial Interaction

4.2.1 Embodied Navigation

TitleYearVenue
GeoNav: Empowering MLLMs with Explicit Geospatial Reasoning Abilities for Language-Goal Aerial Navigation2026PR
OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation2026ICLR
What You See is What You Reach: Towards Spatial Navigation with High-Level Human Instructions2026AAAI
CAUSALNAV: A Long-term Embodied Navigation System for Autonomous Mobile Robots in Dynamic Outdoor Scenarios2026RAL
LOG-Nav: Efficient Layout-Aware Object-Goal Navigation with Hierarchical Planning2026AAAI
PanoNav: Mapless Zero-Shot Object Navigation with Panoramic Scene Parsing and Dynamic Memory2026AAAI
VPN: Visual Prompt Navigation2026AAAI
History-Enhanced Two-Stage Transformer for Aerial Vision-and-Language Navigation2026AAAI
SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical Planning2026AAAI
Agent Journey Beyond RGB: Hierarchical Semantic-Spatial Representation Enrichment for Vision-and-Language Navigation2026AAAI
Stairway to Success: An Online Floor-Aware Zero-Shot Object-Goal Navigation Framework via LLM-Driven Coarse-to-Fine Exploration2026RAL
FSR-VLN: Fast and Slow Reasoning for Vision-Language Navigation with Hierarchical Multi-modal Scene Graph2025arXiv
GraphEQA: Using 3D Semantic Scene Graphs for Real-time Embodied Question Answering2025CoRL
Human-like Navigation in a World Built for Humans2025CoRL
3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning2025CVPR
Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation2025NeurIPS
EfficientNav: Towards On-Device Object-Goal Navigation with Navigation Map Caching and Retrieval2025NeurIPS
TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making2025NeurIPS
Distilling LLM Prior to Flow Model for Generalizable Agent's Imagination in Object Goal Navigation2025NeurIPS
Embodied VSR: Dynamic Scene Graph-Guided Chain-of-Thought Reasoning for Visual Spatial Tasks2025arXiv
CogNav: Cognitive Process Modeling for Object Goal Navigation with LLMs2025ICCV
Scene Map-based Prompt Tuning for Navigation Instruction Generation2025CVPR
BeliefMapNav: 3D Voxel-Based Belief Map for Zero-Shot Object Navigation2025NeurIPS
From reactive to cognitive: brain-inspired spatial intelligence for embodied agents2025arXiv
GC-VLN: Instruction as Graph Constraints for Training-free Vision-and-Language Navigation2025CoRL
MG-Nav: Dual-Scale Visual Navigation via Sparse Spatial Memory2025arXiv
3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model2025arXiv
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding2025ICCV
LagMemo: Language 3D Gaussian Splatting Memory for Multi-modal Open-vocabulary Multi-goal Visual Navigation2025arXiv
Meta-Memory: Retrieving and Integrating Semantic-Spatial Memories for Robot Spatial Reasoning2025arXiv
MrSteve: Instruction-Following Agents in Minecraft with What-Where-When Memory2025ICLR
ReMEmbR: Building and Reasoning Over Long-Horizon Spatio-Temporal Memory for Robot Navigation2025ICRA
RoboMemory: A Brain-inspired Multi-memory Agentic Framework for Interactive Environmental Learning in Physical Embodied Systems2025arXiv
Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation2025NeurIPS
Visuomotor Navigation for Embodied Robots With Spatial Memory and Semantic Reasoning Cognition2025TNNLS
ObjectReact: Learning Object-Relative Control for Visual Navigation2025CoRL
Ali-UI: Enhancing Complex Vision-Language Navigation with Alignment of Unified Map and Instruction Parsing2025MM
MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation2025ACL
Mem4Nav: Boosting Vision-and-Language Navigation in Urban Environments with a Hierarchical Spatial-Cognition Long-Short Memory System2025arXiv
Learning Bird’s Eye View scene graph and knowledge-inspired policy for embodied visual navigation2025KBS
Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation2025ICCV
OVL-MAP: An Online Visual Language Map Approach for Vision-and-Language Navigation in Continuous Environments2025RAL
MemoNav: Working Memory Model for Visual Navigation2024CVPR
Embodied Contrastive Learning with Geometric Consistency and Behavioral Awareness for Object Navigation2024MM
Explore until Confident: Efficient Exploration for Embodied Question Answering2024arXiv
TopoNav: Topological Navigation for Efficient Exploration in Sparse Reward Environments2024IROS
Cog-GA: A Large Language Models-based Generative Agent for Vision-Language Navigation in Continuous Environments2024arXiv
Embodied-RAG: General Non-parametric Embodied Memory for Retrieval and Generation2024arXiv
MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation2024ACL
Visual Language Maps for Robot Navigation2023ICRA
Context-Aware Planning and Environment-Aware Memory for Instruction Following Embodied Agents2023ICCV
BEVBert: Multimodal Map Pre-training for Language-guided Navigation2023ICCV
Bird’s-Eye-View Scene Graph for Vision-Language Navigation2023ICCV
Learning Navigational Visual Representations with Semantic Map Supervision2023ICCV
GridMM: Grid Memory Map for Vision-and-Language Navigation2023ICCV
Scene Graph Contrastive Learning for Embodied Navigation2023ICCV
Topological Semantic Graph Memory for Image-Goal Navigation2023CoRL
Cross-modal Map Learning for Vision and Language Navigation2022CVPR
Weakly-Supervised Multi-Granularity Map Learning for Vision-and-Language Navigation2022NeurIPS
Visual Graph Memory with Unsupervised Representation for Visual Navigation2021ICCV
No RL, No Simulation: Learning to Navigate without Navigating2021NeurIPS
Cognitive Mapping and Planning for Visual Navigation2017CVPR

4.2.2 Embodied Manipulation


📌 Citation

If you find this survey or repository useful for your research, please consider citing:

@article{tian2026spatial,
  title={Spatial Intelligence from a Cognitive Map Perspective: A Survey},
  author={Tian, Yuxuan and Ji, Yuheng and Zheng, Xiaolong and Qin, Ziheng and Wang, Yipu and Zheng, Xinyi and Liu, Yuyang and Bai, Shuanghao and Li, Zhe and Wang, Liang and others},
  year={2026},
  publisher={Preprints}
}