sonia-raychaudhuri/awesome-semantic-maps

23

6 commits

updated Aug 9, 2025

See the code

README

Semantic Mapping in Indoor Embodied AI - A Survey on Advances, Challenges, and Future Directions

Sonia Raychaudhuri, Angel X. Chang


This repository contains the list of papers from our survey which presents the latest advances, open challenges and future directions in semantic mapping for indoor embodied AI.


Table of Contents

2025 · 2024 · 2023 · 2022 · 2021 · 2020 · 2019 · 2018 · 2017 · 2016 · 2015 · 2014 · 2013 · 2012 · 2011 · 2010 · 2009 · 2008 · 2007 · 2006 · 2005 · 2004 · 2003 · 2002 · 2001 · 2000 · 1998 · 1997 · 1996 · 1989 · 1985 · 1984 · 1981

2025

2024

Papers
ETPNav: Evolving topological planning for vision-language navigation in continuous environments
One Map to Find Them All: Real-time Open-Vocabulary Mapping for Zero-shot Multi-Object Navigation
SLAM Handbook
Unimate and Beyond: Exploring the Genesis of Industrial Robotics
Multi-level neural scene graphs for dynamic urban environments
Robohop: Segment-based topological map representation for open-world visual navigation
End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
Collaborative dynamic 3d scene graphs for automated driving
Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning
World models for autonomous driving: An initial survey
Semantic Gaussians: Open-Vocabulary Scene Understanding with 3D Gaussian Splatting
Foundations of spatial perception for robotics: Hierarchical representations and real-time systems
Goat-bench: A benchmark for multi-modal lifelong navigation
GaussNav: Gaussian Splatting for Visual Navigation
Sgs-slam: Semantic gaussian splatting for neural dense slam
LaneSegNet: Map Learning with Lane Segment Perception for Autonomous Driving
Bevformer: learning bird's-eye-view representation from lidar-camera via spatiotemporal transformers
Embodied AI with Large Language Models: A Survey and New HRI Framework
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment
Estimating Map Completeness in Robot Exploration
Clio: Real-time task-driven open-set 3d scene graphs
OpenEQA: Embodied Question Answering in the Era of Foundation Models
QueSTMaps: Queryable Semantic Topological Maps for 3D Scene Understanding
LangSplat: 3D language Gaussian splatting
Learning generalizable feature fields for mobile manipulation
Explore until Confident: Efficient Exploration for Embodied Question Answering
Collaborative Instance Navigation: Leveraging Agent Self-Dialogue to Minimize User Input
Mobileclip: Fast image-text models through multi-modal reinforced training
A survey of visual SLAM in dynamic environment: The evolution from geometric to semantic approaches
Hierarchical Open-Vocabulary 3D Scene Graphs for Language-Grounded Robot Navigation
Embodied navigation with multi-modal information: A survey from tasks to methodology
SED: A simple encoder-decoder for open-vocabulary semantic segmentation
EmbodiedSAM: Online Segment Any 3D Thing in Real Time
3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning
HM3D-OVON: A Dataset and Benchmark for Open-Vocabulary Object Goal Navigation
BEV perception for autonomous driving: State of the art and future perspectives
A Survey of Embodied Learning for Object-Centric Robotic Manipulation
Gaussiangrasper: 3d language gaussian splatting for open-vocabulary robotic grasping

2023

Papers
Active slam: A review on last decade
BEVBert: Multimodal Map Pre-training for Language-guided Navigation
A review of high-definition map creation methods for autonomous driving
S-graphs+: Real-time localization and mapping leveraging hierarchical representations
GOAT: Go to any thing
Open-vocabulary queryable scene representations for real world planning
How To Not Train Your Dragon: Training-free Embodied Object Navigation with Semantic Frontiers
CLIP2Scene: Towards Label-efficient 3D Scene Understanding by CLIP
PaLM-E: An Embodied Multimodal Language Model
Semantic and Topological Mapping using Intersection Identification
Cows on pasture: Baselines and benchmarks for language-driven zero-shot object navigation
3D scene reconstruction and mapping with real time human detection for search and rescue robotics
Navigating to objects in the real world
ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
Visual language maps for robot navigation
VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
ConceptFusion: Open-set Multimodal 3D Mapping
LeRF: Language embedded radiance fields
Topological semantic graph memory for image-goal navigation
Segment Anything
Renderable Neural Radiance Map for Visual Navigation
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Visual instruction tuning
Bevfusion: Multi-task multi-sensor fusion with unified bird's-eye view representation
Feature-realistic neural fusion for real-time, open set scene understanding
GPT-4 Technical Report
Dinov2: Learning robust visual features without supervision
OpenScene: 3D Scene Understanding with Open Vocabularies
Visual SLAM integration with semantic segmentation and deep learning: A review
Constructing maps for autonomous robotics: An introductory conceptual overview
MOPA: Modular Object Navigation with PointGoal Agents
Nerf-slam: Real-time dense monocular slam with neural radiance fields
CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory
LM-Nav: Robotic navigation with large pre-trained models of language, vision, and action
Distilled feature fields enable few-shot language-guided manipulation
Vision-based dirt distribution mapping using deep learning
A systematic literature review on long-term localization and mapping for mobile robots
Language-enhanced RNR-map: Querying renderable neural radiance field maps with natural language
Habitat-Matterport 3D Semantics dataset
VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation
3D-aware object goal navigation via simultaneous exploration and identification
Recognize Anything: A Strong Image Tagging Model
ESC: Exploration with soft commonsense constraints for zero-shot object navigation

2022

Papers
Collaborative mobile robotics for semantic mapping: A survey
Flamingo: a Visual Language Model for Few-Shot Learning
Learning deep sdf maps online for robot navigation and exploration
Fast 3D sparse topological skeleton graph generation for mobile robot global planning
Retrospectives on the Embodied AI Workshop
PointResNet: residual network for 3D point cloud segmentation and classification
A survey of embodied AI: From simulators to research tasks
Panoptic nerf: 3d-to-2d label transfer for panoptic urban scene segmentation
Uncertainty-driven planner for exploration and navigation
Learning to Map for Active Semantic Goal Navigation
Cross-modal map learning for vision and language navigation
Hydra: A real-time spatial perception system for 3D scene graph construction and optimization
Simple but effective: CLIP embeddings for embodied AI
Instance-Specific Image Goal Navigation: Training Embodied Agents to Find Object Instances
Language-driven Semantic Segmentation
A survey of transformers
Stubborn: A Strong Baseline for Indoor Object Navigation
Curiosity-driven exploration via latent bayesian surprise
Stronger together: Air-ground robotic collaboration using semantics
isdf: Real-time neural signed distance fields for robot perception
TEACh: Task-driven Embodied Agents that Chat
Enough is enough: Towards autonomous uncertainty-driven stopping criteria
Towards Accurate Loop Closure Detection in Semantic SLAM With 3D Semantic Covisibility Graphs
Visual slam: What are the current trends and what to expect?
A simple approach for visual room rearrangement: 3D mapping and semantic search
The revisiting problem in simultaneous localization and mapping: A survey on visual loop closure detection
Language Understanding for Field and Service Robots in a Priori Unknown Environments
Clip-nerf: Text-and-image driven manipulation of neural radiance fields
Habitat Challenge 2022
3D-Aware Object Goal Navigation via Simultaneous Exploration and Identification
A survey of visual navigation: From geometry to embodied AI
Extract free dense labels from clip
NICE-SLAM: Neural implicit scalable encoding for SLAM

2021

Papers
An improved visual SLAM based on affine transformation for ORB feature extraction
ORB-SLAM3: An accurate open-source library for visual, visual--inertial, and multimap SLAM
Emerging properties in self-supervised vision transformers
Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric Views
SEAL: Self-supervised embodied active learning using exploration and 3D consistency
Topological planning with transformers for vision-and-language navigation
An improved initialization method for monocular visual-inertial SLAM
Deep reinforcement learning for map-less goal-driven robot navigation
Learning to Map for Active Semantic Goal Navigation
Open-vocabulary Object Detection via Vision and Language Knowledge Distillation
A comparative survey of lidar-slam and lidar based sensor technologies
Visual graph memory with unsupervised representation for visual navigation
Prototypical Contrastive Learning of Unsupervised Representations
SA-LOAM: Semantic-aided LiDAR SLAM with Loop Closure
Ion: Instance-level object navigation
Active mapping and robot exploration: A survey
Nerf: Representing scenes as neural radiance fields for view synthesis
Any way you look at it: Semantic crossview localization and mapping with lidar
Voldor+ slam: For the times when feature-based or direct methods are not good enough
Semantic slam with autonomous object-level data association
Semantic loop closure detection based on graph matching in multi-objects scenes
Learning transferable visual models from natural language supervision
Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI
Language-Aligned Waypoint (LAW) Supervision for Vision-and-Language Navigation in Continuous Environments
Kimera: From SLAM to spatial perception with 3D dynamic scene graphs
imap: Implicit mapping and positioning in real-time
Habitat 2.0: Training Home Assistants to Rearrange their Habitat
SLAM; definition and evolution
Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras
NeSF: Neural semantic fields for generalizable semantic segmentation of 3D scenes
Visual Room Rearrangement
Scenegraphfusion: Isncremental 3d scene graph prediction from rgb-d sequences
A novel informative autonomous exploration strategy with uniform sampling for quadrotors
Point transformer
In-place scene labelling and understanding with implicit scene representation
Deep learning for embodied vision navigation: A survey

2020

Papers
Multimodal estimation and communication of latent semantic knowledge for robust execution of robot instructions
Real-time human-robot communication for manipulation tasks in partially observed environments
Combining optimal control and learning for visual navigation in novel environments
Rearrangement: A challenge for embodied ai
ObjectNav Revisited: On Evaluation of Embodied Agents Navigating to Objects
Learning to plan with uncertain topological maps
Object Goal Navigation using Goal-Oriented Semantic Exploration
Neural Topological SLAM for Visual Navigation
SoundSpaces: Audio-visual navigation in 3D environments
Learning to Set Waypoints for Audio-Visual Navigation
A survey on deep learning for localization and mapping: Towards the age of spatial machine intelligence
Sloam: Semantic lidar odometry and mapping for forest inventory
An image is worth 16x16 words: Transformers for image recognition at scale
Semantics for robotic mapping, perception and interaction: A survey
Integrated Task and Motion Planning
Sim-to-real reinforcement learning applied to end-to-end vehicle control
Learning to See before Learning to Act: Visual Pre-training for Manipulation
Seeing the un-scene: Learning amodal semantic maps for room navigation
Language-guided semantic mapping and mobile manipulation in partially observable environments
Occupancy anticipation for efficient exploration and navigation
3D Dynamic Scene Graphs: Actionable Spatial Perception with Places, Objects, and Humans
Kimera: an open-source library for real-time metric-semantic localization and mapping
Superglue: Learning feature matching with graph neural networks
OrcVIO: Object residual constrained visual-inertial odometry
Lio-sam: Tightly-coupled lidar inertial odometry via smoothing and mapping
ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
MultiON: Benchmarking semantic map memory using multi-object navigation
Safe Robot Navigation Via Multi-Modal Anomaly Detection
A survey of image semantics-based visual simultaneous localization and mapping: Application-oriented solutions to autonomous navigation of mobile robots
Sapien: A simulated part-based interactive environment
Fusionlane: Multi-sensor fusion for lane marking semantic segmentation using deep neural networks
Transporter Networks: Rearranging the Visual World for Robotic Manipulation
Fusion-aware point convolution for online semantic 3d scene segmentation

2019

2018

2017

Papers
A dataset for developing and benchmarking active vision
Probabilistic data association for semantic SLAM
Past, present, and future of simultaneous localization and mapping: Toward the robust-perception age
Matterport3D: Learning from RGB-D Data in Indoor Environments
Bundlefusion: Real-time globally consistent 3d reconstruction using on-the-fly surface reintegration
Factor graphs for robot perception
Cognitive mapping and planning for visual navigation
Mask R-CNN
An iterative closest points algorithm for registration of 3D laser scanner point clouds with geometric features
AI2-THOR: An Interactive 3D Environment for Visual AI
A review of spatial reasoning and interaction for real-world robotics
Semanticfusion: Dense 3d semantic mapping with convolutional neural networks
Visual-inertial monocular SLAM with map reuse
Curiosity-driven exploration by self-supervised prediction
Pointnet: Deep learning on point sets for 3d classification and segmentation
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Co-fusion: Real-time segmentation, tracking and fusion of multiple objects
Virtual-to-real deep reinforcement learning: Continuous control of mobile robots for mapless navigation
Visual SLAM algorithms: A survey from 2010 to 2016
Cnn-slam: Real-time dense monocular slam with learned depth prediction
DA-RNN: Semantic mapping with data associated recurrent neural networks
Direct monocular odometry using points and lines
Keyframe-based monocular SLAM: design, survey, and future directions
Robotic pick-and-place of novel objects in clutter with multi-affordance grasping and cross-domain image matching

2016

2015

2014

2013

2012

2011

2010

2009

2008

2007

2006

2005

2004

2003

2002

2001

2000

1998

1997

1996

1989

1985

1984

1981

sonia-raychaudhuri/awesome-semantic-maps

23

6 commits

updated Aug 9, 2025

See the code

README

Semantic Mapping in Indoor Embodied AI - A Survey on Advances, Challenges, and Future Directions

Sonia Raychaudhuri, Angel X. Chang


This repository contains the list of papers from our survey which presents the latest advances, open challenges and future directions in semantic mapping for indoor embodied AI.


Table of Contents

2025 · 2024 · 2023 · 2022 · 2021 · 2020 · 2019 · 2018 · 2017 · 2016 · 2015 · 2014 · 2013 · 2012 · 2011 · 2010 · 2009 · 2008 · 2007 · 2006 · 2005 · 2004 · 2003 · 2002 · 2001 · 2000 · 1998 · 1997 · 1996 · 1989 · 1985 · 1984 · 1981

2025

2024

Papers
ETPNav: Evolving topological planning for vision-language navigation in continuous environments
One Map to Find Them All: Real-time Open-Vocabulary Mapping for Zero-shot Multi-Object Navigation
SLAM Handbook
Unimate and Beyond: Exploring the Genesis of Industrial Robotics
Multi-level neural scene graphs for dynamic urban environments
Robohop: Segment-based topological map representation for open-world visual navigation
End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
Collaborative dynamic 3d scene graphs for automated driving
Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning
World models for autonomous driving: An initial survey
Semantic Gaussians: Open-Vocabulary Scene Understanding with 3D Gaussian Splatting
Foundations of spatial perception for robotics: Hierarchical representations and real-time systems
Goat-bench: A benchmark for multi-modal lifelong navigation
GaussNav: Gaussian Splatting for Visual Navigation
Sgs-slam: Semantic gaussian splatting for neural dense slam
LaneSegNet: Map Learning with Lane Segment Perception for Autonomous Driving
Bevformer: learning bird's-eye-view representation from lidar-camera via spatiotemporal transformers
Embodied AI with Large Language Models: A Survey and New HRI Framework
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
InstructNav: Zero-shot System for Generic Instruction Navigation in Unexplored Environment
Estimating Map Completeness in Robot Exploration
Clio: Real-time task-driven open-set 3d scene graphs
OpenEQA: Embodied Question Answering in the Era of Foundation Models
QueSTMaps: Queryable Semantic Topological Maps for 3D Scene Understanding
LangSplat: 3D language Gaussian splatting
Learning generalizable feature fields for mobile manipulation
Explore until Confident: Efficient Exploration for Embodied Question Answering
Collaborative Instance Navigation: Leveraging Agent Self-Dialogue to Minimize User Input
Mobileclip: Fast image-text models through multi-modal reinforced training
A survey of visual SLAM in dynamic environment: The evolution from geometric to semantic approaches
Hierarchical Open-Vocabulary 3D Scene Graphs for Language-Grounded Robot Navigation
Embodied navigation with multi-modal information: A survey from tasks to methodology
SED: A simple encoder-decoder for open-vocabulary semantic segmentation
EmbodiedSAM: Online Segment Any 3D Thing in Real Time
3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning
HM3D-OVON: A Dataset and Benchmark for Open-Vocabulary Object Goal Navigation
BEV perception for autonomous driving: State of the art and future perspectives
A Survey of Embodied Learning for Object-Centric Robotic Manipulation
Gaussiangrasper: 3d language gaussian splatting for open-vocabulary robotic grasping

2023

Papers
Active slam: A review on last decade
BEVBert: Multimodal Map Pre-training for Language-guided Navigation
A review of high-definition map creation methods for autonomous driving
S-graphs+: Real-time localization and mapping leveraging hierarchical representations
GOAT: Go to any thing
Open-vocabulary queryable scene representations for real world planning
How To Not Train Your Dragon: Training-free Embodied Object Navigation with Semantic Frontiers
CLIP2Scene: Towards Label-efficient 3D Scene Understanding by CLIP
PaLM-E: An Embodied Multimodal Language Model
Semantic and Topological Mapping using Intersection Identification
Cows on pasture: Baselines and benchmarks for language-driven zero-shot object navigation
3D scene reconstruction and mapping with real time human detection for search and rescue robotics
Navigating to objects in the real world
ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills
Visual language maps for robot navigation
VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
ConceptFusion: Open-set Multimodal 3D Mapping
LeRF: Language embedded radiance fields
Topological semantic graph memory for image-goal navigation
Segment Anything
Renderable Neural Radiance Map for Visual Navigation
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Visual instruction tuning
Bevfusion: Multi-task multi-sensor fusion with unified bird's-eye view representation
Feature-realistic neural fusion for real-time, open set scene understanding
GPT-4 Technical Report
Dinov2: Learning robust visual features without supervision
OpenScene: 3D Scene Understanding with Open Vocabularies
Visual SLAM integration with semantic segmentation and deep learning: A review
Constructing maps for autonomous robotics: An introductory conceptual overview
MOPA: Modular Object Navigation with PointGoal Agents
Nerf-slam: Real-time dense monocular slam with neural radiance fields
CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory
LM-Nav: Robotic navigation with large pre-trained models of language, vision, and action
Distilled feature fields enable few-shot language-guided manipulation
Vision-based dirt distribution mapping using deep learning
A systematic literature review on long-term localization and mapping for mobile robots
Language-enhanced RNR-map: Querying renderable neural radiance field maps with natural language
Habitat-Matterport 3D Semantics dataset
VLFM: Vision-Language Frontier Maps for Zero-Shot Semantic Navigation
3D-aware object goal navigation via simultaneous exploration and identification
Recognize Anything: A Strong Image Tagging Model
ESC: Exploration with soft commonsense constraints for zero-shot object navigation

2022

Papers
Collaborative mobile robotics for semantic mapping: A survey
Flamingo: a Visual Language Model for Few-Shot Learning
Learning deep sdf maps online for robot navigation and exploration
Fast 3D sparse topological skeleton graph generation for mobile robot global planning
Retrospectives on the Embodied AI Workshop
PointResNet: residual network for 3D point cloud segmentation and classification
A survey of embodied AI: From simulators to research tasks
Panoptic nerf: 3d-to-2d label transfer for panoptic urban scene segmentation
Uncertainty-driven planner for exploration and navigation
Learning to Map for Active Semantic Goal Navigation
Cross-modal map learning for vision and language navigation
Hydra: A real-time spatial perception system for 3D scene graph construction and optimization
Simple but effective: CLIP embeddings for embodied AI
Instance-Specific Image Goal Navigation: Training Embodied Agents to Find Object Instances
Language-driven Semantic Segmentation
A survey of transformers
Stubborn: A Strong Baseline for Indoor Object Navigation
Curiosity-driven exploration via latent bayesian surprise
Stronger together: Air-ground robotic collaboration using semantics
isdf: Real-time neural signed distance fields for robot perception
TEACh: Task-driven Embodied Agents that Chat
Enough is enough: Towards autonomous uncertainty-driven stopping criteria
Towards Accurate Loop Closure Detection in Semantic SLAM With 3D Semantic Covisibility Graphs
Visual slam: What are the current trends and what to expect?
A simple approach for visual room rearrangement: 3D mapping and semantic search
The revisiting problem in simultaneous localization and mapping: A survey on visual loop closure detection
Language Understanding for Field and Service Robots in a Priori Unknown Environments
Clip-nerf: Text-and-image driven manipulation of neural radiance fields
Habitat Challenge 2022
3D-Aware Object Goal Navigation via Simultaneous Exploration and Identification
A survey of visual navigation: From geometry to embodied AI
Extract free dense labels from clip
NICE-SLAM: Neural implicit scalable encoding for SLAM

2021

Papers
An improved visual SLAM based on affine transformation for ORB feature extraction
ORB-SLAM3: An accurate open-source library for visual, visual--inertial, and multimap SLAM
Emerging properties in self-supervised vision transformers
Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric Views
SEAL: Self-supervised embodied active learning using exploration and 3D consistency
Topological planning with transformers for vision-and-language navigation
An improved initialization method for monocular visual-inertial SLAM
Deep reinforcement learning for map-less goal-driven robot navigation
Learning to Map for Active Semantic Goal Navigation
Open-vocabulary Object Detection via Vision and Language Knowledge Distillation
A comparative survey of lidar-slam and lidar based sensor technologies
Visual graph memory with unsupervised representation for visual navigation
Prototypical Contrastive Learning of Unsupervised Representations
SA-LOAM: Semantic-aided LiDAR SLAM with Loop Closure
Ion: Instance-level object navigation
Active mapping and robot exploration: A survey
Nerf: Representing scenes as neural radiance fields for view synthesis
Any way you look at it: Semantic crossview localization and mapping with lidar
Voldor+ slam: For the times when feature-based or direct methods are not good enough
Semantic slam with autonomous object-level data association
Semantic loop closure detection based on graph matching in multi-objects scenes
Learning transferable visual models from natural language supervision
Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI
Language-Aligned Waypoint (LAW) Supervision for Vision-and-Language Navigation in Continuous Environments
Kimera: From SLAM to spatial perception with 3D dynamic scene graphs
imap: Implicit mapping and positioning in real-time
Habitat 2.0: Training Home Assistants to Rearrange their Habitat
SLAM; definition and evolution
Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras
NeSF: Neural semantic fields for generalizable semantic segmentation of 3D scenes
Visual Room Rearrangement
Scenegraphfusion: Isncremental 3d scene graph prediction from rgb-d sequences
A novel informative autonomous exploration strategy with uniform sampling for quadrotors
Point transformer
In-place scene labelling and understanding with implicit scene representation
Deep learning for embodied vision navigation: A survey

2020

Papers
Multimodal estimation and communication of latent semantic knowledge for robust execution of robot instructions
Real-time human-robot communication for manipulation tasks in partially observed environments
Combining optimal control and learning for visual navigation in novel environments
Rearrangement: A challenge for embodied ai
ObjectNav Revisited: On Evaluation of Embodied Agents Navigating to Objects
Learning to plan with uncertain topological maps
Object Goal Navigation using Goal-Oriented Semantic Exploration
Neural Topological SLAM for Visual Navigation
SoundSpaces: Audio-visual navigation in 3D environments
Learning to Set Waypoints for Audio-Visual Navigation
A survey on deep learning for localization and mapping: Towards the age of spatial machine intelligence
Sloam: Semantic lidar odometry and mapping for forest inventory
An image is worth 16x16 words: Transformers for image recognition at scale
Semantics for robotic mapping, perception and interaction: A survey
Integrated Task and Motion Planning
Sim-to-real reinforcement learning applied to end-to-end vehicle control
Learning to See before Learning to Act: Visual Pre-training for Manipulation
Seeing the un-scene: Learning amodal semantic maps for room navigation
Language-guided semantic mapping and mobile manipulation in partially observable environments
Occupancy anticipation for efficient exploration and navigation
3D Dynamic Scene Graphs: Actionable Spatial Perception with Places, Objects, and Humans
Kimera: an open-source library for real-time metric-semantic localization and mapping
Superglue: Learning feature matching with graph neural networks
OrcVIO: Object residual constrained visual-inertial odometry
Lio-sam: Tightly-coupled lidar inertial odometry via smoothing and mapping
ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
MultiON: Benchmarking semantic map memory using multi-object navigation
Safe Robot Navigation Via Multi-Modal Anomaly Detection
A survey of image semantics-based visual simultaneous localization and mapping: Application-oriented solutions to autonomous navigation of mobile robots
Sapien: A simulated part-based interactive environment
Fusionlane: Multi-sensor fusion for lane marking semantic segmentation using deep neural networks
Transporter Networks: Rearranging the Visual World for Robotic Manipulation
Fusion-aware point convolution for online semantic 3d scene segmentation

2019

2018

2017

Papers
A dataset for developing and benchmarking active vision
Probabilistic data association for semantic SLAM
Past, present, and future of simultaneous localization and mapping: Toward the robust-perception age
Matterport3D: Learning from RGB-D Data in Indoor Environments
Bundlefusion: Real-time globally consistent 3d reconstruction using on-the-fly surface reintegration
Factor graphs for robot perception
Cognitive mapping and planning for visual navigation
Mask R-CNN
An iterative closest points algorithm for registration of 3D laser scanner point clouds with geometric features
AI2-THOR: An Interactive 3D Environment for Visual AI
A review of spatial reasoning and interaction for real-world robotics
Semanticfusion: Dense 3d semantic mapping with convolutional neural networks
Visual-inertial monocular SLAM with map reuse
Curiosity-driven exploration by self-supervised prediction
Pointnet: Deep learning on point sets for 3d classification and segmentation
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
Co-fusion: Real-time segmentation, tracking and fusion of multiple objects
Virtual-to-real deep reinforcement learning: Continuous control of mobile robots for mapless navigation
Visual SLAM algorithms: A survey from 2010 to 2016
Cnn-slam: Real-time dense monocular slam with learned depth prediction
DA-RNN: Semantic mapping with data associated recurrent neural networks
Direct monocular odometry using points and lines
Keyframe-based monocular SLAM: design, survey, and future directions
Robotic pick-and-place of novel objects in clutter with multi-affordance grasping and cross-domain image matching

2016

2015

2014

2013

2012

2011

2010

2009

2008

2007

2006

2005

2004

2003

2002

2001

2000

1998

1997

1996

1989

1985

1984

1981