waynechu1021/Awesome_Visual_Language_Navigation

End-to-End Visual Language Navigation with Limited Sensing: A Survey

TeX

46

28 commits

updated Sep 17, 2026

See the code

README

๐Ÿงญ Awesome Visual Language Navigation

Awesome list badge BibTeX Paper

๐ŸŒŸ Overview

๐Ÿ“š Preliminaries

Simulation, Data and Control

Toward More Realistic Simulation

Towards More Diverse Datasets&Benchmarks

YearPaperVenueResources
2026HumanoidVLN: A Physics-Grounded Simulator and Benchmark for Vision-Language Navigation Across Diverse Humanoid Embodiments2026 arXivProject
2026DaViNCi: A Dataset Towards Outdoor Vision-and-Language Navigation with Continuous Actions and Dynamic Elements2026 arXivProject
2026Beyond Isolation: A Unified Benchmark for General-Purpose Navigation2026 RSSProject Code
2026MiniVLA-Nav v1: A Multi-Scene Simulation Dataset for Language-Conditioned Robot Navigation2026 arXivโ€”
2026Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration2026 arXivโ€”
2026CoNavBench: Collaborative Long-Horizon Vision-Language Navigation Benchmark2026 ICLRProject
2026Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification2026 arXivโ€”
2026AirNav: A Large-Scale Real-World UAV Vision-and-Language Navigation Dataset with Natural and Diverse Instructions2026 arXivโ€”
2025IndoorUAV: Benchmarking Vision-Language UAV Navigation in Continuous Indoor Environments2025 arXivโ€”
2025VL-LN Bench: Towards Long-horizon Goal-oriented Navigation with Active Dialogs2025 arXivCode Data
2025OpenFly: A Comprehensive Platform for Aerial Vision-Language Navigation2025 arXivCode Data
2025Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent Dialogues2025 ICCVProject Code Data
2025RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation2025 CVPRCode Data
2025Towards long-horizon vision-language navigation: Platform, benchmark and method2025 CVPRCode
2025CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos2025 CVPRCode Data
2025UrbanNav: Learning Language-Guided Urban Navigation from Web-Scale Human Trajectories2025 arXivCode
2025UAV-Flow Colosseo: A Real-World Benchmark for Flying-on-a-Word UAV Imitation Learning2025 arXivโ€”
2024Embodiedcity: A benchmark platform for embodied agent in real-world city environment2024 arXivCode
2024InstruGen: Automatic Instruction Generation for Vision-and-Language Navigation Via Large Multimodal Models2024 arXivCode
2024Towards realistic uav vision-language navigation: Platform, benchmark, and methodology2024 arXivProject Code Data
2024CityNav: Language-goal aerial navigation dataset with geographic information2024 arXivCode
2024Mind the error! detection and localization of instruction errors in vision-and-language navigation2024 IROSProject Code
2024Hazard challenge: Embodied decision making in dynamically changing environments2024 arXivProject
2023AerialVLN: Vision-and-Language Navigation for UAVs2023 ICCVCode
2023Iterative vision-and-language navigation2023 CVPRCode
2023Scaling Data Generation in Vision-and-Language Navigation2023 ICCVโ€”
2022REVE-CE: Remote Embodied Visual Referring Expression in Continuous Environment2022 RA-Lโ€”
2021Talk2Nav: Long-Range Vision-and-Language Navigation with Dual Attention and Spatial Memory2021 IJCVโ€”
2020Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding2020 arXivโ€”
2018Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments2018 CVPRโ€”

Towards More Continuous Control

๐Ÿง  Context Modeling

Context Perception

Spatial Perception

Temporal Perception

Context Memory

Memory Representation

Memory Retrieval and Utilization

Context Efficiency

Reducing Temporal Redundancy

Reducing Spatial Redundancy

Computational Reuse and Acceleration

๐Ÿ”ฎ Imaginative Prediction

Spatial Imagination

Occupancy Prediction

View Completion

3D Representation

Temporal Imagination

Pixel Level

Feature Level

Inverse Dynamics Modeling and Causality

Inverse Dynamics

Causal Learning

๐ŸŽฏ Decision Planning

Structured Decision

Instruction Decomposition

Sub-Goal Formulation

Reasoning-based Decision

LLM as External Reasoner

Integrated Perception-Reasoning-Action Model

YearPaperVenueResources
2026GroundingVLN: Reasoning and Acting with Grounding for Vision-Language Navigation2026 arXivโ€”
2026MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control2026 arXivProject Code
2026ReflectVLN: Training Vision-Language Navigation Agents with Reflective Reasoning2026 arXivCode
2026AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation2026 CVPRProject Code
2026A Deployable Embodied Vision-Language Navigation System with Hierarchical Cognition and Context-Aware Exploration2026 arXivโ€”
2026AsyncShield: A Plug-and-Play Edge Adapter for Asynchronous Cloud-based VLA Navigation2026 arXivโ€”
2026AURA: Multimodal Shared Autonomy for Real-World Urban Navigation2026 arXivโ€”
2026HiRO-Nav: Hybrid Reasoning Enables Efficient Embodied Navigation2026 arXivโ€”
2026AgentVLN: Towards Agentic Vision-and-Language Navigation2026 arXivCode
2026EmergeNav: Structured Embodied Inference for Zero-Shot Vision-and-Language Navigation in Continuous Environments2026 arXivโ€”
2026ABot-N0: Technical Report on the VLA Foundation Model for Versatile Embodied Navigation2026 arXivโ€”
2026VLingNav: Embodied Navigation with Adaptive Reasoning and Visual-Assisted Linguistic Memory2026 arXivโ€”
2025D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation2025 arXivโ€”
2025MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots2025 arXivProject Code
2025CompassNav: Steering from Path Imitation to Decision Understanding in Navigation2025 arXivโ€”
2025AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation2025 arXivCode
2025Nav-r1: Reasoning and navigation in embodied scenes2025 arXivProject Code
2025OctoNav: Towards Generalist Embodied Navigation2025 arXivโ€”
2025Aux-Think: Exploring Reasoning Strategies for Data-Efficient Vision-Language Navigation2025 NeurlPSProject Data
2025NavCoT: Boosting LLM-based Vision-and-Language Navigation via Learning Disentangled Reasoning2025 TPAMIโ€”

Predictive Reasoning Model

Hierarchical Decision

High-Level System

Low-Level System

Hybrid Architecture

โš ๏ธ Reflection from Error

Sources and Taxonomy of Errors in VLN

Instruction Errors

Perceptual Errors

  • Covered by papers already listed above.

Planning and Decision Errors

Error Prevention

Prevent Instruction Errors

Prevent Perception Errors

  • Covered by papers already listed above.

Prevent Planning and Decision Errors

Learn From Error

Distribution Alignment

YearPaperVenueResources
2024Vision-Language Navigation with Energy-Based Policy2024 NeurIPSโ€”

DAgger-based

Exploration-based

Test-time Adaptation & Self Correstion

๐ŸŒ Broader Navigation Literature

Surveys and Evaluation

YearPaperVenueResources
2026A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation2026 IEEE TASEโ€”
2026A comprehensive review of recent advancements in vision-and-language navigation2026 Discover Computingโ€”
2026Robot Navigation via Foundation Language Models: A Review2026 ACM Computing Surveysโ€”
2026Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models2026 arXivโ€”
2026Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap2026 arXivโ€”
2026Mini-BEHAVIOR-Gran: Revealing U-Shaped Effects of Instruction Granularity on Language-Guided Embodied Agents2026 arXivโ€”
2026NavTrust: Benchmarking Trustworthiness for Embodied Navigation2026 arXivProject
2024Vision-and-language navigation today and tomorrow: A survey in the era of foundation models2024 TMLRCode
2024Vision-language navigation: a survey and taxonomy2024 Neural Computing and Applicationsโ€”
2023Advances in embodied navigation using large language models: A survey2023 arXivโ€”
2023A2Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models2023 arXivโ€”
2023Visual language navigation: A survey and open challenges2023 Artificial Intelligence Reviewโ€”
2022Vision-and-language navigation: A survey of tasks, methods, and future directions2022 ACLCode
2019General evaluation for instruction conditioned navigation using dynamic time warping2019 arXivโ€”
2018On evaluation of embodied navigation agents2018 arXivโ€”

General Navigation and Object Navigation

YearPaperVenueResources
2026Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon Navigation2026 arXivโ€”
2026LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation2026 arXivProject Code
2026VoLN: Vision-Only Long-Horizon Navigation---Paradigm, Benchmark, and Method2026 arXivProject Code Data
2026ABot-N1: Toward a General Visual Language Navigation Foundation Model2026 arXivProject
2026Qwen-RobotNav: A Scalable Navigation Model Designed for an Agentic Navigation System2026 arXivProject
2026GN0: Toward a Unified Paradigm for Generation, Evaluation, and Policy Learning in Visual-Language Navigation2026 arXivProject Code Model
2026R2F: Repurposing Ray Frontiers for LLM-free Object Navigation2026 arXivโ€”
2026Hydra-Nav: Object Navigation via Adaptive Dual-Process Reasoning2026 arXivโ€”
2026MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation2026 arXivโ€”
2025RANGER: A Monocular Zero-Shot Semantic Navigation Framework through Contextual Adaptation2025 arXivโ€”
2025SocialNav-Map: Dynamic Mapping with Human Trajectory Prediction for Zero-Shot Social Navigation2025 arXivCode
2025NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions2025 arXivCode
2025C-NAV: Towards self-evolving continual object navigation in open world2025 arXivโ€”
20253D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning2025 arXivโ€”
2024TopV-Nav: Unlocking the top-view spatial reasoning potential of MLLM for zero-shot object navigation2024 arXivโ€”
2023ViNT: A foundation model for visual navigation2023 arXivProject Code
2022GNM: A general navigation model to drive any robot2022 arXivCode

Embodied / VLA Context

Policies, Efficiency, and Applied Navigation

๐Ÿš€ Future Directions

Representative outlook papers explicitly discussed in the manuscript are listed below.

Toward Action-Aware Policies

  • Action grounding beyond static image-text reasoning.

Toward Visual Predictive CoT

  • Reasoning grounded in predicted visual futures.

Towards Unified Visual Geometry Navigation Model

  • Tighter coupling between geometry and policy learning.

Towards Dynamic Environment

Towards Collaborative VLN

waynechu1021/Awesome_Visual_Language_Navigation

End-to-End Visual Language Navigation with Limited Sensing: A Survey

TeX

46

28 commits

updated Sep 17, 2026

See the code

README

๐Ÿงญ Awesome Visual Language Navigation

Awesome list badge BibTeX Paper

๐ŸŒŸ Overview

๐Ÿ“š Preliminaries

Simulation, Data and Control

Toward More Realistic Simulation

Towards More Diverse Datasets&Benchmarks

YearPaperVenueResources
2026HumanoidVLN: A Physics-Grounded Simulator and Benchmark for Vision-Language Navigation Across Diverse Humanoid Embodiments2026 arXivProject
2026DaViNCi: A Dataset Towards Outdoor Vision-and-Language Navigation with Continuous Actions and Dynamic Elements2026 arXivProject
2026Beyond Isolation: A Unified Benchmark for General-Purpose Navigation2026 RSSProject Code
2026MiniVLA-Nav v1: A Multi-Scene Simulation Dataset for Language-Conditioned Robot Navigation2026 arXivโ€”
2026Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration2026 arXivโ€”
2026CoNavBench: Collaborative Long-Horizon Vision-Language Navigation Benchmark2026 ICLRProject
2026Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification2026 arXivโ€”
2026AirNav: A Large-Scale Real-World UAV Vision-and-Language Navigation Dataset with Natural and Diverse Instructions2026 arXivโ€”
2025IndoorUAV: Benchmarking Vision-Language UAV Navigation in Continuous Indoor Environments2025 arXivโ€”
2025VL-LN Bench: Towards Long-horizon Goal-oriented Navigation with Active Dialogs2025 arXivCode Data
2025OpenFly: A Comprehensive Platform for Aerial Vision-Language Navigation2025 arXivCode Data
2025Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent Dialogues2025 ICCVProject Code Data
2025RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation2025 CVPRCode Data
2025Towards long-horizon vision-language navigation: Platform, benchmark and method2025 CVPRCode
2025CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos2025 CVPRCode Data
2025UrbanNav: Learning Language-Guided Urban Navigation from Web-Scale Human Trajectories2025 arXivCode
2025UAV-Flow Colosseo: A Real-World Benchmark for Flying-on-a-Word UAV Imitation Learning2025 arXivโ€”
2024Embodiedcity: A benchmark platform for embodied agent in real-world city environment2024 arXivCode
2024InstruGen: Automatic Instruction Generation for Vision-and-Language Navigation Via Large Multimodal Models2024 arXivCode
2024Towards realistic uav vision-language navigation: Platform, benchmark, and methodology2024 arXivProject Code Data
2024CityNav: Language-goal aerial navigation dataset with geographic information2024 arXivCode
2024Mind the error! detection and localization of instruction errors in vision-and-language navigation2024 IROSProject Code
2024Hazard challenge: Embodied decision making in dynamically changing environments2024 arXivProject
2023AerialVLN: Vision-and-Language Navigation for UAVs2023 ICCVCode
2023Iterative vision-and-language navigation2023 CVPRCode
2023Scaling Data Generation in Vision-and-Language Navigation2023 ICCVโ€”
2022REVE-CE: Remote Embodied Visual Referring Expression in Continuous Environment2022 RA-Lโ€”
2021Talk2Nav: Long-Range Vision-and-Language Navigation with Dual Attention and Spatial Memory2021 IJCVโ€”
2020Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding2020 arXivโ€”
2018Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments2018 CVPRโ€”

Towards More Continuous Control

๐Ÿง  Context Modeling

Context Perception

Spatial Perception

Temporal Perception

Context Memory

Memory Representation

Memory Retrieval and Utilization

Context Efficiency

Reducing Temporal Redundancy

Reducing Spatial Redundancy

Computational Reuse and Acceleration

๐Ÿ”ฎ Imaginative Prediction

Spatial Imagination

Occupancy Prediction

View Completion

3D Representation

Temporal Imagination

Pixel Level

Feature Level

Inverse Dynamics Modeling and Causality

Inverse Dynamics

Causal Learning

๐ŸŽฏ Decision Planning

Structured Decision

Instruction Decomposition

Sub-Goal Formulation

Reasoning-based Decision

LLM as External Reasoner

Integrated Perception-Reasoning-Action Model

YearPaperVenueResources
2026GroundingVLN: Reasoning and Acting with Grounding for Vision-Language Navigation2026 arXivโ€”
2026MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control2026 arXivProject Code
2026ReflectVLN: Training Vision-Language Navigation Agents with Reflective Reasoning2026 arXivCode
2026AwareVLN: Reasoning with Self-awareness for Vision-Language Navigation2026 CVPRProject Code
2026A Deployable Embodied Vision-Language Navigation System with Hierarchical Cognition and Context-Aware Exploration2026 arXivโ€”
2026AsyncShield: A Plug-and-Play Edge Adapter for Asynchronous Cloud-based VLA Navigation2026 arXivโ€”
2026AURA: Multimodal Shared Autonomy for Real-World Urban Navigation2026 arXivโ€”
2026HiRO-Nav: Hybrid Reasoning Enables Efficient Embodied Navigation2026 arXivโ€”
2026AgentVLN: Towards Agentic Vision-and-Language Navigation2026 arXivCode
2026EmergeNav: Structured Embodied Inference for Zero-Shot Vision-and-Language Navigation in Continuous Environments2026 arXivโ€”
2026ABot-N0: Technical Report on the VLA Foundation Model for Versatile Embodied Navigation2026 arXivโ€”
2026VLingNav: Embodied Navigation with Adaptive Reasoning and Visual-Assisted Linguistic Memory2026 arXivโ€”
2025D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation2025 arXivโ€”
2025MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots2025 arXivProject Code
2025CompassNav: Steering from Path Imitation to Decision Understanding in Navigation2025 arXivโ€”
2025AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation2025 arXivCode
2025Nav-r1: Reasoning and navigation in embodied scenes2025 arXivProject Code
2025OctoNav: Towards Generalist Embodied Navigation2025 arXivโ€”
2025Aux-Think: Exploring Reasoning Strategies for Data-Efficient Vision-Language Navigation2025 NeurlPSProject Data
2025NavCoT: Boosting LLM-based Vision-and-Language Navigation via Learning Disentangled Reasoning2025 TPAMIโ€”

Predictive Reasoning Model

Hierarchical Decision

High-Level System

Low-Level System

Hybrid Architecture

โš ๏ธ Reflection from Error

Sources and Taxonomy of Errors in VLN

Instruction Errors

Perceptual Errors

  • Covered by papers already listed above.

Planning and Decision Errors

Error Prevention

Prevent Instruction Errors

Prevent Perception Errors

  • Covered by papers already listed above.

Prevent Planning and Decision Errors

Learn From Error

Distribution Alignment

YearPaperVenueResources
2024Vision-Language Navigation with Energy-Based Policy2024 NeurIPSโ€”

DAgger-based

Exploration-based

Test-time Adaptation & Self Correstion

๐ŸŒ Broader Navigation Literature

Surveys and Evaluation

YearPaperVenueResources
2026A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation2026 IEEE TASEโ€”
2026A comprehensive review of recent advancements in vision-and-language navigation2026 Discover Computingโ€”
2026Robot Navigation via Foundation Language Models: A Review2026 ACM Computing Surveysโ€”
2026Vision-Language Navigation for Aerial Robots: Towards the Era of Large Language Models2026 arXivโ€”
2026Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap2026 arXivโ€”
2026Mini-BEHAVIOR-Gran: Revealing U-Shaped Effects of Instruction Granularity on Language-Guided Embodied Agents2026 arXivโ€”
2026NavTrust: Benchmarking Trustworthiness for Embodied Navigation2026 arXivProject
2024Vision-and-language navigation today and tomorrow: A survey in the era of foundation models2024 TMLRCode
2024Vision-language navigation: a survey and taxonomy2024 Neural Computing and Applicationsโ€”
2023Advances in embodied navigation using large language models: A survey2023 arXivโ€”
2023A2Nav: Action-Aware Zero-Shot Robot Navigation by Exploiting Vision-and-Language Ability of Foundation Models2023 arXivโ€”
2023Visual language navigation: A survey and open challenges2023 Artificial Intelligence Reviewโ€”
2022Vision-and-language navigation: A survey of tasks, methods, and future directions2022 ACLCode
2019General evaluation for instruction conditioned navigation using dynamic time warping2019 arXivโ€”
2018On evaluation of embodied navigation agents2018 arXivโ€”

General Navigation and Object Navigation

YearPaperVenueResources
2026Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon Navigation2026 arXivโ€”
2026LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation2026 arXivProject Code
2026VoLN: Vision-Only Long-Horizon Navigation---Paradigm, Benchmark, and Method2026 arXivProject Code Data
2026ABot-N1: Toward a General Visual Language Navigation Foundation Model2026 arXivProject
2026Qwen-RobotNav: A Scalable Navigation Model Designed for an Agentic Navigation System2026 arXivProject
2026GN0: Toward a Unified Paradigm for Generation, Evaluation, and Policy Learning in Visual-Language Navigation2026 arXivProject Code Model
2026R2F: Repurposing Ray Frontiers for LLM-free Object Navigation2026 arXivโ€”
2026Hydra-Nav: Object Navigation via Adaptive Dual-Process Reasoning2026 arXivโ€”
2026MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation2026 arXivโ€”
2025RANGER: A Monocular Zero-Shot Semantic Navigation Framework through Contextual Adaptation2025 arXivโ€”
2025SocialNav-Map: Dynamic Mapping with Human Trajectory Prediction for Zero-Shot Social Navigation2025 arXivCode
2025NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions2025 arXivCode
2025C-NAV: Towards self-evolving continual object navigation in open world2025 arXivโ€”
20253D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning2025 arXivโ€”
2024TopV-Nav: Unlocking the top-view spatial reasoning potential of MLLM for zero-shot object navigation2024 arXivโ€”
2023ViNT: A foundation model for visual navigation2023 arXivProject Code
2022GNM: A general navigation model to drive any robot2022 arXivCode

Embodied / VLA Context

Policies, Efficiency, and Applied Navigation

๐Ÿš€ Future Directions

Representative outlook papers explicitly discussed in the manuscript are listed below.

Toward Action-Aware Policies

  • Action grounding beyond static image-text reasoning.

Toward Visual Predictive CoT

  • Reasoning grounded in predicted visual futures.

Towards Unified Visual Geometry Navigation Model

  • Tighter coupling between geometry and policy learning.

Towards Dynamic Environment

Towards Collaborative VLN

Languages

TeX

100.0%