End-to-End Visual Language Navigation with Limited Sensing: A Survey
TeX
46
28 commits
updated Sep 17, 2026
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | ActiveVLN: Towards Active Exploration via Multi-Turn RL in Vision-and-Language Navigation | 2026 ICRA | |
| 2026 | StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling | 2026 ICRA | |
| 2025 | NaVILA: Legged Robot Vision-Language-Action Model for Navigation | 2025 RSS | |
| 2021 | VLN-BERT: A Recurrent Vision-and-Language BERT for Navigation | 2021 CVPR |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | DecoVLN: Decoupling Observation, Reasoning, and Correction for Vision-and-Language Navigation | 2026 arXiv | โ |
| 2025 | RANGER: A Monocular Zero-Shot Semantic Navigation Framework through Contextual Adaptation | 2025 arXiv | โ |
| 2025 | C-NAV: Towards self-evolving continual object navigation in open world | 2025 arXiv | โ |
| 2025 | AstraNav-Memory: Contexts Compression for Long Memory | 2025 arXiv | โ |
| 2025 | COSMO: Combination of Selective Memorization for Low-cost Vision-and-Language Navigation | 2025 ICCV | โ |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | Structured Observation Language for Efficient and Generalizable Vision-Language Navigation | 2026 arXiv | โ |
| 2025 | Embodied Navigation Foundation Model | 2025 arXiv | |
| 2025 | Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks | 2025 RSS |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | FreqCache: Accelerating Embodied VLN Models with Adaptive Frequency-Guided Token Caching | 2026 arXiv | โ |
| 2026 | VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness | 2026 arXiv | โ |
| 2026 | History-Conditioned Spatio-Temporal Visual Token Pruning for Efficient Vision-Language Navigation | 2026 arXiv | โ |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | MapDream: Task-Driven Map Learning for Vision-Language Navigation | 2026 arXiv | โ |
| 2026 | MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming | 2026 AAAI | |
| 2025 | PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation | 2025 Neural Networks | โ |
| 2024 | Imagine Before Go: Self-Supervised Generative Map for Object Goal Navigation | 2024 CVPR | |
| 2023 | PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language Navigation | 2023 NeurIPS |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN | 2026 arXiv | โ |
| 2026 | NavDreamer: Video Models as Zero-Shot 3D Navigators | 2026 arXiv | โ |
| 2026 | Sparse Video Generation Propels Real-World Beyond-the-View Vision-Language Navigation | 2026 arXiv | |
| 2025 | AstraNav-World: World Model for Foresight Control and Consistency | 2025 arXiv | โ |
| 2025 | VISTAv2: World Imagination for Indoor Vision-and-Language Navigation | 2025 arXiv | โ |
| 2025 | DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation | 2025 arXiv | โ |
| 2025 | VISTA: Generative Visual Imagination for Vision-and-Language Navigation | 2025 arXiv | โ |
| 2025 | Navigation World Models | 2025 CVPR | โ |
| 2025 | Do Visual Imaginations Improve Vision-and-Language Navigation Agents? | 2025 CVPR | โ |
| 2021 | PathDreamer: A World Model for Indoor Navigation | 2021 ICCV | โ |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2025 | GC-VLN: Instruction as graph constraints for training-free vision-and-language navigation | 2025 arXiv | โ |
| 2025 | Constraint-aware zero-shot vision-language navigation in continuous environments | 2025 TPAMI | โ |
| 2024 | Boosting efficient reinforcement learning for vision-and-language navigation with open-sourced llm | 2024 RA-L | โ |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | VLN-NF: Feasibility-Aware Vision-and-Language Navigation with False-Premise Instructions | 2026 arXiv | |
| 2024 | Mind the error! detection and localization of instruction errors in vision-and-language navigation | 2024 IROS |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | Where Did It Go Wrong? Capability-Oriented Failure Attribution for Vision-and-Language Navigation Agents | 2026 arXiv | โ |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2025 | Seeing with Partial Certainty: Conformal Prediction for Robotic Scene Recognition in Built Environments | 2025 arXiv | โ |
| 2025 | Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent Dialogues | 2025 ICCV | |
| 2024 | I2EDL: Interactive Instruction Error Detection and Localization | 2024 RO-MAN | โ |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical Planning | 2026 arXiv | โ |
| 2026 | EvolveNav: Empowering LLM-Based Vision-Language Navigation via Self-Improving Embodied Reasoning | 2026 TPAMI |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2024 | Vision-Language Navigation with Energy-Based Policy | 2024 NeurIPS | โ |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2025 | MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation | 2025 arXiv | โ |
| 2024 | ฯ0: A Vision-Language-Action Flow Model for General Robot Control | 2024 arXiv | โ |
| 2024 | OpenVLA: An Open-Source Vision-Language-Action Model | 2024 arXiv |
Representative outlook papers explicitly discussed in the manuscript are listed below.
TeX
100.0%
End-to-End Visual Language Navigation with Limited Sensing: A Survey
TeX
46
28 commits
updated Sep 17, 2026
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | ActiveVLN: Towards Active Exploration via Multi-Turn RL in Vision-and-Language Navigation | 2026 ICRA | |
| 2026 | StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling | 2026 ICRA | |
| 2025 | NaVILA: Legged Robot Vision-Language-Action Model for Navigation | 2025 RSS | |
| 2021 | VLN-BERT: A Recurrent Vision-and-Language BERT for Navigation | 2021 CVPR |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | DecoVLN: Decoupling Observation, Reasoning, and Correction for Vision-and-Language Navigation | 2026 arXiv | โ |
| 2025 | RANGER: A Monocular Zero-Shot Semantic Navigation Framework through Contextual Adaptation | 2025 arXiv | โ |
| 2025 | C-NAV: Towards self-evolving continual object navigation in open world | 2025 arXiv | โ |
| 2025 | AstraNav-Memory: Contexts Compression for Long Memory | 2025 arXiv | โ |
| 2025 | COSMO: Combination of Selective Memorization for Low-cost Vision-and-Language Navigation | 2025 ICCV | โ |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | Structured Observation Language for Efficient and Generalizable Vision-Language Navigation | 2026 arXiv | โ |
| 2025 | Embodied Navigation Foundation Model | 2025 arXiv | |
| 2025 | Uni-NaVid: A Video-based Vision-Language-Action Model for Unifying Embodied Navigation Tasks | 2025 RSS |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | FreqCache: Accelerating Embodied VLN Models with Adaptive Frequency-Guided Token Caching | 2026 arXiv | โ |
| 2026 | VLN-Cache: Enabling Token Caching for VLN Models with Visual/Semantic Dynamics Awareness | 2026 arXiv | โ |
| 2026 | History-Conditioned Spatio-Temporal Visual Token Pruning for Efficient Vision-Language Navigation | 2026 arXiv | โ |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | MapDream: Task-Driven Map Learning for Vision-Language Navigation | 2026 arXiv | โ |
| 2026 | MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming | 2026 AAAI | |
| 2025 | PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation | 2025 Neural Networks | โ |
| 2024 | Imagine Before Go: Self-Supervised Generative Map for Object Goal Navigation | 2024 CVPR | |
| 2023 | PanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language Navigation | 2023 NeurIPS |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN | 2026 arXiv | โ |
| 2026 | NavDreamer: Video Models as Zero-Shot 3D Navigators | 2026 arXiv | โ |
| 2026 | Sparse Video Generation Propels Real-World Beyond-the-View Vision-Language Navigation | 2026 arXiv | |
| 2025 | AstraNav-World: World Model for Foresight Control and Consistency | 2025 arXiv | โ |
| 2025 | VISTAv2: World Imagination for Indoor Vision-and-Language Navigation | 2025 arXiv | โ |
| 2025 | DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation | 2025 arXiv | โ |
| 2025 | VISTA: Generative Visual Imagination for Vision-and-Language Navigation | 2025 arXiv | โ |
| 2025 | Navigation World Models | 2025 CVPR | โ |
| 2025 | Do Visual Imaginations Improve Vision-and-Language Navigation Agents? | 2025 CVPR | โ |
| 2021 | PathDreamer: A World Model for Indoor Navigation | 2021 ICCV | โ |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2025 | GC-VLN: Instruction as graph constraints for training-free vision-and-language navigation | 2025 arXiv | โ |
| 2025 | Constraint-aware zero-shot vision-language navigation in continuous environments | 2025 TPAMI | โ |
| 2024 | Boosting efficient reinforcement learning for vision-and-language navigation with open-sourced llm | 2024 RA-L | โ |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | VLN-NF: Feasibility-Aware Vision-and-Language Navigation with False-Premise Instructions | 2026 arXiv | |
| 2024 | Mind the error! detection and localization of instruction errors in vision-and-language navigation | 2024 IROS |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | Where Did It Go Wrong? Capability-Oriented Failure Attribution for Vision-and-Language Navigation Agents | 2026 arXiv | โ |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2025 | Seeing with Partial Certainty: Conformal Prediction for Robotic Scene Recognition in Built Environments | 2025 arXiv | โ |
| 2025 | Collaborative Instance Object Navigation: Leveraging Uncertainty-Awareness to Minimize Human-Agent Dialogues | 2025 ICCV | |
| 2024 | I2EDL: Interactive Instruction Error Detection and Localization | 2024 RO-MAN | โ |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2026 | SeqWalker: Sequential-Horizon Vision-and-Language Navigation with Hierarchical Planning | 2026 arXiv | โ |
| 2026 | EvolveNav: Empowering LLM-Based Vision-Language Navigation via Self-Improving Embodied Reasoning | 2026 TPAMI |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2024 | Vision-Language Navigation with Energy-Based Policy | 2024 NeurIPS | โ |
| Year | Paper | Venue | Resources |
|---|---|---|---|
| 2025 | MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation | 2025 arXiv | โ |
| 2024 | ฯ0: A Vision-Language-Action Flow Model for General Robot Control | 2024 arXiv | โ |
| 2024 | OpenVLA: An Open-Source Vision-Language-Action Model | 2024 arXiv |
Representative outlook papers explicitly discussed in the manuscript are listed below.
TeX
100.0%