[ICLR 2026 Oral] Visual Planning: Let's Think Only with Images
Python
377
41 commits
updated Apr 24, 2026
We introduce Visual Planning, a new reasoning paradigm where planning is conducted entirely through sequences of images, without relying on language. Unlike traditional multimodal models that use visual input but still reason in text, our approach enables models to "think" directly in the visual domain. We propose a reinforcement learning framework, VPRL, which significantly outperforms language-based baselines on spatial navigation tasks.
We propose a novel two-stage reinforcement learning training framework:
We release the following model checkpoints on Hugging Face:
| Environment | Checkpoint |
|---|---|
| MiniBehaviour | VPRL-7B-MiniBehaviour |
| Maze | VPRL-7B-Maze |
| FrozenLake | VPRL-7B-FrozenLake |
Please first create a conda environment:
conda create -n visualplanning python=3.12.3
conda activate visualplanning
Then run:
bash scripts/install.sh
VPFT corresponds to supervised fine-tuning on optimal trajectories:
bash scripts/sft_optimal.sh frozenlake
bash scripts/sft_optimal.sh maze
bash scripts/sft_optimal.sh minibehaviour
Stage 1 performs policy initialization with random trajectory supervision:
bash scripts/sft_random.sh frozenlake
bash scripts/sft_random.sh maze
bash scripts/sft_random.sh minibehaviour
Stage 2 performs reinforcement learning with GRPO:
bash scripts/grpo.sh frozenlake
bash scripts/grpo.sh maze
bash scripts/grpo.sh minibehaviour
We evaluate VPRL across three diverse visual planning environments:
If you find Visual Planning useful for your research and applications, please cite using this BibTeX:
@misc{xu2025visualplanningletsthink,
title={Visual Planning: Let's Think Only with Images},
author={Yi Xu and Chengzu Li and Han Zhou and Xingchen Wan and Caiqi Zhang and Anna Korhonen and Ivan Vulić},
year={2025},
eprint={2505.11409},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2505.11409},
}
4 followers · starred May 2025
[ICLR 2026 Oral] Visual Planning: Let's Think Only with Images
Python
377
41 commits
updated Apr 24, 2026
We introduce Visual Planning, a new reasoning paradigm where planning is conducted entirely through sequences of images, without relying on language. Unlike traditional multimodal models that use visual input but still reason in text, our approach enables models to "think" directly in the visual domain. We propose a reinforcement learning framework, VPRL, which significantly outperforms language-based baselines on spatial navigation tasks.
We propose a novel two-stage reinforcement learning training framework:
We release the following model checkpoints on Hugging Face:
| Environment | Checkpoint |
|---|---|
| MiniBehaviour | VPRL-7B-MiniBehaviour |
| Maze | VPRL-7B-Maze |
| FrozenLake | VPRL-7B-FrozenLake |
Please first create a conda environment:
conda create -n visualplanning python=3.12.3
conda activate visualplanning
Then run:
bash scripts/install.sh
VPFT corresponds to supervised fine-tuning on optimal trajectories:
bash scripts/sft_optimal.sh frozenlake
bash scripts/sft_optimal.sh maze
bash scripts/sft_optimal.sh minibehaviour
Stage 1 performs policy initialization with random trajectory supervision:
bash scripts/sft_random.sh frozenlake
bash scripts/sft_random.sh maze
bash scripts/sft_random.sh minibehaviour
Stage 2 performs reinforcement learning with GRPO:
bash scripts/grpo.sh frozenlake
bash scripts/grpo.sh maze
bash scripts/grpo.sh minibehaviour
We evaluate VPRL across three diverse visual planning environments:
If you find Visual Planning useful for your research and applications, please cite using this BibTeX:
@misc{xu2025visualplanningletsthink,
title={Visual Planning: Let's Think Only with Images},
author={Yi Xu and Chengzu Li and Han Zhou and Xingchen Wan and Caiqi Zhang and Anna Korhonen and Ivan Vulić},
year={2025},
eprint={2505.11409},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2505.11409},
}
4 followers · starred May 2025