2026.2.1 🚀 The test set of our Mazeplanning dataset is now available on Huggingface.2025.10.29 🚀 Our paper Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs is now available on arXiv together with the code release. Latent Sketchpad extends frontier MLLMs (e.g., Gemma3 and Qwen2.5-VL) to interleave text and visual latents generation, incorporating visual thoughts directly into reasoning.
Multimodal Large Language Models (MLLMs) excel at visual understanding, but they face key limitations:
Latent Sketchpad extends frontier MLLMs (e.g., Gemma3, Qwen2.5-VL) with an internal visual scratchpad, enabling models to interleave textual reasoning and visual latent generation within the autoregressive process.
# Clone this repository
git clone https://github.com/hwanyu112/Latent-Sketchpad.git
cd Latent-Sketchpad
# (Optional) Create and activate a virtual environment
conda create -n sketchpad python=3.10 -y
conda activate sketchpad
# Install dependencies
# Note: Please select the appropriate file according to the model you plan to use.
# Gemma3
bash scripts/setup_gemma.sh
# or Qwen2.5-VL
# bash scripts/setup_qwen.sh
We provide training scripts for different backbone MLLMs (e.g., Qwen2.5-VL, Gemma3).
Please make sure to first configure the environment according to Installation.
Before running any training script, ensure that all parameter settings in the script are correctly configured according to your experimental setup (e.g., dataset paths, model names, and checkpoints).
To enable language-output-only pretrained MLLMs to interleave text and visual generatrion, the training process is divided into two stages:
Stage 1 (training_stage1.sh)
Downstream task fine-tuning.
# Stage 1 Training
bash scripts/training_stage1.sh
Stage 2 (training_stage2.sh)
Vision Head training
# Stage 2 Training
bash scripts/training_stage2.sh
We provide evaluation scripts to test Latent Sketchpad on our MazePlanning dataset.
# Evaluate a trained checkpoint on MazePlanning dataset
export GENERATION_TYPE=multimodal # replace with 'text_only' for text-only generation
python evaluate.py --model_path /path/to/model \
--decoder_path /path/to/sketch_decoder.ckpt \
--data_path /path/to/test_data.json \
--image_folder imgs/ \
--output_dir /path/to/output_dir
We also provide an inference script.
export GENERATION_TYPE=multimodal # replace with 'text_only' for text-only generation
python inference.py --model_path /path/to/model \
--decoder_path /path/to/sketch_decoder.ckpt \
--data_path /path/to/test_data.json \
--image_folder imgs/ \
--output_dir /path/to/output_dir
Install dependencies for the Sketch Decoder separately (to avoid conflicts with the MLLM environment):
# Install dependencies
bash decoder/setup.sh
We provide training scripts for pretraining the Sketch Decoder.
Before launching the training, please refer to the configuration file
decoder/configs/vit-only-gemma3-12-224-40G.json
and adjust the settings as needed.
In particular:
"number_per_class": -1 (control the number of samples from each class)"cate_num": -1 (control the number of classes)The default value -1 means that all available samples from the QuickDraw dataset will be used for training.
If you plan to train the Sketch Decoder with a vision encoder other than OpenCLIP, Gemma3, or Qwen2.5-VL,
please modify the implementation in decoder/vision_encoder_wrapper.py
and update the corresponding "input_dim" field in the configuration file to match the feature dimension of your chosen encoder.
After configuring the file, run the following commands to start training:
# Train the Sketch Decoder from scratch
cd decoder/
bash run.sh
We provide pretrained Sketch Decoder weights on Hugging Face — you are welcome to download them before running the demo:
The repository includes pretrained checkpoints corresponding to three different vision encoders from:
You can then launch the demo script to visualize the pretrained Sketch Decoder reconstruction pipeline:
python sketch_decoder/app.py \
--vision_model "google/gemma-3-12b-it" \
--checkpoint_path "path/to/decoder" \
--feature_dim 1152
You can select the appropriate checkpoint for your chosen vision model and specify it via --checkpoint_path.
The --feature_dim should match the output dimension of the corresponding vision encoder (e.g., Gemma3: 1152, OpenCLIP: 1024, Qwen2.5-VL: 1280).
The demo works as follows:
Below we show the reconstruction results of the same input image using different vision encoders.
These results highlight that the pretrained Sketch Decoder not only ensures broad compatibility across diverse vision encoders, but also demonstrates strong generalization.
We provide the script 4o_ls.py for running the GPT-4o + Latent Sketchpad setup. You can execute it directly using:
export GENERATION_TYPE=agent
export OPENAI_API_KEY=YOUR_API_KEY
python 4o_ls.py --model_path /path/to/latent_sketchpad \
--decoder_path /path/to/sketch_decoder_qwen25_vl.ckpt \
--data_path /path/to/data.json \
--image_folder ls_imgs/ \
--output_dir /path/to/output
The MazePlanning dataset is specifically designed to evaluate interleaved multimodal reasoning.
It provides complex mazes paired with step-by-step interleaved text-and-image reasoning sequences, allowing models to demonstrate planning and imagination beyond static perception.
Training Set:
Evaluation Sets:
📢 Availability:
The dataset is coming soon. Stay tuned for the official release! 🚀
Latent Sketchpad enables different MLLMs to perform interleaved text-and-visual reasoning, extending their capabilities beyond text-only deliberation.
👉 For full multimodal showcase videos with interleaved reasoning, please visit our 🌐 project page.
If you find our work helpful for your research, please consider citing our work.
@article{zhang2025latentsketchpad,
title={Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs},
author={Zhang, Huanyu and Wu, Wenshan and Li, Chengzu and Shang, Ning and Xia, Yan and Huang, Yangyu and Zhang, Yifan and Dong, Li and Zhang, Zhang and Wang, Liang and Tan, Tieniu and Wei, Furu},
journal={arXiv preprint arXiv:2510.24514},
year={2025}
}
Explore our related researches:
23 followers · starred Oct 2025
Python
98.0%
Shell
2.0%
2026.2.1 🚀 The test set of our Mazeplanning dataset is now available on Huggingface.2025.10.29 🚀 Our paper Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs is now available on arXiv together with the code release. Latent Sketchpad extends frontier MLLMs (e.g., Gemma3 and Qwen2.5-VL) to interleave text and visual latents generation, incorporating visual thoughts directly into reasoning.
Multimodal Large Language Models (MLLMs) excel at visual understanding, but they face key limitations:
Latent Sketchpad extends frontier MLLMs (e.g., Gemma3, Qwen2.5-VL) with an internal visual scratchpad, enabling models to interleave textual reasoning and visual latent generation within the autoregressive process.
# Clone this repository
git clone https://github.com/hwanyu112/Latent-Sketchpad.git
cd Latent-Sketchpad
# (Optional) Create and activate a virtual environment
conda create -n sketchpad python=3.10 -y
conda activate sketchpad
# Install dependencies
# Note: Please select the appropriate file according to the model you plan to use.
# Gemma3
bash scripts/setup_gemma.sh
# or Qwen2.5-VL
# bash scripts/setup_qwen.sh
We provide training scripts for different backbone MLLMs (e.g., Qwen2.5-VL, Gemma3).
Please make sure to first configure the environment according to Installation.
Before running any training script, ensure that all parameter settings in the script are correctly configured according to your experimental setup (e.g., dataset paths, model names, and checkpoints).
To enable language-output-only pretrained MLLMs to interleave text and visual generatrion, the training process is divided into two stages:
Stage 1 (training_stage1.sh)
Downstream task fine-tuning.
# Stage 1 Training
bash scripts/training_stage1.sh
Stage 2 (training_stage2.sh)
Vision Head training
# Stage 2 Training
bash scripts/training_stage2.sh
We provide evaluation scripts to test Latent Sketchpad on our MazePlanning dataset.
# Evaluate a trained checkpoint on MazePlanning dataset
export GENERATION_TYPE=multimodal # replace with 'text_only' for text-only generation
python evaluate.py --model_path /path/to/model \
--decoder_path /path/to/sketch_decoder.ckpt \
--data_path /path/to/test_data.json \
--image_folder imgs/ \
--output_dir /path/to/output_dir
We also provide an inference script.
export GENERATION_TYPE=multimodal # replace with 'text_only' for text-only generation
python inference.py --model_path /path/to/model \
--decoder_path /path/to/sketch_decoder.ckpt \
--data_path /path/to/test_data.json \
--image_folder imgs/ \
--output_dir /path/to/output_dir
Install dependencies for the Sketch Decoder separately (to avoid conflicts with the MLLM environment):
# Install dependencies
bash decoder/setup.sh
We provide training scripts for pretraining the Sketch Decoder.
Before launching the training, please refer to the configuration file
decoder/configs/vit-only-gemma3-12-224-40G.json
and adjust the settings as needed.
In particular:
"number_per_class": -1 (control the number of samples from each class)"cate_num": -1 (control the number of classes)The default value -1 means that all available samples from the QuickDraw dataset will be used for training.
If you plan to train the Sketch Decoder with a vision encoder other than OpenCLIP, Gemma3, or Qwen2.5-VL,
please modify the implementation in decoder/vision_encoder_wrapper.py
and update the corresponding "input_dim" field in the configuration file to match the feature dimension of your chosen encoder.
After configuring the file, run the following commands to start training:
# Train the Sketch Decoder from scratch
cd decoder/
bash run.sh
We provide pretrained Sketch Decoder weights on Hugging Face — you are welcome to download them before running the demo:
The repository includes pretrained checkpoints corresponding to three different vision encoders from:
You can then launch the demo script to visualize the pretrained Sketch Decoder reconstruction pipeline:
python sketch_decoder/app.py \
--vision_model "google/gemma-3-12b-it" \
--checkpoint_path "path/to/decoder" \
--feature_dim 1152
You can select the appropriate checkpoint for your chosen vision model and specify it via --checkpoint_path.
The --feature_dim should match the output dimension of the corresponding vision encoder (e.g., Gemma3: 1152, OpenCLIP: 1024, Qwen2.5-VL: 1280).
The demo works as follows:
Below we show the reconstruction results of the same input image using different vision encoders.
These results highlight that the pretrained Sketch Decoder not only ensures broad compatibility across diverse vision encoders, but also demonstrates strong generalization.
We provide the script 4o_ls.py for running the GPT-4o + Latent Sketchpad setup. You can execute it directly using:
export GENERATION_TYPE=agent
export OPENAI_API_KEY=YOUR_API_KEY
python 4o_ls.py --model_path /path/to/latent_sketchpad \
--decoder_path /path/to/sketch_decoder_qwen25_vl.ckpt \
--data_path /path/to/data.json \
--image_folder ls_imgs/ \
--output_dir /path/to/output
The MazePlanning dataset is specifically designed to evaluate interleaved multimodal reasoning.
It provides complex mazes paired with step-by-step interleaved text-and-image reasoning sequences, allowing models to demonstrate planning and imagination beyond static perception.
Training Set:
Evaluation Sets:
📢 Availability:
The dataset is coming soon. Stay tuned for the official release! 🚀
Latent Sketchpad enables different MLLMs to perform interleaved text-and-visual reasoning, extending their capabilities beyond text-only deliberation.
👉 For full multimodal showcase videos with interleaved reasoning, please visit our 🌐 project page.
If you find our work helpful for your research, please consider citing our work.
@article{zhang2025latentsketchpad,
title={Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs},
author={Zhang, Huanyu and Wu, Wenshan and Li, Chengzu and Shang, Ning and Xia, Yan and Huang, Yangyu and Zhang, Yifan and Dong, Li and Zhang, Zhang and Wang, Liang and Tan, Tieniu and Wei, Furu},
journal={arXiv preprint arXiv:2510.24514},
year={2025}
}
Explore our related researches:
23 followers · starred Oct 2025
Python
98.0%
Shell
2.0%