Latent Sketchpad – Pretrained Sketch Decoder
2
26 commits
1 linked in READMEs
updated Nov 5, 2025
Latent Sketchpad extends pretrained MLLMs to interleave text and visual latent generation,
allowing them to form visual thoughts as part of the reasoning process.
The framework equips the pretrained MLLM with a Vision Head to generate visual latents autoregressively during inference. A separately pretrained Sketch Decoder visualizes these latents into interpretable sketches.
This repository provides the pretrained sketch decoder component of Latent Sketchpad, a framework that equips Multimodal Large Language Models (MLLMs) with an internal visual scratchpad for visual reasoning and imagination.
For more details, please visit our project homepage and paper.
To use the pretrained Sketch Decoder, please refer to the Code Repository 📘.
The repository contains detailed installation steps, example scripts, and guidance on how to visualize visual latents with the decoder.
You can run the provided demo script from the GitHub repository to visualize the pretrained Sketch Decoder reconstruction pipeline:
python sketch_decoder/app.py \
--vision_model "google/gemma-3-12b-it" \
--checkpoint_path "path/to/decoder" \
--feature_dim 1152
You can select the appropriate checkpoint for your chosen vision model and specify it via --checkpoint_path.
The --feature_dim should match the output dimension of the corresponding vision encoder (e.g., Gemma3: 1152, OpenCLIP: 1024, Qwen2.5-VL: 1280).
This repository provides pretrained Sketch Decoder checkpoints adapted for different vision encoders:
sketch_decoder_gemma3.ckpt – compatible with Gemma3 vision encodersketch_decoder_openclip.ckpt – compatible with OpenCLIP vision encodersketch_decoder_qwen25_vl.ckpt – compatible with Qwen2.5-VL vision encoderWe would like to thank all collaborators and contributors involved in the development of Latent Sketchpad.
If you use this model or the Latent Sketchpad framework in your research, please cite:
@article{zhang2025latentsketchpad,
title={Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs},
author={Zhang, Huanyu and Wu, Wenshan and Li, Chengzu and Shang, Ning and Xia, Yan and Huang, Yangyu and Zhang, Yifan and Dong, Li and Zhang, Zhang and Wang, Liang and Tan, Tieniu and Wei, Furu},
journal={arXiv preprint arXiv:2510.24514},
year={2025}
}
Latent Sketchpad – Pretrained Sketch Decoder
2
26 commits
1 linked in READMEs
updated Nov 5, 2025
Latent Sketchpad extends pretrained MLLMs to interleave text and visual latent generation,
allowing them to form visual thoughts as part of the reasoning process.
The framework equips the pretrained MLLM with a Vision Head to generate visual latents autoregressively during inference. A separately pretrained Sketch Decoder visualizes these latents into interpretable sketches.
This repository provides the pretrained sketch decoder component of Latent Sketchpad, a framework that equips Multimodal Large Language Models (MLLMs) with an internal visual scratchpad for visual reasoning and imagination.
For more details, please visit our project homepage and paper.
To use the pretrained Sketch Decoder, please refer to the Code Repository 📘.
The repository contains detailed installation steps, example scripts, and guidance on how to visualize visual latents with the decoder.
You can run the provided demo script from the GitHub repository to visualize the pretrained Sketch Decoder reconstruction pipeline:
python sketch_decoder/app.py \
--vision_model "google/gemma-3-12b-it" \
--checkpoint_path "path/to/decoder" \
--feature_dim 1152
You can select the appropriate checkpoint for your chosen vision model and specify it via --checkpoint_path.
The --feature_dim should match the output dimension of the corresponding vision encoder (e.g., Gemma3: 1152, OpenCLIP: 1024, Qwen2.5-VL: 1280).
This repository provides pretrained Sketch Decoder checkpoints adapted for different vision encoders:
sketch_decoder_gemma3.ckpt – compatible with Gemma3 vision encodersketch_decoder_openclip.ckpt – compatible with OpenCLIP vision encodersketch_decoder_qwen25_vl.ckpt – compatible with Qwen2.5-VL vision encoderWe would like to thank all collaborators and contributors involved in the development of Latent Sketchpad.
If you use this model or the Latent Sketchpad framework in your research, please cite:
@article{zhang2025latentsketchpad,
title={Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs},
author={Zhang, Huanyu and Wu, Wenshan and Li, Chengzu and Shang, Ning and Xia, Yan and Huang, Yangyu and Zhang, Yifan and Dong, Li and Zhang, Zhang and Wang, Liang and Tan, Tieniu and Wei, Furu},
journal={arXiv preprint arXiv:2510.24514},
year={2025}
}