inclusionAI/ZwZ-RL-VQA

Dataset

ZwZ-RL-VQA: Region-to-Image Distilled Training Data for Fine-Grained Perception

17

23 commits

1 linked in READMEs

updated May 28, 2026

See the code

README

ZwZ-RL-VQA: Region-to-Image Distilled Training Data for Fine-Grained Perception

This synthetic dataset is generated via Region-to-Image Distillation (R2I) for training multimodal large language models (MLLMs) on fine-grained perception tasks without test-time tool use.

📖 Overview

The Zooming without Zooming (ZwZ) method transforms "zooming" from an inference-time tool into a training-time primitive:

  1. Zoom-in Synthesis: Strong teacher models (Qwen3-VL-235B, GLM-4.5V) generate questions and answers on micro-cropped regions where fine details are unambiguous
  2. Zoom-out Distillation: Region-grounded supervision is distilled back to full images with explicit bounding-box overlays
  3. Single-Pass Inference: Trained models internalize zooming benefits, achieving fine-grained perception in one forward pass

📊 Dataset Statistics

AttributeValue
Total Samples37k
Source ImagesSA-1B, LAION, MetaCLIP, Visual Genome, CC12M, STPLS3D
Image ResolutionMostly > 1000×1000 (high-resolution)
Crop Ratiomostly < 10% of full image area (fine-grained focus)
Question TypesCounting, OCR, Color, Structure, Material, Identification
Consensus Filter>6/8 agreement among teacher ensembles

🏗️ Data Generation Pipeline

Teachers Used

RoleModel
Question GeneratorQwen3-VL-235B-A22B-Instruct
Answer Generator 1Qwen3-VL-235B-A22B-Instruct
Answer Generator 2GLM-4.5V

Quality Control

  • ✅ Consensus Filtering: Only retain QA pairs with >75% teacher agreement (6/8 votes)
  • ✅ Difficulty Filtering: Reject samples that baseline Qwen3-VL-8B answers correctly >50% of the time
  • ✅ Visual Grounding: Bounding boxes overlaid on images to resolve referential ambiguity

📂 Data Structure & Extraction

The image data is provided in multiple split compressed files to ensure reliable downloading.

1. Extract Training Images

After downloading all images.tar.gz.* parts, use the following command to merge and extract them:

cd images/
# Merge split files and extract to the current directory
cat images.tar.gz* | tar -xvf - -C ./

2. Original Data & Synthesis (Optional)

If you are interested in how the training data images.tar.gz.* was synthesized, you can refer to the data synthesis script.

The synthesis process uses the original images. To extract the source data, follow these steps:

cd original_images/
# Merge split files and extract to the current directory
cat original_images.tar.gz* | tar -xvf - -C ./

Once extracted, you can use the script mentioned above to reproduce the dataset from these original images.

🎯 Intended Use

This dataset is designed for:

  • Reinforcement Learning on MLLMs (e.g., with DAPO/GRPO)
  • Research on distilling tool-use capabilities into single-pass models

📈 Training Results

Models trained on this dataset (ZwZ-4B/7B/8B) achieve:

ModelZoomBenchHR-Bench-4KHR-Bench-8KVStar
ZwZ-4B55.7481.7579.5092.67
ZwZ-7B55.6275.3873.2588.48
ZwZ-8B58.1184.3882.0091.10

vs. Qwen3-VL-8B baseline: 37.87 / 78.88 / 74.63 / 86.39

📄 Citation

@article{wei2026zooming,
  title={Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception},
  author={Wei, Lai and He, Liangbo and Lan, Jun and Dong, Lingzhong and Cai, Yutong and Li, Siyuan and Zhu, Huijia and Wang, Weiqiang and Kong, Linghe and Wang, Yue and Zhang, Zhuosheng and Huang, Weiran},
  journal={arXiv preprint arXiv:2602.11858},
  year={2026}
}

📝 License

Apache-2.0 License

fine-grained-perception
multimodal
region-to-image-distillation
vision-language-model

Contributors

langdaohlb

11 commits

WaltonFuture

11 commits

m1ngcheng

1 commits

inclusionAI/ZwZ-RL-VQA

Dataset

ZwZ-RL-VQA: Region-to-Image Distilled Training Data for Fine-Grained Perception

17

23 commits

1 linked in READMEs

updated May 28, 2026

See the code

README

ZwZ-RL-VQA: Region-to-Image Distilled Training Data for Fine-Grained Perception

This synthetic dataset is generated via Region-to-Image Distillation (R2I) for training multimodal large language models (MLLMs) on fine-grained perception tasks without test-time tool use.

📖 Overview

The Zooming without Zooming (ZwZ) method transforms "zooming" from an inference-time tool into a training-time primitive:

  1. Zoom-in Synthesis: Strong teacher models (Qwen3-VL-235B, GLM-4.5V) generate questions and answers on micro-cropped regions where fine details are unambiguous
  2. Zoom-out Distillation: Region-grounded supervision is distilled back to full images with explicit bounding-box overlays
  3. Single-Pass Inference: Trained models internalize zooming benefits, achieving fine-grained perception in one forward pass

📊 Dataset Statistics

AttributeValue
Total Samples37k
Source ImagesSA-1B, LAION, MetaCLIP, Visual Genome, CC12M, STPLS3D
Image ResolutionMostly > 1000×1000 (high-resolution)
Crop Ratiomostly < 10% of full image area (fine-grained focus)
Question TypesCounting, OCR, Color, Structure, Material, Identification
Consensus Filter>6/8 agreement among teacher ensembles

🏗️ Data Generation Pipeline

Teachers Used

RoleModel
Question GeneratorQwen3-VL-235B-A22B-Instruct
Answer Generator 1Qwen3-VL-235B-A22B-Instruct
Answer Generator 2GLM-4.5V

Quality Control

  • ✅ Consensus Filtering: Only retain QA pairs with >75% teacher agreement (6/8 votes)
  • ✅ Difficulty Filtering: Reject samples that baseline Qwen3-VL-8B answers correctly >50% of the time
  • ✅ Visual Grounding: Bounding boxes overlaid on images to resolve referential ambiguity

📂 Data Structure & Extraction

The image data is provided in multiple split compressed files to ensure reliable downloading.

1. Extract Training Images

After downloading all images.tar.gz.* parts, use the following command to merge and extract them:

cd images/
# Merge split files and extract to the current directory
cat images.tar.gz* | tar -xvf - -C ./

2. Original Data & Synthesis (Optional)

If you are interested in how the training data images.tar.gz.* was synthesized, you can refer to the data synthesis script.

The synthesis process uses the original images. To extract the source data, follow these steps:

cd original_images/
# Merge split files and extract to the current directory
cat original_images.tar.gz* | tar -xvf - -C ./

Once extracted, you can use the script mentioned above to reproduce the dataset from these original images.

🎯 Intended Use

This dataset is designed for:

  • Reinforcement Learning on MLLMs (e.g., with DAPO/GRPO)
  • Research on distilling tool-use capabilities into single-pass models

📈 Training Results

Models trained on this dataset (ZwZ-4B/7B/8B) achieve:

ModelZoomBenchHR-Bench-4KHR-Bench-8KVStar
ZwZ-4B55.7481.7579.5092.67
ZwZ-7B55.6275.3873.2588.48
ZwZ-8B58.1184.3882.0091.10

vs. Qwen3-VL-8B baseline: 37.87 / 78.88 / 74.63 / 86.39

📄 Citation

@article{wei2026zooming,
  title={Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perception},
  author={Wei, Lai and He, Liangbo and Lan, Jun and Dong, Lingzhong and Cai, Yutong and Li, Siyuan and Zhu, Huijia and Wang, Weiqiang and Kong, Linghe and Wang, Yue and Zhang, Zhuosheng and Huang, Weiran},
  journal={arXiv preprint arXiv:2602.11858},
  year={2026}
}

📝 License

Apache-2.0 License

fine-grained-perception
multimodal
region-to-image-distillation
vision-language-model

Contributors

langdaohlb

11 commits

WaltonFuture

11 commits

m1ngcheng

1 commits