JuntaoTang/TriGap

Dataset

TriGap: Multimodal Continual Instruction Tuning Benchmark

2

9 commits

5 linked in READMEs

updated Sep 15, 2026

See the code

README


TriGap: Multimodal Continual Instruction Tuning Benchmark

TriGap is a challenging benchmark designed for Multimodal Large Language Models (MLLMs) in the context of Continual Instruction Tuning. It comprises instruction data derived from 10 diverse VQA datasets, covering domains such as medical imaging, scientific documents, autonomous driving, and abstract reasoning.

πŸ“š Citation

If you use this benchmark in your research, please cite the following works:

@article{tang2026prism,
  title={Prism: A Plug-in Reproducible Infrastructure for Scalable Multimodal Continual Instruction Tuning},
  author={Tang, Jun-Tao and Shi, Yu-Cheng and Xie, Zhen-Hao and Zhou, Da-Wei},
  journal={arXiv preprint arXiv:2605.26110},
  year={2026}
}

@inproceedings{xie2026same,
  title={SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning},
  author={Xie, Zhen-Hao and Tang, Jun-Tao and Shi, Yu-Cheng and Ye, Han-Jia and Zhan, De-Chuan and Zhou, Da-Wei},
  booktitle={International Conference on Machine Learning (ICML)},
  year={2026}
}

πŸ“₯ Dataset Download

The TriGap benchmark relies on image data from 10 underlying VQA datasets. After downloading the instruction data, you must download the corresponding images from their original sources.

Please refer to the Download Link column below to access the official repositories or project pages for each dataset.

Dataset NamePaper LinkDownload Link (Source)
PMCVQAArXivGitHub
DocVQAArXivProject Page
ChartQAArXivHF
IconQAArXivProject Page
InfographicVQAArXivProject Page
ArxivQAArXivProject Page
RoadsideArXivGitHub
ChemVQA-HF
FloodNetVQAIEEE XploreGitHub
CLEVRArXivGitHub

πŸ“‚ Directory Structure

To ensure compatibility with the TriGap benchmark loader, please organize your downloaded data into the following directory structure.

1. Image Data Structure

Place all downloaded images into an images folder, maintaining the specific sub-directory structure required by each dataset:

images
β”œβ”€β”€ ArxivQA
β”œβ”€β”€ CLEVR
β”‚   β”œβ”€β”€ test
β”‚   β”œβ”€β”€ train
β”‚   └── val
β”œβ”€β”€ ChartQA
β”œβ”€β”€ ChemVQA
β”‚   β”œβ”€β”€ test
β”‚   └── train
β”œβ”€β”€ DocVQA
β”‚   └── spdocvqa_images
β”œβ”€β”€ FloodNetVQA
β”‚   β”œβ”€β”€ test_images
β”‚   β”œβ”€β”€ train_images
β”‚   └── valid_images
β”œβ”€β”€ IconQA
β”‚   β”œβ”€β”€ test
β”‚   β”‚   β”œβ”€β”€ choose_img
β”‚   β”‚   β”œβ”€β”€ choose_txt
β”‚   β”‚   └── fill_in_blank
β”‚   β”œβ”€β”€ train
β”‚   β”‚   β”œβ”€β”€ choose_img
β”‚   β”‚   β”œβ”€β”€ choose_txt
β”‚   β”‚   └── fill_in_blank
β”‚   └── val
β”‚       β”œβ”€β”€ choose_img
β”‚       β”œβ”€β”€ choose_txt
β”‚       └── fill_in_blank
β”œβ”€β”€ InfographicVQA
β”œβ”€β”€ PMCVQA
└── Roadside
    β”œβ”€β”€ train_img
    └── val_img

2. Final Benchmark Structure

The final root directory for the TriGap benchmark should look like this:

TriGap/
β”œβ”€β”€ images/          # Contains all the image sub-folders listed above
└── instructions/    # Contains the TriGap instruction JSON files
continual-learning
instruction-tuning
multimodal
vision-language

Contributors

JuntaoTang

9 commits

JuntaoTang/TriGap

Dataset

TriGap: Multimodal Continual Instruction Tuning Benchmark

2

9 commits

5 linked in READMEs

updated Sep 15, 2026

See the code

README


TriGap: Multimodal Continual Instruction Tuning Benchmark

TriGap is a challenging benchmark designed for Multimodal Large Language Models (MLLMs) in the context of Continual Instruction Tuning. It comprises instruction data derived from 10 diverse VQA datasets, covering domains such as medical imaging, scientific documents, autonomous driving, and abstract reasoning.

πŸ“š Citation

If you use this benchmark in your research, please cite the following works:

@article{tang2026prism,
  title={Prism: A Plug-in Reproducible Infrastructure for Scalable Multimodal Continual Instruction Tuning},
  author={Tang, Jun-Tao and Shi, Yu-Cheng and Xie, Zhen-Hao and Zhou, Da-Wei},
  journal={arXiv preprint arXiv:2605.26110},
  year={2026}
}

@inproceedings{xie2026same,
  title={SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning},
  author={Xie, Zhen-Hao and Tang, Jun-Tao and Shi, Yu-Cheng and Ye, Han-Jia and Zhan, De-Chuan and Zhou, Da-Wei},
  booktitle={International Conference on Machine Learning (ICML)},
  year={2026}
}

πŸ“₯ Dataset Download

The TriGap benchmark relies on image data from 10 underlying VQA datasets. After downloading the instruction data, you must download the corresponding images from their original sources.

Please refer to the Download Link column below to access the official repositories or project pages for each dataset.

Dataset NamePaper LinkDownload Link (Source)
PMCVQAArXivGitHub
DocVQAArXivProject Page
ChartQAArXivHF
IconQAArXivProject Page
InfographicVQAArXivProject Page
ArxivQAArXivProject Page
RoadsideArXivGitHub
ChemVQA-HF
FloodNetVQAIEEE XploreGitHub
CLEVRArXivGitHub

πŸ“‚ Directory Structure

To ensure compatibility with the TriGap benchmark loader, please organize your downloaded data into the following directory structure.

1. Image Data Structure

Place all downloaded images into an images folder, maintaining the specific sub-directory structure required by each dataset:

images
β”œβ”€β”€ ArxivQA
β”œβ”€β”€ CLEVR
β”‚   β”œβ”€β”€ test
β”‚   β”œβ”€β”€ train
β”‚   └── val
β”œβ”€β”€ ChartQA
β”œβ”€β”€ ChemVQA
β”‚   β”œβ”€β”€ test
β”‚   └── train
β”œβ”€β”€ DocVQA
β”‚   └── spdocvqa_images
β”œβ”€β”€ FloodNetVQA
β”‚   β”œβ”€β”€ test_images
β”‚   β”œβ”€β”€ train_images
β”‚   └── valid_images
β”œβ”€β”€ IconQA
β”‚   β”œβ”€β”€ test
β”‚   β”‚   β”œβ”€β”€ choose_img
β”‚   β”‚   β”œβ”€β”€ choose_txt
β”‚   β”‚   └── fill_in_blank
β”‚   β”œβ”€β”€ train
β”‚   β”‚   β”œβ”€β”€ choose_img
β”‚   β”‚   β”œβ”€β”€ choose_txt
β”‚   β”‚   └── fill_in_blank
β”‚   └── val
β”‚       β”œβ”€β”€ choose_img
β”‚       β”œβ”€β”€ choose_txt
β”‚       └── fill_in_blank
β”œβ”€β”€ InfographicVQA
β”œβ”€β”€ PMCVQA
└── Roadside
    β”œβ”€β”€ train_img
    └── val_img

2. Final Benchmark Structure

The final root directory for the TriGap benchmark should look like this:

TriGap/
β”œβ”€β”€ images/          # Contains all the image sub-folders listed above
└── instructions/    # Contains the TriGap instruction JSON files
continual-learning
instruction-tuning
multimodal
vision-language

Contributors

JuntaoTang

9 commits