JefferyZhan/Griffon-G-CCMD-8M

Dataset

Griffon-G-CCMD-8M Dataset Card

0

4 commits

1 linked in READMEs

updated Aug 12, 2025

See the code

README

Griffon-G-CCMD-8M Dataset Card

News

** [2025/08/12] ** We are glad to release the data of the Griffon-G CCMD-8M.

** [2025/08/10] ** We are happy to announce that Griffon v2 is accepted to ICCV 2025.

Dataset details

We provide the training data for Stage 2: Paradigm Pre-Adaptation Pre-training and Stage 3: Comprehensive Instruction Tuning. For the Stage 1 Alignment, please follow the official guideline of ShareGPT-4V. We prodive the stage 2 data in the pretrain folder and stage 3 data in the SFT folder. For the alignment of the annotations and the images, please refer to the Github Repo.

Pretrain Data

After downloading the repo, please download the images from the open-source data. The source images include: Object365-2023, COCO(train2017 & train2014), V3Det, Visual Genemo, Flickrs30K Entities.

Instruction Tuning Data

We provide the visual tokenizer training data with both images and annotations. For the others, we provide the processed annotations, and please download the images from the source datasets. As the general instruction data contains millions of images from different sources, we list them below. When the quota is ready, we will provide the used images directly. If any potential image is missing, please contact us and we can provide the missing ones.

  • general_instructions.json: ai2d, allava_laion, allava_vflan, a-okvqa, ChartQA, docvqa, DUE_Benchmark, dvqa, gqa, infoQA, llava_pretrain, ocr_vqa, scienceqa, sam, sharegpt4v, share_textvqa, synthdog-en, textvqa, visual genome, VisualMRC, vizwiz, web-celebrity, web-landmark, and wikiart.
  • CT-datasetv2.tar.gz: contains both the images and annotations.
  • Other annotations: the same as pretrain data images.

License

Attribution-NonCommercial 4.0 International. It should abide by the policy of the original data sources.

Citation

To use this data, please cite

@article{zhan2024griffon-G,
  title={Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models},
  author={Zhan, Yufei and Zhao, Hongyin and Zhu, Yousong and Yang, Fan and Tang, Ming and Wang, Jinqiao},
  journal={arXiv preprint arXiv:2410.16163},
  year={2024}
}

Contributors

JefferyZhan

4 commits

JefferyZhan/Griffon-G-CCMD-8M

Dataset

Griffon-G-CCMD-8M Dataset Card

0

4 commits

1 linked in READMEs

updated Aug 12, 2025

See the code

README

Griffon-G-CCMD-8M Dataset Card

News

** [2025/08/12] ** We are glad to release the data of the Griffon-G CCMD-8M.

** [2025/08/10] ** We are happy to announce that Griffon v2 is accepted to ICCV 2025.

Dataset details

We provide the training data for Stage 2: Paradigm Pre-Adaptation Pre-training and Stage 3: Comprehensive Instruction Tuning. For the Stage 1 Alignment, please follow the official guideline of ShareGPT-4V. We prodive the stage 2 data in the pretrain folder and stage 3 data in the SFT folder. For the alignment of the annotations and the images, please refer to the Github Repo.

Pretrain Data

After downloading the repo, please download the images from the open-source data. The source images include: Object365-2023, COCO(train2017 & train2014), V3Det, Visual Genemo, Flickrs30K Entities.

Instruction Tuning Data

We provide the visual tokenizer training data with both images and annotations. For the others, we provide the processed annotations, and please download the images from the source datasets. As the general instruction data contains millions of images from different sources, we list them below. When the quota is ready, we will provide the used images directly. If any potential image is missing, please contact us and we can provide the missing ones.

  • general_instructions.json: ai2d, allava_laion, allava_vflan, a-okvqa, ChartQA, docvqa, DUE_Benchmark, dvqa, gqa, infoQA, llava_pretrain, ocr_vqa, scienceqa, sam, sharegpt4v, share_textvqa, synthdog-en, textvqa, visual genome, VisualMRC, vizwiz, web-celebrity, web-landmark, and wikiart.
  • CT-datasetv2.tar.gz: contains both the images and annotations.
  • Other annotations: the same as pretrain data images.

License

Attribution-NonCommercial 4.0 International. It should abide by the policy of the original data sources.

Citation

To use this data, please cite

@article{zhan2024griffon-G,
  title={Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models},
  author={Zhan, Yufei and Zhao, Hongyin and Zhu, Yousong and Yang, Fan and Tang, Ming and Wang, Jinqiao},
  journal={arXiv preprint arXiv:2410.16163},
  year={2024}
}

Contributors

JefferyZhan

4 commits