Vision-language conversation in English, Chinese, French, Spanish, Russian, Japanese, Arabic, Hindi, Bengali and Urdu
This repository contains the multilingual, multimodal dataset used to train PALO. The dataset includes 665K English instructions from LLaVA-v1.5 and translations of LLaVA-Instruct-150K into Chinese, French, Spanish, Russian, Japanese, Arabic, Hindi, Bengali, and Urdu, totaling nearly 2.1M instructions. Please refer to the Section # 3.1 of our paper for details.
Please download the images from constituting datasets,
.jpgAfter downloading all of them, organize the data as follows in PALO/data,
βββ coco
β βββ train2017
βββ gqa
β βββ images
βββ ocr_vqa
β βββ images
βββ textvqa
β βββ train_images
βββ vg
βββ VG_100K
βββ VG_100K_2
4 commits
Vision-language conversation in English, Chinese, French, Spanish, Russian, Japanese, Arabic, Hindi, Bengali and Urdu
This repository contains the multilingual, multimodal dataset used to train PALO. The dataset includes 665K English instructions from LLaVA-v1.5 and translations of LLaVA-Instruct-150K into Chinese, French, Spanish, Russian, Japanese, Arabic, Hindi, Bengali, and Urdu, totaling nearly 2.1M instructions. Please refer to the Section # 3.1 of our paper for details.
Please download the images from constituting datasets,
.jpgAfter downloading all of them, organize the data as follows in PALO/data,
βββ coco
β βββ train2017
βββ gqa
β βββ images
βββ ocr_vqa
β βββ images
βββ textvqa
β βββ train_images
βββ vg
βββ VG_100K
βββ VG_100K_2
4 commits