neulab/PangeaInstruct

Dataset

87

stars

16

commits

1

linked in READMEs

Feb 2, 2025

updated

multilingual
multimodal

README

PangeaInstruct

Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages

๐Ÿ‡ช๐Ÿ‡น ๐Ÿ‡ธ๐Ÿ‡ฆ ๐Ÿ‡ง๐Ÿ‡ฌ ๐Ÿ‡ง๐Ÿ‡ฉ ๐Ÿ‡จ๐Ÿ‡ฟ ๐Ÿ‡ฉ๐Ÿ‡ช ๐Ÿ‡ฌ๐Ÿ‡ท ๐Ÿ‡ฌ๐Ÿ‡ง ๐Ÿ‡บ๐Ÿ‡ธ ๐Ÿ‡ช๐Ÿ‡ธ ๐Ÿ‡ฎ๐Ÿ‡ท ๐Ÿ‡ซ๐Ÿ‡ท ๐Ÿ‡ฎ๐Ÿ‡ช ๐Ÿ‡ฎ๐Ÿ‡ณ ๐Ÿ‡ฎ๐Ÿ‡ฉ ๐Ÿ‡ณ๐Ÿ‡ฌ ๐Ÿ‡ฎ๐Ÿ‡น ๐Ÿ‡ฎ๐Ÿ‡ฑ ๐Ÿ‡ฏ๐Ÿ‡ต ๐Ÿ‡ฎ๐Ÿ‡ฉ ๐Ÿ‡ฐ๐Ÿ‡ท ๐Ÿ‡ณ๐Ÿ‡ฑ ๐Ÿ‡ฒ๐Ÿ‡ณ ๐Ÿ‡ฒ๐Ÿ‡พ ๐Ÿ‡ณ๐Ÿ‡ด ๐Ÿ‡ต๐Ÿ‡ฑ ๐Ÿ‡ต๐Ÿ‡น ๐Ÿ‡ง๐Ÿ‡ท ๐Ÿ‡ท๐Ÿ‡ด ๐Ÿ‡ท๐Ÿ‡บ ๐Ÿ‡ฑ๐Ÿ‡ฐ ๐Ÿ‡ฎ๐Ÿ‡ฉ ๐Ÿ‡ฐ๐Ÿ‡ช ๐Ÿ‡น๐Ÿ‡ฟ ๐Ÿ‡ฑ๐Ÿ‡ฐ ๐Ÿ‡ฎ๐Ÿ‡ณ ๐Ÿ‡ฎ๐Ÿ‡ณ ๐Ÿ‡น๐Ÿ‡ญ ๐Ÿ‡น๐Ÿ‡ท ๐Ÿ‡บ๐Ÿ‡ฆ ๐Ÿ‡ต๐Ÿ‡ฐ ๐Ÿ‡ฎ๐Ÿ‡ณ ๐Ÿ‡ป๐Ÿ‡ณ ๐Ÿ‡จ๐Ÿ‡ณ ๐Ÿ‡น๐Ÿ‡ผ

๐Ÿ  Homepage | ๐Ÿค– Pangea-7B | ๐Ÿ“Š PangeaIns | ๐Ÿงช PangeaBench | ๐Ÿ’ป Github | ๐Ÿ“„ Arxiv | ๐Ÿ“• PDF | ๐Ÿ–ฅ๏ธ Demo

description

This README provides comprehensive details on the PangeaIns dataset, which was utilized during the instruction tuning phase for Pangea-7B.

Description of PangeaIns

PangeaIns is a 6M multilingual multicultural multimodal instruction tuning dataset spanning 39 languages.

PangeaIns Data Source

PangeaIns data path: PangeaIns.json (# samples: 6450624) PangeaIns data source:

Dataset NameDataset Path# Samples
ALLAVA-4Vgeneral/ALLAVA-4V/data.json621327
allava_vflangeneral/allava_vflan/data.json325122
Cambrian737kgeneral/cambrian/data.json736934
ChartQAdoc+chart/ChartQA/data.json28299
Code-Feedbacktext-only/Code-Feedback/data.json20000
doc-vqadoc+chart/doc-vqa/data.json9665
gpt4v-datasetcaption/gpt4v-dataset/data.json10822
GQA-rugeneral/GQA-ru/data.json40000
laion-1M-qacultural/laion-multi-1M/captions-1M-generated-qas-llava.json1028791
laion-300K-captioncultural/laion-multi-1M/laion-300K-caption-llava.json300000
llava-en-zh-300kgeneral/llava-en-zh-300k/data.json50000
LLaVA-Finetunecultural/laion-cultural-150k/laion-cultural-150k.json151072
Llava-JP-Instruct-108Kgeneral/LLaVA-JP-Instruct-108K/data.json108855
llava-med-zh-instruct-60Kgeneral/llava-med-zh-instruct-60k/data.json56649
LLaVA-NeXtgeneral/LLaVA-NeXt-Data/data.json119853
LVIS-Instruct4Vgeneral/LVIS-Instruct4V/data.json222697
MTVQAgeneral/MTVQA/data.json6678
nvlr2-llavageneral/nvlr2-llava/data.json86373
NuminaMath-CoTtext-only/NuminaMath-CoT/data.json100000
OpenHermes-2.5text-only/Openhermes-2.5/data.json399900
palo_multilingual_datasetgeneral/palo_multilingual_dataset/urdu-100k.json99992
ShareGPT-4ogeneral/ShareGPT-4o/data.json57289
ShareGPT4Vgeneral/ShareGPT4V/data.json91021
STAIR-Captionscaption/STAIR-Captions/data.json82783
table-vqadoc+chart/table-vqa/data.json16408
Viet-Doc-VQAdoc+chart/Viet-Doc-VQA/data.json12000
Viet-DOC-VQA-IIdoc+chart/Viet-DOC-VQA-II/data.json14998
Viet-OCR-VQAdoc+chart/Viet-OCR-VQA/data.json30000
Viet-ShareGPT-4o-Text-VQAgeneral/Viet-ShareGPT-4o-Text-VQA/data.json42678
webui_multilingual_ocrocr/webui_multilingual_ocr/data.json300000
translationtranslation/data.json1280328

Applications

PangeaIns was designed specifically for training the Pangea-7B model.

Code Instructions

The dataset follows the LLaVA data format. To retrieve all files from PangeaIns, use the following script:

from huggingface_hub import HfApi, hf_hub_download
import json

# Initialize the API client
api = HfApi()
dataset_name = "neulab/PangeaInstruct"

# Retrieve and download all files in the dataset
files = api.list_repo_files(repo_id=dataset_name, repo_type="dataset")

for file in files:
    hf_hub_download(repo_id=dataset_name, filename=file, repo_type="dataset")
    print(f"File downloaded: {file}")

# Load the complete PangeaIns dataset
with open('PangeaIns.json') as f:
  data = json.load(f)

Please note that image data is provided in compressed formats such as .tar or .zip. After downloading, you may need to extract these files to access the images. For images.tar files, you could untar them by running

tar -xvf images.tar

For images.zip files, you could unzip them by running

unzip images.zip

For some large tar files, we uploaded tar files splitted using the split command, such as split -n 4 -d images.tar part_. For example, in the cultural/laion-multi-1M subset, we splitted the images.tar file into 4 parts, part_00, part_01, part_02, and part_03. In such cases, you would need to first combine the splits and then extract the tar file.

cat part_* > images.tar
tar -xvf images.tar

Each subset within the PangeaIns dataset (e.g., ChartQA) contains a .json file for metadata and a corresponding .tar/.zip file for the images.

Citing the Dataset

BibTeX Citation:

@article{yue2024pangeafullyopenmultilingual,
  title={Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages},
  author={Xiang Yue and Yueqi Song and Akari Asai and Seungone Kim and Jean de Dieu Nyandwi and Simran Khanuja and Anjali Kantharuban and Lintang Sutawika and Sathyanarayanan Ramamoorthy and Graham Neubig},
  year={2024},
  journal={arXiv preprint arXiv:2410.16153},
  url={https://arxiv.org/abs/2410.16153}
}

Contact

Corresponding to: {xyue2,yueqis,gneubig}@cs.cmu.edu

Contributors

yueqis

9 commits

yuexiang96

6 commits

lbourdois

1 commits

neulab/PangeaInstruct

Dataset

87

stars

16

commits

1

linked in READMEs

Feb 2, 2025

updated

multilingual
multimodal

README

PangeaInstruct

Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages

๐Ÿ‡ช๐Ÿ‡น ๐Ÿ‡ธ๐Ÿ‡ฆ ๐Ÿ‡ง๐Ÿ‡ฌ ๐Ÿ‡ง๐Ÿ‡ฉ ๐Ÿ‡จ๐Ÿ‡ฟ ๐Ÿ‡ฉ๐Ÿ‡ช ๐Ÿ‡ฌ๐Ÿ‡ท ๐Ÿ‡ฌ๐Ÿ‡ง ๐Ÿ‡บ๐Ÿ‡ธ ๐Ÿ‡ช๐Ÿ‡ธ ๐Ÿ‡ฎ๐Ÿ‡ท ๐Ÿ‡ซ๐Ÿ‡ท ๐Ÿ‡ฎ๐Ÿ‡ช ๐Ÿ‡ฎ๐Ÿ‡ณ ๐Ÿ‡ฎ๐Ÿ‡ฉ ๐Ÿ‡ณ๐Ÿ‡ฌ ๐Ÿ‡ฎ๐Ÿ‡น ๐Ÿ‡ฎ๐Ÿ‡ฑ ๐Ÿ‡ฏ๐Ÿ‡ต ๐Ÿ‡ฎ๐Ÿ‡ฉ ๐Ÿ‡ฐ๐Ÿ‡ท ๐Ÿ‡ณ๐Ÿ‡ฑ ๐Ÿ‡ฒ๐Ÿ‡ณ ๐Ÿ‡ฒ๐Ÿ‡พ ๐Ÿ‡ณ๐Ÿ‡ด ๐Ÿ‡ต๐Ÿ‡ฑ ๐Ÿ‡ต๐Ÿ‡น ๐Ÿ‡ง๐Ÿ‡ท ๐Ÿ‡ท๐Ÿ‡ด ๐Ÿ‡ท๐Ÿ‡บ ๐Ÿ‡ฑ๐Ÿ‡ฐ ๐Ÿ‡ฎ๐Ÿ‡ฉ ๐Ÿ‡ฐ๐Ÿ‡ช ๐Ÿ‡น๐Ÿ‡ฟ ๐Ÿ‡ฑ๐Ÿ‡ฐ ๐Ÿ‡ฎ๐Ÿ‡ณ ๐Ÿ‡ฎ๐Ÿ‡ณ ๐Ÿ‡น๐Ÿ‡ญ ๐Ÿ‡น๐Ÿ‡ท ๐Ÿ‡บ๐Ÿ‡ฆ ๐Ÿ‡ต๐Ÿ‡ฐ ๐Ÿ‡ฎ๐Ÿ‡ณ ๐Ÿ‡ป๐Ÿ‡ณ ๐Ÿ‡จ๐Ÿ‡ณ ๐Ÿ‡น๐Ÿ‡ผ

๐Ÿ  Homepage | ๐Ÿค– Pangea-7B | ๐Ÿ“Š PangeaIns | ๐Ÿงช PangeaBench | ๐Ÿ’ป Github | ๐Ÿ“„ Arxiv | ๐Ÿ“• PDF | ๐Ÿ–ฅ๏ธ Demo

description

This README provides comprehensive details on the PangeaIns dataset, which was utilized during the instruction tuning phase for Pangea-7B.

Description of PangeaIns

PangeaIns is a 6M multilingual multicultural multimodal instruction tuning dataset spanning 39 languages.

PangeaIns Data Source

PangeaIns data path: PangeaIns.json (# samples: 6450624) PangeaIns data source:

Dataset NameDataset Path# Samples
ALLAVA-4Vgeneral/ALLAVA-4V/data.json621327
allava_vflangeneral/allava_vflan/data.json325122
Cambrian737kgeneral/cambrian/data.json736934
ChartQAdoc+chart/ChartQA/data.json28299
Code-Feedbacktext-only/Code-Feedback/data.json20000
doc-vqadoc+chart/doc-vqa/data.json9665
gpt4v-datasetcaption/gpt4v-dataset/data.json10822
GQA-rugeneral/GQA-ru/data.json40000
laion-1M-qacultural/laion-multi-1M/captions-1M-generated-qas-llava.json1028791
laion-300K-captioncultural/laion-multi-1M/laion-300K-caption-llava.json300000
llava-en-zh-300kgeneral/llava-en-zh-300k/data.json50000
LLaVA-Finetunecultural/laion-cultural-150k/laion-cultural-150k.json151072
Llava-JP-Instruct-108Kgeneral/LLaVA-JP-Instruct-108K/data.json108855
llava-med-zh-instruct-60Kgeneral/llava-med-zh-instruct-60k/data.json56649
LLaVA-NeXtgeneral/LLaVA-NeXt-Data/data.json119853
LVIS-Instruct4Vgeneral/LVIS-Instruct4V/data.json222697
MTVQAgeneral/MTVQA/data.json6678
nvlr2-llavageneral/nvlr2-llava/data.json86373
NuminaMath-CoTtext-only/NuminaMath-CoT/data.json100000
OpenHermes-2.5text-only/Openhermes-2.5/data.json399900
palo_multilingual_datasetgeneral/palo_multilingual_dataset/urdu-100k.json99992
ShareGPT-4ogeneral/ShareGPT-4o/data.json57289
ShareGPT4Vgeneral/ShareGPT4V/data.json91021
STAIR-Captionscaption/STAIR-Captions/data.json82783
table-vqadoc+chart/table-vqa/data.json16408
Viet-Doc-VQAdoc+chart/Viet-Doc-VQA/data.json12000
Viet-DOC-VQA-IIdoc+chart/Viet-DOC-VQA-II/data.json14998
Viet-OCR-VQAdoc+chart/Viet-OCR-VQA/data.json30000
Viet-ShareGPT-4o-Text-VQAgeneral/Viet-ShareGPT-4o-Text-VQA/data.json42678
webui_multilingual_ocrocr/webui_multilingual_ocr/data.json300000
translationtranslation/data.json1280328

Applications

PangeaIns was designed specifically for training the Pangea-7B model.

Code Instructions

The dataset follows the LLaVA data format. To retrieve all files from PangeaIns, use the following script:

from huggingface_hub import HfApi, hf_hub_download
import json

# Initialize the API client
api = HfApi()
dataset_name = "neulab/PangeaInstruct"

# Retrieve and download all files in the dataset
files = api.list_repo_files(repo_id=dataset_name, repo_type="dataset")

for file in files:
    hf_hub_download(repo_id=dataset_name, filename=file, repo_type="dataset")
    print(f"File downloaded: {file}")

# Load the complete PangeaIns dataset
with open('PangeaIns.json') as f:
  data = json.load(f)

Please note that image data is provided in compressed formats such as .tar or .zip. After downloading, you may need to extract these files to access the images. For images.tar files, you could untar them by running

tar -xvf images.tar

For images.zip files, you could unzip them by running

unzip images.zip

For some large tar files, we uploaded tar files splitted using the split command, such as split -n 4 -d images.tar part_. For example, in the cultural/laion-multi-1M subset, we splitted the images.tar file into 4 parts, part_00, part_01, part_02, and part_03. In such cases, you would need to first combine the splits and then extract the tar file.

cat part_* > images.tar
tar -xvf images.tar

Each subset within the PangeaIns dataset (e.g., ChartQA) contains a .json file for metadata and a corresponding .tar/.zip file for the images.

Citing the Dataset

BibTeX Citation:

@article{yue2024pangeafullyopenmultilingual,
  title={Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages},
  author={Xiang Yue and Yueqi Song and Akari Asai and Seungone Kim and Jean de Dieu Nyandwi and Simran Khanuja and Anjali Kantharuban and Lintang Sutawika and Sathyanarayanan Ramamoorthy and Graham Neubig},
  year={2024},
  journal={arXiv preprint arXiv:2410.16153},
  url={https://arxiv.org/abs/2410.16153}
}

Contact

Corresponding to: {xyue2,yueqis,gneubig}@cs.cmu.edu

Contributors

yueqis

9 commits

yuexiang96

6 commits

lbourdois

1 commits