Muennighoff/xP3x

Dataset

95

stars

342

commits

2

linked in READMEs

May 23, 2025

updated

Browse cluster: Multilingual NLP Datasets and Corpora

README

Dataset Card for xP3x

Table of Contents

Dataset Description

Dataset Summary

xP3x (Crosslingual Public Pool of Prompts eXtended) is a collection of prompts & datasets across 277 languages & 16 NLP tasks. It contains all of xP3 + much more! It is used for training future contenders of mT0 & BLOOMZ at project Aya @Cohere Labs 🧡

  • Creation: The dataset can be recreated using instructions available here together with the file in this repository named xp3x_create.py. We provide this version to save processing time.
  • Languages: 277
  • xP3 Dataset Family:
NameExplanationExample models
xP3x Mixture of 17 tasks in 277 languages with English promptsWIP - Join us at Project Aya @C4AI to help!
xP3 Mixture of 13 training tasks in 46 languages with English promptsbloomz & mt0-xxl
xP3mt Mixture of 13 training tasks in 46 languages with prompts in 20 languages (machine-translated from English)bloomz-mt & mt0-xxl-mt
xP3all xP3 + evaluation datasets adding an additional 3 tasks for a total of 16 tasks in 46 languages with English prompts
xP3megds Megatron-DeepSpeed processed version of xP3bloomz
P3 Repreprocessed version of the English-only P3 with 8 training tasksbloomz-p3 & mt0-xxl-p3

Dataset Structure

Data Instances

An example looks as follows:

{
  'inputs': '11月、遂にクロームはファイヤーフォックスを引き離し始めた。_はインターネットユーザーの評価が高まったのだ。\nReplace the _ in the above sentence with the correct option: \n- ファイヤーフォックス\n- クローム',
  'targets': 'クローム',
  'language': 'jpn_Jpan',
  'split': 'test',
  'template': 'Replace',
  'dataset': 'Muennighoff/xwinograd',
  'config': 'jp'
}

Data Fields

The data fields are the same among all splits:

  • inputs: the natural language input fed to the model
  • targets: the natural language target that the model has to generate
  • language: The language code. The codes are an extension of the FLORES-200 codes, where the first part is the language code and the second part the script code.
  • template: The name of the prompt used.
  • dataset: The Hugging Face dataset identifier of where the data stems from.
  • config: The config of the Hugging Face dataset.

Usage

The dataset has 680 gigabytes and 530 million samples. You may want to filter it and then deduplicate depending on your needs.

Loading by language:

# pip install -q datasets
from datasets import load_dataset
ds = load_dataset("Muennighoff/xP3x", "zho_Hans", streaming=True) # Use streaming to not download all at once
for x in ds["train"]:
    print(x)
    break

You can then filter down by the data fields to e.g. only get certain configs or datasets. As every dataset-config-template is its own jsonl file, you can also decide on the datasets, configs and templates you want and only download them. For example, to download all Japanese xwinograd samples, you could do:

# pip install -q datasets
from datasets import load_dataset
import multiprocessing
# pip install --upgrade huggingface-hub
from huggingface_hub import HfFileSystem, hf_hub_url

fs = HfFileSystem()
fps = fs.glob(f"datasets/CohereLabs/xP3x/data/jpn_Jpan/*xwinograd*")
resolved_paths = [fs.resolve_path(file) for file in fps]
data_files = [hf_hub_url(resolved_path.repo_id, resolved_path.path_in_repo, repo_type=resolved_path.repo_type) for resolved_path in resolved_paths]

ds = load_dataset("json", data_files=data_files, num_proc=8)["train"]

Sometimes it may be faster to clone the entire repo. To download all English files, you could do e.g.

GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/CohereLabs/xP3x
cd xP3x
git lfs pull --include="data/eng_Latn/*"

Data Splits

LanguageCodeKilobytes%Samples%
Emilianegl_Latn1040.04020.0
Swiss Germangsw_Latn1040.04080.0
Novialnov_Latn1160.04320.0
Ainu (Latin script)ain_Latn1200.04100.0
Chamorrocha_Latn1200.04520.0
Gothicgot_Goth1200.04020.0
Prussianprg_Latn1200.04240.0
Picardpcd_Latn1400.05300.0
Northern Frisianfrr_Latn1560.05540.0
Uzbek (Latin script)uzb_Latn1560.06000.0
Ottoman Turkish (Latin script)ota_Latn1880.06320.0
Swahili (macrolanguage)swa_Latn2120.07720.0
Talossantzl_Latn2200.08360.0
Kven Finnishfkv_Latn2600.09100.0
Zazazza_Latn2600.01,0560.0
Frisianfry_Latn2680.09560.0
Piemontesepms_Latn2760.09980.0
Kalmykxal_Cyrl2880.09760.0
Hunsrikhrx_Latn3520.01,3800.0
Romanyrom_Latn3640.01,4100.0
Ancient Greek (to 1453)grc_Grek3920.01,2260.0
Tase Naganst_Latn4240.01,6080.0
Albaniansqi_Latn5960.02,2160.0
Guadeloupean Creole Frenchgcf_Latn6080.02,3260.0
Yakutsah_Cyrl6080.01,9860.0
Ho (Latin script)hoc_Latn6320.02,6340.0
Khasikha_Latn6760.02,6640.0
Algerian Arabicarq_Arab6880.02,2780.0
Lower Sorbiandsb_Latn6920.02,5960.0
Chuvashchv_Cyrl7160.02,4460.0
Old Russianorv_Cyrl7520.02,5860.0
Pampangapam_Latn7840.02,9840.0
Kurdish (Latin script)kur_Latn7960.03,0500.0
Ottoman Turkishota_Arab8320.02,7720.0
Kotavaavk_Latn8640.03,1180.0
Upper Sorbianhsb_Latn9000.03,4740.0
Buryatbua_Cyrl9240.03,2180.0
Swabianswg_Latn9960.03,3660.0
Coastal Kadazankzj_Latn1,1360.03,7660.0
Chavacanocbk_Latn1,3520.04,9940.0
Quechuaque_Latn1,7040.05,3120.0
Lingua Franca Nova (Cyrillic script)lfn_Cyrl1,7400.05,4580.0
Groningsgos_Latn1,8640.07,4620.0
Volapükvol_Latn1,9480.07,7120.0
Yue Chinese (Simplified)yue_Hans2,3000.07,8720.0
Mari (Russia)chm_Cyrl2,5400.07,4960.0
Kadazan Dusundtp_Latn2,5480.08,8920.0
Bretonbre_Latn3,0480.011,8680.0
Ladinolad_Latn3,2240.011,9160.0
Cornishcor_Latn3,4920.013,8800.0
Interlingueile_Latn3,7000.014,4680.0
Wu Chinesewuu_Hans3,7840.013,0620.0
Japanese (Katakana)jpn_Kana4,2080.013,9420.0
Idoido_Latn6,1800.023,7420.0
Yiddishiyid_Hebr9,8960.034,4120.01
Klingontlh_Latn11,7160.046,0100.01
Lingua Franca Novalfn_Latn13,3280.046,8260.01
Lojbanjbo_Latn17,4680.066,6940.01
Low Germannds_Latn18,3640.068,0980.01
Interlingua (International Auxiliary Language Association)ina_Latn25,7000.076,5840.01
Javajava25,9040.013,5510.0
Japanese (Kanji)jpn_Hani26,2920.089,9780.02
Norwegiannor_Latn26,7240.093,1160.02
Toki Ponatoki_Latn26,8080.097,1700.02
Latinlat_Latn28,9000.0101,3900.02
Serbo-Croatianhbs_Latn29,4520.0105,7480.02
Nigerian Pidginpcm_Latn145,8720.0288,9920.02
Azerbaijani (South or North; Latin script)aze_Latn147,5640.0277,8750.01
Serbian (Latin script)srp_Latn179,0720.03131,1010.02
Japanese (Hiragana)jpn_Hira188,9440.03628,7580.12
Berber (Latin script)ber_Latn201,4640.03693,6020.13
Jupyter Notebookjupyter_notebook416,0560.06400,0000.08
Yue Chineseyue_Hant613,3520.091,227,4290.23
Haitian Creolehat_Latn629,4200.091,228,2810.23
Mossimos_Latn630,4160.091,223,4810.23
Pangasinanpag_Latn630,6840.091,223,4810.23
Twitwi_Latn631,1720.091,223,4810.23
Bosnianbos_Latn633,0160.091,224,4790.23
Eweewe_Latn633,2920.091,223,4810.23
Bambarabam_Latn634,5200.091,223,4810.23
Javanesejav_Latn635,2480.091,224,0030.23
Southwestern Dinkadik_Latn635,4160.091,223,4810.23
Kabuverdianukea_Latn636,1440.091,223,4810.23
Dyuladyu_Latn636,4640.091,223,4810.23
Venetianvec_Latn637,4120.091,223,4810.23
Chokwecjk_Latn637,5320.091,223,4810.23
Latgalianltg_Latn637,6120.091,223,4810.23
Sundanesesun_Latn638,1200.091,223,4810.23
Asturianast_Latn638,7080.091,223,4810.23
Akanaka_Latn639,6480.091,223,4810.23
Mizolus_Latn639,6800.091,223,4810.23
Guaranigrn_Latn641,5400.091,225,6470.23
Limburgishlim_Latn642,3680.091,223,4810.23
Faroesefao_Latn642,4320.091,224,0670.23
Buginesebug_Latn643,4720.091,223,4810.23
Sangosag_Latn643,5960.091,223,4810.23
Luba-Kasailua_Latn643,6400.091,223,4810.23
Papiamentopap_Latn643,6480.091,223,4810.23
Silesianszl_Latn644,6080.091,223,4810.23
Sicilianscn_Latn645,6360.11,223,4810.23
Kimbundukmb_Latn645,9640.11,223,4810.23
Basqueeus_Latn646,0840.11,246,8770.23
Balineseban_Latn646,4080.11,223,4810.23
Norwegian Nynorsknno_Latn646,9960.11,229,6990.23
Central Aymaraayr_Latn647,2360.11,223,4810.23
Tamasheq (Latin script)taq_Latn648,6560.11,223,4810.23
Kikongokon_Latn648,9920.11,223,4810.23
Friulianfur_Latn649,2720.11,223,4810.23
Ayacucho Quechuaquy_Latn649,9920.11,223,4810.23
Maorimri_Latn650,3360.11,224,2110.23
Icelandicisl_Latn650,3720.11,246,6230.23
Galicianglg_Latn652,0880.11,233,2910.23
Catalancat_Latn652,1160.11,241,3810.23
Lombardlmo_Latn652,1200.11,223,4810.23
Banjar (Latin script)bjn_Latn652,3720.11,223,4810.23
Fijianfij_Latn652,7960.11,223,4810.23
Crimean Tatarcrh_Latn653,9200.11,223,8950.23
Northern Kurdishkmr_Latn654,1080.11,223,4810.23
Ligurianlij_Latn654,4320.11,223,4810.23
Occitanoci_Latn655,6760.11,227,9450.23
Turkmentuk_Latn658,6720.11,241,2050.23
Luxembourgishltz_Latn658,7680.11,225,3390.23
Cebuanoceb_Latn659,1240.11,226,0390.23
Samoansmo_Latn659,7040.11,223,4810.23
Sardiniansrd_Latn660,0000.11,223,4810.23
Bembabem_Latn660,5040.11,223,4810.23
Minangkabau (Latin script)min_Latn660,6720.11,223,4810.23
Acehnese (Latin script)ace_Latn661,0840.11,223,4810.23
Ilocanoilo_Latn661,1840.11,227,6630.23
Irishgle_Latn661,6600.11,227,3570.23
Fonfon_Latn663,1240.11,223,4810.23
Waraywar_Latn664,1200.11,226,5030.23
Norwegian Bokmålnob_Latn666,2400.11,300,6070.24
Tosk Albanianals_Latn666,6920.11,223,4810.23
Standard Malayzsm_Latn667,0880.11,270,7150.24
Southern Sothosot_Latn667,7280.11,223,4810.23
Kabylekab_Latn668,1280.11,346,6050.25
Jingphokac_Latn669,4640.11,223,4810.23
Lingalalin_Latn670,4280.11,323,4810.25
Wolofwol_Latn670,5680.11,373,4810.26
Central Kanuri (Latin script)knc_Latn670,8000.11,223,4810.23
Kikuyukik_Latn672,0960.11,223,4810.23
Tok Pisintpi_Latn672,9160.11,223,4810.23
Nuernus_Latn673,6320.11,223,4810.23
Tagalogtgl_Latn673,6840.11,247,4170.23
Tumbukatum_Latn676,9480.11,223,4810.23
Plateau Malagasyplt_Latn677,8520.11,223,4810.23
Afrikaansafr_Latn679,1640.11,337,0910.25
North Azerbaijaniazj_Latn679,8200.11,223,4810.23
Kabiyèkbp_Latn684,8800.11,223,4810.23
Modern Standard Arabic (Romanized)arb_Latn685,4080.11,223,4810.23
Scottish Gaelicgla_Latn708,6200.11,243,6270.23
Sindhisnd_Arab718,6800.111,223,4810.23
North Levantine Arabicapc_Arab720,0480.111,223,4810.23
Tunisian Arabicaeb_Arab720,3600.111,223,4810.23
South Levantine Arabicajp_Arab720,4880.111,223,4810.23
Dariprs_Arab720,5000.111,223,4810.23
Moroccan Arabicary_Arab722,9040.111,223,4810.23
Egyptian Arabicarz_Arab723,3560.111,223,4810.23
Najdi Arabicars_Arab725,7840.111,223,4810.23
Acehnese (Arabic script)ace_Arab726,2720.111,223,4810.23
Mesopotamian Arabicacm_Arab728,4720.111,223,4810.23
Ta’izzi-Adeni Arabicacq_Arab734,7800.111,223,4810.23
South Azerbaijaniazb_Arab735,7280.111,223,4810.23
Central Kanuri (Arabic script)knc_Arab746,9360.111,223,4810.23
Rundirun_Latn749,7920.111,296,1110.24
Banjar (Arabic script)bjn_Arab751,1120.111,223,4810.23
Central Kurdishckb_Arab756,8040.111,223,4810.23
Bashkirbak_Cyrl758,8160.111,223,4810.23
Kashmiri (Arabic script)kas_Arab759,1400.111,223,4810.23
Tatartat_Cyrl764,2120.111,247,6850.23
Minangkabau (Arabic script)min_Arab765,3840.111,223,4810.23
Kazakhkaz_Cyrl766,1760.111,232,6970.23
Halh Mongoliankhk_Cyrl776,3840.111,224,3530.23
Tajiktgk_Cyrl780,4520.111,223,4810.23
Eastern Yiddishydd_Hebr781,4520.121,223,4810.23
Uyghuruig_Arab785,4440.121,256,9990.24
Armenianhye_Armn789,9520.121,228,1710.23
Hebrewheb_Hebr793,1440.121,604,3650.3
Belarusianbel_Cyrl806,5880.121,261,1970.24
Macedonianmkd_Cyrl813,4360.121,384,5670.26
Welshcym_Latn821,0360.121,321,4550.25
Northern Uzbekuzn_Latn835,5600.121,273,4040.24
Central Atlas Tamazighttzm_Tfng843,5080.121,223,4810.23
Tamasheq (Tifinagh script)taq_Tfng848,1040.121,223,4810.23
Magahimag_Deva851,3600.131,223,4810.23
Bhojpuribho_Deva854,8480.131,223,4810.23
Awadhiawa_Deva857,0960.131,224,0370.23
Chhattisgarhihne_Deva859,3320.131,223,4810.23
Kyrgyzkir_Cyrl860,7000.131,250,1630.23
Maithilimai_Deva863,4760.131,223,4810.23
Assameseasm_Beng865,9040.131,223,4810.23
Kashmiri (Devanagari script)kas_Deva867,2320.131,223,4810.23
Sanskritsan_Deva879,2360.131,223,4810.23
Laolao_Laoo888,2400.131,223,4810.23
Odiaory_Orya890,5080.131,223,4810.23
Santalisat_Olck902,3000.131,223,4810.23
Kannadakan_Knda909,2600.131,223,4810.23
Meitei (Bengali script)mni_Beng917,9840.141,223,4810.23
Georgiankat_Geor928,7120.141,226,7290.23
Kambakam_Latn936,4680.142,136,6150.4
Tigrinyatir_Ethi949,6080.141,276,5360.24
Swatissw_Latn950,5640.142,195,0020.41
Malayalammal_Mlym953,9840.141,225,0830.23
Nigerian Fulfuldefuv_Latn956,3280.142,126,6520.4
Umbunduumb_Latn974,1040.142,264,5530.43
Gandalug_Latn975,7800.142,273,4810.43
Northern Sothonso_Latn978,4840.142,250,9710.42
Khmerkhm_Khmr984,7560.141,227,8250.23
Luoluo_Latn993,0680.152,249,2420.42
Standard Tibetanbod_Tibt993,7320.151,223,4810.23
Tswanatsn_Latn1,009,3280.152,323,4810.44
Kinyarwandakin_Latn1,010,7520.152,273,4810.43
Sinhalasin_Sinh1,012,0120.151,256,5820.24
Xhosaxho_Latn1,019,8040.152,323,4810.44
Shonasna_Latn1,026,3200.152,273,4810.43
Esperantoepo_Latn1,029,4440.152,612,0830.49
Tsongatso_Latn1,031,8560.152,323,4810.44
Dzongkhadzo_Tibt1,033,5520.151,223,4810.23
Zuluzul_Latn1,039,2960.152,323,4810.44
Serbiansrp_Cyrl1,040,0240.151,362,5980.26
Nyanjanya_Latn1,061,7800.162,323,4810.44
Shanshn_Mymr1,074,9400.161,223,4810.23
Igboibo_Latn1,095,3000.162,282,3010.43
Hausahau_Latn1,112,2720.162,335,7380.44
West Central Oromogaz_Latn1,115,6000.162,343,2600.44
Nepalinpi_Deva1,144,6760.171,281,4300.24
Yorubayor_Latn1,164,5400.172,334,8010.44
Southern Pashtopbt_Arab1,170,8400.171,365,5330.26
Somalisom_Latn1,198,3200.182,482,4370.47
Burmesemya_Mymr1,228,1960.181,279,8820.24
Amharicamh_Ethi1,261,1280.191,980,2150.37
Eastern Panjabipan_Guru1,305,6360.191,307,8970.25
Gujaratiguj_Gujr1,331,7800.21,317,3140.25
Marathimar_Deva1,494,0240.221,443,9500.27
Bengaliben_Beng1,650,2720.241,411,5140.27
Chinese (Traditional)zho_Hant1,778,7360.261,956,1890.37
Tamiltam_Taml1,833,3280.271,394,4730.26
Swahiliswh_Latn1,970,7840.294,185,6080.79
Telugutel_Telu2,224,4800.331,573,3250.3
Ukrainianukr_Cyrl2,227,6160.332,216,1190.42
Western Persianpes_Arab2,389,3400.351,811,1210.34
Turkishtur_Latn3,106,6000.464,146,1530.78
Urduurd_Arab3,553,9600.523,513,2180.66
Koreankor_Hang4,642,4680.683,415,9200.64
Pythonpython4,728,5040.73,142,9620.59
Japanesejpn_Jpan5,079,7880.754,193,5700.79
Thaitha_Thai6,860,7041.014,666,2990.88
Chinese (Simplified)zho_Hans8,063,6841.197,355,5091.38
Vietnamesevie_Latn8,398,8241.246,194,9251.16
Indonesianind_Latn9,380,1441.385,301,8121.0
Hindihin_Deva9,914,3281.465,612,1761.05
Croatianhrv_Latn10,028,0281.485,583,9751.05
Modern Standard Arabicarb_Arab11,051,0641.637,232,5511.36
Romanianron_Latn11,441,6361.685,594,9271.05
Maltesemlt_Latn11,614,4881.715,513,8851.04
Slovenianslv_Latn12,014,9121.775,533,6891.04
Estonianest_Latn12,126,2121.795,584,0571.05
Lithuanianlit_Latn12,253,9761.85,603,0471.05
Slovakslk_Latn12,286,3001.815,513,4811.04
Standard Latvianlvs_Latn12,298,5841.815,517,2871.04
Polishpol_Latn12,409,6841.835,868,6311.1
Hungarianhun_Latn12,607,4201.866,086,6211.14
Russianrus_Cyrl13,110,9081.938,798,9271.65
Czechces_Latn14,316,0522.116,418,4621.21
Bulgarianbul_Cyrl14,615,4682.157,265,8851.37
Swedishswe_Latn14,646,6562.165,634,3631.06
Finnishfin_Latn15,011,4642.216,077,5011.14
Danishdan_Latn16,136,6122.385,831,1091.1
Dutchnld_Latn22,387,0203.38,992,8641.69
Greekell_Grek23,144,2963.417,224,0011.36
Italianita_Latn23,952,8243.539,967,7381.87
Portuguesepor_Latn27,297,2524.0211,242,8082.11
Germandeu_Latn27,909,8084.1115,806,9692.97
Frenchfra_Latn28,428,6084.1816,365,9843.08
Spanishspa_Latn30,969,5804.5616,315,9283.07
Englisheng_Latn69,530,38410.2453,015,6909.96
Total-679,318,704100532,107,156100

Language specifics

  • Japanese: Data in jpn_Hira, jpn_Kana, jpn_Hani is guaranteed to have Hiragana, Katakana or Kanji, respectively in each sample. However, they may still include other styles. So while all samples in jpn_Kana are guaranteed to have Katakana, there may still be Hiragana or Kanji.

Dataset Creation

Source Data

Training datasets

Dataset specifics

  • Flores-200: There are three prompts for Flores: continuation, question, command, which represent three commonly used prompting styles, i.e. making a prompt seem like a natural continuation, turning it into a question or commanding the model to do something.
  • tatoeba_mt: Contains duplicates. For example, it has data that is both classified as jpn_Kana and jpn_Jpan, so you may want to deduplicate.

Additional Information

Licensing Information

The dataset collection is released under Apache 2.0. Note that individual datasets may have different licenses.

Citation Information

@article{muennighoff2022crosslingual,
  title={Crosslingual generalization through multitask finetuning},
  author={Muennighoff, Niklas and Wang, Thomas and Sutawika, Lintang and Roberts, Adam and Biderman, Stella and Scao, Teven Le and Bari, M Saiful and Shen, Sheng and Yong, Zheng-Xin and Schoelkopf, Hailey and others},
  journal={arXiv preprint arXiv:2211.01786},
  year={2022}
}

Contributions

Thanks to the contributors of promptsource for adding many prompts used in this dataset. Thanks to the Aya team @Cohere Labs 🧡

Contributors

Muennighoff

335 commits

alexrs

3 commits

dushrntic

1 commits

Muennighoff/xP3x

Dataset

95

stars

342

commits

2

linked in READMEs

May 23, 2025

updated

Browse cluster: Multilingual NLP Datasets and Corpora

README

Dataset Card for xP3x

Table of Contents

Dataset Description

Dataset Summary

xP3x (Crosslingual Public Pool of Prompts eXtended) is a collection of prompts & datasets across 277 languages & 16 NLP tasks. It contains all of xP3 + much more! It is used for training future contenders of mT0 & BLOOMZ at project Aya @Cohere Labs 🧡

  • Creation: The dataset can be recreated using instructions available here together with the file in this repository named xp3x_create.py. We provide this version to save processing time.
  • Languages: 277
  • xP3 Dataset Family:
NameExplanationExample models
xP3x Mixture of 17 tasks in 277 languages with English promptsWIP - Join us at Project Aya @C4AI to help!
xP3 Mixture of 13 training tasks in 46 languages with English promptsbloomz & mt0-xxl
xP3mt Mixture of 13 training tasks in 46 languages with prompts in 20 languages (machine-translated from English)bloomz-mt & mt0-xxl-mt
xP3all xP3 + evaluation datasets adding an additional 3 tasks for a total of 16 tasks in 46 languages with English prompts
xP3megds Megatron-DeepSpeed processed version of xP3bloomz
P3 Repreprocessed version of the English-only P3 with 8 training tasksbloomz-p3 & mt0-xxl-p3

Dataset Structure

Data Instances

An example looks as follows:

{
  'inputs': '11月、遂にクロームはファイヤーフォックスを引き離し始めた。_はインターネットユーザーの評価が高まったのだ。\nReplace the _ in the above sentence with the correct option: \n- ファイヤーフォックス\n- クローム',
  'targets': 'クローム',
  'language': 'jpn_Jpan',
  'split': 'test',
  'template': 'Replace',
  'dataset': 'Muennighoff/xwinograd',
  'config': 'jp'
}

Data Fields

The data fields are the same among all splits:

  • inputs: the natural language input fed to the model
  • targets: the natural language target that the model has to generate
  • language: The language code. The codes are an extension of the FLORES-200 codes, where the first part is the language code and the second part the script code.
  • template: The name of the prompt used.
  • dataset: The Hugging Face dataset identifier of where the data stems from.
  • config: The config of the Hugging Face dataset.

Usage

The dataset has 680 gigabytes and 530 million samples. You may want to filter it and then deduplicate depending on your needs.

Loading by language:

# pip install -q datasets
from datasets import load_dataset
ds = load_dataset("Muennighoff/xP3x", "zho_Hans", streaming=True) # Use streaming to not download all at once
for x in ds["train"]:
    print(x)
    break

You can then filter down by the data fields to e.g. only get certain configs or datasets. As every dataset-config-template is its own jsonl file, you can also decide on the datasets, configs and templates you want and only download them. For example, to download all Japanese xwinograd samples, you could do:

# pip install -q datasets
from datasets import load_dataset
import multiprocessing
# pip install --upgrade huggingface-hub
from huggingface_hub import HfFileSystem, hf_hub_url

fs = HfFileSystem()
fps = fs.glob(f"datasets/CohereLabs/xP3x/data/jpn_Jpan/*xwinograd*")
resolved_paths = [fs.resolve_path(file) for file in fps]
data_files = [hf_hub_url(resolved_path.repo_id, resolved_path.path_in_repo, repo_type=resolved_path.repo_type) for resolved_path in resolved_paths]

ds = load_dataset("json", data_files=data_files, num_proc=8)["train"]

Sometimes it may be faster to clone the entire repo. To download all English files, you could do e.g.

GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/CohereLabs/xP3x
cd xP3x
git lfs pull --include="data/eng_Latn/*"

Data Splits

LanguageCodeKilobytes%Samples%
Emilianegl_Latn1040.04020.0
Swiss Germangsw_Latn1040.04080.0
Novialnov_Latn1160.04320.0
Ainu (Latin script)ain_Latn1200.04100.0
Chamorrocha_Latn1200.04520.0
Gothicgot_Goth1200.04020.0
Prussianprg_Latn1200.04240.0
Picardpcd_Latn1400.05300.0
Northern Frisianfrr_Latn1560.05540.0
Uzbek (Latin script)uzb_Latn1560.06000.0
Ottoman Turkish (Latin script)ota_Latn1880.06320.0
Swahili (macrolanguage)swa_Latn2120.07720.0
Talossantzl_Latn2200.08360.0
Kven Finnishfkv_Latn2600.09100.0
Zazazza_Latn2600.01,0560.0
Frisianfry_Latn2680.09560.0
Piemontesepms_Latn2760.09980.0
Kalmykxal_Cyrl2880.09760.0
Hunsrikhrx_Latn3520.01,3800.0
Romanyrom_Latn3640.01,4100.0
Ancient Greek (to 1453)grc_Grek3920.01,2260.0
Tase Naganst_Latn4240.01,6080.0
Albaniansqi_Latn5960.02,2160.0
Guadeloupean Creole Frenchgcf_Latn6080.02,3260.0
Yakutsah_Cyrl6080.01,9860.0
Ho (Latin script)hoc_Latn6320.02,6340.0
Khasikha_Latn6760.02,6640.0
Algerian Arabicarq_Arab6880.02,2780.0
Lower Sorbiandsb_Latn6920.02,5960.0
Chuvashchv_Cyrl7160.02,4460.0
Old Russianorv_Cyrl7520.02,5860.0
Pampangapam_Latn7840.02,9840.0
Kurdish (Latin script)kur_Latn7960.03,0500.0
Ottoman Turkishota_Arab8320.02,7720.0
Kotavaavk_Latn8640.03,1180.0
Upper Sorbianhsb_Latn9000.03,4740.0
Buryatbua_Cyrl9240.03,2180.0
Swabianswg_Latn9960.03,3660.0
Coastal Kadazankzj_Latn1,1360.03,7660.0
Chavacanocbk_Latn1,3520.04,9940.0
Quechuaque_Latn1,7040.05,3120.0
Lingua Franca Nova (Cyrillic script)lfn_Cyrl1,7400.05,4580.0
Groningsgos_Latn1,8640.07,4620.0
Volapükvol_Latn1,9480.07,7120.0
Yue Chinese (Simplified)yue_Hans2,3000.07,8720.0
Mari (Russia)chm_Cyrl2,5400.07,4960.0
Kadazan Dusundtp_Latn2,5480.08,8920.0
Bretonbre_Latn3,0480.011,8680.0
Ladinolad_Latn3,2240.011,9160.0
Cornishcor_Latn3,4920.013,8800.0
Interlingueile_Latn3,7000.014,4680.0
Wu Chinesewuu_Hans3,7840.013,0620.0
Japanese (Katakana)jpn_Kana4,2080.013,9420.0
Idoido_Latn6,1800.023,7420.0
Yiddishiyid_Hebr9,8960.034,4120.01
Klingontlh_Latn11,7160.046,0100.01
Lingua Franca Novalfn_Latn13,3280.046,8260.01
Lojbanjbo_Latn17,4680.066,6940.01
Low Germannds_Latn18,3640.068,0980.01
Interlingua (International Auxiliary Language Association)ina_Latn25,7000.076,5840.01
Javajava25,9040.013,5510.0
Japanese (Kanji)jpn_Hani26,2920.089,9780.02
Norwegiannor_Latn26,7240.093,1160.02
Toki Ponatoki_Latn26,8080.097,1700.02
Latinlat_Latn28,9000.0101,3900.02
Serbo-Croatianhbs_Latn29,4520.0105,7480.02
Nigerian Pidginpcm_Latn145,8720.0288,9920.02
Azerbaijani (South or North; Latin script)aze_Latn147,5640.0277,8750.01
Serbian (Latin script)srp_Latn179,0720.03131,1010.02
Japanese (Hiragana)jpn_Hira188,9440.03628,7580.12
Berber (Latin script)ber_Latn201,4640.03693,6020.13
Jupyter Notebookjupyter_notebook416,0560.06400,0000.08
Yue Chineseyue_Hant613,3520.091,227,4290.23
Haitian Creolehat_Latn629,4200.091,228,2810.23
Mossimos_Latn630,4160.091,223,4810.23
Pangasinanpag_Latn630,6840.091,223,4810.23
Twitwi_Latn631,1720.091,223,4810.23
Bosnianbos_Latn633,0160.091,224,4790.23
Eweewe_Latn633,2920.091,223,4810.23
Bambarabam_Latn634,5200.091,223,4810.23
Javanesejav_Latn635,2480.091,224,0030.23
Southwestern Dinkadik_Latn635,4160.091,223,4810.23
Kabuverdianukea_Latn636,1440.091,223,4810.23
Dyuladyu_Latn636,4640.091,223,4810.23
Venetianvec_Latn637,4120.091,223,4810.23
Chokwecjk_Latn637,5320.091,223,4810.23
Latgalianltg_Latn637,6120.091,223,4810.23
Sundanesesun_Latn638,1200.091,223,4810.23
Asturianast_Latn638,7080.091,223,4810.23
Akanaka_Latn639,6480.091,223,4810.23
Mizolus_Latn639,6800.091,223,4810.23
Guaranigrn_Latn641,5400.091,225,6470.23
Limburgishlim_Latn642,3680.091,223,4810.23
Faroesefao_Latn642,4320.091,224,0670.23
Buginesebug_Latn643,4720.091,223,4810.23
Sangosag_Latn643,5960.091,223,4810.23
Luba-Kasailua_Latn643,6400.091,223,4810.23
Papiamentopap_Latn643,6480.091,223,4810.23
Silesianszl_Latn644,6080.091,223,4810.23
Sicilianscn_Latn645,6360.11,223,4810.23
Kimbundukmb_Latn645,9640.11,223,4810.23
Basqueeus_Latn646,0840.11,246,8770.23
Balineseban_Latn646,4080.11,223,4810.23
Norwegian Nynorsknno_Latn646,9960.11,229,6990.23
Central Aymaraayr_Latn647,2360.11,223,4810.23
Tamasheq (Latin script)taq_Latn648,6560.11,223,4810.23
Kikongokon_Latn648,9920.11,223,4810.23
Friulianfur_Latn649,2720.11,223,4810.23
Ayacucho Quechuaquy_Latn649,9920.11,223,4810.23
Maorimri_Latn650,3360.11,224,2110.23
Icelandicisl_Latn650,3720.11,246,6230.23
Galicianglg_Latn652,0880.11,233,2910.23
Catalancat_Latn652,1160.11,241,3810.23
Lombardlmo_Latn652,1200.11,223,4810.23
Banjar (Latin script)bjn_Latn652,3720.11,223,4810.23
Fijianfij_Latn652,7960.11,223,4810.23
Crimean Tatarcrh_Latn653,9200.11,223,8950.23
Northern Kurdishkmr_Latn654,1080.11,223,4810.23
Ligurianlij_Latn654,4320.11,223,4810.23
Occitanoci_Latn655,6760.11,227,9450.23
Turkmentuk_Latn658,6720.11,241,2050.23
Luxembourgishltz_Latn658,7680.11,225,3390.23
Cebuanoceb_Latn659,1240.11,226,0390.23
Samoansmo_Latn659,7040.11,223,4810.23
Sardiniansrd_Latn660,0000.11,223,4810.23
Bembabem_Latn660,5040.11,223,4810.23
Minangkabau (Latin script)min_Latn660,6720.11,223,4810.23
Acehnese (Latin script)ace_Latn661,0840.11,223,4810.23
Ilocanoilo_Latn661,1840.11,227,6630.23
Irishgle_Latn661,6600.11,227,3570.23
Fonfon_Latn663,1240.11,223,4810.23
Waraywar_Latn664,1200.11,226,5030.23
Norwegian Bokmålnob_Latn666,2400.11,300,6070.24
Tosk Albanianals_Latn666,6920.11,223,4810.23
Standard Malayzsm_Latn667,0880.11,270,7150.24
Southern Sothosot_Latn667,7280.11,223,4810.23
Kabylekab_Latn668,1280.11,346,6050.25
Jingphokac_Latn669,4640.11,223,4810.23
Lingalalin_Latn670,4280.11,323,4810.25
Wolofwol_Latn670,5680.11,373,4810.26
Central Kanuri (Latin script)knc_Latn670,8000.11,223,4810.23
Kikuyukik_Latn672,0960.11,223,4810.23
Tok Pisintpi_Latn672,9160.11,223,4810.23
Nuernus_Latn673,6320.11,223,4810.23
Tagalogtgl_Latn673,6840.11,247,4170.23
Tumbukatum_Latn676,9480.11,223,4810.23
Plateau Malagasyplt_Latn677,8520.11,223,4810.23
Afrikaansafr_Latn679,1640.11,337,0910.25
North Azerbaijaniazj_Latn679,8200.11,223,4810.23
Kabiyèkbp_Latn684,8800.11,223,4810.23
Modern Standard Arabic (Romanized)arb_Latn685,4080.11,223,4810.23
Scottish Gaelicgla_Latn708,6200.11,243,6270.23
Sindhisnd_Arab718,6800.111,223,4810.23
North Levantine Arabicapc_Arab720,0480.111,223,4810.23
Tunisian Arabicaeb_Arab720,3600.111,223,4810.23
South Levantine Arabicajp_Arab720,4880.111,223,4810.23
Dariprs_Arab720,5000.111,223,4810.23
Moroccan Arabicary_Arab722,9040.111,223,4810.23
Egyptian Arabicarz_Arab723,3560.111,223,4810.23
Najdi Arabicars_Arab725,7840.111,223,4810.23
Acehnese (Arabic script)ace_Arab726,2720.111,223,4810.23
Mesopotamian Arabicacm_Arab728,4720.111,223,4810.23
Ta’izzi-Adeni Arabicacq_Arab734,7800.111,223,4810.23
South Azerbaijaniazb_Arab735,7280.111,223,4810.23
Central Kanuri (Arabic script)knc_Arab746,9360.111,223,4810.23
Rundirun_Latn749,7920.111,296,1110.24
Banjar (Arabic script)bjn_Arab751,1120.111,223,4810.23
Central Kurdishckb_Arab756,8040.111,223,4810.23
Bashkirbak_Cyrl758,8160.111,223,4810.23
Kashmiri (Arabic script)kas_Arab759,1400.111,223,4810.23
Tatartat_Cyrl764,2120.111,247,6850.23
Minangkabau (Arabic script)min_Arab765,3840.111,223,4810.23
Kazakhkaz_Cyrl766,1760.111,232,6970.23
Halh Mongoliankhk_Cyrl776,3840.111,224,3530.23
Tajiktgk_Cyrl780,4520.111,223,4810.23
Eastern Yiddishydd_Hebr781,4520.121,223,4810.23
Uyghuruig_Arab785,4440.121,256,9990.24
Armenianhye_Armn789,9520.121,228,1710.23
Hebrewheb_Hebr793,1440.121,604,3650.3
Belarusianbel_Cyrl806,5880.121,261,1970.24
Macedonianmkd_Cyrl813,4360.121,384,5670.26
Welshcym_Latn821,0360.121,321,4550.25
Northern Uzbekuzn_Latn835,5600.121,273,4040.24
Central Atlas Tamazighttzm_Tfng843,5080.121,223,4810.23
Tamasheq (Tifinagh script)taq_Tfng848,1040.121,223,4810.23
Magahimag_Deva851,3600.131,223,4810.23
Bhojpuribho_Deva854,8480.131,223,4810.23
Awadhiawa_Deva857,0960.131,224,0370.23
Chhattisgarhihne_Deva859,3320.131,223,4810.23
Kyrgyzkir_Cyrl860,7000.131,250,1630.23
Maithilimai_Deva863,4760.131,223,4810.23
Assameseasm_Beng865,9040.131,223,4810.23
Kashmiri (Devanagari script)kas_Deva867,2320.131,223,4810.23
Sanskritsan_Deva879,2360.131,223,4810.23
Laolao_Laoo888,2400.131,223,4810.23
Odiaory_Orya890,5080.131,223,4810.23
Santalisat_Olck902,3000.131,223,4810.23
Kannadakan_Knda909,2600.131,223,4810.23
Meitei (Bengali script)mni_Beng917,9840.141,223,4810.23
Georgiankat_Geor928,7120.141,226,7290.23
Kambakam_Latn936,4680.142,136,6150.4
Tigrinyatir_Ethi949,6080.141,276,5360.24
Swatissw_Latn950,5640.142,195,0020.41
Malayalammal_Mlym953,9840.141,225,0830.23
Nigerian Fulfuldefuv_Latn956,3280.142,126,6520.4
Umbunduumb_Latn974,1040.142,264,5530.43
Gandalug_Latn975,7800.142,273,4810.43
Northern Sothonso_Latn978,4840.142,250,9710.42
Khmerkhm_Khmr984,7560.141,227,8250.23
Luoluo_Latn993,0680.152,249,2420.42
Standard Tibetanbod_Tibt993,7320.151,223,4810.23
Tswanatsn_Latn1,009,3280.152,323,4810.44
Kinyarwandakin_Latn1,010,7520.152,273,4810.43
Sinhalasin_Sinh1,012,0120.151,256,5820.24
Xhosaxho_Latn1,019,8040.152,323,4810.44
Shonasna_Latn1,026,3200.152,273,4810.43
Esperantoepo_Latn1,029,4440.152,612,0830.49
Tsongatso_Latn1,031,8560.152,323,4810.44
Dzongkhadzo_Tibt1,033,5520.151,223,4810.23
Zuluzul_Latn1,039,2960.152,323,4810.44
Serbiansrp_Cyrl1,040,0240.151,362,5980.26
Nyanjanya_Latn1,061,7800.162,323,4810.44
Shanshn_Mymr1,074,9400.161,223,4810.23
Igboibo_Latn1,095,3000.162,282,3010.43
Hausahau_Latn1,112,2720.162,335,7380.44
West Central Oromogaz_Latn1,115,6000.162,343,2600.44
Nepalinpi_Deva1,144,6760.171,281,4300.24
Yorubayor_Latn1,164,5400.172,334,8010.44
Southern Pashtopbt_Arab1,170,8400.171,365,5330.26
Somalisom_Latn1,198,3200.182,482,4370.47
Burmesemya_Mymr1,228,1960.181,279,8820.24
Amharicamh_Ethi1,261,1280.191,980,2150.37
Eastern Panjabipan_Guru1,305,6360.191,307,8970.25
Gujaratiguj_Gujr1,331,7800.21,317,3140.25
Marathimar_Deva1,494,0240.221,443,9500.27
Bengaliben_Beng1,650,2720.241,411,5140.27
Chinese (Traditional)zho_Hant1,778,7360.261,956,1890.37
Tamiltam_Taml1,833,3280.271,394,4730.26
Swahiliswh_Latn1,970,7840.294,185,6080.79
Telugutel_Telu2,224,4800.331,573,3250.3
Ukrainianukr_Cyrl2,227,6160.332,216,1190.42
Western Persianpes_Arab2,389,3400.351,811,1210.34
Turkishtur_Latn3,106,6000.464,146,1530.78
Urduurd_Arab3,553,9600.523,513,2180.66
Koreankor_Hang4,642,4680.683,415,9200.64
Pythonpython4,728,5040.73,142,9620.59
Japanesejpn_Jpan5,079,7880.754,193,5700.79
Thaitha_Thai6,860,7041.014,666,2990.88
Chinese (Simplified)zho_Hans8,063,6841.197,355,5091.38
Vietnamesevie_Latn8,398,8241.246,194,9251.16
Indonesianind_Latn9,380,1441.385,301,8121.0
Hindihin_Deva9,914,3281.465,612,1761.05
Croatianhrv_Latn10,028,0281.485,583,9751.05
Modern Standard Arabicarb_Arab11,051,0641.637,232,5511.36
Romanianron_Latn11,441,6361.685,594,9271.05
Maltesemlt_Latn11,614,4881.715,513,8851.04
Slovenianslv_Latn12,014,9121.775,533,6891.04
Estonianest_Latn12,126,2121.795,584,0571.05
Lithuanianlit_Latn12,253,9761.85,603,0471.05
Slovakslk_Latn12,286,3001.815,513,4811.04
Standard Latvianlvs_Latn12,298,5841.815,517,2871.04
Polishpol_Latn12,409,6841.835,868,6311.1
Hungarianhun_Latn12,607,4201.866,086,6211.14
Russianrus_Cyrl13,110,9081.938,798,9271.65
Czechces_Latn14,316,0522.116,418,4621.21
Bulgarianbul_Cyrl14,615,4682.157,265,8851.37
Swedishswe_Latn14,646,6562.165,634,3631.06
Finnishfin_Latn15,011,4642.216,077,5011.14
Danishdan_Latn16,136,6122.385,831,1091.1
Dutchnld_Latn22,387,0203.38,992,8641.69
Greekell_Grek23,144,2963.417,224,0011.36
Italianita_Latn23,952,8243.539,967,7381.87
Portuguesepor_Latn27,297,2524.0211,242,8082.11
Germandeu_Latn27,909,8084.1115,806,9692.97
Frenchfra_Latn28,428,6084.1816,365,9843.08
Spanishspa_Latn30,969,5804.5616,315,9283.07
Englisheng_Latn69,530,38410.2453,015,6909.96
Total-679,318,704100532,107,156100

Language specifics

  • Japanese: Data in jpn_Hira, jpn_Kana, jpn_Hani is guaranteed to have Hiragana, Katakana or Kanji, respectively in each sample. However, they may still include other styles. So while all samples in jpn_Kana are guaranteed to have Katakana, there may still be Hiragana or Kanji.

Dataset Creation

Source Data

Training datasets

Dataset specifics

  • Flores-200: There are three prompts for Flores: continuation, question, command, which represent three commonly used prompting styles, i.e. making a prompt seem like a natural continuation, turning it into a question or commanding the model to do something.
  • tatoeba_mt: Contains duplicates. For example, it has data that is both classified as jpn_Kana and jpn_Jpan, so you may want to deduplicate.

Additional Information

Licensing Information

The dataset collection is released under Apache 2.0. Note that individual datasets may have different licenses.

Citation Information

@article{muennighoff2022crosslingual,
  title={Crosslingual generalization through multitask finetuning},
  author={Muennighoff, Niklas and Wang, Thomas and Sutawika, Lintang and Roberts, Adam and Biderman, Stella and Scao, Teven Le and Bari, M Saiful and Shen, Sheng and Yong, Zheng-Xin and Schoelkopf, Hailey and others},
  journal={arXiv preprint arXiv:2211.01786},
  year={2022}
}

Contributions

Thanks to the contributors of promptsource for adding many prompts used in this dataset. Thanks to the Aya team @Cohere Labs 🧡

Contributors

Muennighoff

335 commits

alexrs

3 commits

dushrntic

1 commits