CohereForAI/aya-101

Model

665

stars

23

commits

9

repos using this model

2

linked in READMEs

Sep 10, 2025

updated

afr
amh
ara
aze
bel
ben
bul
cat
ceb
ces
cym
dan
deu
ell
endpoints_compatible
eng
epo
est
eus
fil
fin
fra
fry
gla
gle
glg
guj
hat
hau
heb
hin
hun
hye
ibo
ind
isl
ita
jav
jpn
kan
kat
kaz
khm
kir
kor
kur
lao
lat
lav
lit
ltz
mal
mar
mkd
mlg
mlt
mon
mri
msa
mya
nep
nld
nor
nso
nya
ory
pan
pes
pol
por
pus
ron
rus
safetensors
sin
slk
slv
smo
sna
snd
som
sot
spa
sqi
srp
sun
swa
swe
t5
tam
tel
text2text-generation
text-generation-inference
tgk
tha
transformers
tur
twi
ukr
urd
uzb
vie
xho
yid
yor
zho
zul

README

Aya model summary image

Model Card for Aya 101

Model Summary

The Aya model is a massively multilingual generative language model that follows instructions in 101 languages. Aya outperforms mT0 and BLOOMZ a wide variety of automatic and human evaluations despite covering double the number of languages. The Aya model is trained using xP3x, Aya Dataset, Aya Collection, a subset of DataProvenance collection and ShareGPT-Command. We release the checkpoints under a Apache-2.0 license to further our mission of multilingual technologies empowering a multilingual world.

Use

# pip install -q transformers
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer

checkpoint = "CohereLabs/aya-101"

tokenizer = AutoTokenizer.from_pretrained(checkpoint)
aya_model = AutoModelForSeq2SeqLM.from_pretrained(checkpoint)

# Turkish to English translation
tur_inputs = tokenizer.encode("Translate to English: Aya cok dilli bir dil modelidir.", return_tensors="pt")
tur_outputs = aya_model.generate(tur_inputs, max_new_tokens=128)
print(tokenizer.decode(tur_outputs[0]))
# Aya is a multi-lingual language model

# Q: Why are there so many languages in India?
hin_inputs = tokenizer.encode("भारत में इतनी सारी भाषाएँ क्यों हैं?", return_tensors="pt")
hin_outputs = aya_model.generate(hin_inputs, max_new_tokens=128)
print(tokenizer.decode(hin_outputs[0]))
# Expected output: भारत में कई भाषाएँ हैं और विभिन्न भाषाओं के बोली जाने वाले लोग हैं। यह विभिन्नता भाषाई विविधता और सांस्कृतिक विविधता का परिणाम है। Translates to "India has many languages and people speaking different languages. This diversity is the result of linguistic diversity and cultural diversity."

Model Details

Finetuning

  • Architecture: Same as mt5-xxl
  • Number of Samples seen during Finetuning: 25M
  • Batch size: 256
  • Hardware: TPUv4-128
  • Software: T5X, Jax

Data Sources

The Aya model is trained on the following datasets:

All datasets are subset to the 101 languages supported by mT5. See the paper for details about filtering and pruning.

Evaluation

We refer to Section 5 from our paper for multilingual eval across 99 languages – including discriminative and generative tasks, human evaluation, and simulated win rates that cover both held-out tasks and in-distribution performance.

Bias, Risks, and Limitations

For a detailed overview of our effort at safety mitigation and benchmarking toxicity and bias across multiple languages, we refer to Sections 6 and 7 of our paper: Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model.

We hope that the release of the Aya model will make community-based redteaming efforts possible, by exposing an open-source massively-multilingual model for community research.

Citation

BibTeX:

@article{üstün2024aya,
  title={Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model},
  author={Ahmet Üstün and Viraat Aryabumi and Zheng-Xin Yong and Wei-Yin Ko and Daniel D'souza and Gbemileke Onilude and Neel Bhandari and Shivalika Singh and Hui-Lee Ooi and Amr Kayid and Freddie Vargus and Phil Blunsom and Shayne Longpre and Niklas Muennighoff and Marzieh Fadaee and Julia Kreutzer and Sara Hooker},
  journal={arXiv preprint arXiv:2402.07827},
  year={2024}
}

Languages Covered

Click to see Languages Covered

Below is the list of languages used in finetuning the Aya Model. We group languages into higher-, mid-, and lower-resourcedness based on a language classification by Joshi et. al, 2020. For further details, we refer to our paper

ISO CodeLanguage NameScriptFamilySubgroupingResourcedness
afrAfrikaansLatinIndo-EuropeanGermanicMid
amhAmharicGe'ezAfro-AsiaticSemiticLow
araArabicArabicAfro-AsiaticSemiticHigh
azeAzerbaijaniArabic/LatinTurkicCommon TurkicLow
belBelarusianCyrillicIndo-EuropeanBalto-SlavicMid
benBengaliBengaliIndo-EuropeanIndo-AryanMid
bulBulgarianCyrillicIndo-EuropeanBalto-SlavicMid
catCatalanLatinIndo-EuropeanItalicHigh
cebCebuanoLatinAustronesianMalayo-PolynesianMid
cesCzechLatinIndo-EuropeanBalto-SlavicHigh
cymWelshLatinIndo-EuropeanCelticLow
danDanishLatinIndo-EuropeanGermanicMid
deuGermanLatinIndo-EuropeanGermanicHigh
ellGreekGreekIndo-EuropeanGraeco-PhrygianMid
engEnglishLatinIndo-EuropeanGermanicHigh
epoEsperantoLatinConstructedEsperanticLow
estEstonianLatinUralicFinnicMid
eusBasqueLatinBasque-High
finFinnishLatinUralicFinnicHigh
filTagalogLatinAustronesianMalayo-PolynesianMid
fraFrenchLatinIndo-EuropeanItalicHigh
fryWestern FrisianLatinIndo-EuropeanGermanicLow
glaScottish GaelicLatinIndo-EuropeanCelticLow
gleIrishLatinIndo-EuropeanCelticLow
glgGalicianLatinIndo-EuropeanItalicMid
gujGujaratiGujaratiIndo-EuropeanIndo-AryanLow
hatHaitian CreoleLatinIndo-EuropeanItalicLow
hauHausaLatinAfro-AsiaticChadicLow
hebHebrewHebrewAfro-AsiaticSemiticMid
hinHindiDevanagariIndo-EuropeanIndo-AryanHigh
hunHungarianLatinUralic-High
hyeArmenianArmenianIndo-EuropeanArmenicLow
iboIgboLatinAtlantic-CongoBenue-CongoLow
indIndonesianLatinAustronesianMalayo-PolynesianMid
islIcelandicLatinIndo-EuropeanGermanicLow
itaItalianLatinIndo-EuropeanItalicHigh
javJavaneseLatinAustronesianMalayo-PolynesianLow
jpnJapaneseJapaneseJaponicJapanesicHigh
kanKannadaKannadaDravidianSouth DravidianLow
katGeorgianGeorgianKartvelianGeorgian-ZanMid
kazKazakhCyrillicTurkicCommon TurkicMid
khmKhmerKhmerAustroasiaticKhmericLow
kirKyrgyzCyrillicTurkicCommon TurkicLow
korKoreanHangulKoreanicKoreanHigh
kurKurdishLatinIndo-EuropeanIranianLow
laoLaoLaoTai-KadaiKam-TaiLow
lavLatvianLatinIndo-EuropeanBalto-SlavicMid
latLatinLatinIndo-EuropeanItalicMid
litLithuanianLatinIndo-EuropeanBalto-SlavicMid
ltzLuxembourgishLatinIndo-EuropeanGermanicLow
malMalayalamMalayalamDravidianSouth DravidianLow
marMarathiDevanagariIndo-EuropeanIndo-AryanLow
mkdMacedonianCyrillicIndo-EuropeanBalto-SlavicLow
mlgMalagasyLatinAustronesianMalayo-PolynesianLow
mltMalteseLatinAfro-AsiaticSemiticLow
monMongolianCyrillicMongolic-KhitanMongolicLow
mriMaoriLatinAustronesianMalayo-PolynesianLow
msaMalayLatinAustronesianMalayo-PolynesianMid
myaBurmeseMyanmarSino-TibetanBurmo-QiangicLow
nepNepaliDevanagariIndo-EuropeanIndo-AryanLow
nldDutchLatinIndo-EuropeanGermanicHigh
norNorwegianLatinIndo-EuropeanGermanicLow
nsoNorthern SothoLatinAtlantic-CongoBenue-CongoLow
nyaChichewaLatinAtlantic-CongoBenue-CongoLow
oryOriyaOriyaIndo-EuropeanIndo-AryanLow
panPunjabiGurmukhiIndo-EuropeanIndo-AryanLow
pesPersianArabicIndo-EuropeanIranianHigh
polPolishLatinIndo-EuropeanBalto-SlavicHigh
porPortugueseLatinIndo-EuropeanItalicHigh
pusPashtoArabicIndo-EuropeanIranianLow
ronRomanianLatinIndo-EuropeanItalicMid
rusRussianCyrillicIndo-EuropeanBalto-SlavicHigh
sinSinhalaSinhalaIndo-EuropeanIndo-AryanLow
slkSlovakLatinIndo-EuropeanBalto-SlavicMid
slvSlovenianLatinIndo-EuropeanBalto-SlavicMid
smoSamoanLatinAustronesianMalayo-PolynesianLow
snaShonaLatinIndo-EuropeanIndo-AryanLow
sndSindhiArabicIndo-EuropeanIndo-AryanLow
somSomaliLatinAfro-AsiaticCushiticLow
sotSouthern SothoLatinAtlantic-CongoBenue-CongoLow
spaSpanishLatinIndo-EuropeanItalicHigh
sqiAlbanianLatinIndo-EuropeanAlbanianLow
srpSerbianCyrillicIndo-EuropeanBalto-SlavicHigh
sunSundaneseLatinAustronesianMalayo-PolynesianLow
swaSwahiliLatinAtlantic-CongoBenue-CongoLow
sweSwedishLatinIndo-EuropeanGermanicHigh
tamTamilTamilDravidianSouth DravidianMid
telTeluguTeluguDravidianSouth DravidianLow
tgkTajikCyrillicIndo-EuropeanIranianLow
thaThaiThaiTai-KadaiKam-TaiMid
turTurkishLatinTurkicCommon TurkicHigh
twiTwiLatinAtlantic-CongoNiger-CongoLow
ukrUkrainianCyrillicIndo-EuropeanBalto-SlavicMid
urdUrduArabicIndo-EuropeanIndo-AryanMid
uzbUzbekLatinTurkicCommon TurkicMid
vieVietnameseLatinAustroasiaticVieticHigh
xhoXhosaLatinAtlantic-CongoBenue-CongoLow
yidYiddishHebrewIndo-EuropeanGermanicLow
yorYorubaLatinAtlantic-CongoBenue-CongoLow
zhoChineseHanSino-TibetanSiniticHigh
zulZuluLatinAtlantic-CongoBenue-CongoLow

Model Card Contact

For errors in this model card, contact Ahmet or Viraat, {ahmet, viraat} at cohere dot com.

Contributors

AU
Ahmet Ustun

10 commits

viraat

5 commits

alexrs

3 commits

coherecode

2 commits

CohereForAI/aya-101

Model

665

stars

23

commits

9

repos using this model

2

linked in READMEs

Sep 10, 2025

updated

afr
amh
ara
aze
bel
ben
bul
cat
ceb
ces
cym
dan
deu
ell
endpoints_compatible
eng
epo
est
eus
fil
fin
fra
fry
gla
gle
glg
guj
hat
hau
heb
hin
hun
hye
ibo
ind
isl
ita
jav
jpn
kan
kat
kaz
khm
kir
kor
kur
lao
lat
lav
lit
ltz
mal
mar
mkd
mlg
mlt
mon
mri
msa
mya
nep
nld
nor
nso
nya
ory
pan
pes
pol
por
pus
ron
rus
safetensors
sin
slk
slv
smo
sna
snd
som
sot
spa
sqi
srp
sun
swa
swe
t5
tam
tel
text2text-generation
text-generation-inference
tgk
tha
transformers
tur
twi
ukr
urd
uzb
vie
xho
yid
yor
zho
zul

README

Aya model summary image

Model Card for Aya 101

Model Summary

The Aya model is a massively multilingual generative language model that follows instructions in 101 languages. Aya outperforms mT0 and BLOOMZ a wide variety of automatic and human evaluations despite covering double the number of languages. The Aya model is trained using xP3x, Aya Dataset, Aya Collection, a subset of DataProvenance collection and ShareGPT-Command. We release the checkpoints under a Apache-2.0 license to further our mission of multilingual technologies empowering a multilingual world.

Use

# pip install -q transformers
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer

checkpoint = "CohereLabs/aya-101"

tokenizer = AutoTokenizer.from_pretrained(checkpoint)
aya_model = AutoModelForSeq2SeqLM.from_pretrained(checkpoint)

# Turkish to English translation
tur_inputs = tokenizer.encode("Translate to English: Aya cok dilli bir dil modelidir.", return_tensors="pt")
tur_outputs = aya_model.generate(tur_inputs, max_new_tokens=128)
print(tokenizer.decode(tur_outputs[0]))
# Aya is a multi-lingual language model

# Q: Why are there so many languages in India?
hin_inputs = tokenizer.encode("भारत में इतनी सारी भाषाएँ क्यों हैं?", return_tensors="pt")
hin_outputs = aya_model.generate(hin_inputs, max_new_tokens=128)
print(tokenizer.decode(hin_outputs[0]))
# Expected output: भारत में कई भाषाएँ हैं और विभिन्न भाषाओं के बोली जाने वाले लोग हैं। यह विभिन्नता भाषाई विविधता और सांस्कृतिक विविधता का परिणाम है। Translates to "India has many languages and people speaking different languages. This diversity is the result of linguistic diversity and cultural diversity."

Model Details

Finetuning

  • Architecture: Same as mt5-xxl
  • Number of Samples seen during Finetuning: 25M
  • Batch size: 256
  • Hardware: TPUv4-128
  • Software: T5X, Jax

Data Sources

The Aya model is trained on the following datasets:

All datasets are subset to the 101 languages supported by mT5. See the paper for details about filtering and pruning.

Evaluation

We refer to Section 5 from our paper for multilingual eval across 99 languages – including discriminative and generative tasks, human evaluation, and simulated win rates that cover both held-out tasks and in-distribution performance.

Bias, Risks, and Limitations

For a detailed overview of our effort at safety mitigation and benchmarking toxicity and bias across multiple languages, we refer to Sections 6 and 7 of our paper: Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model.

We hope that the release of the Aya model will make community-based redteaming efforts possible, by exposing an open-source massively-multilingual model for community research.

Citation

BibTeX:

@article{üstün2024aya,
  title={Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model},
  author={Ahmet Üstün and Viraat Aryabumi and Zheng-Xin Yong and Wei-Yin Ko and Daniel D'souza and Gbemileke Onilude and Neel Bhandari and Shivalika Singh and Hui-Lee Ooi and Amr Kayid and Freddie Vargus and Phil Blunsom and Shayne Longpre and Niklas Muennighoff and Marzieh Fadaee and Julia Kreutzer and Sara Hooker},
  journal={arXiv preprint arXiv:2402.07827},
  year={2024}
}

Languages Covered

Click to see Languages Covered

Below is the list of languages used in finetuning the Aya Model. We group languages into higher-, mid-, and lower-resourcedness based on a language classification by Joshi et. al, 2020. For further details, we refer to our paper

ISO CodeLanguage NameScriptFamilySubgroupingResourcedness
afrAfrikaansLatinIndo-EuropeanGermanicMid
amhAmharicGe'ezAfro-AsiaticSemiticLow
araArabicArabicAfro-AsiaticSemiticHigh
azeAzerbaijaniArabic/LatinTurkicCommon TurkicLow
belBelarusianCyrillicIndo-EuropeanBalto-SlavicMid
benBengaliBengaliIndo-EuropeanIndo-AryanMid
bulBulgarianCyrillicIndo-EuropeanBalto-SlavicMid
catCatalanLatinIndo-EuropeanItalicHigh
cebCebuanoLatinAustronesianMalayo-PolynesianMid
cesCzechLatinIndo-EuropeanBalto-SlavicHigh
cymWelshLatinIndo-EuropeanCelticLow
danDanishLatinIndo-EuropeanGermanicMid
deuGermanLatinIndo-EuropeanGermanicHigh
ellGreekGreekIndo-EuropeanGraeco-PhrygianMid
engEnglishLatinIndo-EuropeanGermanicHigh
epoEsperantoLatinConstructedEsperanticLow
estEstonianLatinUralicFinnicMid
eusBasqueLatinBasque-High
finFinnishLatinUralicFinnicHigh
filTagalogLatinAustronesianMalayo-PolynesianMid
fraFrenchLatinIndo-EuropeanItalicHigh
fryWestern FrisianLatinIndo-EuropeanGermanicLow
glaScottish GaelicLatinIndo-EuropeanCelticLow
gleIrishLatinIndo-EuropeanCelticLow
glgGalicianLatinIndo-EuropeanItalicMid
gujGujaratiGujaratiIndo-EuropeanIndo-AryanLow
hatHaitian CreoleLatinIndo-EuropeanItalicLow
hauHausaLatinAfro-AsiaticChadicLow
hebHebrewHebrewAfro-AsiaticSemiticMid
hinHindiDevanagariIndo-EuropeanIndo-AryanHigh
hunHungarianLatinUralic-High
hyeArmenianArmenianIndo-EuropeanArmenicLow
iboIgboLatinAtlantic-CongoBenue-CongoLow
indIndonesianLatinAustronesianMalayo-PolynesianMid
islIcelandicLatinIndo-EuropeanGermanicLow
itaItalianLatinIndo-EuropeanItalicHigh
javJavaneseLatinAustronesianMalayo-PolynesianLow
jpnJapaneseJapaneseJaponicJapanesicHigh
kanKannadaKannadaDravidianSouth DravidianLow
katGeorgianGeorgianKartvelianGeorgian-ZanMid
kazKazakhCyrillicTurkicCommon TurkicMid
khmKhmerKhmerAustroasiaticKhmericLow
kirKyrgyzCyrillicTurkicCommon TurkicLow
korKoreanHangulKoreanicKoreanHigh
kurKurdishLatinIndo-EuropeanIranianLow
laoLaoLaoTai-KadaiKam-TaiLow
lavLatvianLatinIndo-EuropeanBalto-SlavicMid
latLatinLatinIndo-EuropeanItalicMid
litLithuanianLatinIndo-EuropeanBalto-SlavicMid
ltzLuxembourgishLatinIndo-EuropeanGermanicLow
malMalayalamMalayalamDravidianSouth DravidianLow
marMarathiDevanagariIndo-EuropeanIndo-AryanLow
mkdMacedonianCyrillicIndo-EuropeanBalto-SlavicLow
mlgMalagasyLatinAustronesianMalayo-PolynesianLow
mltMalteseLatinAfro-AsiaticSemiticLow
monMongolianCyrillicMongolic-KhitanMongolicLow
mriMaoriLatinAustronesianMalayo-PolynesianLow
msaMalayLatinAustronesianMalayo-PolynesianMid
myaBurmeseMyanmarSino-TibetanBurmo-QiangicLow
nepNepaliDevanagariIndo-EuropeanIndo-AryanLow
nldDutchLatinIndo-EuropeanGermanicHigh
norNorwegianLatinIndo-EuropeanGermanicLow
nsoNorthern SothoLatinAtlantic-CongoBenue-CongoLow
nyaChichewaLatinAtlantic-CongoBenue-CongoLow
oryOriyaOriyaIndo-EuropeanIndo-AryanLow
panPunjabiGurmukhiIndo-EuropeanIndo-AryanLow
pesPersianArabicIndo-EuropeanIranianHigh
polPolishLatinIndo-EuropeanBalto-SlavicHigh
porPortugueseLatinIndo-EuropeanItalicHigh
pusPashtoArabicIndo-EuropeanIranianLow
ronRomanianLatinIndo-EuropeanItalicMid
rusRussianCyrillicIndo-EuropeanBalto-SlavicHigh
sinSinhalaSinhalaIndo-EuropeanIndo-AryanLow
slkSlovakLatinIndo-EuropeanBalto-SlavicMid
slvSlovenianLatinIndo-EuropeanBalto-SlavicMid
smoSamoanLatinAustronesianMalayo-PolynesianLow
snaShonaLatinIndo-EuropeanIndo-AryanLow
sndSindhiArabicIndo-EuropeanIndo-AryanLow
somSomaliLatinAfro-AsiaticCushiticLow
sotSouthern SothoLatinAtlantic-CongoBenue-CongoLow
spaSpanishLatinIndo-EuropeanItalicHigh
sqiAlbanianLatinIndo-EuropeanAlbanianLow
srpSerbianCyrillicIndo-EuropeanBalto-SlavicHigh
sunSundaneseLatinAustronesianMalayo-PolynesianLow
swaSwahiliLatinAtlantic-CongoBenue-CongoLow
sweSwedishLatinIndo-EuropeanGermanicHigh
tamTamilTamilDravidianSouth DravidianMid
telTeluguTeluguDravidianSouth DravidianLow
tgkTajikCyrillicIndo-EuropeanIranianLow
thaThaiThaiTai-KadaiKam-TaiMid
turTurkishLatinTurkicCommon TurkicHigh
twiTwiLatinAtlantic-CongoNiger-CongoLow
ukrUkrainianCyrillicIndo-EuropeanBalto-SlavicMid
urdUrduArabicIndo-EuropeanIndo-AryanMid
uzbUzbekLatinTurkicCommon TurkicMid
vieVietnameseLatinAustroasiaticVieticHigh
xhoXhosaLatinAtlantic-CongoBenue-CongoLow
yidYiddishHebrewIndo-EuropeanGermanicLow
yorYorubaLatinAtlantic-CongoBenue-CongoLow
zhoChineseHanSino-TibetanSiniticHigh
zulZuluLatinAtlantic-CongoBenue-CongoLow

Model Card Contact

For errors in this model card, contact Ahmet or Viraat, {ahmet, viraat} at cohere dot com.

Contributors

AU
Ahmet Ustun

10 commits

viraat

5 commits

alexrs

3 commits

coherecode

2 commits