We present Monolingual-Quechua-IIC, a monolingual corpus of Southern Quechua, which can be used to build language models using Transformers models. This corpus also includes the Wiki and OSCAR corpora. We used this corpus to build Llama-RoBERTa-Quechua, the first language model for Southern Quechua using Transformers.
Southern Quechua
Apache-2.0
@inproceedings{zevallos2022introducing,
title={Introducing QuBERT: A Large Monolingual Corpus and BERT Model for Southern Quechua},
author={Zevallos, Rodolfo and Ortega, John and Chen, William and Castro, Richard and Bel, Nuria and Toshio, Cesar and Venturas, Renzo and Aradiel, Hilario and Melgarejo, Nelsi},
booktitle={Proceedings of the Third Workshop on Deep Learning for Low-Resource Natural Language Processing},
pages={1--13},
year={2022}
}
Thanks to @rjzevallos for adding this dataset.
27 commits
2 commits
We present Monolingual-Quechua-IIC, a monolingual corpus of Southern Quechua, which can be used to build language models using Transformers models. This corpus also includes the Wiki and OSCAR corpora. We used this corpus to build Llama-RoBERTa-Quechua, the first language model for Southern Quechua using Transformers.
Southern Quechua
Apache-2.0
@inproceedings{zevallos2022introducing,
title={Introducing QuBERT: A Large Monolingual Corpus and BERT Model for Southern Quechua},
author={Zevallos, Rodolfo and Ortega, John and Chen, William and Castro, Richard and Bel, Nuria and Toshio, Cesar and Venturas, Renzo and Aradiel, Hilario and Melgarejo, Nelsi},
booktitle={Proceedings of the Third Workshop on Deep Learning for Low-Resource Natural Language Processing},
pages={1--13},
year={2022}
}
Thanks to @rjzevallos for adding this dataset.
27 commits
2 commits