[NeurIPS 2024] 🕸 GlotCC Dataset and Pipline
21
stars
14
commits
Jupyter Notebook
primary language
Apr 6, 2025
updated
GlotCC is a multilingual corpus built by the GlotLID language identification and cisnlp/Ungoliant pipeline from CommonCrawl.
Lastest version supports more than 1000 languages and is filtered based on adopted filters from C4, CCNet, MADLAD-400, RedPajama-Data-v2, OSCAR, Gopher, RefinedWeb, FineWeb, Datatrove, Dolma, Pile-CC, Pretrainer's Guide, and GlotScript. ™ The logo features a llama with the style of C.C. from the Code Geass anime reading a book.
GlotCC Dataset, Version 1: https://huggingface.co/datasets/cis-lmu/GlotCC-V1
We forked oscar-project/ungoliant to cisnlp/ungoliant and made the necessary changes to integrate it with the GlotLID language identification model.
For detailed instructions on running the pipeline, refer to the cisnlp/ungoliant repository. The README is up-to-date.
If you find our repo and data useful for your research, please cite:
@article{kargaran2024glotcc,
title = {Glot{CC}: An Open Broad-Coverage CommonCrawl Corpus and Pipeline for Minority Languages},
author = {Kargaran, Amir Hossein and Yvon, Fran{\c{c}}ois and Sch{\"u}tze, Hinrich},
journal = {Advances in Neural Information Processing Systems},
year = {2024},
url = {https://arxiv.org/abs/2410.23825}
}
14 commits
Jupyter Notebook
100.0%
[NeurIPS 2024] 🕸 GlotCC Dataset and Pipline
21
stars
14
commits
Jupyter Notebook
primary language
Apr 6, 2025
updated
GlotCC is a multilingual corpus built by the GlotLID language identification and cisnlp/Ungoliant pipeline from CommonCrawl.
Lastest version supports more than 1000 languages and is filtered based on adopted filters from C4, CCNet, MADLAD-400, RedPajama-Data-v2, OSCAR, Gopher, RefinedWeb, FineWeb, Datatrove, Dolma, Pile-CC, Pretrainer's Guide, and GlotScript. ™ The logo features a llama with the style of C.C. from the Code Geass anime reading a book.
GlotCC Dataset, Version 1: https://huggingface.co/datasets/cis-lmu/GlotCC-V1
We forked oscar-project/ungoliant to cisnlp/ungoliant and made the necessary changes to integrate it with the GlotLID language identification model.
For detailed instructions on running the pipeline, refer to the cisnlp/ungoliant repository. The README is up-to-date.
If you find our repo and data useful for your research, please cite:
@article{kargaran2024glotcc,
title = {Glot{CC}: An Open Broad-Coverage CommonCrawl Corpus and Pipeline for Minority Languages},
author = {Kargaran, Amir Hossein and Yvon, Fran{\c{c}}ois and Sch{\"u}tze, Hinrich},
journal = {Advances in Neural Information Processing Systems},
year = {2024},
url = {https://arxiv.org/abs/2410.23825}
}
14 commits
Jupyter Notebook
100.0%