11 repos
Large-scale text datasets and language resources designed for training and evaluating multilingual natural language processing models. These repositories provide curated collections spanning hundreds of languages, from common to low-resource variants, enabling researchers to build and benchmark cross-lingual systems. The cluster includes corpus collection projects, dataset processing pipelines, and foundational resources for multilingual language model training.