16 repos
Multilingual translations and localization of the Open Assistant (OASST) dataset, a crowdsourced collection of conversational data for training open-source language models. The cluster centers on coordinating translation efforts across multiple languages including Bengali, Portuguese, Hindi, Dutch, Spanish, Turkish, and others, making dialogue datasets accessible for non-English AI training and evaluation. These repositories primarily serve as translation hubs and language-specific data repositories rather than as distinct software tools or frameworks.