10 repos
Models and systems for translating between and embedding text across multiple languages, with emphasis on many-to-many translation and cross-lingual understanding. The cluster centers on large pretrained transformer models (mBART, mT5, M2M-100) that handle dozens of languages simultaneously, alongside multilingual embedding models like Nomic. These repositories represent both the model architectures themselves and their applications in production translation and semantic search across diverse language pairs.