Multilingual Text Embeddings

24 repos

Models and tools for generating vector embeddings from text across multiple languages, enabling semantic search, similarity matching, and downstream NLP tasks in language-agnostic ways. The cluster centers on pretrained embedding models like Snowflake Arctic Embed, multilingual E5, and XLM-RoBERTa variants, along with frameworks and utilities for deploying and using these embeddings in production systems. Developers working on cross-lingual search, recommendation systems, or multilingual machine learning applications would find both the foundational models and integration tooling here.

en ·4,854
de ·4,854
fr ·4,850
es ·4,847
ar ·4,835
it ·4,739
pt ·4,732
nl ·4,675
hi ·4,615
ru ·4,615