238 repos across 11 sub-areas
Multilingual Text-to-Speech Systems
37 repos
Text-to-speech (TTS) synthesis systems supporting multiple languages including English, German, Chinese, Korean, and Spanish. The cluster centers on open-source TTS implementations, from real-time inference systems to specialized models like MOSS-TTS variants and Bark, enabling developers to generate natural-sounding speech across diverse linguistic contexts. Repositories here cover model architectures, inference optimization, and language-specific adaptations for practical speech generation applications.
Cluster 462303
36 repos
Large Language Models with PyTorch
33 repos
PyTorch-based implementations and tooling for large language models, with multilingual support (French and English). The cluster centers on transformer-based models of varying scales, from billions to hundreds of millions of parameters, using the safetensors format for efficient model serialization and distribution. Repositories here provide model weights, training utilities, inference frameworks, and supporting infrastructure for working with modern LLMs in production and research contexts.
Multilingual Large Language Models
31 repos
Open-source and commercial large language models optimized for instruction-following and multilingual support across English, Spanish, French, German, and Italian. This cluster includes compact efficient models like Mistral Nemo and Ministral variants alongside larger vision-capable models like Pixtral, representing a focus on practical deployment of performant LLMs across diverse languages and resource constraints. The repos here emphasize model weights, quantized variants (GGUF formats), and inference-ready implementations rather than training frameworks.
Multilingual Speech Recognition and ASR
30 repos
Automatic speech recognition (ASR) systems and models supporting multiple languages including Spanish, German, English, French, and Portuguese. The cluster centers on accessible speech-to-text implementations, with prominent models like OpenAI's Whisper series (large, base, medium, small, tiny) providing open-source foundations for building multilingual voice interfaces. Repositories here cover model implementations, fine-tuning approaches, and practical ASR applications across diverse language communities.
Large Language Models & Transformers
21 repos
Transformer-based language models and the supporting infrastructure for working with them. This cluster centers on Meta's Llama model series across multiple scales (1B through 405B parameters), along with the broader ecosystem of transformer libraries and model serialization formats like safetensors. Repositories here cover model implementations, inference optimization, and interoperability across different frameworks.
Cluster 462312
16 repos
Multilingual NLP and Legal Language Models
13 repos
Pre-trained transformer models and tools for natural language processing across multiple languages (French, English, German, Spanish) with a focus on semantic similarity, sentence representation, and legal domain applications. The cluster centers on PyTorch-based models like legal-xlm variants and multilingual sentence embeddings, designed for cross-lingual understanding, paraphrase detection, and specialized tasks like sentence boundary detection. This is a resource collection for practitioners building multilingual NLP systems and legal tech applications.
Multilingual Machine Translation & Embeddings
10 repos
Models and systems for translating between and embedding text across multiple languages, with emphasis on many-to-many translation and cross-lingual understanding. The cluster centers on large pretrained transformer models (mBART, mT5, M2M-100) that handle dozens of languages simultaneously, alongside multilingual embedding models like Nomic. These repositories represent both the model architectures themselves and their applications in production translation and semantic search across diverse language pairs.
Cluster 462307
7 repos
Liquid language and edge deployment
4 repos
Multi-language support for the Liquid template engine and related edge computing tools, with a focus on cross-language implementations and runtime optimization. The cluster centers on Liquid-based projects across multiple programming languages, alongside edge deployment and inference optimization frameworks—reflected in the prevalence of 'liquid', 'edge', and multilingual topic tags (de, en, es).