Cluster 448828

238 repos across 11 sub-areas

Multilingual Text-to-Speech Systems

37 repos

Text-to-speech (TTS) synthesis systems supporting multiple languages including English, German, Chinese, Korean, and Spanish. The cluster centers on open-source TTS implementations, from real-time inference systems to specialized models like MOSS-TTS variants and Bark, enabling developers to generate natural-sounding speech across diverse linguistic contexts. Repositories here cover model architectures, inference optimization, and language-specific adaptations for practical speech generation applications.

Cluster 462303

36 repos

Large Language Models with PyTorch

33 repos

PyTorch-based implementations and tooling for large language models, with multilingual support (French and English). The cluster centers on transformer-based models of varying scales, from billions to hundreds of millions of parameters, using the safetensors format for efficient model serialization and distribution. Repositories here provide model weights, training utilities, inference frameworks, and supporting infrastructure for working with modern LLMs in production and research contexts.

Multilingual Large Language Models

31 repos

Open-source and commercial large language models optimized for instruction-following and multilingual support across English, Spanish, French, German, and Italian. This cluster includes compact efficient models like Mistral Nemo and Ministral variants alongside larger vision-capable models like Pixtral, representing a focus on practical deployment of performant LLMs across diverse languages and resource constraints. The repos here emphasize model weights, quantized variants (GGUF formats), and inference-ready implementations rather than training frameworks.

Multilingual Speech Recognition and ASR

30 repos

Automatic speech recognition (ASR) systems and models supporting multiple languages including Spanish, German, English, French, and Portuguese. The cluster centers on accessible speech-to-text implementations, with prominent models like OpenAI's Whisper series (large, base, medium, small, tiny) providing open-source foundations for building multilingual voice interfaces. Repositories here cover model implementations, fine-tuning approaches, and practical ASR applications across diverse language communities.

Large Language Models & Transformers

21 repos

Transformer-based language models and the supporting infrastructure for working with them. This cluster centers on Meta's Llama model series across multiple scales (1B through 405B parameters), along with the broader ecosystem of transformer libraries and model serialization formats like safetensors. Repositories here cover model implementations, inference optimization, and interoperability across different frameworks.

Cluster 462312

16 repos

Multilingual NLP and Legal Language Models

13 repos

Pre-trained transformer models and tools for natural language processing across multiple languages (French, English, German, Spanish) with a focus on semantic similarity, sentence representation, and legal domain applications. The cluster centers on PyTorch-based models like legal-xlm variants and multilingual sentence embeddings, designed for cross-lingual understanding, paraphrase detection, and specialized tasks like sentence boundary detection. This is a resource collection for practitioners building multilingual NLP systems and legal tech applications.

Multilingual Machine Translation & Embeddings

10 repos

Models and systems for translating between and embedding text across multiple languages, with emphasis on many-to-many translation and cross-lingual understanding. The cluster centers on large pretrained transformer models (mBART, mT5, M2M-100) that handle dozens of languages simultaneously, alongside multilingual embedding models like Nomic. These repositories represent both the model architectures themselves and their applications in production translation and semantic search across diverse language pairs.

Cluster 462307

7 repos

Liquid language and edge deployment

mixed

4 repos

Multi-language support for the Liquid template engine and related edge computing tools, with a focus on cross-language implementations and runtime optimization. The cluster centers on Liquid-based projects across multiple programming languages, alongside edge deployment and inference optimization frameworks—reflected in the prevalence of 'liquid', 'edge', and multilingual topic tags (de, en, es).