Multilingual LLMs and Speech Models

238 repos across 11 sub-areas

Large language models and speech synthesis systems designed for multiple languages, particularly non-English languages like Spanish, French, German, Portuguese, and others. The cluster centers on open-source model implementations and fine-tuned variants optimized for multilingual understanding, generation, and audio synthesis, with notable models like Aya (focused on non-English languages), RWKV, and Fish Speech representing different approaches to scaling language capabilities across linguistic boundaries.

Multilingual Text-to-Speech Systems

37 repos

Text-to-speech (TTS) synthesis systems supporting multiple languages including English, German, Chinese, Korean, and Spanish. The cluster centers on open-source TTS implementations, from real-time inference systems to specialized models like MOSS-TTS variants and Bark, enabling developers to generate natural-sounding speech across diverse linguistic contexts. Repositories here cover model architectures, inference optimization, and language-specific adaptations for practical speech generation applications.

Multilingual Model Conversion & ONNX Optimization

36 repos

Tools and model artifacts for converting large language models to ONNX format for efficient inference across multiple languages. The cluster centers on pre-converted model repositories (LFM2, LFM2.5 series) optimized for deployment, alongside supporting infrastructure for safetensors handling and cross-language (English, French, German, Spanish) model serving. This area addresses the practical challenge of taking large pretrained models and preparing them for production inference with reduced computational requirements.

Large Language Models with PyTorch

33 repos

PyTorch-based implementations and tooling for large language models, with multilingual support (French and English). The cluster centers on transformer-based models of varying scales, from billions to hundreds of millions of parameters, using the safetensors format for efficient model serialization and distribution. Repositories here provide model weights, training utilities, inference frameworks, and supporting infrastructure for working with modern LLMs in production and research contexts.

Multilingual Large Language Models

31 repos

Open-source and commercial large language models optimized for instruction-following and multilingual support across English, Spanish, French, German, and Italian. This cluster includes compact efficient models like Mistral Nemo and Ministral variants alongside larger vision-capable models like Pixtral, representing a focus on practical deployment of performant LLMs across diverse languages and resource constraints. The repos here emphasize model weights, quantized variants (GGUF formats), and inference-ready implementations rather than training frameworks.

Multilingual Speech Recognition and ASR

30 repos

Automatic speech recognition (ASR) systems and models supporting multiple languages including Spanish, German, English, French, and Portuguese. The cluster centers on accessible speech-to-text implementations, with prominent models like OpenAI's Whisper series (large, base, medium, small, tiny) providing open-source foundations for building multilingual voice interfaces. Repositories here cover model implementations, fine-tuning approaches, and practical ASR applications across diverse language communities.

Large Language Models & Transformers

21 repos

Transformer-based language models and the supporting infrastructure for working with them. This cluster centers on Meta's Llama model series across multiple scales (1B through 405B parameters), along with the broader ecosystem of transformer libraries and model serialization formats like safetensors. Repositories here cover model implementations, inference optimization, and interoperability across different frameworks.

Multilingual Named Entity Recognition & PII Detection

16 repos

Tools and models for extracting named entities and detecting personally identifiable information (PII) across multiple languages including French, English, German, Italian, and Spanish. The cluster centers on GLiNER-based models and frameworks that provide language-agnostic or multilingual capabilities for information extraction tasks, with particular emphasis on privacy-focused PII filtering and detection pipelines. These repositories offer both base models and specialized variants designed for production guardrails and data protection workflows.

Multilingual NLP and Legal Language Models

13 repos

Pre-trained transformer models and tools for natural language processing across multiple languages (French, English, German, Spanish) with a focus on semantic similarity, sentence representation, and legal domain applications. The cluster centers on PyTorch-based models like legal-xlm variants and multilingual sentence embeddings, designed for cross-lingual understanding, paraphrase detection, and specialized tasks like sentence boundary detection. This is a resource collection for practitioners building multilingual NLP systems and legal tech applications.

Multilingual Machine Translation & Embeddings

10 repos

Models and systems for translating between and embedding text across multiple languages, with emphasis on many-to-many translation and cross-lingual understanding. The cluster centers on large pretrained transformer models (mBART, mT5, M2M-100) that handle dozens of languages simultaneously, alongside multilingual embedding models like Nomic. These repositories represent both the model architectures themselves and their applications in production translation and semantic search across diverse language pairs.

Text-to-Speech Models & Synthesis

7 repos

Open-source text-to-speech (TTS) models and implementations across multiple languages including English, Chinese, Italian, and Spanish. The cluster features efficient model architectures like Qwen3-TTS (0.6B parameters) and Darwin-TTS (1.7B), alongside voice cloning and customization tools. Repositories here span from base model implementations to cross-tokenizer variants and practical voice synthesis applications, representing both research models and production-ready TTS systems.

Liquid language and edge deployment

mixed

4 repos

Multi-language support for the Liquid template engine and related edge computing tools, with a focus on cross-language implementations and runtime optimization. The cluster centers on Liquid-based projects across multiple programming languages, alongside edge deployment and inference optimization frameworks—reflected in the prevalence of 'liquid', 'edge', and multilingual topic tags (de, en, es).