242 repos across 12 sub-areas
Open-source language models and speech synthesis systems with support for multiple languages including English, French, Spanish, German, and Portuguese. The cluster includes compact models optimized for efficiency (like the S1-Mini variants), large instruction-tuned models (Aya series), and speech-to-text/text-to-speech systems (Fish Speech). These projects emphasize making multilingual AI accessible through open weights and community-driven development.
Liquid Finance & LFM2 Models
34 repos
Machine learning models and implementations centered on Liquid Finance (LFM2) in multiple scales, packaged primarily as ONNX-format models for efficient inference. The cluster includes base model variants (350M through 2.6B parameters), specialized versions like RAG-augmented and experimental configurations, and supporting infrastructure for deployment. Repositories span model artifacts, conversion utilities, and integration examples across Japanese and English language support.
Text-to-Speech and Voice Generation
33 repos
Neural text-to-speech (TTS) systems and voice synthesis models that convert written text into spoken audio. The cluster centers on efficient TTS implementations including fish-speech, Qwen3-TTS, MOSS-TTS, and Bark—ranging from lightweight mobile-friendly models to customizable voice generation systems. These repositories span multilingual support (English, Italian, French, Korean, Spanish) and cover both pre-trained models and frameworks for building voice synthesis applications.
Cluster 647053
29 repos
Large Language Models and Text Generation
28 repos
Libraries, models, and infrastructure for building and deploying large language models (LLMs) with a focus on transformer-based architectures and efficient text generation. The cluster centers on the BLOOM model family across multiple scales (1B to 7B parameters) and related tooling for model serialization (safetensors), inference optimization (text-generation-inference), and PyTorch-based implementations. Repositories here provide both the model weights and the underlying frameworks needed to understand, fine-tune, and serve modern language models at various scales.
Multilingual Text Embeddings
24 repos
Models and tools for generating vector embeddings from text across multiple languages, enabling semantic search, similarity matching, and downstream NLP tasks in language-agnostic ways. The cluster centers on pretrained embedding models like Snowflake Arctic Embed, multilingual E5, and XLM-RoBERTa variants, along with frameworks and utilities for deploying and using these embeddings in production systems. Developers working on cross-lingual search, recommendation systems, or multilingual machine learning applications would find both the foundational models and integration tooling here.
Cluster 647049
21 repos
Multilingual Text-to-Speech Systems
19 repos
Text-to-speech (TTS) synthesis systems supporting multiple languages including English, German, Chinese, Korean, and Spanish. The cluster centers on open-source TTS implementations, from real-time inference systems to specialized models like MOSS-TTS variants and Bark, enabling developers to generate natural-sounding speech across diverse linguistic contexts. Repositories here cover model architectures, inference optimization, and language-specific adaptations for practical speech generation applications.
Cluster 647056
18 repos
Cluster 647047
16 repos
Large Language Models and Llama
9 repos
Meta's Llama model family across multiple scales and variants, from 1B to 405B parameters. These repositories contain trained model weights, instruction-tuned versions, and implementations designed for deployment and inference of state-of-the-art open-source language models. Developers using this cluster will find model artifacts, configuration details, and integration points for building applications with Llama models.
Multilingual Open-Source Language Models
6 repos
Large language models released as open-source alternatives to proprietary systems, with native or optimized support for multiple languages including French, English, German, and Spanish. The cluster centers on instruction-tuned model variants across different scales and architectures—from the dense Mistral Nemo and StableLM families to the sparse mixture-of-experts Mixtral models—designed for practical deployment and fine-tuning. These repositories contain model weights in safetensors format, documentation, and inference code that enable developers to run capable, multilingual reasoning systems without reliance on closed APIs.
Cluster 647045
5 repos