129 repos across 6 sub-areas
Libraries, tools, and applications for automatic speech-to-text conversion, with heavy focus on OpenAI's Whisper model and its deployment across multiple platforms. The cluster spans Python implementations for training and inference, Swift/iOS integrations, Rust-based optimizations for edge deployment, and specialized tools like speaker diarization and real-time transcription. Developers here will find production-ready inference engines, model optimization techniques, and cross-platform wrappers enabling Whisper deployment from cloud services to mobile and embedded devices.
Speech-to-Text Applications & Tools
31 repos
Cross-platform applications and libraries for converting spoken audio to text, with a strong emphasis on accessibility features and lightweight implementations. The cluster spans multiple platforms—primarily Swift for macOS/iOS clients, Python for backends and scripting, and Rust for performance-critical components—often built around OpenAI's Whisper model or similar speech recognition engines. Projects range from system-level HUD overlays and accessibility integrations to standalone transcription utilities, many using modern frameworks like Tauri for cross-platform desktop deployment.
Cluster 635234
27 repos
Cluster 635236
26 repos
Cluster 635238
18 repos
Cluster 635235
15 repos
Speech Recognition and Audio Processing SDKs
12 repos
Language-specific SDKs and libraries for speech-to-text and audio processing, with a focus on offline speech recognition and machine learning inference. The cluster centers on the Deepgram SDK ecosystem (Node.js, JavaScript, Python, Go, .NET) alongside supporting ML frameworks and audio processing tools. This collection is valuable for developers building voice-enabled applications that need cross-language SDK support or exploring offline speech recognition techniques.