1,759
stars
0
commits
10
repos using this model
17
linked in READMEs
May 10, 2024
updated
pyannote/segmentation
695
pyannote/speaker-diarization-3.1
3,523
pyannote/speaker-diarization
1,326
pyannote/speaker-diarization-community-1
1,594
pyannote/overlapped-speech-detection
64
pyannote/voice-activity-detection
241
pyannote/brouhaha
31
pyannote/embedding
236
modelscope/3D-Speaker
A Repository for Single- and Multi-modal Speaker Verification, Speaker Recognition and Speaker…
3,132
debpalash/VoiceStudio
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design,…
21,995
QuentinFuxa/WhisperLiveKit
Real-time, local speech-to-text with streaming ASR, speaker diarization, translation, and…
11,013
salute-developers/GigaAM
Foundational Model for Speech Recognition Tasks
790
revdotcom/reverb
Open source inference code for Rev's model
440
huggingface/diarizers
331
ai-bot-pro/achatbot
An open source chat bot architecture for voice/vision (and multimodal) assistants, local(CPU/GPU…
89
JaesungHuh/SimpleDiarization
Simple diarization model
53
PaulKinlan/web-ai-showcase
Every AI model you can run locally in a browser — interactive explainers, live controls, see-inside…
4
xolra0d/auto_audio
1
juanmc2005/diart
A python package to build AI-powered real-time audio applications
2,023
kadirnar/whisper-plus
WhisperPlus: Faster, Smarter, and More Capable 🚀
1,957
john-rocky/CoreML-Models
Core ML model zoo for iOS/macOS — PyTorch models converted to ready-to-use .mlpackage, each with a…
1,863
asiff00/On-Device-Speech-to-Speech-Conversational-AI
This is an on-CPU real-time conversational system for two-way speech communication with AI models,…
257
ReisCook/Voice_Extractor
Automated speech dataset creator
225
thainph/node-trans
Real-time audio translation app
125
murtaza-nasir/whisperx-asr-service
99
Gr122lyBr/voicetag
Speaker identification powered by pyannote and resemblyzer
MatiasDiBernardo/Lowcost-ITW-curation
Low-cost, CPU-friendly preprocessing pipeline and metric toolkit to curate and evaluate in-the-wild…
7
biyachuev/yt-transcriber
AI-powered audio/video processing: transcription, speaker diarization, LLM refinement, translation…
Possum/CogFluency
Measure Cognitive Friction of speeches using competitive Toastmasters derived model
D4X-max/Speaker-Diarization-VAD-CAS
VAD & Clustering Audio Segments
frederik-ai/diart-quantization
vinnyvinnyvinny/local-transcribe
Source code for local transcribe