audio-cpp/audio.cpp-gguf

Model

118

stars

51

commits

1

linked in READMEs

Sep 9, 2026

updated

audio.cpp
audio-to-audio
automatic-speech-recognition
gguf
quantized
source-separation
speaker-diarization
speech
text-to-audio
text-to-speech
voice-conversion
Browse cluster: Speech Recognition and Audio Processing

README

audio.cpp GGUF Model Packages

This directory contains audio.cpp-native GGUF conversions of multiple speech models. These files are intended for use with audio.cpp.

If you enjoy the project, please star audio.cpp on GitHub and this Hugging Face repository.

For conversion details, supported layouts, direct-file loading, sidecar embedding, and the latest compatibility notes, see the audio.cpp GGUF guide:

!!! Converted and quantized packages are checked with automated metrics, but perceived quality can still differ for human listeners. Please validate the exact package, backend, and route to confirm the output is acceptable for your use case.

Files

The table lists the GGUF packages currently provided by this repository.

Directoryaudio.cpp familyGGUF providedOriginal model license
ACE-Step1.5-GGUFace_stepBF16 + Q8MIT
AudioSR-GGUFaudiosrF32MIT
BS-RoFormer-ep368-GGUFbs_roformerQ8Apache-2.0
Chatterbox-GGUFchatterboxF16 + Q8MIT
Citrinet-ASR-GGUFcitrinet_asrQ8CC-BY-4.0
Confucius4-TTS-GGUFconfucius4_ttsoriginalApache-2.0
ControlFoley-GGUFcontrolfoleyF32Apache-2.0
DotTTS-Edit-GGUFdots_ttsBF16 + Q8Apache-2.0
DotTTS-MF-GGUFdots_ttsBF16Apache-2.0
DotTTS-SOAR-GGUFdots_ttsoriginal + BF16Apache-2.0
DramaBox-GGUFdramaboxQ8LTX-2 Community License
Fish-Audio-S2-Pro-GGUFfish_audioBF16 + Q8Fish Audio Research License
Fun-ASR-Nano-2512-GGUFfun_asr_nanoF16 + Q8FunASR Model Open Source License Agreement v1.1
HeartMuLa-GGUFheartmulaF16 + Q8Apache-2.0
HTDemucs-GGUFhtdemucsF16 + Q8MIT
Higgs-Audio-v3-STT-GGUFhiggs_audio_sttF16 + Q8Apache-2.0
Higgs-Audio-v3-TTS-4B-GGUFhiggs_audio_ttsBF16 + Q8Boson Higgs TTS 3 Research and Non-Commercial License
Hviske-v5.3-GGUFhviske_asrQ8CC-BY-NC-4.0
IndexTTS2-GGUFindex_tts2original + F16 + Q8bilibili Model Use License Agreement
IndexTTS2.5-GGUFindex_tts2original + F16 + Q8bilibili Model Use License Agreement
Inflect-Micro-v2-GGUFinflect_v2originalApache-2.0
Irodori-TTS-500M-v3-GGUFirodori_ttsF16 + Q8MIT
Irodori-TTS-600M-v3-VoiceDesign-GGUFirodori_ttsF16 + Q8MIT
Irodori-TTS-v4-Small-GGUFirodori_ttsF16 + Q8MIT
Kroko-ASR-GGUFkroko_asrQ8CC-BY-SA community model license
MMS-Forced-Aligner-GGUFmms_forced_alignerF16CC-BY-NC-4.0
MOSS-TTS-Local-v1.5-GGUFmoss_tts_localBF16 + Q8Apache-2.0
MOSS-TTS-Nano-100M-GGUFmoss_tts_nanoBF16 + Q8Apache-2.0
MOSS-VoiceGenerator-GGUFmoss_voicegenBF16 + F16 codec decodeApache-2.0
MagpieTTS-Multilingual-357M-GGUFmagpie_ttsoriginalNVIDIA Open Model License
MeanVC2-GGUFmeanvc2F32 + Q4_KApache-2.0
Mel-Band-RoFormer-GGUFmel_band_roformerF16 + Q8MIT
MiniMax-H3-Q4-GGUFminimax_h3Q4_K + INT8 DiT optionMiniMax-H3 Community License
MioCodec-25Hz-44.1kHz-v2-GGUFmiocodecoriginal + F16 + Q8MIT
MioTTS-1.7B-GGUFmiottsoriginal + BF16 + Q8Apache-2.0
MuScriptor-Small-GGUFmuscriptorF32CC-BY-NC-4.0
Nemotron-3.5-ASR-Streaming-0.6B-GGUFnemotron_asrF16 + Q8OpenMDW-1.1
NeuTTS-2E-GGUFneuttsoriginalApache-2.0
OmniVoice-GGUFomnivoiceBF16 + F16 + Q8Apache-2.0
Parakeet-TDT-0.6B-v3-GGUFparakeet_tdtF16 + Q8CC-BY-4.0
PersonaPlex-GGUFpersonaplexQ4_K + Q8Apache-2.0
PocketTTS-GGUFpocket_ttsBF16 + Q8CC-BY-4.0
Qwen3-ASR-0.6B-GGUFqwen3_asrF16 + Q8Apache-2.0
Qwen3-ASR-1.7B-GGUFqwen3_asrF16 + Q8Apache-2.0
Qwen3-ForcedAligner-0.6B-GGUFqwen3_forced_alignerF16 + Q8Apache-2.0
Qwen3-TTS-12Hz-0.6B-Base-GGUFqwen3_ttsBF16 + Q8Apache-2.0
Qwen3-TTS-12Hz-1.7B-Base-GGUFqwen3_ttsoriginal + BF16 + Q8Apache-2.0
Qwen3-TTS-12Hz-1.7B-CustomVoice-GGUFqwen3_ttsBF16 + Q8Apache-2.0
Qwen3-TTS-12Hz-1.7B-VoiceDesign-GGUFqwen3_ttsBF16 + Q8Apache-2.0
RVC-GGUFrvcF16MIT
SeedVC-MLX-GGUFseed_vcoriginal + F16 + Q8GPL-3.0
Sortformer-Diar-4spk-v1-GGUFsortformer_diarF16 + Q8CC-BY-NC-4.0
Stable-Audio-3-Medium-GGUFstable_audioF16 + Q8Stability AI Community License
Stable-Audio-3-Small-Music-GGUFstable_audioF16 + Q8Stability AI Community License
Stable-Audio-3-Small-SFX-GGUFstable_audioF16 + Q8Stability AI Community License
Supertonic-3-GGUFsupertonicoriginal + F16 + Q8BigScience Open RAIL-M
Vevo2-GGUFvevo2original + F16 + Q8CC-BY-NC-ND-4.0
VibeVoice-1.5B-GGUFvibevoiceBF16 + Q8 + Q4MIT
VibeVoice-ASR-GGUFvibevoice_asrF16 + Q8MIT
VoxCPM2-GGUFvoxcpm2original + BF16 + Q8Apache-2.0
Voxtral-Mini-4B-Realtime-2602-GGUFvoxtral_realtimeBF16 + Q8 + Q4_KApache-2.0

Usage

Pass a GGUF file directly as --model:

audiocpp_cli --task tts --family supertonic --model Supertonic-3-GGUF/supertonic-3-orig.gguf --backend cuda --language en --text "Hello." --voice-id M1 --out out.wav

For ASR:

audiocpp_cli --task asr --family qwen3_asr --model Qwen3-ASR-0.6B-GGUF/qwen3-asr-0.6b-f16.gguf --backend cuda --audio speech.wav --text "" --text-out transcript.txt

License

Each GGUF file is a converted form of its original model. Use and redistribution are governed by the corresponding original model license listed above. Please review the original model card and license terms before using or redistributing any converted weights.

Contributors

audio-cpp

51 commits

audio-cpp/audio.cpp-gguf

Model

118

stars

51

commits

1

linked in READMEs

Sep 9, 2026

updated

audio.cpp
audio-to-audio
automatic-speech-recognition
gguf
quantized
source-separation
speaker-diarization
speech
text-to-audio
text-to-speech
voice-conversion
Browse cluster: Speech Recognition and Audio Processing

README

audio.cpp GGUF Model Packages

This directory contains audio.cpp-native GGUF conversions of multiple speech models. These files are intended for use with audio.cpp.

If you enjoy the project, please star audio.cpp on GitHub and this Hugging Face repository.

For conversion details, supported layouts, direct-file loading, sidecar embedding, and the latest compatibility notes, see the audio.cpp GGUF guide:

!!! Converted and quantized packages are checked with automated metrics, but perceived quality can still differ for human listeners. Please validate the exact package, backend, and route to confirm the output is acceptable for your use case.

Files

The table lists the GGUF packages currently provided by this repository.

Directoryaudio.cpp familyGGUF providedOriginal model license
ACE-Step1.5-GGUFace_stepBF16 + Q8MIT
AudioSR-GGUFaudiosrF32MIT
BS-RoFormer-ep368-GGUFbs_roformerQ8Apache-2.0
Chatterbox-GGUFchatterboxF16 + Q8MIT
Citrinet-ASR-GGUFcitrinet_asrQ8CC-BY-4.0
Confucius4-TTS-GGUFconfucius4_ttsoriginalApache-2.0
ControlFoley-GGUFcontrolfoleyF32Apache-2.0
DotTTS-Edit-GGUFdots_ttsBF16 + Q8Apache-2.0
DotTTS-MF-GGUFdots_ttsBF16Apache-2.0
DotTTS-SOAR-GGUFdots_ttsoriginal + BF16Apache-2.0
DramaBox-GGUFdramaboxQ8LTX-2 Community License
Fish-Audio-S2-Pro-GGUFfish_audioBF16 + Q8Fish Audio Research License
Fun-ASR-Nano-2512-GGUFfun_asr_nanoF16 + Q8FunASR Model Open Source License Agreement v1.1
HeartMuLa-GGUFheartmulaF16 + Q8Apache-2.0
HTDemucs-GGUFhtdemucsF16 + Q8MIT
Higgs-Audio-v3-STT-GGUFhiggs_audio_sttF16 + Q8Apache-2.0
Higgs-Audio-v3-TTS-4B-GGUFhiggs_audio_ttsBF16 + Q8Boson Higgs TTS 3 Research and Non-Commercial License
Hviske-v5.3-GGUFhviske_asrQ8CC-BY-NC-4.0
IndexTTS2-GGUFindex_tts2original + F16 + Q8bilibili Model Use License Agreement
IndexTTS2.5-GGUFindex_tts2original + F16 + Q8bilibili Model Use License Agreement
Inflect-Micro-v2-GGUFinflect_v2originalApache-2.0
Irodori-TTS-500M-v3-GGUFirodori_ttsF16 + Q8MIT
Irodori-TTS-600M-v3-VoiceDesign-GGUFirodori_ttsF16 + Q8MIT
Irodori-TTS-v4-Small-GGUFirodori_ttsF16 + Q8MIT
Kroko-ASR-GGUFkroko_asrQ8CC-BY-SA community model license
MMS-Forced-Aligner-GGUFmms_forced_alignerF16CC-BY-NC-4.0
MOSS-TTS-Local-v1.5-GGUFmoss_tts_localBF16 + Q8Apache-2.0
MOSS-TTS-Nano-100M-GGUFmoss_tts_nanoBF16 + Q8Apache-2.0
MOSS-VoiceGenerator-GGUFmoss_voicegenBF16 + F16 codec decodeApache-2.0
MagpieTTS-Multilingual-357M-GGUFmagpie_ttsoriginalNVIDIA Open Model License
MeanVC2-GGUFmeanvc2F32 + Q4_KApache-2.0
Mel-Band-RoFormer-GGUFmel_band_roformerF16 + Q8MIT
MiniMax-H3-Q4-GGUFminimax_h3Q4_K + INT8 DiT optionMiniMax-H3 Community License
MioCodec-25Hz-44.1kHz-v2-GGUFmiocodecoriginal + F16 + Q8MIT
MioTTS-1.7B-GGUFmiottsoriginal + BF16 + Q8Apache-2.0
MuScriptor-Small-GGUFmuscriptorF32CC-BY-NC-4.0
Nemotron-3.5-ASR-Streaming-0.6B-GGUFnemotron_asrF16 + Q8OpenMDW-1.1
NeuTTS-2E-GGUFneuttsoriginalApache-2.0
OmniVoice-GGUFomnivoiceBF16 + F16 + Q8Apache-2.0
Parakeet-TDT-0.6B-v3-GGUFparakeet_tdtF16 + Q8CC-BY-4.0
PersonaPlex-GGUFpersonaplexQ4_K + Q8Apache-2.0
PocketTTS-GGUFpocket_ttsBF16 + Q8CC-BY-4.0
Qwen3-ASR-0.6B-GGUFqwen3_asrF16 + Q8Apache-2.0
Qwen3-ASR-1.7B-GGUFqwen3_asrF16 + Q8Apache-2.0
Qwen3-ForcedAligner-0.6B-GGUFqwen3_forced_alignerF16 + Q8Apache-2.0
Qwen3-TTS-12Hz-0.6B-Base-GGUFqwen3_ttsBF16 + Q8Apache-2.0
Qwen3-TTS-12Hz-1.7B-Base-GGUFqwen3_ttsoriginal + BF16 + Q8Apache-2.0
Qwen3-TTS-12Hz-1.7B-CustomVoice-GGUFqwen3_ttsBF16 + Q8Apache-2.0
Qwen3-TTS-12Hz-1.7B-VoiceDesign-GGUFqwen3_ttsBF16 + Q8Apache-2.0
RVC-GGUFrvcF16MIT
SeedVC-MLX-GGUFseed_vcoriginal + F16 + Q8GPL-3.0
Sortformer-Diar-4spk-v1-GGUFsortformer_diarF16 + Q8CC-BY-NC-4.0
Stable-Audio-3-Medium-GGUFstable_audioF16 + Q8Stability AI Community License
Stable-Audio-3-Small-Music-GGUFstable_audioF16 + Q8Stability AI Community License
Stable-Audio-3-Small-SFX-GGUFstable_audioF16 + Q8Stability AI Community License
Supertonic-3-GGUFsupertonicoriginal + F16 + Q8BigScience Open RAIL-M
Vevo2-GGUFvevo2original + F16 + Q8CC-BY-NC-ND-4.0
VibeVoice-1.5B-GGUFvibevoiceBF16 + Q8 + Q4MIT
VibeVoice-ASR-GGUFvibevoice_asrF16 + Q8MIT
VoxCPM2-GGUFvoxcpm2original + BF16 + Q8Apache-2.0
Voxtral-Mini-4B-Realtime-2602-GGUFvoxtral_realtimeBF16 + Q8 + Q4_KApache-2.0

Usage

Pass a GGUF file directly as --model:

audiocpp_cli --task tts --family supertonic --model Supertonic-3-GGUF/supertonic-3-orig.gguf --backend cuda --language en --text "Hello." --voice-id M1 --out out.wav

For ASR:

audiocpp_cli --task asr --family qwen3_asr --model Qwen3-ASR-0.6B-GGUF/qwen3-asr-0.6b-f16.gguf --backend cuda --audio speech.wav --text "" --text-out transcript.txt

License

Each GGUF file is a converted form of its original model. Use and redistribution are governed by the corresponding original model license listed above. Please review the original model card and license terms before using or redistributing any converted weights.

Contributors

audio-cpp

51 commits