Benchmarking open-source and text-to-speech models on intelligibility (WER), speed (RTFx, batched inference on a fixed GPU) and β for voice-cloning runs β speaker similarity (SIM), across two multilingual eval sets.
Benchmarking open-source and text-to-speech models on intelligibility (WER), speed (RTFx, batched inference on a fixed GPU) and β for voice-cloning runs β speaker similarity (SIM), across two multilingual eval sets.