CosyVoice_DPO_NOTES: Supercharge Your Cosyvoice model with Cutting-Edge DPO Fine-Tuning!
Python
129
39 commits
updated Aug 8, 2025
🎙️ CosyVoice_DPO_NOTES: Supercharge Your TTS with Cutting-Edge DPO Fine-Tuning! 🎙️
Welcome to CosyVoice_DPO_NOTES, the go-to resource for TTS practitioners looking to push the boundaries of speech synthesis with Direct Preference Optimization (DPO)! Built on the powerful CosyVoice framework by FunAudioLLM, this repository is your treasure trove of practical insights, code snippets, and detailed notes on fine-tuning CosyVoice models using DPO to achieve unparalleled speaker similarity, pronunciation accuracy, and naturalness in multilingual and zero-shot scenarios. 🌍🎵
"Speaker A<|endofprompt|>" in multi-speaker SFT, with workarounds for reducing errors.Clone this repo, dive into the code, and start experimenting with CosyVoice’s pretrained models (CosyVoice2-0.5B, CosyVoice-300M-SFT, etc.) available on Hugging Face or ModelScope. Join the conversation on GitHub Issues to share your tweaks or ask for help. Whether you’re crafting a multilingual chatbot or a next-gen audiobook narrator, CosyVoice_DPO_NOTES is your shortcut to TTS excellence! 🚀
🔗 Repo
📚 Inspired by: CosyVoice 2 & CosyVoice 3
🎧 Demos: Check out CosyVoice’s demos
Let’s make synthetic voices sound more human than ever! 🎙️
39 commits
Python
100.0%
CosyVoice_DPO_NOTES: Supercharge Your Cosyvoice model with Cutting-Edge DPO Fine-Tuning!
Python
129
39 commits
updated Aug 8, 2025
🎙️ CosyVoice_DPO_NOTES: Supercharge Your TTS with Cutting-Edge DPO Fine-Tuning! 🎙️
Welcome to CosyVoice_DPO_NOTES, the go-to resource for TTS practitioners looking to push the boundaries of speech synthesis with Direct Preference Optimization (DPO)! Built on the powerful CosyVoice framework by FunAudioLLM, this repository is your treasure trove of practical insights, code snippets, and detailed notes on fine-tuning CosyVoice models using DPO to achieve unparalleled speaker similarity, pronunciation accuracy, and naturalness in multilingual and zero-shot scenarios. 🌍🎵
"Speaker A<|endofprompt|>" in multi-speaker SFT, with workarounds for reducing errors.Clone this repo, dive into the code, and start experimenting with CosyVoice’s pretrained models (CosyVoice2-0.5B, CosyVoice-300M-SFT, etc.) available on Hugging Face or ModelScope. Join the conversation on GitHub Issues to share your tweaks or ask for help. Whether you’re crafting a multilingual chatbot or a next-gen audiobook narrator, CosyVoice_DPO_NOTES is your shortcut to TTS excellence! 🚀
🔗 Repo
📚 Inspired by: CosyVoice 2 & CosyVoice 3
🎧 Demos: Check out CosyVoice’s demos
Let’s make synthetic voices sound more human than ever! 🎙️
39 commits
Python
100.0%