CosyVoiceNet is a C#/.NET port and integration layer for the FunAudioLLM/CosyVoice text-to-speech pipeline. It is based on the original CosyVoice Python implementation, but has been adjusted for .NET usage, lazy model downloads, high-level application integration, CPU/CUDA backend selection, saved voice reuse, and runtime optimization profiles.
The goal is simple: external apps should be able to request CosyVoice capability through one C# facade, without requiring Python at runtime.
CosyVoiceRuntimeOptions for logging, profiling, thread counts, cache
strategy, affinity, and optimization behavior.| Model alias | Local model | Upstream model | Main modes |
|---|---|---|---|
cosyvoice3 | Fun-CosyVoice3-0.5B | FunAudioLLM/Fun-CosyVoice3-0.5B-2512 | zero-shot, cross-lingual, instruct2, saved voices |
cosyvoice2 | CosyVoice2-0.5B | FunAudioLLM/CosyVoice2-0.5B | zero-shot, cross-lingual, instruct2, saved voices |
cosyvoice | CosyVoice-300M | FunAudioLLM/CosyVoice-300M | zero-shot, cross-lingual, saved voices |
sft | CosyVoice-300M-SFT | FunAudioLLM/CosyVoice-300M-SFT | built-in SFT speakers |
instruct | CosyVoice-300M-Instruct | FunAudioLLM/CosyVoice-300M-Instruct | original 300M instruct route |
The model registry lives in CosyVoiceNet.cli.CosyVoiceModels. The downloader is
lazy: asking for cosyvoice3 does not download every other model.
CosyVoiceNet targets net10.0.
dotnet restore CosyVoiceNet\CosyVoiceNet.csproj
dotnet build CosyVoiceNet\CosyVoiceNet.csproj
The default Windows build references the CUDA TorchSharp runtime package. A
caller can still force CPU at runtime with CosyVoiceBackend.Cpu. CUDA-capable
Windows builds also reference ONNX Runtime GPU so the prompt-audio ONNX models
can try CUDA when the resolved backend is CUDA. If an ONNX CUDA provider cannot
initialize on a machine, CosyVoiceNet falls back to CPU for that ONNX session.
For a CPU-only build:
dotnet build CosyVoiceNet\CosyVoiceNet.csproj -p:TorchUseCuda=false
To explicitly disable only ONNX CUDA binaries while keeping the Torch CUDA runtime:
dotnet build CosyVoiceNet\CosyVoiceNet.csproj -p:OnnxUseCuda=false
Generation benchmarks and repeat-run stability checks live in TIMINGS.md. The README keeps the project overview and usage examples focused, while the timing file can grow as more hardware, profiles, and models are tested.
External applications should normally use CosyVoiceReturner or the
ICosyVoiceReturner interface instead of directly constructing the lower-level
runtime classes.
using CosyVoiceNet;
using CosyVoiceNet.cli;
ICosyVoiceReturner tts = CosyVoiceReturner.Shared;
This is static capability information and does not initialize model weights.
foreach (var model in tts.GetModels())
{
Console.WriteLine($"{model.LocalName} downloaded={model.IsDownloaded} features={model.Features}");
}
var progress = new Progress<CosyVoiceDownloadProgress>(item =>
{
var percent = item.ModelPercent?.ToString("0.0") ?? "?";
Console.WriteLine($"{item.Stage}: {item.Message} ({percent}%)");
});
var modelDirectory = tts.EnsureModelDownloaded("cosyvoice3", progress);
Console.WriteLine(modelDirectory);
The global backend defaults to Auto, which attempts CUDA and falls back to CPU
when CUDA is unavailable. You can make the choice explicit globally:
tts.SetGlobalBackend(CosyVoiceBackend.Cuda);
Or per model/load/generation request:
var loaded = tts.LoadModel(
model: "cosyvoice3",
backend: CosyVoiceBackend.Cpu,
runtimeOptions: new CosyVoiceRuntimeOptions
{
OptimizationProfile = CosyVoiceOptimizationProfile.Balanced,
CpuThreads = 8,
CpuInteropThreads = 1
});
Console.WriteLine($"{loaded.LocalName} active backend: {loaded.ActiveBackend}");
Zero-shot generation needs both a prompt WAV and the transcript of that WAV.
var result = tts.Generate(new CosyVoiceTtsRequest
{
Model = "cosyvoice3",
Text = "A quick CosyVoiceNet test from the C# pipeline.",
PromptText = "This is the exact transcript of the prompt audio.",
PromptWav = @"D:\voices\prompt.wav",
Backend = CosyVoiceBackend.Cuda,
RuntimeOptions = new CosyVoiceRuntimeOptions
{
OptimizationProfile = CosyVoiceOptimizationProfile.Throughput
}
});
File.WriteAllBytes(@"D:\voices\out.wav", result.WavBytes);
Console.WriteLine($"Generated {result.DurationSeconds:0.00}s in {result.InferenceTime.TotalSeconds:0.00}s");
CosyVoiceTtsResult.WavBytes is the playable WAV payload. RawFloat32Bytes
contains the raw mono float32 samples used before WAV packaging.
var savedVoice = tts.CloneAndSaveVoice(new CosyVoiceCloneRequest(
Model: "cosyvoice3",
VoiceName: "my_voice",
PromptText: "This is the exact transcript of the prompt audio.",
PromptWav: @"D:\voices\prompt.wav",
Backend: CosyVoiceBackend.Cuda));
var result = tts.Generate(new CosyVoiceTtsRequest
{
Model = "cosyvoice3",
Text = "This uses the saved voice without reprocessing the prompt WAV.",
Voice = savedVoice,
Backend = CosyVoiceBackend.Cuda
});
Saved voices are model-specific. First-time generation from a provided WAV voice selector can clone that WAV automatically, then later calls reuse the saved voice name.
var voices = tts.GetVoices(
model: "cosyvoice3",
providedWavs: new[] { @"D:\voices\guest.wav" },
ensureDownloaded: false);
foreach (var voice in voices)
{
Console.WriteLine($"{voice.Id} kind={voice.Kind} cloneOnFirstUse={voice.RequiresClone}");
}
For clone-capable models, provided WAVs and integrated prompt WAVs are exposed as voice options without forcing model initialization. Built-in SFT voices are reported for SFT models.
Leave Mode as CosyVoiceTtsMode.Auto for normal use. The facade chooses the
route from the fields you provide:
InstructText chooses Instruct2 for CosyVoice2/CosyVoice3 or Instruct
for the 300M instruct model.CrossLingual = true chooses the cross-lingual route.Voice chooses SFT for SFT-only models, otherwise a saved/provided voice.PromptWav without a voice chooses zero-shot.var result = tts.Generate(new CosyVoiceTtsRequest
{
Model = "cosyvoice3",
Text = "Say this with a calm, clear streaming voice.",
Voice = "my_voice",
InstructText = "Use a relaxed and natural delivery.",
Backend = CosyVoiceBackend.Cuda
});
Runtime behavior is controlled through objects passed to the API, not through environment variables. Useful knobs include:
OptimizationProfile: Compatibility, Balanced, Throughput, or
LowMemory.CpuThreads and CpuInteropThreads: Torch CPU worker settings.CpuProcessorAffinityMask: optional process affinity mask for CPU pinning.QwenKvCacheBackend and LegacyTransformerCacheBackend: cache strategy.QwenAttentionBackend, QwenMlpBackend, and SamplingBackend: lower-level
generation implementation choices.Logger and Profiler: opt-in diagnostics sinks.TraceTextInput, TracePromptTrim, TraceLlmInputShapes, and
TraceGeneratedTokens: debugging traces that should stay off during normal
generation.CosyVoiceNet is intentionally close to the original CosyVoice behavior where it matters for model logic, token preparation, prompt handling, flow generation, and vocoder output. It is still not a direct packaging of the Python code. The C# runtime has been adapted and optimized for:
Huge thanks to the FunAudioLLM team for CosyVoice and the research/model releases that made this port possible:
The upstream CosyVoice project also acknowledges major work from:
Please keep upstream license and attribution requirements in mind when redistributing models or derived code. The bundled upstream CosyVoice source tree uses the Apache License 2.0.
If you use this project in published work, cite the original CosyVoice papers:
@article{du2024cosyvoice,
title={Cosyvoice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic tokens},
author={Du, Zhihao and Chen, Qian and Zhang, Shiliang and Hu, Kai and Lu, Heng and Yang, Yexin and Hu, Hangrui and Zheng, Siqi and Gu, Yue and Ma, Ziyang and others},
journal={arXiv preprint arXiv:2407.05407},
year={2024}
}
@article{du2024cosyvoice2,
title={Cosyvoice 2: Scalable streaming speech synthesis with large language models},
author={Du, Zhihao and Wang, Yuxuan and Chen, Qian and Shi, Xian and Lv, Xiang and Zhao, Tianyu and Gao, Zhifu and Yang, Yexin and Gao, Changfeng and Wang, Hui and others},
journal={arXiv preprint arXiv:2412.10117},
year={2024}
}
@article{du2025cosyvoice,
title={CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training},
author={Du, Zhihao and Gao, Changfeng and Wang, Yuxuan and Yu, Fan and Zhao, Tianyu and Wang, Hao and Lv, Xiang and Wang, Hui and Shi, Xian and An, Keyu and others},
journal={arXiv preprint arXiv:2505.17589},
year={2025}
}
@inproceedings{lyu2025build,
title={Build LLM-Based Zero-Shot Streaming TTS System with Cosyvoice},
author={Lyu, Xiang and Wang, Yuxuan and Zhao, Tianyu and Wang, Hao and Liu, Huadai and Du, Zhihao},
booktitle={ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
pages={1--2},
year={2025},
organization={IEEE}
}
If CosyVoiceNet saves you time, please consider tipping or supporting(links available in linktree) development through the ArnieTW Linktree. It helps keep the port maintained, tested, and improving.
C#
100.0%
CosyVoiceNet is a C#/.NET port and integration layer for the FunAudioLLM/CosyVoice text-to-speech pipeline. It is based on the original CosyVoice Python implementation, but has been adjusted for .NET usage, lazy model downloads, high-level application integration, CPU/CUDA backend selection, saved voice reuse, and runtime optimization profiles.
The goal is simple: external apps should be able to request CosyVoice capability through one C# facade, without requiring Python at runtime.
CosyVoiceRuntimeOptions for logging, profiling, thread counts, cache
strategy, affinity, and optimization behavior.| Model alias | Local model | Upstream model | Main modes |
|---|---|---|---|
cosyvoice3 | Fun-CosyVoice3-0.5B | FunAudioLLM/Fun-CosyVoice3-0.5B-2512 | zero-shot, cross-lingual, instruct2, saved voices |
cosyvoice2 | CosyVoice2-0.5B | FunAudioLLM/CosyVoice2-0.5B | zero-shot, cross-lingual, instruct2, saved voices |
cosyvoice | CosyVoice-300M | FunAudioLLM/CosyVoice-300M | zero-shot, cross-lingual, saved voices |
sft | CosyVoice-300M-SFT | FunAudioLLM/CosyVoice-300M-SFT | built-in SFT speakers |
instruct | CosyVoice-300M-Instruct | FunAudioLLM/CosyVoice-300M-Instruct | original 300M instruct route |
The model registry lives in CosyVoiceNet.cli.CosyVoiceModels. The downloader is
lazy: asking for cosyvoice3 does not download every other model.
CosyVoiceNet targets net10.0.
dotnet restore CosyVoiceNet\CosyVoiceNet.csproj
dotnet build CosyVoiceNet\CosyVoiceNet.csproj
The default Windows build references the CUDA TorchSharp runtime package. A
caller can still force CPU at runtime with CosyVoiceBackend.Cpu. CUDA-capable
Windows builds also reference ONNX Runtime GPU so the prompt-audio ONNX models
can try CUDA when the resolved backend is CUDA. If an ONNX CUDA provider cannot
initialize on a machine, CosyVoiceNet falls back to CPU for that ONNX session.
For a CPU-only build:
dotnet build CosyVoiceNet\CosyVoiceNet.csproj -p:TorchUseCuda=false
To explicitly disable only ONNX CUDA binaries while keeping the Torch CUDA runtime:
dotnet build CosyVoiceNet\CosyVoiceNet.csproj -p:OnnxUseCuda=false
Generation benchmarks and repeat-run stability checks live in TIMINGS.md. The README keeps the project overview and usage examples focused, while the timing file can grow as more hardware, profiles, and models are tested.
External applications should normally use CosyVoiceReturner or the
ICosyVoiceReturner interface instead of directly constructing the lower-level
runtime classes.
using CosyVoiceNet;
using CosyVoiceNet.cli;
ICosyVoiceReturner tts = CosyVoiceReturner.Shared;
This is static capability information and does not initialize model weights.
foreach (var model in tts.GetModels())
{
Console.WriteLine($"{model.LocalName} downloaded={model.IsDownloaded} features={model.Features}");
}
var progress = new Progress<CosyVoiceDownloadProgress>(item =>
{
var percent = item.ModelPercent?.ToString("0.0") ?? "?";
Console.WriteLine($"{item.Stage}: {item.Message} ({percent}%)");
});
var modelDirectory = tts.EnsureModelDownloaded("cosyvoice3", progress);
Console.WriteLine(modelDirectory);
The global backend defaults to Auto, which attempts CUDA and falls back to CPU
when CUDA is unavailable. You can make the choice explicit globally:
tts.SetGlobalBackend(CosyVoiceBackend.Cuda);
Or per model/load/generation request:
var loaded = tts.LoadModel(
model: "cosyvoice3",
backend: CosyVoiceBackend.Cpu,
runtimeOptions: new CosyVoiceRuntimeOptions
{
OptimizationProfile = CosyVoiceOptimizationProfile.Balanced,
CpuThreads = 8,
CpuInteropThreads = 1
});
Console.WriteLine($"{loaded.LocalName} active backend: {loaded.ActiveBackend}");
Zero-shot generation needs both a prompt WAV and the transcript of that WAV.
var result = tts.Generate(new CosyVoiceTtsRequest
{
Model = "cosyvoice3",
Text = "A quick CosyVoiceNet test from the C# pipeline.",
PromptText = "This is the exact transcript of the prompt audio.",
PromptWav = @"D:\voices\prompt.wav",
Backend = CosyVoiceBackend.Cuda,
RuntimeOptions = new CosyVoiceRuntimeOptions
{
OptimizationProfile = CosyVoiceOptimizationProfile.Throughput
}
});
File.WriteAllBytes(@"D:\voices\out.wav", result.WavBytes);
Console.WriteLine($"Generated {result.DurationSeconds:0.00}s in {result.InferenceTime.TotalSeconds:0.00}s");
CosyVoiceTtsResult.WavBytes is the playable WAV payload. RawFloat32Bytes
contains the raw mono float32 samples used before WAV packaging.
var savedVoice = tts.CloneAndSaveVoice(new CosyVoiceCloneRequest(
Model: "cosyvoice3",
VoiceName: "my_voice",
PromptText: "This is the exact transcript of the prompt audio.",
PromptWav: @"D:\voices\prompt.wav",
Backend: CosyVoiceBackend.Cuda));
var result = tts.Generate(new CosyVoiceTtsRequest
{
Model = "cosyvoice3",
Text = "This uses the saved voice without reprocessing the prompt WAV.",
Voice = savedVoice,
Backend = CosyVoiceBackend.Cuda
});
Saved voices are model-specific. First-time generation from a provided WAV voice selector can clone that WAV automatically, then later calls reuse the saved voice name.
var voices = tts.GetVoices(
model: "cosyvoice3",
providedWavs: new[] { @"D:\voices\guest.wav" },
ensureDownloaded: false);
foreach (var voice in voices)
{
Console.WriteLine($"{voice.Id} kind={voice.Kind} cloneOnFirstUse={voice.RequiresClone}");
}
For clone-capable models, provided WAVs and integrated prompt WAVs are exposed as voice options without forcing model initialization. Built-in SFT voices are reported for SFT models.
Leave Mode as CosyVoiceTtsMode.Auto for normal use. The facade chooses the
route from the fields you provide:
InstructText chooses Instruct2 for CosyVoice2/CosyVoice3 or Instruct
for the 300M instruct model.CrossLingual = true chooses the cross-lingual route.Voice chooses SFT for SFT-only models, otherwise a saved/provided voice.PromptWav without a voice chooses zero-shot.var result = tts.Generate(new CosyVoiceTtsRequest
{
Model = "cosyvoice3",
Text = "Say this with a calm, clear streaming voice.",
Voice = "my_voice",
InstructText = "Use a relaxed and natural delivery.",
Backend = CosyVoiceBackend.Cuda
});
Runtime behavior is controlled through objects passed to the API, not through environment variables. Useful knobs include:
OptimizationProfile: Compatibility, Balanced, Throughput, or
LowMemory.CpuThreads and CpuInteropThreads: Torch CPU worker settings.CpuProcessorAffinityMask: optional process affinity mask for CPU pinning.QwenKvCacheBackend and LegacyTransformerCacheBackend: cache strategy.QwenAttentionBackend, QwenMlpBackend, and SamplingBackend: lower-level
generation implementation choices.Logger and Profiler: opt-in diagnostics sinks.TraceTextInput, TracePromptTrim, TraceLlmInputShapes, and
TraceGeneratedTokens: debugging traces that should stay off during normal
generation.CosyVoiceNet is intentionally close to the original CosyVoice behavior where it matters for model logic, token preparation, prompt handling, flow generation, and vocoder output. It is still not a direct packaging of the Python code. The C# runtime has been adapted and optimized for:
Huge thanks to the FunAudioLLM team for CosyVoice and the research/model releases that made this port possible:
The upstream CosyVoice project also acknowledges major work from:
Please keep upstream license and attribution requirements in mind when redistributing models or derived code. The bundled upstream CosyVoice source tree uses the Apache License 2.0.
If you use this project in published work, cite the original CosyVoice papers:
@article{du2024cosyvoice,
title={Cosyvoice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic tokens},
author={Du, Zhihao and Chen, Qian and Zhang, Shiliang and Hu, Kai and Lu, Heng and Yang, Yexin and Hu, Hangrui and Zheng, Siqi and Gu, Yue and Ma, Ziyang and others},
journal={arXiv preprint arXiv:2407.05407},
year={2024}
}
@article{du2024cosyvoice2,
title={Cosyvoice 2: Scalable streaming speech synthesis with large language models},
author={Du, Zhihao and Wang, Yuxuan and Chen, Qian and Shi, Xian and Lv, Xiang and Zhao, Tianyu and Gao, Zhifu and Yang, Yexin and Gao, Changfeng and Wang, Hui and others},
journal={arXiv preprint arXiv:2412.10117},
year={2024}
}
@article{du2025cosyvoice,
title={CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training},
author={Du, Zhihao and Gao, Changfeng and Wang, Yuxuan and Yu, Fan and Zhao, Tianyu and Wang, Hao and Lv, Xiang and Wang, Hui and Shi, Xian and An, Keyu and others},
journal={arXiv preprint arXiv:2505.17589},
year={2025}
}
@inproceedings{lyu2025build,
title={Build LLM-Based Zero-Shot Streaming TTS System with Cosyvoice},
author={Lyu, Xiang and Wang, Yuxuan and Zhao, Tianyu and Wang, Hao and Liu, Huadai and Du, Zhihao},
booktitle={ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
pages={1--2},
year={2025},
organization={IEEE}
}
If CosyVoiceNet saves you time, please consider tipping or supporting(links available in linktree) development through the ArnieTW Linktree. It helps keep the port maintained, tested, and improving.
C#
100.0%