ArnieTW/CosyVoiceNet

FunAudio/CosyVoice port to dotNet

C#

1

10 commits

updated Jun 16, 2026

See the code

README

CosyVoiceNet

CosyVoiceNet is a C#/.NET port and integration layer for the FunAudioLLM/CosyVoice text-to-speech pipeline. It is based on the original CosyVoice Python implementation, but has been adjusted for .NET usage, lazy model downloads, high-level application integration, CPU/CUDA backend selection, saved voice reuse, and runtime optimization profiles.

The goal is simple: external apps should be able to request CosyVoice capability through one C# facade, without requiring Python at runtime.

What It Does

  • Runs CosyVoice models from .NET through TorchSharp and the local C# pipeline.
  • Supports CPU and CUDA backend selection from a single CUDA-capable build.
  • Lets callers set a global backend and override backend per model/request.
  • Auto-downloads a model only when that specific model is requested.
  • Exposes model capabilities without initializing model weights.
  • Supports zero-shot, cross-lingual, instruct, instruct2, SFT, and saved voices depending on the selected model.
  • Can clone a prompt WAV into a named saved voice, then generate future TTS by voice name.
  • Returns generated audio as bytes. CosyVoiceNet does not decide where external apps store WAV files.
  • Uses CosyVoiceRuntimeOptions for logging, profiling, thread counts, cache strategy, affinity, and optimization behavior.

Supported Models

Model aliasLocal modelUpstream modelMain modes
cosyvoice3Fun-CosyVoice3-0.5BFunAudioLLM/Fun-CosyVoice3-0.5B-2512zero-shot, cross-lingual, instruct2, saved voices
cosyvoice2CosyVoice2-0.5BFunAudioLLM/CosyVoice2-0.5Bzero-shot, cross-lingual, instruct2, saved voices
cosyvoiceCosyVoice-300MFunAudioLLM/CosyVoice-300Mzero-shot, cross-lingual, saved voices
sftCosyVoice-300M-SFTFunAudioLLM/CosyVoice-300M-SFTbuilt-in SFT speakers
instructCosyVoice-300M-InstructFunAudioLLM/CosyVoice-300M-Instructoriginal 300M instruct route

The model registry lives in CosyVoiceNet.cli.CosyVoiceModels. The downloader is lazy: asking for cosyvoice3 does not download every other model.

Build

CosyVoiceNet targets net10.0.

dotnet restore CosyVoiceNet\CosyVoiceNet.csproj
dotnet build CosyVoiceNet\CosyVoiceNet.csproj

The default Windows build references the CUDA TorchSharp runtime package. A caller can still force CPU at runtime with CosyVoiceBackend.Cpu. CUDA-capable Windows builds also reference ONNX Runtime GPU so the prompt-audio ONNX models can try CUDA when the resolved backend is CUDA. If an ONNX CUDA provider cannot initialize on a machine, CosyVoiceNet falls back to CPU for that ONNX session.

For a CPU-only build:

dotnet build CosyVoiceNet\CosyVoiceNet.csproj -p:TorchUseCuda=false

To explicitly disable only ONNX CUDA binaries while keeping the Torch CUDA runtime:

dotnet build CosyVoiceNet\CosyVoiceNet.csproj -p:OnnxUseCuda=false

Timings

Generation benchmarks and repeat-run stability checks live in TIMINGS.md. The README keeps the project overview and usage examples focused, while the timing file can grow as more hardware, profiles, and models are tested.

High-Level API

External applications should normally use CosyVoiceReturner or the ICosyVoiceReturner interface instead of directly constructing the lower-level runtime classes.

using CosyVoiceNet;
using CosyVoiceNet.cli;

ICosyVoiceReturner tts = CosyVoiceReturner.Shared;

List Models

This is static capability information and does not initialize model weights.

foreach (var model in tts.GetModels())
{
    Console.WriteLine($"{model.LocalName} downloaded={model.IsDownloaded} features={model.Features}");
}

Download A Model With Progress

var progress = new Progress<CosyVoiceDownloadProgress>(item =>
{
    var percent = item.ModelPercent?.ToString("0.0") ?? "?";
    Console.WriteLine($"{item.Stage}: {item.Message} ({percent}%)");
});

var modelDirectory = tts.EnsureModelDownloaded("cosyvoice3", progress);
Console.WriteLine(modelDirectory);

Backend Selection

The global backend defaults to Auto, which attempts CUDA and falls back to CPU when CUDA is unavailable. You can make the choice explicit globally:

tts.SetGlobalBackend(CosyVoiceBackend.Cuda);

Or per model/load/generation request:

var loaded = tts.LoadModel(
    model: "cosyvoice3",
    backend: CosyVoiceBackend.Cpu,
    runtimeOptions: new CosyVoiceRuntimeOptions
    {
        OptimizationProfile = CosyVoiceOptimizationProfile.Balanced,
        CpuThreads = 8,
        CpuInteropThreads = 1
    });

Console.WriteLine($"{loaded.LocalName} active backend: {loaded.ActiveBackend}");

Generate From A Prompt WAV

Zero-shot generation needs both a prompt WAV and the transcript of that WAV.

var result = tts.Generate(new CosyVoiceTtsRequest
{
    Model = "cosyvoice3",
    Text = "A quick CosyVoiceNet test from the C# pipeline.",
    PromptText = "This is the exact transcript of the prompt audio.",
    PromptWav = @"D:\voices\prompt.wav",
    Backend = CosyVoiceBackend.Cuda,
    RuntimeOptions = new CosyVoiceRuntimeOptions
    {
        OptimizationProfile = CosyVoiceOptimizationProfile.Throughput
    }
});

File.WriteAllBytes(@"D:\voices\out.wav", result.WavBytes);
Console.WriteLine($"Generated {result.DurationSeconds:0.00}s in {result.InferenceTime.TotalSeconds:0.00}s");

CosyVoiceTtsResult.WavBytes is the playable WAV payload. RawFloat32Bytes contains the raw mono float32 samples used before WAV packaging.

Clone And Reuse A Voice

var savedVoice = tts.CloneAndSaveVoice(new CosyVoiceCloneRequest(
    Model: "cosyvoice3",
    VoiceName: "my_voice",
    PromptText: "This is the exact transcript of the prompt audio.",
    PromptWav: @"D:\voices\prompt.wav",
    Backend: CosyVoiceBackend.Cuda));

var result = tts.Generate(new CosyVoiceTtsRequest
{
    Model = "cosyvoice3",
    Text = "This uses the saved voice without reprocessing the prompt WAV.",
    Voice = savedVoice,
    Backend = CosyVoiceBackend.Cuda
});

Saved voices are model-specific. First-time generation from a provided WAV voice selector can clone that WAV automatically, then later calls reuse the saved voice name.

List Voices

var voices = tts.GetVoices(
    model: "cosyvoice3",
    providedWavs: new[] { @"D:\voices\guest.wav" },
    ensureDownloaded: false);

foreach (var voice in voices)
{
    Console.WriteLine($"{voice.Id} kind={voice.Kind} cloneOnFirstUse={voice.RequiresClone}");
}

For clone-capable models, provided WAVs and integrated prompt WAVs are exposed as voice options without forcing model initialization. Built-in SFT voices are reported for SFT models.

Instruct And Cross-Lingual

Leave Mode as CosyVoiceTtsMode.Auto for normal use. The facade chooses the route from the fields you provide:

  • InstructText chooses Instruct2 for CosyVoice2/CosyVoice3 or Instruct for the 300M instruct model.
  • CrossLingual = true chooses the cross-lingual route.
  • Voice chooses SFT for SFT-only models, otherwise a saved/provided voice.
  • PromptWav without a voice chooses zero-shot.
var result = tts.Generate(new CosyVoiceTtsRequest
{
    Model = "cosyvoice3",
    Text = "Say this with a calm, clear streaming voice.",
    Voice = "my_voice",
    InstructText = "Use a relaxed and natural delivery.",
    Backend = CosyVoiceBackend.Cuda
});

Runtime Options

Runtime behavior is controlled through objects passed to the API, not through environment variables. Useful knobs include:

  • OptimizationProfile: Compatibility, Balanced, Throughput, or LowMemory.
  • CpuThreads and CpuInteropThreads: Torch CPU worker settings.
  • CpuProcessorAffinityMask: optional process affinity mask for CPU pinning.
  • QwenKvCacheBackend and LegacyTransformerCacheBackend: cache strategy.
  • QwenAttentionBackend, QwenMlpBackend, and SamplingBackend: lower-level generation implementation choices.
  • Logger and Profiler: opt-in diagnostics sinks.
  • TraceTextInput, TracePromptTrim, TraceLlmInputShapes, and TraceGeneratedTokens: debugging traces that should stay off during normal generation.

Notes On The Port

CosyVoiceNet is intentionally close to the original CosyVoice behavior where it matters for model logic, token preparation, prompt handling, flow generation, and vocoder output. It is still not a direct packaging of the Python code. The C# runtime has been adapted and optimized for:

  • .NET application embedding.
  • No Python dependency at runtime.
  • Lazy model acquisition.
  • Reusable high-level request/response types.
  • CPU/CUDA backend management.
  • Saved voice workflows.
  • Per-request and per-model runtime options.
  • Profiling and logging through explicit API hooks.

Kudos

Huge thanks to the FunAudioLLM team for CosyVoice and the research/model releases that made this port possible:

The upstream CosyVoice project also acknowledges major work from:

Please keep upstream license and attribution requirements in mind when redistributing models or derived code. The bundled upstream CosyVoice source tree uses the Apache License 2.0.

Citations

If you use this project in published work, cite the original CosyVoice papers:

@article{du2024cosyvoice,
  title={Cosyvoice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic tokens},
  author={Du, Zhihao and Chen, Qian and Zhang, Shiliang and Hu, Kai and Lu, Heng and Yang, Yexin and Hu, Hangrui and Zheng, Siqi and Gu, Yue and Ma, Ziyang and others},
  journal={arXiv preprint arXiv:2407.05407},
  year={2024}
}

@article{du2024cosyvoice2,
  title={Cosyvoice 2: Scalable streaming speech synthesis with large language models},
  author={Du, Zhihao and Wang, Yuxuan and Chen, Qian and Shi, Xian and Lv, Xiang and Zhao, Tianyu and Gao, Zhifu and Yang, Yexin and Gao, Changfeng and Wang, Hui and others},
  journal={arXiv preprint arXiv:2412.10117},
  year={2024}
}

@article{du2025cosyvoice,
  title={CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training},
  author={Du, Zhihao and Gao, Changfeng and Wang, Yuxuan and Yu, Fan and Zhao, Tianyu and Wang, Hao and Lv, Xiang and Wang, Hui and Shi, Xian and An, Keyu and others},
  journal={arXiv preprint arXiv:2505.17589},
  year={2025}
}

@inproceedings{lyu2025build,
  title={Build LLM-Based Zero-Shot Streaming TTS System with Cosyvoice},
  author={Lyu, Xiang and Wang, Yuxuan and Zhao, Tianyu and Wang, Hao and Liu, Huadai and Du, Zhihao},
  booktitle={ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
  pages={1--2},
  year={2025},
  organization={IEEE}
}

Socials And Support

If CosyVoiceNet saves you time, please consider tipping or supporting(links available in linktree) development through the ArnieTW Linktree. It helps keep the port maintained, tested, and improving.

ArnieTW/CosyVoiceNet

FunAudio/CosyVoice port to dotNet

C#

1

10 commits

updated Jun 16, 2026

See the code

README

CosyVoiceNet

CosyVoiceNet is a C#/.NET port and integration layer for the FunAudioLLM/CosyVoice text-to-speech pipeline. It is based on the original CosyVoice Python implementation, but has been adjusted for .NET usage, lazy model downloads, high-level application integration, CPU/CUDA backend selection, saved voice reuse, and runtime optimization profiles.

The goal is simple: external apps should be able to request CosyVoice capability through one C# facade, without requiring Python at runtime.

What It Does

  • Runs CosyVoice models from .NET through TorchSharp and the local C# pipeline.
  • Supports CPU and CUDA backend selection from a single CUDA-capable build.
  • Lets callers set a global backend and override backend per model/request.
  • Auto-downloads a model only when that specific model is requested.
  • Exposes model capabilities without initializing model weights.
  • Supports zero-shot, cross-lingual, instruct, instruct2, SFT, and saved voices depending on the selected model.
  • Can clone a prompt WAV into a named saved voice, then generate future TTS by voice name.
  • Returns generated audio as bytes. CosyVoiceNet does not decide where external apps store WAV files.
  • Uses CosyVoiceRuntimeOptions for logging, profiling, thread counts, cache strategy, affinity, and optimization behavior.

Supported Models

Model aliasLocal modelUpstream modelMain modes
cosyvoice3Fun-CosyVoice3-0.5BFunAudioLLM/Fun-CosyVoice3-0.5B-2512zero-shot, cross-lingual, instruct2, saved voices
cosyvoice2CosyVoice2-0.5BFunAudioLLM/CosyVoice2-0.5Bzero-shot, cross-lingual, instruct2, saved voices
cosyvoiceCosyVoice-300MFunAudioLLM/CosyVoice-300Mzero-shot, cross-lingual, saved voices
sftCosyVoice-300M-SFTFunAudioLLM/CosyVoice-300M-SFTbuilt-in SFT speakers
instructCosyVoice-300M-InstructFunAudioLLM/CosyVoice-300M-Instructoriginal 300M instruct route

The model registry lives in CosyVoiceNet.cli.CosyVoiceModels. The downloader is lazy: asking for cosyvoice3 does not download every other model.

Build

CosyVoiceNet targets net10.0.

dotnet restore CosyVoiceNet\CosyVoiceNet.csproj
dotnet build CosyVoiceNet\CosyVoiceNet.csproj

The default Windows build references the CUDA TorchSharp runtime package. A caller can still force CPU at runtime with CosyVoiceBackend.Cpu. CUDA-capable Windows builds also reference ONNX Runtime GPU so the prompt-audio ONNX models can try CUDA when the resolved backend is CUDA. If an ONNX CUDA provider cannot initialize on a machine, CosyVoiceNet falls back to CPU for that ONNX session.

For a CPU-only build:

dotnet build CosyVoiceNet\CosyVoiceNet.csproj -p:TorchUseCuda=false

To explicitly disable only ONNX CUDA binaries while keeping the Torch CUDA runtime:

dotnet build CosyVoiceNet\CosyVoiceNet.csproj -p:OnnxUseCuda=false

Timings

Generation benchmarks and repeat-run stability checks live in TIMINGS.md. The README keeps the project overview and usage examples focused, while the timing file can grow as more hardware, profiles, and models are tested.

High-Level API

External applications should normally use CosyVoiceReturner or the ICosyVoiceReturner interface instead of directly constructing the lower-level runtime classes.

using CosyVoiceNet;
using CosyVoiceNet.cli;

ICosyVoiceReturner tts = CosyVoiceReturner.Shared;

List Models

This is static capability information and does not initialize model weights.

foreach (var model in tts.GetModels())
{
    Console.WriteLine($"{model.LocalName} downloaded={model.IsDownloaded} features={model.Features}");
}

Download A Model With Progress

var progress = new Progress<CosyVoiceDownloadProgress>(item =>
{
    var percent = item.ModelPercent?.ToString("0.0") ?? "?";
    Console.WriteLine($"{item.Stage}: {item.Message} ({percent}%)");
});

var modelDirectory = tts.EnsureModelDownloaded("cosyvoice3", progress);
Console.WriteLine(modelDirectory);

Backend Selection

The global backend defaults to Auto, which attempts CUDA and falls back to CPU when CUDA is unavailable. You can make the choice explicit globally:

tts.SetGlobalBackend(CosyVoiceBackend.Cuda);

Or per model/load/generation request:

var loaded = tts.LoadModel(
    model: "cosyvoice3",
    backend: CosyVoiceBackend.Cpu,
    runtimeOptions: new CosyVoiceRuntimeOptions
    {
        OptimizationProfile = CosyVoiceOptimizationProfile.Balanced,
        CpuThreads = 8,
        CpuInteropThreads = 1
    });

Console.WriteLine($"{loaded.LocalName} active backend: {loaded.ActiveBackend}");

Generate From A Prompt WAV

Zero-shot generation needs both a prompt WAV and the transcript of that WAV.

var result = tts.Generate(new CosyVoiceTtsRequest
{
    Model = "cosyvoice3",
    Text = "A quick CosyVoiceNet test from the C# pipeline.",
    PromptText = "This is the exact transcript of the prompt audio.",
    PromptWav = @"D:\voices\prompt.wav",
    Backend = CosyVoiceBackend.Cuda,
    RuntimeOptions = new CosyVoiceRuntimeOptions
    {
        OptimizationProfile = CosyVoiceOptimizationProfile.Throughput
    }
});

File.WriteAllBytes(@"D:\voices\out.wav", result.WavBytes);
Console.WriteLine($"Generated {result.DurationSeconds:0.00}s in {result.InferenceTime.TotalSeconds:0.00}s");

CosyVoiceTtsResult.WavBytes is the playable WAV payload. RawFloat32Bytes contains the raw mono float32 samples used before WAV packaging.

Clone And Reuse A Voice

var savedVoice = tts.CloneAndSaveVoice(new CosyVoiceCloneRequest(
    Model: "cosyvoice3",
    VoiceName: "my_voice",
    PromptText: "This is the exact transcript of the prompt audio.",
    PromptWav: @"D:\voices\prompt.wav",
    Backend: CosyVoiceBackend.Cuda));

var result = tts.Generate(new CosyVoiceTtsRequest
{
    Model = "cosyvoice3",
    Text = "This uses the saved voice without reprocessing the prompt WAV.",
    Voice = savedVoice,
    Backend = CosyVoiceBackend.Cuda
});

Saved voices are model-specific. First-time generation from a provided WAV voice selector can clone that WAV automatically, then later calls reuse the saved voice name.

List Voices

var voices = tts.GetVoices(
    model: "cosyvoice3",
    providedWavs: new[] { @"D:\voices\guest.wav" },
    ensureDownloaded: false);

foreach (var voice in voices)
{
    Console.WriteLine($"{voice.Id} kind={voice.Kind} cloneOnFirstUse={voice.RequiresClone}");
}

For clone-capable models, provided WAVs and integrated prompt WAVs are exposed as voice options without forcing model initialization. Built-in SFT voices are reported for SFT models.

Instruct And Cross-Lingual

Leave Mode as CosyVoiceTtsMode.Auto for normal use. The facade chooses the route from the fields you provide:

  • InstructText chooses Instruct2 for CosyVoice2/CosyVoice3 or Instruct for the 300M instruct model.
  • CrossLingual = true chooses the cross-lingual route.
  • Voice chooses SFT for SFT-only models, otherwise a saved/provided voice.
  • PromptWav without a voice chooses zero-shot.
var result = tts.Generate(new CosyVoiceTtsRequest
{
    Model = "cosyvoice3",
    Text = "Say this with a calm, clear streaming voice.",
    Voice = "my_voice",
    InstructText = "Use a relaxed and natural delivery.",
    Backend = CosyVoiceBackend.Cuda
});

Runtime Options

Runtime behavior is controlled through objects passed to the API, not through environment variables. Useful knobs include:

  • OptimizationProfile: Compatibility, Balanced, Throughput, or LowMemory.
  • CpuThreads and CpuInteropThreads: Torch CPU worker settings.
  • CpuProcessorAffinityMask: optional process affinity mask for CPU pinning.
  • QwenKvCacheBackend and LegacyTransformerCacheBackend: cache strategy.
  • QwenAttentionBackend, QwenMlpBackend, and SamplingBackend: lower-level generation implementation choices.
  • Logger and Profiler: opt-in diagnostics sinks.
  • TraceTextInput, TracePromptTrim, TraceLlmInputShapes, and TraceGeneratedTokens: debugging traces that should stay off during normal generation.

Notes On The Port

CosyVoiceNet is intentionally close to the original CosyVoice behavior where it matters for model logic, token preparation, prompt handling, flow generation, and vocoder output. It is still not a direct packaging of the Python code. The C# runtime has been adapted and optimized for:

  • .NET application embedding.
  • No Python dependency at runtime.
  • Lazy model acquisition.
  • Reusable high-level request/response types.
  • CPU/CUDA backend management.
  • Saved voice workflows.
  • Per-request and per-model runtime options.
  • Profiling and logging through explicit API hooks.

Kudos

Huge thanks to the FunAudioLLM team for CosyVoice and the research/model releases that made this port possible:

The upstream CosyVoice project also acknowledges major work from:

Please keep upstream license and attribution requirements in mind when redistributing models or derived code. The bundled upstream CosyVoice source tree uses the Apache License 2.0.

Citations

If you use this project in published work, cite the original CosyVoice papers:

@article{du2024cosyvoice,
  title={Cosyvoice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic tokens},
  author={Du, Zhihao and Chen, Qian and Zhang, Shiliang and Hu, Kai and Lu, Heng and Yang, Yexin and Hu, Hangrui and Zheng, Siqi and Gu, Yue and Ma, Ziyang and others},
  journal={arXiv preprint arXiv:2407.05407},
  year={2024}
}

@article{du2024cosyvoice2,
  title={Cosyvoice 2: Scalable streaming speech synthesis with large language models},
  author={Du, Zhihao and Wang, Yuxuan and Chen, Qian and Shi, Xian and Lv, Xiang and Zhao, Tianyu and Gao, Zhifu and Yang, Yexin and Gao, Changfeng and Wang, Hui and others},
  journal={arXiv preprint arXiv:2412.10117},
  year={2024}
}

@article{du2025cosyvoice,
  title={CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training},
  author={Du, Zhihao and Gao, Changfeng and Wang, Yuxuan and Yu, Fan and Zhao, Tianyu and Wang, Hao and Lv, Xiang and Wang, Hui and Shi, Xian and An, Keyu and others},
  journal={arXiv preprint arXiv:2505.17589},
  year={2025}
}

@inproceedings{lyu2025build,
  title={Build LLM-Based Zero-Shot Streaming TTS System with Cosyvoice},
  author={Lyu, Xiang and Wang, Yuxuan and Zhao, Tianyu and Wang, Hao and Liu, Huadai and Du, Zhihao},
  booktitle={ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
  pages={1--2},
  year={2025},
  organization={IEEE}
}

Socials And Support

If CosyVoiceNet saves you time, please consider tipping or supporting(links available in linktree) development through the ArnieTW Linktree. It helps keep the port maintained, tested, and improving.

Languages

C#

100.0%