Chunbo Hao1,2 · Junjie Zheng2 · Guobin Ma1 · Yuepeng Jiang1 · Huakang Chen1 · Wenjie Tian1 · Gongyu Chen2 · Zihao Chen2 · Lei Xie1
1 Audio, Speech and Language Processing Group (ASLP@NPU), School of Computer Science, Northwestern Polytechnical University, China
2 AI Lab, GiantNetwork, China
Overall architecture of YingMusic-Singer-Plus. Left: SFT training pipeline. Right: GRPO training pipeline.
YingMusic-Singer-Plus is a fully diffusion-based singing voice synthesis model that enables melody-controllable singing voice editing with flexible lyric manipulation, requiring no manual alignment or precise phoneme annotation.
Given only three inputs — an optional timbre reference, a melody-providing singing clip, and modified lyrics — YingMusic-Singer-Plus synthesizes high-fidelity singing voices at 44.1 kHz while faithfully preserving the original melody.
conda create -n YingMusic-Singer-Plus python=3.10
conda activate YingMusic-Singer-Plus
# uv is much faster than pip
pip install uv
uv pip install -r requirements.txt
conda --version.envs/ and create a folder named YingMusic-Singer-Plus.tar -xvf <package_name>.| CPU Architecture | GPU | OS | Download |
|---|---|---|---|
| ARM | NVIDIA | Linux | Coming soon |
| AMD64 | NVIDIA | Linux | Coming soon |
| AMD64 | NVIDIA | Windows | Coming soon |
Build the image:
docker build -t YingMusic-Singer-Plus .
Run inference:
docker run --gpus all -it YingMusic-Singer-Plus
Visit https://huggingface.co/spaces/ASLP-lab/YingMusic-Singer-Plus to try the model instantly in your browser.
python app_local.py
python infer_api.py \
--ref_audio path/to/ref.wav \
--melody_audio path/to/melody.wav \
--ref_text "该体谅的不执着|如果那天我" \
--target_text "好多天|看不完你" \
--output output.wav
Enable vocal separation and accompaniment mixing:
python infer_api.py \
--ref_audio ref.wav \
--melody_audio melody.wav \
--ref_text "..." \
--target_text "..." \
--separate_vocals \ # separate vocals from the input before processing
--mix_accompaniment \ # mix the synthesized vocal back with the accompaniment
--output mixed_output.wav
Note: All audio fed to the model must be pure vocal tracks (no accompaniment). If your inputs contain accompaniment, run vocal separation first using
src/third_party/MusicSourceSeparationTraining/inference_api.py.
The input JSONL file should contain one JSON object per line, formatted as follows:
{"id": "1", "melody_ref_path": "XXX", "gen_text": "好多天|看不完你", "timbre_ref_path": "XXX", "timbre_ref_text": "该体谅的不执着|如果那天我"}
python batch_infer.py \
--input_type jsonl \
--input_path /path/to/input.jsonl \
--output_dir /path/to/output \
--ckpt_path /path/to/ckpts \
--num_gpus 4
Multi-process inference on LyricEditBench (melody control) — the test set will be downloaded automatically:
python inference_mp.py \
--input_type lyric_edit_bench_melody_control \
--output_dir path/to/LyricEditBench_melody_control \
--ckpt_path ASLP-lab/YingMusic-Singer-Plus \
--num_gpus 8
Multi-process inference on LyricEditBench (singing edit):
python inference_mp.py \
--input_type lyric_edit_bench_sing_edit \
--output_dir path/to/LyricEditBench_sing_edit \
--ckpt_path ASLP-lab/YingMusic-Singer-Plus \
--num_gpus 8
YingMusic-Singer-Plus consists of four core components:
| Component | Description |
|---|---|
| VAE | Stable Audio 2 encoder/decoder; downsamples stereo 44.1 kHz audio by 2048× |
| Melody Extractor | Encoder of a pretrained MIDI extraction model (SOME); captures disentangled melody information |
| IPA Tokenizer | Converts Chinese & English lyrics into a unified phoneme sequence with sentence-level alignment |
| DiT-based CFM | Conditional flow matching backbone following F5-TTS (22 layers, 16 heads, hidden dim 1024) |
Total parameters: ~727.3M (453.6M CFM + 156.1M VAE + 117.6M Melody Extractor)
We introduce LyricEditBench, the first benchmark for melody-preserving lyric modification evaluation, built on GTSinger. The dataset is available on HuggingFace at https://huggingface.co/datasets/ASLP-lab/LyricEditBench.
Comparison with baseline models on LyricEditBench across task types (Table 1) and languages. Metrics — P: PER, S: SIM, F: F0-CORR, V: VS — are detailed in Section 3. Best results in bold.
This work builds upon the following open-source projects:
The code and model weights in this project are licensed under CC BY 4.0, except for the following:
The VAE model weights and inference code (in src/YingMusic-Singer-Plus/utils/stable-audio-tools) are derived from Stable Audio Open by Stability AI, and are licensed under the Stability AI Community License.
Chunbo Hao1,2 · Junjie Zheng2 · Guobin Ma1 · Yuepeng Jiang1 · Huakang Chen1 · Wenjie Tian1 · Gongyu Chen2 · Zihao Chen2 · Lei Xie1
1 Audio, Speech and Language Processing Group (ASLP@NPU), School of Computer Science, Northwestern Polytechnical University, China
2 AI Lab, GiantNetwork, China
Overall architecture of YingMusic-Singer-Plus. Left: SFT training pipeline. Right: GRPO training pipeline.
YingMusic-Singer-Plus is a fully diffusion-based singing voice synthesis model that enables melody-controllable singing voice editing with flexible lyric manipulation, requiring no manual alignment or precise phoneme annotation.
Given only three inputs — an optional timbre reference, a melody-providing singing clip, and modified lyrics — YingMusic-Singer-Plus synthesizes high-fidelity singing voices at 44.1 kHz while faithfully preserving the original melody.
conda create -n YingMusic-Singer-Plus python=3.10
conda activate YingMusic-Singer-Plus
# uv is much faster than pip
pip install uv
uv pip install -r requirements.txt
conda --version.envs/ and create a folder named YingMusic-Singer-Plus.tar -xvf <package_name>.| CPU Architecture | GPU | OS | Download |
|---|---|---|---|
| ARM | NVIDIA | Linux | Coming soon |
| AMD64 | NVIDIA | Linux | Coming soon |
| AMD64 | NVIDIA | Windows | Coming soon |
Build the image:
docker build -t YingMusic-Singer-Plus .
Run inference:
docker run --gpus all -it YingMusic-Singer-Plus
Visit https://huggingface.co/spaces/ASLP-lab/YingMusic-Singer-Plus to try the model instantly in your browser.
python app_local.py
python infer_api.py \
--ref_audio path/to/ref.wav \
--melody_audio path/to/melody.wav \
--ref_text "该体谅的不执着|如果那天我" \
--target_text "好多天|看不完你" \
--output output.wav
Enable vocal separation and accompaniment mixing:
python infer_api.py \
--ref_audio ref.wav \
--melody_audio melody.wav \
--ref_text "..." \
--target_text "..." \
--separate_vocals \ # separate vocals from the input before processing
--mix_accompaniment \ # mix the synthesized vocal back with the accompaniment
--output mixed_output.wav
Note: All audio fed to the model must be pure vocal tracks (no accompaniment). If your inputs contain accompaniment, run vocal separation first using
src/third_party/MusicSourceSeparationTraining/inference_api.py.
The input JSONL file should contain one JSON object per line, formatted as follows:
{"id": "1", "melody_ref_path": "XXX", "gen_text": "好多天|看不完你", "timbre_ref_path": "XXX", "timbre_ref_text": "该体谅的不执着|如果那天我"}
python batch_infer.py \
--input_type jsonl \
--input_path /path/to/input.jsonl \
--output_dir /path/to/output \
--ckpt_path /path/to/ckpts \
--num_gpus 4
Multi-process inference on LyricEditBench (melody control) — the test set will be downloaded automatically:
python inference_mp.py \
--input_type lyric_edit_bench_melody_control \
--output_dir path/to/LyricEditBench_melody_control \
--ckpt_path ASLP-lab/YingMusic-Singer-Plus \
--num_gpus 8
Multi-process inference on LyricEditBench (singing edit):
python inference_mp.py \
--input_type lyric_edit_bench_sing_edit \
--output_dir path/to/LyricEditBench_sing_edit \
--ckpt_path ASLP-lab/YingMusic-Singer-Plus \
--num_gpus 8
YingMusic-Singer-Plus consists of four core components:
| Component | Description |
|---|---|
| VAE | Stable Audio 2 encoder/decoder; downsamples stereo 44.1 kHz audio by 2048× |
| Melody Extractor | Encoder of a pretrained MIDI extraction model (SOME); captures disentangled melody information |
| IPA Tokenizer | Converts Chinese & English lyrics into a unified phoneme sequence with sentence-level alignment |
| DiT-based CFM | Conditional flow matching backbone following F5-TTS (22 layers, 16 heads, hidden dim 1024) |
Total parameters: ~727.3M (453.6M CFM + 156.1M VAE + 117.6M Melody Extractor)
We introduce LyricEditBench, the first benchmark for melody-preserving lyric modification evaluation, built on GTSinger. The dataset is available on HuggingFace at https://huggingface.co/datasets/ASLP-lab/LyricEditBench.
Comparison with baseline models on LyricEditBench across task types (Table 1) and languages. Metrics — P: PER, S: SIM, F: F0-CORR, V: VS — are detailed in Section 3. Best results in bold.
This work builds upon the following open-source projects:
The code and model weights in this project are licensed under CC BY 4.0, except for the following:
The VAE model weights and inference code (in src/YingMusic-Singer-Plus/utils/stable-audio-tools) are derived from Stable Audio Open by Stability AI, and are licensed under the Stability AI Community License.