GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
17
22 commits
2 linked in READMEs
updated Jul 24, 2025
Dataset of GTSinger (NeurIPS 2024 Spotlight): A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks.
We introduce GTSinger, a large Global, multi-Technique, free-to-use, high-quality singing corpus with realistic music scores, designed for all singing tasks, along with its benchmarks.
We provide the full corpus for free in this repository.
And metadata.json and phone_set.json are also offered for each language in processed. Note: you should change the wav_fn for each segment to your own absolute path! And you can use metadata of multiple languages by concat their data! We will provide the metadata for other languages soon!
Besides, we also provide our dataset on Google Drive.
Moreover, you can visit our Demo Page for the audio samples of our dataset as well as the results of our benchmarks.
Through this repo you can access our full dataset (audio along with TextGrid, json, musicxml) and processed data (metadata.json, phone_set.json, spker_set.json) on Hugging Face for free! Hope our data is helpful for your research.
Besides, we also provide our dataset on .
Please note that, if you are using GTSinger, it means that you have accepted the terms of license.
Our dataset is organized hierarchically.
It presents nine top-level folders, each corresponding to a distinct language.
Within each language folder, there are five sub-folders, each representing a specific singing technique.
These technique folders contain numerous song entries, with each song further divided into several controlled comparison groups: a control group (natural singing without the specific technique), and a technique group (densely employing the specific technique).
Our singing voices and speech are recorded at a 48kHz sampling rate with 24-bit resolution in WAV format.
Alignments and annotations are provided in TextGrid files, including word boundaries, phoneme boundaries, phoneme-level annotations for six techniques, and global style labels (singing method, emotion, pace, and range).
We also provide realistic music scores in musicxml format.
Notably, we provide an additional JSON file for each singing voice, facilitating data parsing and processing for singing models.
Here is the data structure of our dataset:
.
βββ Chinese
βΒ Β βββ ZH-Alto-1
βΒ Β βββ ZH-Tenor-1
βββ English
βΒ Β βββ EN-Alto-1
βΒ Β βΒ Β βββ Breathy
βΒ Β βΒ Β βββ Glissando
βΒ Β βΒ Β β βββ my love
βΒ Β βΒ Β β βββ Control_Group
βΒ Β βΒ Β β βββ Glissando_Group
βΒ Β βΒ Β β βββ Paired_Speech_Group
βΒ Β βΒ Β βββ Mixed_Voice_and_Falsetto
βΒ Β βΒ Β βββ Pharyngeal
βΒ Β βΒ Β βββ Vibrato
βΒ Β βββ EN-Alto-2
βΒ Β βΒ Β βββ Breathy
βΒ Β βΒ Β βββ Glissando
βΒ Β βΒ Β βββ Mixed_Voice_and_Falsetto
βΒ Β βΒ Β βββ Pharyngeal
βΒ Β βΒ Β βββ Vibrato
βΒ Β βββ EN-Tenor-1
βΒ Β Β Β βββ Breathy
βΒ Β Β Β βββ Glissando
βΒ Β Β Β βββ Mixed_Voice_and_Falsetto
βΒ Β Β Β βββ Pharyngeal
βΒ Β Β Β βββ Vibrato
βββ French
βΒ Β βββ FR-Soprano-1
βΒ Β βββ FR-Tenor-1
βββ German
βΒ Β βββ DE-Soprano-1
βΒ Β βββ DE-Tenor-1
βββ Italian
βΒ Β βββ IT-Bass-1
βΒ Β βββ IT-Bass-2
βΒ Β βββ IT-Soprano-1
βββ Japanese
βΒ Β βββ JA-Soprano-1
βΒ Β βββ JA-Tenor-1
βββ Korean
βΒ Β βββ KO-Soprano-1
βΒ Β βββ KO-Soprano-2
βΒ Β βββ KO-Tenor-1
βββ Russian
βΒ Β βββ RU-Alto-1
βββ Spanish
βββ ES-Bass-1
βββ ES-Soprano-1
If you find this code useful in your research, please cite our work:
@article{zhang2024gtsinger,
title={Gtsinger: A global multi-technique singing corpus with realistic music scores for all singing tasks},
author={Zhang, Yu and Pan, Changhao and Guo, Wenxiang and Li, Ruiqi and Zhu, Zhiyuan and Wang, Jialei and Xu, Wenhao and Lu, Jingyu and Hong, Zhiqing and Wang, Chuxin and others},
journal={arXiv preprint arXiv:2409.13832},
year={2024}
}
Any organization or individual is prohibited from using any technology mentioned in this paper to generate someone's singing without his/her consent, including but not limited to government leaders, political figures, and celebrities. If you do not comply with this item, you could be in violation of copyright laws.
22 commits
GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
17
22 commits
2 linked in READMEs
updated Jul 24, 2025
Dataset of GTSinger (NeurIPS 2024 Spotlight): A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks.
We introduce GTSinger, a large Global, multi-Technique, free-to-use, high-quality singing corpus with realistic music scores, designed for all singing tasks, along with its benchmarks.
We provide the full corpus for free in this repository.
And metadata.json and phone_set.json are also offered for each language in processed. Note: you should change the wav_fn for each segment to your own absolute path! And you can use metadata of multiple languages by concat their data! We will provide the metadata for other languages soon!
Besides, we also provide our dataset on Google Drive.
Moreover, you can visit our Demo Page for the audio samples of our dataset as well as the results of our benchmarks.
Through this repo you can access our full dataset (audio along with TextGrid, json, musicxml) and processed data (metadata.json, phone_set.json, spker_set.json) on Hugging Face for free! Hope our data is helpful for your research.
Besides, we also provide our dataset on .
Please note that, if you are using GTSinger, it means that you have accepted the terms of license.
Our dataset is organized hierarchically.
It presents nine top-level folders, each corresponding to a distinct language.
Within each language folder, there are five sub-folders, each representing a specific singing technique.
These technique folders contain numerous song entries, with each song further divided into several controlled comparison groups: a control group (natural singing without the specific technique), and a technique group (densely employing the specific technique).
Our singing voices and speech are recorded at a 48kHz sampling rate with 24-bit resolution in WAV format.
Alignments and annotations are provided in TextGrid files, including word boundaries, phoneme boundaries, phoneme-level annotations for six techniques, and global style labels (singing method, emotion, pace, and range).
We also provide realistic music scores in musicxml format.
Notably, we provide an additional JSON file for each singing voice, facilitating data parsing and processing for singing models.
Here is the data structure of our dataset:
.
βββ Chinese
βΒ Β βββ ZH-Alto-1
βΒ Β βββ ZH-Tenor-1
βββ English
βΒ Β βββ EN-Alto-1
βΒ Β βΒ Β βββ Breathy
βΒ Β βΒ Β βββ Glissando
βΒ Β βΒ Β β βββ my love
βΒ Β βΒ Β β βββ Control_Group
βΒ Β βΒ Β β βββ Glissando_Group
βΒ Β βΒ Β β βββ Paired_Speech_Group
βΒ Β βΒ Β βββ Mixed_Voice_and_Falsetto
βΒ Β βΒ Β βββ Pharyngeal
βΒ Β βΒ Β βββ Vibrato
βΒ Β βββ EN-Alto-2
βΒ Β βΒ Β βββ Breathy
βΒ Β βΒ Β βββ Glissando
βΒ Β βΒ Β βββ Mixed_Voice_and_Falsetto
βΒ Β βΒ Β βββ Pharyngeal
βΒ Β βΒ Β βββ Vibrato
βΒ Β βββ EN-Tenor-1
βΒ Β Β Β βββ Breathy
βΒ Β Β Β βββ Glissando
βΒ Β Β Β βββ Mixed_Voice_and_Falsetto
βΒ Β Β Β βββ Pharyngeal
βΒ Β Β Β βββ Vibrato
βββ French
βΒ Β βββ FR-Soprano-1
βΒ Β βββ FR-Tenor-1
βββ German
βΒ Β βββ DE-Soprano-1
βΒ Β βββ DE-Tenor-1
βββ Italian
βΒ Β βββ IT-Bass-1
βΒ Β βββ IT-Bass-2
βΒ Β βββ IT-Soprano-1
βββ Japanese
βΒ Β βββ JA-Soprano-1
βΒ Β βββ JA-Tenor-1
βββ Korean
βΒ Β βββ KO-Soprano-1
βΒ Β βββ KO-Soprano-2
βΒ Β βββ KO-Tenor-1
βββ Russian
βΒ Β βββ RU-Alto-1
βββ Spanish
βββ ES-Bass-1
βββ ES-Soprano-1
If you find this code useful in your research, please cite our work:
@article{zhang2024gtsinger,
title={Gtsinger: A global multi-technique singing corpus with realistic music scores for all singing tasks},
author={Zhang, Yu and Pan, Changhao and Guo, Wenxiang and Li, Ruiqi and Zhu, Zhiyuan and Wang, Jialei and Xu, Wenhao and Lu, Jingyu and Hong, Zhiqing and Wang, Chuxin and others},
journal={arXiv preprint arXiv:2409.13832},
year={2024}
}
Any organization or individual is prohibited from using any technology mentioned in this paper to generate someone's singing without his/her consent, including but not limited to government leaders, political figures, and celebrities. If you do not comply with this item, you could be in violation of copyright laws.
22 commits