The MusicBench dataset is a music audio-text pair dataset that was designed for text-to-music generation purpose and released along with Mustango text-to-music model. MusicBench is based on the MusicCaps dataset, which it expands from 5,521 samples to 52,768 training and 400 test samples!
MusicBench expands MusicCaps by:
Train set size = 52,768 samples Test set size = 400
MusicBench consists of 3 .json files and attached audio files in .tar.gz form.
The train set contains audio augmented samples and enhanced captions. Additionally, it offers ChatGPT rephrased captions for all the audio samples. Both TestA and TestB sets contain the same audio content, but TestB has all 4 possible control sentences (related to 4 music features) in captions of all samples, while TestA has no control sentences in the captions.
For more details, see Figure 1 in our paper.
Each row of a .json file has:
Hereby, we also present you the FMACaps evaluation dataset which consists of 1000 samples extracted from the Free Music Archive (FMA) and pseudocaptioned through extracting tags from audio and then utilizing ChatGPT in-context learning. More information is available in our paper!
Most of the samples are 10 second long, exceptions are between 5 to 10 seconds long.
Data size: 1,000 samples Sampling rate: 16 kHz
Files included:
The structure of each .json file is identical to MusicBench, as described in the previous section, with the exception of "alt_caption" column being empty. All captions are in the "main_caption" column!
@inproceedings{melechovsky2024mustango,
title={Mustango: Toward Controllable Text-to-Music Generation},
author={Melechovsky, Jan and Guo, Zixun and Ghosal, Deepanway and Majumder, Navonil and Herremans, Dorien and Poria, Soujanya},
booktitle={Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)},
pages={8286--8309},
year={2024}
}
License: cc-by-sa-3.0
The MusicBench dataset is a music audio-text pair dataset that was designed for text-to-music generation purpose and released along with Mustango text-to-music model. MusicBench is based on the MusicCaps dataset, which it expands from 5,521 samples to 52,768 training and 400 test samples!
MusicBench expands MusicCaps by:
Train set size = 52,768 samples Test set size = 400
MusicBench consists of 3 .json files and attached audio files in .tar.gz form.
The train set contains audio augmented samples and enhanced captions. Additionally, it offers ChatGPT rephrased captions for all the audio samples. Both TestA and TestB sets contain the same audio content, but TestB has all 4 possible control sentences (related to 4 music features) in captions of all samples, while TestA has no control sentences in the captions.
For more details, see Figure 1 in our paper.
Each row of a .json file has:
Hereby, we also present you the FMACaps evaluation dataset which consists of 1000 samples extracted from the Free Music Archive (FMA) and pseudocaptioned through extracting tags from audio and then utilizing ChatGPT in-context learning. More information is available in our paper!
Most of the samples are 10 second long, exceptions are between 5 to 10 seconds long.
Data size: 1,000 samples Sampling rate: 16 kHz
Files included:
The structure of each .json file is identical to MusicBench, as described in the previous section, with the exception of "alt_caption" column being empty. All captions are in the "main_caption" column!
@inproceedings{melechovsky2024mustango,
title={Mustango: Toward Controllable Text-to-Music Generation},
author={Melechovsky, Jan and Guo, Zixun and Ghosal, Deepanway and Majumder, Navonil and Herremans, Dorien and Poria, Soujanya},
booktitle={Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)},
pages={8286--8309},
year={2024}
}
License: cc-by-sa-3.0