mulab-mir/muchomusic

Dataset

MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models

8

3 commits

1 linked in READMEs

updated Aug 5, 2024

See the code

README

MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models

MuChoMusic is a benchmark designed to evaluate music understanding in multimodal language models focused on audio. It includes 1,187 multiple-choice questions validated by human annotators, based on 644 music tracks from two publicly available music datasets. These questions cover a wide variety of genres and assess knowledge and reasoning across several musical concepts and their cultural and functional contexts. The benchmark provides a holistic evaluation of five open-source models, revealing challenges such as over-reliance on the language modality and highlighting the need for better multimodal integration.

Note on Audio Files

This dataset comes without audio files. The audio files can be downloaded from two datasets: SongDescriberDataset (SDD) and MusicCaps. Please see the code repository for more information on how to download the audio.

Citation

If you use this dataset, please cite our paper:

@inproceedings{weck2024muchomusic,
   title={MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models},
   author={Weck, Benno and Manco, Ilaria and Benetos, Emmanouil and Quinton, Elio and Fazekas, György and Bogdanov, Dmitry},
   booktitle = {Proceedings of the 25th International Society for Music Information Retrieval Conference (ISMIR)},
   year={2024}
}
multimodal
music

mulab-mir/muchomusic

Dataset

MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models

8

3 commits

1 linked in READMEs

updated Aug 5, 2024

See the code

README

MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models

MuChoMusic is a benchmark designed to evaluate music understanding in multimodal language models focused on audio. It includes 1,187 multiple-choice questions validated by human annotators, based on 644 music tracks from two publicly available music datasets. These questions cover a wide variety of genres and assess knowledge and reasoning across several musical concepts and their cultural and functional contexts. The benchmark provides a holistic evaluation of five open-source models, revealing challenges such as over-reliance on the language modality and highlighting the need for better multimodal integration.

Note on Audio Files

This dataset comes without audio files. The audio files can be downloaded from two datasets: SongDescriberDataset (SDD) and MusicCaps. Please see the code repository for more information on how to download the audio.

Citation

If you use this dataset, please cite our paper:

@inproceedings{weck2024muchomusic,
   title={MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models},
   author={Weck, Benno and Manco, Ilaria and Benetos, Emmanouil and Quinton, Elio and Fazekas, György and Bogdanov, Dmitry},
   booktitle = {Proceedings of the 25th International Society for Music Information Retrieval Conference (ISMIR)},
   year={2024}
}
multimodal
music