The MidiCaps dataset [1] is a large-scale dataset of 168,385 midi music files with descriptive text captions, and a set of extracted musical features.
The captions have been produced through a captioning pipeline incorporating MIR feature extraction and LLM Claude 3 to caption the data from extracted features with an in-context learning task. The framework used to extract the captions is available open source on github. The original MIDI files originate from the Lakh MIDI Dataset [2,3] and are creative commons licenced.
Listen to a few example synthesized midi files with their captions here.
If you use this dataset, please cite the paper in which it is presented: Jan Melechovsky, Abhinaba Roy, Dorien Herremans, 2024, MidiCaps: A large-scale MIDI dataset with text captions.
We provide all the midi files in a .tar.gz form. Captions are provided as .json files. The "short" version contains the midi file name and the associated caption.
The dataset file contains these main columns:
Additionally, the file contains the following features that were used for captioning:
Last, the file contains the following additional features:
If you use this dataset, please cite the paper that presents it:
BibTeX:
@article{Melechovsky2024,
author = {Jan Melechovsky and Abhinaba Roy and Dorien Herremans},
title = {MidiCaps: A Large-scale MIDI Dataset with Text Captions},
year = {2024},
journal = {arXiv:2406.02255}
}
APA: Jan Melechovsky, Abhinaba Roy, Dorien Herremans, 2024, MidiCaps: A large-scale MIDI dataset with text captions. arXiv:2406.02255.
GitHub: https://github.com/AMAAI-Lab/MidiCaps
[1] Jan Melechovsky, Abhinaba Roy, Dorien Herremans. 2024. MidiCaps: A large-scale MIDI dataset with text captions. arXiv:2406.02255.
[2] Raffel, Colin. Learning-based methods for comparing sequences, with applications to audio-to-midi alignment and matching. Columbia University, 2016.
The MidiCaps dataset [1] is a large-scale dataset of 168,385 midi music files with descriptive text captions, and a set of extracted musical features.
The captions have been produced through a captioning pipeline incorporating MIR feature extraction and LLM Claude 3 to caption the data from extracted features with an in-context learning task. The framework used to extract the captions is available open source on github. The original MIDI files originate from the Lakh MIDI Dataset [2,3] and are creative commons licenced.
Listen to a few example synthesized midi files with their captions here.
If you use this dataset, please cite the paper in which it is presented: Jan Melechovsky, Abhinaba Roy, Dorien Herremans, 2024, MidiCaps: A large-scale MIDI dataset with text captions.
We provide all the midi files in a .tar.gz form. Captions are provided as .json files. The "short" version contains the midi file name and the associated caption.
The dataset file contains these main columns:
Additionally, the file contains the following features that were used for captioning:
Last, the file contains the following additional features:
If you use this dataset, please cite the paper that presents it:
BibTeX:
@article{Melechovsky2024,
author = {Jan Melechovsky and Abhinaba Roy and Dorien Herremans},
title = {MidiCaps: A Large-scale MIDI Dataset with Text Captions},
year = {2024},
journal = {arXiv:2406.02255}
}
APA: Jan Melechovsky, Abhinaba Roy, Dorien Herremans, 2024, MidiCaps: A large-scale MIDI dataset with text captions. arXiv:2406.02255.
GitHub: https://github.com/AMAAI-Lab/MidiCaps
[1] Jan Melechovsky, Abhinaba Roy, Dorien Herremans. 2024. MidiCaps: A large-scale MIDI dataset with text captions. arXiv:2406.02255.
[2] Raffel, Colin. Learning-based methods for comparing sequences, with applications to audio-to-midi alignment and matching. Columbia University, 2016.