CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages
13
25 commits
1 linked in READMEs
updated Feb 24, 2025
CLaMP 3 is a state-of-the-art framework for music information retrieval (MIR) across multiple modalities (βοΈ text, πΌ sheet music, π΅ audio, πΉ MIDI, and πΌοΈ images) and languages (π 27 trained, 100 supported). It leverages contrastive learning to align diverse music modalities into a shared representation space, enabling seamless cross-modal retrieval. You can think of it as a more comprehensive version of CLAP or MuLanβwith much stronger performance, support for all major music modalities, and global language coverage.
π Why CLaMP 3?
β
Multimodal: Works with βοΈ text, πΌ sheet music, π΅ audio, πΉ MIDI, and πΌοΈ images
β
Multilingual: Supports 27 trained & generalizes to π 100 languages
β
SOTA Performance: Significantly outperforms previous strong baselines across modalities and languages π
π‘ Text-to-Music Retrieval: Search music with text (100 languages!)
πΈ Image-to-Music Retrieval: Match music to images π¨
π Cross-Modal Retrieval: Find related music across different modalities
π οΈ Zero-Shot Classification: Identify genre, mood, style, & more π·οΈ
πΌ Semantic Similarity: Measure semantic similarity between generated & reference music
π Check it out: CLaMP 3 Homepage
For users who want to get started quickly with CLaMP3, follow these steps:
Run the following commands:
conda create -n clamp3 python=3.10.16 -y
conda activate clamp3
conda install pytorch torchvision torchaudio pytorch-cuda=11.8 -c pytorch -c nvidia -y
pip install -r requirements.txt
clamp3_*.py ScriptsCLaMP 3 provides scripts for semantic search, semantic similarity calculation, retrieval performance evaluation, and feature extraction across five modalities. Simply provide the file path, and the script will automatically detect the modality and extract the relevant features.
Supported formats include:
.mp3, .wav.mid, .midi.mxl, .musicxml, .xml.png, .jpg.txt (in 100 languages)cache/ directory and reused in future runs to avoid recomputation.temp/ and cleaned up after each run.Note: All files in a folder must belong to the same modality for processing.
clamp3_search.py - Semantic SearchRun retrieval tasks by comparing a query file to reference files in ref_dir. The query and ref_dir can be any modality, so there are 25 possible retrieval combinations, e.g., text-to-music, image-to-music, music-to-music, music-to-text (zero-shot music classification), etc.
python clamp3_search.py <query_file> <ref_dir> [--top_k TOP_K]
clamp3_score.py - Semantic Similarity CalculationThis script calculates semantic similarity between query and reference files. By default, it uses pairwise mode, but you can switch to group mode using the --group flag.
python clamp3_score.py <query_dir> <ref_dir> [--group]
Pairwise Mode (default):
Compares files with matching prefixes and identical folder structures.
Folder structure example:
query_dir/
βββ en/
β βββ sample1.wav
βββ zh/
β βββ sample1.1.wav
β βββ sample1.2.wav
β βββ sample2.wav
ref_dir/
βββ en/
β βββ sample1.txt
βββ zh/
β βββ sample1.txt
β βββ sample2.txt
query_dir/en/sample1.wav and ref_dir/en/sample1.txt).query_dir/zh/sample1.1.wav, query_dir/zh/sample1.2.wav) can correspond to one reference file (e.g., ref_dir/zh/sample1.txt).Important:
Group Mode:
Compares all query files to all reference files and calculates the average similarity.
Enable Group Mode:
python clamp3_score.py query_dir ref_dir --group
clamp3_eval.py - Retrieval Performance EvaluationEvaluates CLaMP3's retrieval performance on a paired dataset using metrics like MRR and Hit@K. Works the same way as pairwise mode in clamp3_score.pyβrequiring matching folder structure and filenames between query_dir and ref_dir.
python clamp3_eval.py <query_dir> <ref_dir>
clamp3_embd.py - Feature ExtractionIf other scripts don't meet your needs, use clamp3_embd.py to extract features.
python clamp3_embd.py <input_dir_path> <output_dir_path> [--get_global]
Feature Output:
--get_global β Shape: (1, T, 768) (T = time steps). Uses last hidden states before avg pooling, ideal for applications needing temporal info. Fine-tuning recommended.--get_global β Shape: (1, 768). Uses avg pooled features, suitable for applications needing global info, can be used directly.Note: Ensure the model weights are placed in the
code/folder, and verify the configuration hyperparameters before use.
Before using CLaMP 3, preprocess MusicXML files into Interleaved ABC, MIDI files into MTF, and audio files into MERT-extracted features.
CLaMP 3 requires Interleaved ABC notation for sheet music. Follow these steps:
Convert MusicXML (.mxl, .xml, .musicxml) to standard ABC using batch_xml2abc.py:
python batch_xml2abc.py <input_dir> <output_dir>
.mxl, .xml, .musicxml files.abc (Standard ABC) files will be savedConvert Standard ABC into Interleaved ABC using batch_interleaved_abc.py:
python batch_interleaved_abc.py <input_dir> <output_dir>
.abc (Standard ABC) filesCLaMP 3 processes performance signals in MIDI Text Format (MTF). Convert MIDI files (.mid, .midi) into MTF format using batch_midi2mtf.py:
python batch_midi2mtf.py <input_dir> <output_dir> --m3_compatible
.mid, .midi files.mtf files will be saved (MTF format for CLaMP 3)--m3_compatible flag must be included to ensure the output format is compatible with CLaMP 3. Without this flag, the extracted MTF files will not work correctly in the pipeline.For audio processing, CLaMP 3 uses MERT-extracted features instead of raw waveforms. Extract MERT-based features from raw audio (.mp3, .wav) using extract_mert.py:
python extract_mert.py --input_path <input_path> --output_path <output_path> --model_path m-a-p/MERT-v1-95M --mean_features
.mp3, .wav.npy (Processed audio features for CLaMP 3)CLaMP 3 is the most powerful music retrieval model, and in most cases, retraining is not needed. However, if necessary, follow these steps.
Modify config.py to adjust hyperparameters and data paths.
Train on your own data.
To train CLaMP 3 on symbolic music (e.g., sheet music, MIDI), run:
python -m torch.distributed.launch --nproc_per_node=<GPUs> --use_env train_clamp3_symbolic.py
For audio data, use:
python -m torch.distributed.launch --nproc_per_node=<GPUs> --use_env train_clamp3_audio.py
For most use cases, it's best to use pre-trained weights instead of training from scratch.
| Version | Best for | Download Link |
|---|---|---|
| CLaMP 3 SAAS | Audio-based retrieval (Recommended) | Download SAAS |
| CLaMP 3 C2 | Symbolic music retrieval (Sheet music, MIDI) | Download C2 |
By default, CLaMP 3 is configured for the SAAS version (optimized for audio).
config.py from "saas" to "c2".After training (or using pre-trained weights), extract features using extract_clamp3.py:
accelerate launch extract_clamp3.py --epoch <epoch> <input_dir> <output_dir> --get_global
--epoch <epoch>: (Optional) Specify the checkpoint epoch.<input_dir>: Directory containing the input files.<output_dir>: Destination folder for the output .npy features.--get_global: (Required for retrieval!) Extracts a global semantic vector for each input.All extracted features are stored as .npy files.
Note: For retrieval,
--get_globalmust be used. Without it, CLaMP 3 will not work correctly for retrieval tasks. You only omit--get_globalif you are performing downstream fine-tuning or need raw feature extraction for custom tasks.
If you find CLaMP 3 useful in your work, please consider citing our paper:
@misc{wu2025clamp3universalmusic,
title={CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages},
author={Shangda Wu and Zhancheng Guo and Ruibin Yuan and Junyan Jiang and Seungheon Doh and Gus Xia and Juhan Nam and Xiaobing Li and Feng Yu and Maosong Sun},
year={2025},
eprint={2502.10362},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2502.10362}
}
CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages
13
25 commits
1 linked in READMEs
updated Feb 24, 2025
CLaMP 3 is a state-of-the-art framework for music information retrieval (MIR) across multiple modalities (βοΈ text, πΌ sheet music, π΅ audio, πΉ MIDI, and πΌοΈ images) and languages (π 27 trained, 100 supported). It leverages contrastive learning to align diverse music modalities into a shared representation space, enabling seamless cross-modal retrieval. You can think of it as a more comprehensive version of CLAP or MuLanβwith much stronger performance, support for all major music modalities, and global language coverage.
π Why CLaMP 3?
β
Multimodal: Works with βοΈ text, πΌ sheet music, π΅ audio, πΉ MIDI, and πΌοΈ images
β
Multilingual: Supports 27 trained & generalizes to π 100 languages
β
SOTA Performance: Significantly outperforms previous strong baselines across modalities and languages π
π‘ Text-to-Music Retrieval: Search music with text (100 languages!)
πΈ Image-to-Music Retrieval: Match music to images π¨
π Cross-Modal Retrieval: Find related music across different modalities
π οΈ Zero-Shot Classification: Identify genre, mood, style, & more π·οΈ
πΌ Semantic Similarity: Measure semantic similarity between generated & reference music
π Check it out: CLaMP 3 Homepage
For users who want to get started quickly with CLaMP3, follow these steps:
Run the following commands:
conda create -n clamp3 python=3.10.16 -y
conda activate clamp3
conda install pytorch torchvision torchaudio pytorch-cuda=11.8 -c pytorch -c nvidia -y
pip install -r requirements.txt
clamp3_*.py ScriptsCLaMP 3 provides scripts for semantic search, semantic similarity calculation, retrieval performance evaluation, and feature extraction across five modalities. Simply provide the file path, and the script will automatically detect the modality and extract the relevant features.
Supported formats include:
.mp3, .wav.mid, .midi.mxl, .musicxml, .xml.png, .jpg.txt (in 100 languages)cache/ directory and reused in future runs to avoid recomputation.temp/ and cleaned up after each run.Note: All files in a folder must belong to the same modality for processing.
clamp3_search.py - Semantic SearchRun retrieval tasks by comparing a query file to reference files in ref_dir. The query and ref_dir can be any modality, so there are 25 possible retrieval combinations, e.g., text-to-music, image-to-music, music-to-music, music-to-text (zero-shot music classification), etc.
python clamp3_search.py <query_file> <ref_dir> [--top_k TOP_K]
clamp3_score.py - Semantic Similarity CalculationThis script calculates semantic similarity between query and reference files. By default, it uses pairwise mode, but you can switch to group mode using the --group flag.
python clamp3_score.py <query_dir> <ref_dir> [--group]
Pairwise Mode (default):
Compares files with matching prefixes and identical folder structures.
Folder structure example:
query_dir/
βββ en/
β βββ sample1.wav
βββ zh/
β βββ sample1.1.wav
β βββ sample1.2.wav
β βββ sample2.wav
ref_dir/
βββ en/
β βββ sample1.txt
βββ zh/
β βββ sample1.txt
β βββ sample2.txt
query_dir/en/sample1.wav and ref_dir/en/sample1.txt).query_dir/zh/sample1.1.wav, query_dir/zh/sample1.2.wav) can correspond to one reference file (e.g., ref_dir/zh/sample1.txt).Important:
Group Mode:
Compares all query files to all reference files and calculates the average similarity.
Enable Group Mode:
python clamp3_score.py query_dir ref_dir --group
clamp3_eval.py - Retrieval Performance EvaluationEvaluates CLaMP3's retrieval performance on a paired dataset using metrics like MRR and Hit@K. Works the same way as pairwise mode in clamp3_score.pyβrequiring matching folder structure and filenames between query_dir and ref_dir.
python clamp3_eval.py <query_dir> <ref_dir>
clamp3_embd.py - Feature ExtractionIf other scripts don't meet your needs, use clamp3_embd.py to extract features.
python clamp3_embd.py <input_dir_path> <output_dir_path> [--get_global]
Feature Output:
--get_global β Shape: (1, T, 768) (T = time steps). Uses last hidden states before avg pooling, ideal for applications needing temporal info. Fine-tuning recommended.--get_global β Shape: (1, 768). Uses avg pooled features, suitable for applications needing global info, can be used directly.Note: Ensure the model weights are placed in the
code/folder, and verify the configuration hyperparameters before use.
Before using CLaMP 3, preprocess MusicXML files into Interleaved ABC, MIDI files into MTF, and audio files into MERT-extracted features.
CLaMP 3 requires Interleaved ABC notation for sheet music. Follow these steps:
Convert MusicXML (.mxl, .xml, .musicxml) to standard ABC using batch_xml2abc.py:
python batch_xml2abc.py <input_dir> <output_dir>
.mxl, .xml, .musicxml files.abc (Standard ABC) files will be savedConvert Standard ABC into Interleaved ABC using batch_interleaved_abc.py:
python batch_interleaved_abc.py <input_dir> <output_dir>
.abc (Standard ABC) filesCLaMP 3 processes performance signals in MIDI Text Format (MTF). Convert MIDI files (.mid, .midi) into MTF format using batch_midi2mtf.py:
python batch_midi2mtf.py <input_dir> <output_dir> --m3_compatible
.mid, .midi files.mtf files will be saved (MTF format for CLaMP 3)--m3_compatible flag must be included to ensure the output format is compatible with CLaMP 3. Without this flag, the extracted MTF files will not work correctly in the pipeline.For audio processing, CLaMP 3 uses MERT-extracted features instead of raw waveforms. Extract MERT-based features from raw audio (.mp3, .wav) using extract_mert.py:
python extract_mert.py --input_path <input_path> --output_path <output_path> --model_path m-a-p/MERT-v1-95M --mean_features
.mp3, .wav.npy (Processed audio features for CLaMP 3)CLaMP 3 is the most powerful music retrieval model, and in most cases, retraining is not needed. However, if necessary, follow these steps.
Modify config.py to adjust hyperparameters and data paths.
Train on your own data.
To train CLaMP 3 on symbolic music (e.g., sheet music, MIDI), run:
python -m torch.distributed.launch --nproc_per_node=<GPUs> --use_env train_clamp3_symbolic.py
For audio data, use:
python -m torch.distributed.launch --nproc_per_node=<GPUs> --use_env train_clamp3_audio.py
For most use cases, it's best to use pre-trained weights instead of training from scratch.
| Version | Best for | Download Link |
|---|---|---|
| CLaMP 3 SAAS | Audio-based retrieval (Recommended) | Download SAAS |
| CLaMP 3 C2 | Symbolic music retrieval (Sheet music, MIDI) | Download C2 |
By default, CLaMP 3 is configured for the SAAS version (optimized for audio).
config.py from "saas" to "c2".After training (or using pre-trained weights), extract features using extract_clamp3.py:
accelerate launch extract_clamp3.py --epoch <epoch> <input_dir> <output_dir> --get_global
--epoch <epoch>: (Optional) Specify the checkpoint epoch.<input_dir>: Directory containing the input files.<output_dir>: Destination folder for the output .npy features.--get_global: (Required for retrieval!) Extracts a global semantic vector for each input.All extracted features are stored as .npy files.
Note: For retrieval,
--get_globalmust be used. Without it, CLaMP 3 will not work correctly for retrieval tasks. You only omit--get_globalif you are performing downstream fine-tuning or need raw feature extraction for custom tasks.
If you find CLaMP 3 useful in your work, please consider citing our paper:
@misc{wu2025clamp3universalmusic,
title={CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages},
author={Shangda Wu and Zhancheng Guo and Ruibin Yuan and Junyan Jiang and Seungheon Doh and Gus Xia and Juhan Nam and Xiaobing Li and Feng Yu and Maosong Sun},
year={2025},
eprint={2502.10362},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2502.10362}
}