hengtsune1024/AnchorSteer

[KDD 2026] Official implementation of AnchorSteer: Self-Discovered Concept Injection for Structure-Preserving Music Editing

Python

9

6 commits

updated Aug 6, 2026

See the code

README

AnchorSteer: Self-Discovered Concept Injection for Structure-Preserving Music Editing

arXiv Hugging Face Demo

Chih-Heng Chang1, Keng-Seng Ho1, Chih-Yu Tsai1, Kuan-Lin Chen1, Yi-Hsuan Yang2, Jian-Jiun Ding1,*

1Graduate Institute of Communication Engineering, National Taiwan University
2Artificial Intelligence Center of Research Excellence, National Taiwan University
*Corresponding author

Abstract

Controllable music editing is to modify high-level attributes while strictly preserving rhythmic and melodic structures. However, this task is challenged by a semantic-structural entanglement: steering methods often degrade structure to achieve editing performance, while structural adaptors suppress semantic responsiveness. We propose AnchorSteer, a framework that disentangles this tension by coupling structural anchoring with self-discovered semantic steering. The proposed approach probes internal representations to extract interpretable, label-free concept vectors via a self-supervised reconstruction objective, isolating attributes without curated data. During editing, these portable, plug-and-play concept vectors are injected into diffusion hidden manifolds while a structural adaptor enforces consistency. Variants for unconditioned and conditioned injections are provided to balance robustness and semantic strength. Experiments on ZoME-Bench and subjective tests show that the proposed framework outperforms both steering-only and anchoring-only baselines, enabling significant semantic transformations with high-fidelity structural preservation.


Installation

Hardware requirements

  • Python 3.10
  • CUDA 12.x (tested on 12.1 and 12.8)
  • Single NVIDIA GPU with at least 16 GB VRAM for training; 12 GB is sufficient for inference only
  • Reference hardware: RTX 4080 (16 GB) — approximately 30 minutes per concept for 20 epochs

1. Create and activate the conda environment

conda env create -f environment.yaml
conda activate env_anchorsteer

2. Install the patched local diffusers fork

pip install -e ./diffusers

WARNING: Do not run pip install diffusers afterward; this will overwrite the patched local fork and cause concept injection to silently fail. See diffusers/PATCHED_FILES.md for details.

3. Log in to HuggingFace

hf auth login

Before running any script you must accept the Stability AI license agreement on the model page at https://huggingface.co/stabilityai/stable-audio-open-1.0. The Stable Audio Open weights are downloaded automatically from HuggingFace on the first run and cached at ~/.cache/huggingface/.

4. Download MuseControlLite checkpoints

The structure-preserving editing pipeline (edit.py) requires MuseControlLite adapter checkpoints. Download them from the MuseControlLite project:

# Run from the AnchorSteer root directory
gdown --folder 1Q9B333jcq1czA11JKTbM-DHANJ8YqGbP -O MuseControlLite/checkpoints

This step is only required for edit.py. Plain concept steering via generate.py does not use these checkpoints.


Quick start: Steering-only inference

1. Download pre-trained concept checkpoints

The pre-trained concept checkpoints are hosted on Hugging Face: heng1024/AnchorSteer-weights (DOI: 10.57967/hf/8920)

From the root directory of this repository, download the checkpoints:

hf download heng1024/AnchorSteer-weights --repo-type model --include "exps/*" --local-dir .

This places the exps/ folder directly into the repo root, matching the expected layout.

We provide 5 pre-trained concept checkpoints: rock, jazz, piano, harp, ambient. In the following example, we use rock as an example, but it can be replaced with any of the other four concepts.

2. Run inference

python generate.py --exp_dir exps/rock-transformer

Defaults: 3 samples, prompt "a music piece". Override with --num_sample N and --prompt "your prompt".

3. Expected output

Six WAV files will be written to exps/rock-transformer/best/:

  • a music piece_0_orig.wav … a music piece_2_orig.wav — original (unsteered) generations
  • a music piece_0_edited.wav … a music piece_2_edited.wav — concept-steered generations

Quick start: AnchorSteer inference with MuseControlLite

To apply concept steering to an existing audio file while preserving its musical structure, use edit.py:

python edit.py \
    --source_audio path/to/source.wav \
    --concept_dir exps/rock-transformer

Prerequisites: complete Installation steps 1–3, including the MuseControlLite checkpoint download.

The edited WAV is written to outputs/<source_stem>_rock-transformer.wav.

Options:

FlagDefaultDescription
--prompt"a music piece"Text prompt describing the desired output
--output_diroutputsDirectory where the edited WAV is saved
--negative_prompt""Negative text prompt
--num_steps50Number of diffusion sampling steps
--seed0Random seed for reproducibility
--use_lastoffUse adaptor.pth (final epoch) instead of best.pth

Training a custom concept

The following example trains a rock Controller from scratch. Replace rock with any concept name.

Step 1: Create a dataset config

Create dataset_config/rock.json with the following schema:

{
    "root_dir": "datasets/rock",
    "num_samples": 1000,
    "audio_len": -1,
    "reference_prompt": [
        "rock music with electric guitar and heavy drums",
        "Low quality"
    ],
    "training_labels": [
        [["a music piece", ["rock"]]]
    ]
}

Field reference

FieldTypeDescription
root_dirstringDirectory where generated WAV files are saved (relative to the AnchorSteer root).
num_samplesintNumber of reference audio clips to generate. 1000 is recommended for full training; 10–50 is enough for a quick smoke test.
audio_lenfloatDuration in seconds of each generated clip. -1 generates the full 47-second audio (pair with --train_full_len during training). A positive value such as 5 produces shorter clips that train faster but capture less temporal structure.
reference_prompt[positive, negative]The prompt pair passed to Stable Audio Open when generating the reference audio. The positive prompt steers the style toward the target concept; the negative prompt suppresses unwanted qualities (e.g. "Low quality").
training_labelslistMaps each audio file to the (prompt, concept) pair used during Controller training. Format: [[["text_prompt", ["concept_name"]]]]. The text prompt is the neutral conditioning text the Controller sees at training time — keep it generic (e.g. "a music piece") so the Controller learns a concept delta that is independent of any particular text input.

reference_prompt vs training_labels text prompt: these serve different roles. reference_prompt controls what audio gets generated (concept-specific, e.g. "rock music with electric guitar"). The text prompt inside training_labels is what the Controller is conditioned on during training (concept-neutral, e.g. "a music piece"). Keeping them separate prevents the Controller from entangling concept knowledge with text-prompt knowledge.

Step 2: Generate reference audio

python data_creation.py --config dataset_config/rock.json

Expected output: datasets/rock/0000.wav through datasets/rock/0999.wav plus labels.json and test.json.

Step 3: Train the Controller

python train.py \
    --train_data_dir datasets/rock \
    --output_dir exps/rock-transformer \
    --train_full_len

Expected training time: approximately 30 minutes on an RTX 4080.

Expected outputs in exps/rock-transformer/:

  • best.pth — best checkpoint (lowest validation loss)
  • adaptor.pth — final-epoch checkpoint
  • config.yaml — full argument snapshot for reproducibility
  • loss_history.png — training / validation loss curves

Key flags reference

FlagDefaultDescription
--control_typetransformerController architecture: transformer (bottleneck TransformerEncoder capturing temporal context) or vector (static learnable delta, ~1–3 M params, faster)
--train_full_lenFalseTrain on full 47-second audio; omit to use 5-second crops
--edit_start0Index of the first SAO DiT transformer block to inject the Controller delta
--edit_end23Index of the last SAO DiT transformer block to inject the Controller delta (inclusive)
--learning_rate5e-4AdamW learning rate
--num_train_epochs20Number of training epochs
--mini_batch_size2Mini-batch size per device; effective batch size = mini_batch_size × gradient_accumulation_steps
--encode_batch_size4Batch size for VAE-encoding the dataset during preprocessing; increase for faster encoding on high-VRAM GPUs
--gradient_accumulation_steps2Number of gradient accumulation steps before an optimizer step

Evaluation

See eval/README.md for setup and usage. The evaluator computes CLAP, LPIPS, and Chroma scores for original and edited audio.


Repository layout

AnchorSteer/
├── README.md
├── LICENSE
├── environment.yaml
├── requirements.txt
├── diffusers/              # Pinned local fork of huggingface/diffusers (0.37.0.dev0)
├── dataset_config/         # Concept JSON configs
├── eval/                   # Evaluation scripts (see eval/README.md)
├── models/                 # AnchorSteer core modules
│   ├── __init__.py
│   ├── controller.py       # Vector / Transformer Controller
│   ├── stable_audio.py     # MyStableAudioPipeline + MyStableAudioDiTModel
│   ├── utils.py            # load_pipeline() helper
│   └── musecontrollite/    # AnchorSteer-MuseControlLite integration layer
├── MuseControlLite/        # fundwotsai2001/MuseControlLite — structure-anchoring backbone
│   └── checkpoints/        # MuseControlLite adapter weights (download separately, see Installation)
├── data_creation.py        # Generate reference audio dataset for a concept
├── train.py                # Train Controller
├── generate.py             # Inference: generate original + concept-steered audio pairs from a text prompt
├── edit.py                 # Inference: structure-preserving concept editing of a source audio file
├── config.py               # Argument parser for train.py
├── utils_data.py           # DataLoader for pre-encoded latents
├── misc.py                 # set_seed and small utilities
└── doc/                    # Internal planning and change-log documents

Acknowledgements

We thank the authors of the following open-source projects:


Citation

If you use AnchorSteer in your research, please cite:

@inproceedings{anchosteer2026,
  title     = {AnchorSteer: Self-Discovered Concept Injection for Structure-Preserving Music Editing},
  author    = {Chih-Heng Chang, Keng-Seng Ho, Chih-Yu Tsai, Kuan-Lin Chen, Yi-Hsuan Yang, Jian-Jiun Ding},
  booktitle = {Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2},
  year      = {2026}
}

License

This project is released under the MIT License (see LICENSE). The MuseControlLite component (MuseControlLite/) is included under its own license; see MuseControlLite/LICENSE for details.

hengtsune1024/AnchorSteer

[KDD 2026] Official implementation of AnchorSteer: Self-Discovered Concept Injection for Structure-Preserving Music Editing

Python

9

6 commits

updated Aug 6, 2026

See the code

README

AnchorSteer: Self-Discovered Concept Injection for Structure-Preserving Music Editing

arXiv Hugging Face Demo

Chih-Heng Chang1, Keng-Seng Ho1, Chih-Yu Tsai1, Kuan-Lin Chen1, Yi-Hsuan Yang2, Jian-Jiun Ding1,*

1Graduate Institute of Communication Engineering, National Taiwan University
2Artificial Intelligence Center of Research Excellence, National Taiwan University
*Corresponding author

Abstract

Controllable music editing is to modify high-level attributes while strictly preserving rhythmic and melodic structures. However, this task is challenged by a semantic-structural entanglement: steering methods often degrade structure to achieve editing performance, while structural adaptors suppress semantic responsiveness. We propose AnchorSteer, a framework that disentangles this tension by coupling structural anchoring with self-discovered semantic steering. The proposed approach probes internal representations to extract interpretable, label-free concept vectors via a self-supervised reconstruction objective, isolating attributes without curated data. During editing, these portable, plug-and-play concept vectors are injected into diffusion hidden manifolds while a structural adaptor enforces consistency. Variants for unconditioned and conditioned injections are provided to balance robustness and semantic strength. Experiments on ZoME-Bench and subjective tests show that the proposed framework outperforms both steering-only and anchoring-only baselines, enabling significant semantic transformations with high-fidelity structural preservation.


Installation

Hardware requirements

  • Python 3.10
  • CUDA 12.x (tested on 12.1 and 12.8)
  • Single NVIDIA GPU with at least 16 GB VRAM for training; 12 GB is sufficient for inference only
  • Reference hardware: RTX 4080 (16 GB) — approximately 30 minutes per concept for 20 epochs

1. Create and activate the conda environment

conda env create -f environment.yaml
conda activate env_anchorsteer

2. Install the patched local diffusers fork

pip install -e ./diffusers

WARNING: Do not run pip install diffusers afterward; this will overwrite the patched local fork and cause concept injection to silently fail. See diffusers/PATCHED_FILES.md for details.

3. Log in to HuggingFace

hf auth login

Before running any script you must accept the Stability AI license agreement on the model page at https://huggingface.co/stabilityai/stable-audio-open-1.0. The Stable Audio Open weights are downloaded automatically from HuggingFace on the first run and cached at ~/.cache/huggingface/.

4. Download MuseControlLite checkpoints

The structure-preserving editing pipeline (edit.py) requires MuseControlLite adapter checkpoints. Download them from the MuseControlLite project:

# Run from the AnchorSteer root directory
gdown --folder 1Q9B333jcq1czA11JKTbM-DHANJ8YqGbP -O MuseControlLite/checkpoints

This step is only required for edit.py. Plain concept steering via generate.py does not use these checkpoints.


Quick start: Steering-only inference

1. Download pre-trained concept checkpoints

The pre-trained concept checkpoints are hosted on Hugging Face: heng1024/AnchorSteer-weights (DOI: 10.57967/hf/8920)

From the root directory of this repository, download the checkpoints:

hf download heng1024/AnchorSteer-weights --repo-type model --include "exps/*" --local-dir .

This places the exps/ folder directly into the repo root, matching the expected layout.

We provide 5 pre-trained concept checkpoints: rock, jazz, piano, harp, ambient. In the following example, we use rock as an example, but it can be replaced with any of the other four concepts.

2. Run inference

python generate.py --exp_dir exps/rock-transformer

Defaults: 3 samples, prompt "a music piece". Override with --num_sample N and --prompt "your prompt".

3. Expected output

Six WAV files will be written to exps/rock-transformer/best/:

  • a music piece_0_orig.wav … a music piece_2_orig.wav — original (unsteered) generations
  • a music piece_0_edited.wav … a music piece_2_edited.wav — concept-steered generations

Quick start: AnchorSteer inference with MuseControlLite

To apply concept steering to an existing audio file while preserving its musical structure, use edit.py:

python edit.py \
    --source_audio path/to/source.wav \
    --concept_dir exps/rock-transformer

Prerequisites: complete Installation steps 1–3, including the MuseControlLite checkpoint download.

The edited WAV is written to outputs/<source_stem>_rock-transformer.wav.

Options:

FlagDefaultDescription
--prompt"a music piece"Text prompt describing the desired output
--output_diroutputsDirectory where the edited WAV is saved
--negative_prompt""Negative text prompt
--num_steps50Number of diffusion sampling steps
--seed0Random seed for reproducibility
--use_lastoffUse adaptor.pth (final epoch) instead of best.pth

Training a custom concept

The following example trains a rock Controller from scratch. Replace rock with any concept name.

Step 1: Create a dataset config

Create dataset_config/rock.json with the following schema:

{
    "root_dir": "datasets/rock",
    "num_samples": 1000,
    "audio_len": -1,
    "reference_prompt": [
        "rock music with electric guitar and heavy drums",
        "Low quality"
    ],
    "training_labels": [
        [["a music piece", ["rock"]]]
    ]
}

Field reference

FieldTypeDescription
root_dirstringDirectory where generated WAV files are saved (relative to the AnchorSteer root).
num_samplesintNumber of reference audio clips to generate. 1000 is recommended for full training; 10–50 is enough for a quick smoke test.
audio_lenfloatDuration in seconds of each generated clip. -1 generates the full 47-second audio (pair with --train_full_len during training). A positive value such as 5 produces shorter clips that train faster but capture less temporal structure.
reference_prompt[positive, negative]The prompt pair passed to Stable Audio Open when generating the reference audio. The positive prompt steers the style toward the target concept; the negative prompt suppresses unwanted qualities (e.g. "Low quality").
training_labelslistMaps each audio file to the (prompt, concept) pair used during Controller training. Format: [[["text_prompt", ["concept_name"]]]]. The text prompt is the neutral conditioning text the Controller sees at training time — keep it generic (e.g. "a music piece") so the Controller learns a concept delta that is independent of any particular text input.

reference_prompt vs training_labels text prompt: these serve different roles. reference_prompt controls what audio gets generated (concept-specific, e.g. "rock music with electric guitar"). The text prompt inside training_labels is what the Controller is conditioned on during training (concept-neutral, e.g. "a music piece"). Keeping them separate prevents the Controller from entangling concept knowledge with text-prompt knowledge.

Step 2: Generate reference audio

python data_creation.py --config dataset_config/rock.json

Expected output: datasets/rock/0000.wav through datasets/rock/0999.wav plus labels.json and test.json.

Step 3: Train the Controller

python train.py \
    --train_data_dir datasets/rock \
    --output_dir exps/rock-transformer \
    --train_full_len

Expected training time: approximately 30 minutes on an RTX 4080.

Expected outputs in exps/rock-transformer/:

  • best.pth — best checkpoint (lowest validation loss)
  • adaptor.pth — final-epoch checkpoint
  • config.yaml — full argument snapshot for reproducibility
  • loss_history.png — training / validation loss curves

Key flags reference

FlagDefaultDescription
--control_typetransformerController architecture: transformer (bottleneck TransformerEncoder capturing temporal context) or vector (static learnable delta, ~1–3 M params, faster)
--train_full_lenFalseTrain on full 47-second audio; omit to use 5-second crops
--edit_start0Index of the first SAO DiT transformer block to inject the Controller delta
--edit_end23Index of the last SAO DiT transformer block to inject the Controller delta (inclusive)
--learning_rate5e-4AdamW learning rate
--num_train_epochs20Number of training epochs
--mini_batch_size2Mini-batch size per device; effective batch size = mini_batch_size × gradient_accumulation_steps
--encode_batch_size4Batch size for VAE-encoding the dataset during preprocessing; increase for faster encoding on high-VRAM GPUs
--gradient_accumulation_steps2Number of gradient accumulation steps before an optimizer step

Evaluation

See eval/README.md for setup and usage. The evaluator computes CLAP, LPIPS, and Chroma scores for original and edited audio.


Repository layout

AnchorSteer/
├── README.md
├── LICENSE
├── environment.yaml
├── requirements.txt
├── diffusers/              # Pinned local fork of huggingface/diffusers (0.37.0.dev0)
├── dataset_config/         # Concept JSON configs
├── eval/                   # Evaluation scripts (see eval/README.md)
├── models/                 # AnchorSteer core modules
│   ├── __init__.py
│   ├── controller.py       # Vector / Transformer Controller
│   ├── stable_audio.py     # MyStableAudioPipeline + MyStableAudioDiTModel
│   ├── utils.py            # load_pipeline() helper
│   └── musecontrollite/    # AnchorSteer-MuseControlLite integration layer
├── MuseControlLite/        # fundwotsai2001/MuseControlLite — structure-anchoring backbone
│   └── checkpoints/        # MuseControlLite adapter weights (download separately, see Installation)
├── data_creation.py        # Generate reference audio dataset for a concept
├── train.py                # Train Controller
├── generate.py             # Inference: generate original + concept-steered audio pairs from a text prompt
├── edit.py                 # Inference: structure-preserving concept editing of a source audio file
├── config.py               # Argument parser for train.py
├── utils_data.py           # DataLoader for pre-encoded latents
├── misc.py                 # set_seed and small utilities
└── doc/                    # Internal planning and change-log documents

Acknowledgements

We thank the authors of the following open-source projects:


Citation

If you use AnchorSteer in your research, please cite:

@inproceedings{anchosteer2026,
  title     = {AnchorSteer: Self-Discovered Concept Injection for Structure-Preserving Music Editing},
  author    = {Chih-Heng Chang, Keng-Seng Ho, Chih-Yu Tsai, Kuan-Lin Chen, Yi-Hsuan Yang, Jian-Jiun Ding},
  booktitle = {Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2},
  year      = {2026}
}

License

This project is released under the MIT License (see LICENSE). The MuseControlLite component (MuseControlLite/) is included under its own license; see MuseControlLite/LICENSE for details.

Languages

Python

100.0%