[AACL-IJCNLP] A survey of foundation-model-based audio editing across speech, music, and general audio, covering task taxonomies, training-based and training-free methods, datasets, and evaluation.
See the codeAACL-IJCNLP 2026
Changhao Pan1,*, Yifei Fan1,*, Fan Zhuo1,*, Yifu Chen1, Wenxiang Guo1,
Yu Zhang2, Ruiqi Li2, Zhiyuan Zhu1, Rui Yang1, Shengpeng Ji3,
Chenyuhao Wen1, Jiayang Xu1, Ke Lei1, Xiaoda Yang1, Jingyu Lu1, Zhou Zhao1,†
1Zhejiang University · 2ByteDance · 3Hunyuan Team, Tencent
*Equal contribution · †Corresponding author
This repository is the official repository for Audio Editing in the Era of Foundation Models: A Survey, Which is accepted by AACL-IJCNLP 2026.
MM-Speech/AudioEditSurvey for better management.This is the official repository for Audio Editing in the Era of Foundation Models: A Survey, accepted to AACL-IJCNLP 2026. It is maintained by MM-Speech and collects papers and resources for foundation-model-based audio editing.
Abstract
Audio editing aims to modify a given synthetic or real-world audio signal to meet users' specific needs. As a promising yet challenging direction in AIGC, it has attracted increasing attention in recent years. With the rapid progress of text-to-audio and text-to-speech generation, powerful audio generation models have become the primary foundation for modern audio editing systems. In this survey, we provide a comprehensive review of foundation-model-based audio editing. We first define the scope of audio editing from a unified perspective and present a detailed taxonomy of existing editing tasks. We then summarize the major foundation-model paradigms for audio editing, and review representative approaches from both training-based and training-free perspectives. In addition, we systematically discuss related resources, including datasets, data construction tools, and evaluation protocols. Finally, we identify open challenges in this field and outline promising directions for future research.
In this survey, we focus on works that make direct contributions to audio editing in the era of foundation models. To ensure a precise and focused discussion, we adopt two main inclusion criteria: (1) the task should center on audio editing, which we define as modifying the acoustic attributes, instances, or content of an existing audio recording, without transformations so substantial that they amount to generating an entirely new audio sample; (2) the method should rely on mainstream audio foundation model paradigms. Accordingly, we do not cover works primarily focused on audio generation, nor do we provide an extensive discussion of signal-processing-based audio editing methods. In addition, to maintain a focused scope, spatial audio (multi-channel formats such as binaural stereo and first-order Ambisonics (FOA)) and related editing techniques are beyond the main scope of this survey.

Figure 1: Taxonomy of audio editing tasks.
| Category | Definition | Representative Editing Goals |
|---|---|---|
| Acoustic Editing | Modifies low-level perceptual attributes while preserving the overall structure and source characteristics of the original audio. | restoration, reverberation editing, loudness/mixing control, equalization, spectral texture editing |
| Semantic Editing | Modifies high-level interpretable information conveyed by audio while maintaining task-irrelevant properties. | linguistic editing, expressive editing, stylistic editing |
| Instance Editing | Manipulates identifiable audio entities while preserving the remaining scene and source relationships. | replacement, deletion/extraction, insertion, overlay/remixing |
Representative editors with publicly released implementations and model weights, subject to each project’s license. Editing types follow our taxonomy; Unified groups editors supporting multiple audio domains, with their supported domains listed in the table. Base links the pretrained backbone used by an editor, while Adapter links its additional learned weights.
| Model | Audio Domain | Editing Types | Model Architecture | Paper | Code | Model |
|---|---|---|---|---|---|---|
| Audio-Omni | Speech; Music; Audio | Instance: addition, removal, extraction, source transformation | MLLM + rectified-flow DiT | 🤗 Weights | ||
| AudioMorphix | Speech; Music; Audio | Semantic: pitch / time stretching Instance: addition, removal, replacement, time shifting | Diffusion U-Net (Tango / AudioLDM) | 🤗 Base (Tango 2) 🤗 Base (AudioLDM) | ||
| AuK / AuK-Flash | Speech; Music | Acoustic: restoration, loudness Semantic: words, lyrics, expression Instance: timbre, source extraction | MLLM + rectified-flow DiT | 🤗 AuK 🤗 Flash | ||
| Vevo2 | Speech; Music | Semantic: content, lyrics, prosody, style Instance: voice / singer conversion | Codec LM + flow-matching decoder | 🤗 Weights | ||
| DirectAudioEdit | Music; Audio | Instance: text-guided event replacement / addition / removal | Diffusion U-Net (Tango 2 / AudioLDM2) | 🤗 Base (Tango 2) 🤗 Base (audio) | ||
| DDPM Inversion (ZETA) | Music; Audio | Semantic: musical style Instance: instrument / sound-event changes | Diffusion U-Net (AudioLDM2) | 🤗 Base (audio) 🤗 Base (music) |
| Model | Editing Types | Model Architecture | Paper | Code | Model |
|---|---|---|---|---|---|
| Ming-UniAudio-Edit | Acoustic: denoising, loudness Semantic: content, prosody, emotion, dialect | Continuous-token LM + diffusion head | 🤗 Weights | ||
| Step-Audio-EditX | Semantic: emotion, speaking style, paralinguistics, pronunciation | Codec LM + flow-matching decoder | 🤗 Weights | ||
| CosyEdit | Semantic: word insertion, deletion, replacement | Codec LM + flow-matching decoder | 🤗 Weights | ||
| VoiceCraft-X | Semantic: multilingual content editing | Codec LM (autoregressive infilling) | 🤗 Weights | ||
| VoiceCraft | Semantic: word insertion, deletion, replacement | Codec LM (autoregressive infilling) | 🤗 Weights | ||
| SSR-Speech | Semantic: word insertion, deletion, replacement | Codec LM (autoregressive infilling) | 🤗 English 🤗 Mandarin | ||
| F5-TTS | Semantic: local content replacement / infilling | Flow-matching DiT | 🤗 Weights | ||
| FluentSpeech | Semantic: content editing, disfluency correction | Diffusion (context-aware denoiser) | 📁 Weights | ||
| EdiTTS | Semantic: content / pitch edits in synthesized speech | Score-based diffusion (Grad-TTS) | 📦 Base |
| Model | Editing Types | Model Architecture | Paper | Code | Model |
|---|---|---|---|---|---|
| YingMusic-Singer-Plus | Semantic: lyrics Instance: singer timbre replacement | Flow-matching DiT | 🤗 Weights | ||
| ACE-Step 1.5 | Semantic: style / local repainting Instance: track extraction / addition (base variant) | LM + flow-matching DiT | 🤗 Turbo 🤗 Base variant | ||
| Instruct-MusicGen | Instance: stem addition, removal, extraction | Codec LM (MusicGen) + adapters | 🤗 Public-data retraining | ||
| MusicGen-Stem | Instance: stem replacement / addition (bass, drums, other) | Multi-stream codec LM | 🤗 Weights | ||
| MelodyFlow | Semantic: genre, mood, style Instance: instrumentation | Flow-matching DiT | 🤗 Weights | ||
| AP-Adapter | Semantic: genre / style transfer Instance: instrument replacement | Diffusion U-Net + audio-prompt adapter | 📁 Adapter 🤗 Base | ||
| AnchorSteer | Semantic: genre / style Instance: instrument changes | Diffusion DiT + structural/concept adapters | 🤗 Concept weights 📁 Structure adapter 🤗 Base (access terms) |
| Model | Editing Types | Model Architecture | Paper | Code | Model |
|---|---|---|---|---|---|
| MMEdit | Acoustic: loudness Instance: event addition, removal, replacement, reordering | ALM + diffusion MMDiT | 🤗 Weights | ||
| SAO-Instruct | Acoustic: filtering, denoising, restoration Semantic: pitch / rate Instance: event manipulation | Diffusion DiT (Stable Audio Open) | 🤗 Weights | ||
| SmartDJ-Editor | Acoustic: volume, reverb, spectral coloration Instance: event addition, removal, extraction, relocation | Diffusion Transformer (U-DiT) | 🤗 Editor weights | ||
| AudioEditor | Instance: event addition, deletion, replacement | Diffusion U-Net (Auffusion) | 🤗 Base | ||
| CoherentAVEdit | Instance: video-conditioned sound-event replacement | Flow-matching Transformer (MMAudio) | 🤗 Weights |
Before the foundation-model era, early neural audio editing methods mainly explored task-specific generative models for local reconstruction and attribute control.
Token-based codec language models cast audio editing as conditional generation over discrete audio tokens. After continuous audio is converted into compact discrete token sequences, target regions are edited through autoregressive continuation, infilling, or selective regeneration conditioned on context, prompts, or task controls.
Diffusion and flow-matching models formulate audio editing as conditional transformation in continuous acoustic spaces, such as mel-spectrograms or audio latents. Instead of infilling discrete tokens, they modify audio through conditional denoising, latent inversion, or continuous flow transformation, making them suitable for high-fidelity reconstruction, region-level refinement, and fine-grained acoustic control in complex scenarios.
Instruction-conditioned and multimodal interfaces for audio editing provide high-level control for foundation-model-based audio editing. They allow users to specify editing intents through natural language instructions, task prompts, reference audio, temporal regions, or visual cues, which are shifted into target spans, task embeddings, event locations, speaker references, or preservation constraints.
Training-based approaches refer to audio editing methods that learn editing behaviors from supervised pairs, pseudo-pairs, or instruction-based triplets before inference. These methods explicitly optimize editing objectives, condition following, and preservation constraints, enabling stable and controllable editing. We group existing works into three categories based on their supervision and conditioning mechanisms, and discuss their core methods and functional scopes.
Figure 2: Overview of training-based audio editing methods.
| Paradigm | Description | Representative Scope |
|---|---|---|
| Task-specific Training | Optimizes models for predefined editing functions or domains. | text-based speech editing, prosody correction, source separation, music stem separation |
| Reference- and Attribute-based Training | Specifies the editing direction through reference audio, style examples, or attribute labels. | voice conversion, timbre transfer, emotion editing, mixing style transfer |
| Instruction-conditioned Training | Learns from instruction-input-output triplets to follow natural-language editing requests. | addition, deletion, replacement, inpainting, super-resolution, music remixing, expressive refinement |
Training-free approaches adapt pretrained audio generative models to editing without parameter updates. They operate by manipulating inference-time mechanisms, such as inversion, attention control, prompt or guidance adjustment, and mask-based constraints. We group existing methods into three common categories, which are often combined to improve localization, preservation, and controllability. Since token-based autoregressive models are less naturally suited to training-free editing, this section mainly focuses on non-autoregressive paradigms, especially diffusion-based foundation models.
Figure 3: Overview of training-free audio editing methods.
| Paradigm | Description | Representative Scope |
|---|---|---|
| Inversion-Based Editing | Maps source audio back into the latent, noise, or trajectory space of a pretrained generative model, then edits it by modifying conditions or sampling trajectories. | DDPM/DDIM inversion, latent inversion, flow-based inversion, speech or music reconstruction and editing |
| Attention-Controlled Editing | Guides pretrained generative models by modifying or reusing internal attention patterns without parameter updates. | cross-attention event localization, self-attention preservation, prompt-level manipulation |
| Mask- and Region-Guided Editing | Specifies where to edit and where to preserve the source audio in waveform, spectrogram, latent, or source-component spaces. | localized editing, inpainting, restoration, source-level manipulation |
| Token-Level Editing with Codec Models | Manipulates discrete audio tokens through masking, infilling, continuation, or selective regeneration at inference time. | speech infilling, localized resynthesis, codec-token editing |
Public datasets for audio editing and controllable audio generation, grouped by their primary audio domain.
This non-exhaustive list highlights datasets suited to audio editing or widely used in the community, with availability verified by the repository maintainers for every entry.
Paired indicates released source–target audio, mixture–stem correspondence, or explicitly matched control/technique takes (✅ / ❌); shared transcripts, audio–text alignment, or audio–MIDI alignment alone do not count. † marks an editing use that requires task construction or adaptation, rather than native editing supervision. Editing types follow our Acoustic / Instance / Semantic taxonomy.
Durations are approximate, without adding together alternate modalities or mixture stems. Text refers to transcripts, captions or instructions; label-only metadata are described in Annotation.
These corpora combine speech, music, and general sounds.
Open-source tools for constructing editing data and annotating existing recordings. Supported Task Type follows our Acoustic / Semantic / Instance taxonomy and indicates the editing supervision that each tool can help construct. Unified covers tools applicable across speech, music and general audio.
Synthesis, source separation, mixing and signal processing for constructing audio examples and source–target pairs.
Tools for extracting or creating content, attribute and temporal annotations from existing audio.
Public evaluation resources for audio editing. Editing categories follow this survey's taxonomy: Acoustic / Semantic / Instance / Composite, where Composite combines different categories within one request. Evaluation method describes the scorer: Expert models, MLLM, or Hybrid (including agent-based evaluation).
Metrics are grouped by the four evaluation dimensions used in this survey. ↑ / ↓ indicate higher / lower is better. Reference / Inputs lists the information needed alongside the edited output.
For local edits, compare the regions or sources that should remain unchanged.
Reusable models and toolkits for multi-dimensional assessment of editing results and audio aesthetics.
Foundation-model-based audio editing still faces several system-level challenges:
If You find this survey or repository useful, please cite our paper:
@article{pan2026audio,
title={Audio Editing in the Era of Foundation Models: A Survey},
author={Pan, Changhao and Fan, Yifei and Zhuo, Fan and Chen, Yifu and Guo, Wenxiang and Zhang, Yu and Li, Ruiqi and Zhu, Zhiyuan and Yang, Rui and Ji, Shengpeng and others},
journal={arXiv preprint arXiv:2606.23139},
year={2026}
}
This repo is meant to keep growing. If an audio editing model, dataset, or benchmark is missing, please feel free to open an issue or a pull request.
Unless otherwise noted below, original content created for this repository is licensed under the MIT License.
The survey paper and content reproduced or adapted from it, including assets/taxonomy_overview.png, assets/train-based.png, and assets/train-free.png, remain under CC BY-NC-SA 4.0. The MIT license does not relicense these materials.
Linked third-party papers, code, models, model weights, datasets, and tools are governed by their respective licenses.
[AACL-IJCNLP] A survey of foundation-model-based audio editing across speech, music, and general audio, covering task taxonomies, training-based and training-free methods, datasets, and evaluation.
See the codeAACL-IJCNLP 2026
Changhao Pan1,*, Yifei Fan1,*, Fan Zhuo1,*, Yifu Chen1, Wenxiang Guo1,
Yu Zhang2, Ruiqi Li2, Zhiyuan Zhu1, Rui Yang1, Shengpeng Ji3,
Chenyuhao Wen1, Jiayang Xu1, Ke Lei1, Xiaoda Yang1, Jingyu Lu1, Zhou Zhao1,†
1Zhejiang University · 2ByteDance · 3Hunyuan Team, Tencent
*Equal contribution · †Corresponding author
This repository is the official repository for Audio Editing in the Era of Foundation Models: A Survey, Which is accepted by AACL-IJCNLP 2026.
MM-Speech/AudioEditSurvey for better management.This is the official repository for Audio Editing in the Era of Foundation Models: A Survey, accepted to AACL-IJCNLP 2026. It is maintained by MM-Speech and collects papers and resources for foundation-model-based audio editing.
Abstract
Audio editing aims to modify a given synthetic or real-world audio signal to meet users' specific needs. As a promising yet challenging direction in AIGC, it has attracted increasing attention in recent years. With the rapid progress of text-to-audio and text-to-speech generation, powerful audio generation models have become the primary foundation for modern audio editing systems. In this survey, we provide a comprehensive review of foundation-model-based audio editing. We first define the scope of audio editing from a unified perspective and present a detailed taxonomy of existing editing tasks. We then summarize the major foundation-model paradigms for audio editing, and review representative approaches from both training-based and training-free perspectives. In addition, we systematically discuss related resources, including datasets, data construction tools, and evaluation protocols. Finally, we identify open challenges in this field and outline promising directions for future research.
In this survey, we focus on works that make direct contributions to audio editing in the era of foundation models. To ensure a precise and focused discussion, we adopt two main inclusion criteria: (1) the task should center on audio editing, which we define as modifying the acoustic attributes, instances, or content of an existing audio recording, without transformations so substantial that they amount to generating an entirely new audio sample; (2) the method should rely on mainstream audio foundation model paradigms. Accordingly, we do not cover works primarily focused on audio generation, nor do we provide an extensive discussion of signal-processing-based audio editing methods. In addition, to maintain a focused scope, spatial audio (multi-channel formats such as binaural stereo and first-order Ambisonics (FOA)) and related editing techniques are beyond the main scope of this survey.

Figure 1: Taxonomy of audio editing tasks.
| Category | Definition | Representative Editing Goals |
|---|---|---|
| Acoustic Editing | Modifies low-level perceptual attributes while preserving the overall structure and source characteristics of the original audio. | restoration, reverberation editing, loudness/mixing control, equalization, spectral texture editing |
| Semantic Editing | Modifies high-level interpretable information conveyed by audio while maintaining task-irrelevant properties. | linguistic editing, expressive editing, stylistic editing |
| Instance Editing | Manipulates identifiable audio entities while preserving the remaining scene and source relationships. | replacement, deletion/extraction, insertion, overlay/remixing |
Representative editors with publicly released implementations and model weights, subject to each project’s license. Editing types follow our taxonomy; Unified groups editors supporting multiple audio domains, with their supported domains listed in the table. Base links the pretrained backbone used by an editor, while Adapter links its additional learned weights.
| Model | Audio Domain | Editing Types | Model Architecture | Paper | Code | Model |
|---|---|---|---|---|---|---|
| Audio-Omni | Speech; Music; Audio | Instance: addition, removal, extraction, source transformation | MLLM + rectified-flow DiT | 🤗 Weights | ||
| AudioMorphix | Speech; Music; Audio | Semantic: pitch / time stretching Instance: addition, removal, replacement, time shifting | Diffusion U-Net (Tango / AudioLDM) | 🤗 Base (Tango 2) 🤗 Base (AudioLDM) | ||
| AuK / AuK-Flash | Speech; Music | Acoustic: restoration, loudness Semantic: words, lyrics, expression Instance: timbre, source extraction | MLLM + rectified-flow DiT | 🤗 AuK 🤗 Flash | ||
| Vevo2 | Speech; Music | Semantic: content, lyrics, prosody, style Instance: voice / singer conversion | Codec LM + flow-matching decoder | 🤗 Weights | ||
| DirectAudioEdit | Music; Audio | Instance: text-guided event replacement / addition / removal | Diffusion U-Net (Tango 2 / AudioLDM2) | 🤗 Base (Tango 2) 🤗 Base (audio) | ||
| DDPM Inversion (ZETA) | Music; Audio | Semantic: musical style Instance: instrument / sound-event changes | Diffusion U-Net (AudioLDM2) | 🤗 Base (audio) 🤗 Base (music) |
| Model | Editing Types | Model Architecture | Paper | Code | Model |
|---|---|---|---|---|---|
| Ming-UniAudio-Edit | Acoustic: denoising, loudness Semantic: content, prosody, emotion, dialect | Continuous-token LM + diffusion head | 🤗 Weights | ||
| Step-Audio-EditX | Semantic: emotion, speaking style, paralinguistics, pronunciation | Codec LM + flow-matching decoder | 🤗 Weights | ||
| CosyEdit | Semantic: word insertion, deletion, replacement | Codec LM + flow-matching decoder | 🤗 Weights | ||
| VoiceCraft-X | Semantic: multilingual content editing | Codec LM (autoregressive infilling) | 🤗 Weights | ||
| VoiceCraft | Semantic: word insertion, deletion, replacement | Codec LM (autoregressive infilling) | 🤗 Weights | ||
| SSR-Speech | Semantic: word insertion, deletion, replacement | Codec LM (autoregressive infilling) | 🤗 English 🤗 Mandarin | ||
| F5-TTS | Semantic: local content replacement / infilling | Flow-matching DiT | 🤗 Weights | ||
| FluentSpeech | Semantic: content editing, disfluency correction | Diffusion (context-aware denoiser) | 📁 Weights | ||
| EdiTTS | Semantic: content / pitch edits in synthesized speech | Score-based diffusion (Grad-TTS) | 📦 Base |
| Model | Editing Types | Model Architecture | Paper | Code | Model |
|---|---|---|---|---|---|
| YingMusic-Singer-Plus | Semantic: lyrics Instance: singer timbre replacement | Flow-matching DiT | 🤗 Weights | ||
| ACE-Step 1.5 | Semantic: style / local repainting Instance: track extraction / addition (base variant) | LM + flow-matching DiT | 🤗 Turbo 🤗 Base variant | ||
| Instruct-MusicGen | Instance: stem addition, removal, extraction | Codec LM (MusicGen) + adapters | 🤗 Public-data retraining | ||
| MusicGen-Stem | Instance: stem replacement / addition (bass, drums, other) | Multi-stream codec LM | 🤗 Weights | ||
| MelodyFlow | Semantic: genre, mood, style Instance: instrumentation | Flow-matching DiT | 🤗 Weights | ||
| AP-Adapter | Semantic: genre / style transfer Instance: instrument replacement | Diffusion U-Net + audio-prompt adapter | 📁 Adapter 🤗 Base | ||
| AnchorSteer | Semantic: genre / style Instance: instrument changes | Diffusion DiT + structural/concept adapters | 🤗 Concept weights 📁 Structure adapter 🤗 Base (access terms) |
| Model | Editing Types | Model Architecture | Paper | Code | Model |
|---|---|---|---|---|---|
| MMEdit | Acoustic: loudness Instance: event addition, removal, replacement, reordering | ALM + diffusion MMDiT | 🤗 Weights | ||
| SAO-Instruct | Acoustic: filtering, denoising, restoration Semantic: pitch / rate Instance: event manipulation | Diffusion DiT (Stable Audio Open) | 🤗 Weights | ||
| SmartDJ-Editor | Acoustic: volume, reverb, spectral coloration Instance: event addition, removal, extraction, relocation | Diffusion Transformer (U-DiT) | 🤗 Editor weights | ||
| AudioEditor | Instance: event addition, deletion, replacement | Diffusion U-Net (Auffusion) | 🤗 Base | ||
| CoherentAVEdit | Instance: video-conditioned sound-event replacement | Flow-matching Transformer (MMAudio) | 🤗 Weights |
Before the foundation-model era, early neural audio editing methods mainly explored task-specific generative models for local reconstruction and attribute control.
Token-based codec language models cast audio editing as conditional generation over discrete audio tokens. After continuous audio is converted into compact discrete token sequences, target regions are edited through autoregressive continuation, infilling, or selective regeneration conditioned on context, prompts, or task controls.
Diffusion and flow-matching models formulate audio editing as conditional transformation in continuous acoustic spaces, such as mel-spectrograms or audio latents. Instead of infilling discrete tokens, they modify audio through conditional denoising, latent inversion, or continuous flow transformation, making them suitable for high-fidelity reconstruction, region-level refinement, and fine-grained acoustic control in complex scenarios.
Instruction-conditioned and multimodal interfaces for audio editing provide high-level control for foundation-model-based audio editing. They allow users to specify editing intents through natural language instructions, task prompts, reference audio, temporal regions, or visual cues, which are shifted into target spans, task embeddings, event locations, speaker references, or preservation constraints.
Training-based approaches refer to audio editing methods that learn editing behaviors from supervised pairs, pseudo-pairs, or instruction-based triplets before inference. These methods explicitly optimize editing objectives, condition following, and preservation constraints, enabling stable and controllable editing. We group existing works into three categories based on their supervision and conditioning mechanisms, and discuss their core methods and functional scopes.
Figure 2: Overview of training-based audio editing methods.
| Paradigm | Description | Representative Scope |
|---|---|---|
| Task-specific Training | Optimizes models for predefined editing functions or domains. | text-based speech editing, prosody correction, source separation, music stem separation |
| Reference- and Attribute-based Training | Specifies the editing direction through reference audio, style examples, or attribute labels. | voice conversion, timbre transfer, emotion editing, mixing style transfer |
| Instruction-conditioned Training | Learns from instruction-input-output triplets to follow natural-language editing requests. | addition, deletion, replacement, inpainting, super-resolution, music remixing, expressive refinement |
Training-free approaches adapt pretrained audio generative models to editing without parameter updates. They operate by manipulating inference-time mechanisms, such as inversion, attention control, prompt or guidance adjustment, and mask-based constraints. We group existing methods into three common categories, which are often combined to improve localization, preservation, and controllability. Since token-based autoregressive models are less naturally suited to training-free editing, this section mainly focuses on non-autoregressive paradigms, especially diffusion-based foundation models.
Figure 3: Overview of training-free audio editing methods.
| Paradigm | Description | Representative Scope |
|---|---|---|
| Inversion-Based Editing | Maps source audio back into the latent, noise, or trajectory space of a pretrained generative model, then edits it by modifying conditions or sampling trajectories. | DDPM/DDIM inversion, latent inversion, flow-based inversion, speech or music reconstruction and editing |
| Attention-Controlled Editing | Guides pretrained generative models by modifying or reusing internal attention patterns without parameter updates. | cross-attention event localization, self-attention preservation, prompt-level manipulation |
| Mask- and Region-Guided Editing | Specifies where to edit and where to preserve the source audio in waveform, spectrogram, latent, or source-component spaces. | localized editing, inpainting, restoration, source-level manipulation |
| Token-Level Editing with Codec Models | Manipulates discrete audio tokens through masking, infilling, continuation, or selective regeneration at inference time. | speech infilling, localized resynthesis, codec-token editing |
Public datasets for audio editing and controllable audio generation, grouped by their primary audio domain.
This non-exhaustive list highlights datasets suited to audio editing or widely used in the community, with availability verified by the repository maintainers for every entry.
Paired indicates released source–target audio, mixture–stem correspondence, or explicitly matched control/technique takes (✅ / ❌); shared transcripts, audio–text alignment, or audio–MIDI alignment alone do not count. † marks an editing use that requires task construction or adaptation, rather than native editing supervision. Editing types follow our Acoustic / Instance / Semantic taxonomy.
Durations are approximate, without adding together alternate modalities or mixture stems. Text refers to transcripts, captions or instructions; label-only metadata are described in Annotation.
These corpora combine speech, music, and general sounds.
Open-source tools for constructing editing data and annotating existing recordings. Supported Task Type follows our Acoustic / Semantic / Instance taxonomy and indicates the editing supervision that each tool can help construct. Unified covers tools applicable across speech, music and general audio.
Synthesis, source separation, mixing and signal processing for constructing audio examples and source–target pairs.
Tools for extracting or creating content, attribute and temporal annotations from existing audio.
Public evaluation resources for audio editing. Editing categories follow this survey's taxonomy: Acoustic / Semantic / Instance / Composite, where Composite combines different categories within one request. Evaluation method describes the scorer: Expert models, MLLM, or Hybrid (including agent-based evaluation).
Metrics are grouped by the four evaluation dimensions used in this survey. ↑ / ↓ indicate higher / lower is better. Reference / Inputs lists the information needed alongside the edited output.
For local edits, compare the regions or sources that should remain unchanged.
Reusable models and toolkits for multi-dimensional assessment of editing results and audio aesthetics.
Foundation-model-based audio editing still faces several system-level challenges:
If You find this survey or repository useful, please cite our paper:
@article{pan2026audio,
title={Audio Editing in the Era of Foundation Models: A Survey},
author={Pan, Changhao and Fan, Yifei and Zhuo, Fan and Chen, Yifu and Guo, Wenxiang and Zhang, Yu and Li, Ruiqi and Zhu, Zhiyuan and Yang, Rui and Ji, Shengpeng and others},
journal={arXiv preprint arXiv:2606.23139},
year={2026}
}
This repo is meant to keep growing. If an audio editing model, dataset, or benchmark is missing, please feel free to open an issue or a pull request.
Unless otherwise noted below, original content created for this repository is licensed under the MIT License.
The survey paper and content reproduced or adapted from it, including assets/taxonomy_overview.png, assets/train-based.png, and assets/train-free.png, remain under CC BY-NC-SA 4.0. The MIT license does not relicense these materials.
Linked third-party papers, code, models, model weights, datasets, and tools are governed by their respective licenses.