MM-Speech/AudioEditSurvey

[AACL-IJCNLP] A survey of foundation-model-based audio editing across speech, music, and general audio, covering task taxonomies, training-based and training-free methods, datasets, and evaluation.

32

11 commits

updated Sep 29, 2026

See the code

README

Awesome Audio Editing

Audio Editing in the Era of Foundation Models: A Survey

AACL-IJCNLP 2026

Changhao Pan1,*, Yifei Fan1,*, Fan Zhuo1,*, Yifu Chen1, Wenxiang Guo1,
Yu Zhang2, Ruiqi Li2, Zhiyuan Zhu1, Rui Yang1, Shengpeng Ji3,
Chenyuhao Wen1, Jiayang Xu1, Ke Lei1, Xiaoda Yang1, Jingyu Lu1, Zhou Zhao1,†

1Zhejiang University  ·  2ByteDance  ·  3Hunyuan Team, Tencent

*Equal contribution  ·  †Corresponding author

arXiv AACL-IJCNLP 2026 Project Page GitHub stars

🌐 Languages

English · 简体中文 · 한국어

🚀 Quick Start

This repository is the official repository for Audio Editing in the Era of Foundation Models: A Survey, Which is accepted by AACL-IJCNLP 2026.

  • We establish a unified taxonomy of acoustic, semantic, and instance editing across speech, music, and general audio, clarifying what each task changes and what it should preserve to support consistent comparisons across editing goals.
  • We review mainstream audio editing techniques through foundation-model architectures and learning paradigms, with an emphasis on their core mechanisms and suitability for different editing scenarios.
  • (Updated Recently) We curate publicly available audio editing models, and summarize their supported task categories and key strengths to help readers identify suitable models.
  • (Updated Recently) We organize publicly available datasets, data construction tools, evaluation benchmarks, and metrics for audio editing, summarizing the audio domains, editing categories, and evaluation dimensions they cover.

🔥What's new

  • 📦 [2026/09] This repository has moved to MM-Speech/AudioEditSurvey for better management.
  • 🏆 [2026/09] Our paper has been accepted to the AACL-IJCNLP 2026!
  • 🎉 [2026/06] We have officially released this survey repository for Audio Editing Models, with the preprint available on arXiv.

Contents

  1. Introduction
  2. Scope
  3. Overall
  4. Foundation Models for Audio Editing
  5. Training-based Audio Editing
  6. Training-free Audio Editing
  7. Resources
  8. Challenges and Future Directions
  9. Citation
  10. Contributing

📌 Introduction

This is the official repository for Audio Editing in the Era of Foundation Models: A Survey, accepted to AACL-IJCNLP 2026. It is maintained by MM-Speech and collects papers and resources for foundation-model-based audio editing.

Abstract
Audio editing aims to modify a given synthetic or real-world audio signal to meet users' specific needs. As a promising yet challenging direction in AIGC, it has attracted increasing attention in recent years. With the rapid progress of text-to-audio and text-to-speech generation, powerful audio generation models have become the primary foundation for modern audio editing systems. In this survey, we provide a comprehensive review of foundation-model-based audio editing. We first define the scope of audio editing from a unified perspective and present a detailed taxonomy of existing editing tasks. We then summarize the major foundation-model paradigms for audio editing, and review representative approaches from both training-based and training-free perspectives. In addition, we systematically discuss related resources, including datasets, data construction tools, and evaluation protocols. Finally, we identify open challenges in this field and outline promising directions for future research.


🎯 Scope

In this survey, we focus on works that make direct contributions to audio editing in the era of foundation models. To ensure a precise and focused discussion, we adopt two main inclusion criteria: (1) the task should center on audio editing, which we define as modifying the acoustic attributes, instances, or content of an existing audio recording, without transformations so substantial that they amount to generating an entirely new audio sample; (2) the method should rely on mainstream audio foundation model paradigms. Accordingly, we do not cover works primarily focused on audio generation, nor do we provide an extensive discussion of signal-processing-based audio editing methods. In addition, to maintain a focused scope, spatial audio (multi-channel formats such as binaural stereo and first-order Ambisonics (FOA)) and related editing techniques are beyond the main scope of this survey.


🧭 Overall

🗂️ Taxonomy Overview

Taxonomy of Audio Editing Tasks

Figure 1: Taxonomy of audio editing tasks.

🧩 Taxonomy Details

CategoryDefinitionRepresentative Editing Goals
Acoustic EditingModifies low-level perceptual attributes while preserving the overall structure and source characteristics of the original audio.restoration, reverberation editing, loudness/mixing control, equalization, spectral texture editing
Semantic EditingModifies high-level interpretable information conveyed by audio while maintaining task-irrelevant properties.linguistic editing, expressive editing, stylistic editing
Instance EditingManipulates identifiable audio entities while preserving the remaining scene and source relationships.replacement, deletion/extraction, insertion, overlay/remixing

📚 Representative Audio Editing Methods

Representative editors with publicly released implementations and model weights, subject to each project’s license. Editing types follow our taxonomy; Unified groups editors supporting multiple audio domains, with their supported domains listed in the table. Base links the pretrained backbone used by an editor, while Adapter links its additional learned weights.

Unified Models

ModelAudio DomainEditing TypesModel ArchitecturePaperCodeModel
Audio-OmniSpeech; Music; AudioInstance: addition, removal, extraction, source transformationMLLM + rectified-flow DiTarXiv PaperGitHub Code🤗 Weights
AudioMorphixSpeech; Music; AudioSemantic: pitch / time stretching
Instance: addition, removal, replacement, time shifting
Diffusion U-Net (Tango / AudioLDM)arXiv PaperHuggingFace Code🤗 Base (Tango 2)
🤗 Base (AudioLDM)
AuK / AuK-FlashSpeech; MusicAcoustic: restoration, loudness
Semantic: words, lyrics, expression
Instance: timbre, source extraction
MLLM + rectified-flow DiTarXiv PaperGitHub Code🤗 AuK
🤗 Flash
Vevo2Speech; MusicSemantic: content, lyrics, prosody, style
Instance: voice / singer conversion
Codec LM + flow-matching decoderarXiv PaperGitHub Code🤗 Weights
DirectAudioEditMusic; AudioInstance: text-guided event replacement / addition / removalDiffusion U-Net (Tango 2 / AudioLDM2)arXiv PaperGitHub Code🤗 Base (Tango 2)
🤗 Base (audio)
DDPM Inversion (ZETA)Music; AudioSemantic: musical style
Instance: instrument / sound-event changes
Diffusion U-Net (AudioLDM2)arXiv PaperGitHub Code🤗 Base (audio)
🤗 Base (music)

Speech Models

ModelEditing TypesModel ArchitecturePaperCodeModel
Ming-UniAudio-EditAcoustic: denoising, loudness
Semantic: content, prosody, emotion, dialect
Continuous-token LM + diffusion headarXiv PaperGitHub Code🤗 Weights
Step-Audio-EditXSemantic: emotion, speaking style, paralinguistics, pronunciationCodec LM + flow-matching decoderarXiv PaperGitHub Code🤗 Weights
CosyEditSemantic: word insertion, deletion, replacementCodec LM + flow-matching decoderarXiv PaperGitHub Code🤗 Weights
VoiceCraft-XSemantic: multilingual content editingCodec LM (autoregressive infilling)arXiv PaperGitHub Code🤗 Weights
VoiceCraftSemantic: word insertion, deletion, replacementCodec LM (autoregressive infilling)arXiv PaperGitHub Code🤗 Weights
SSR-SpeechSemantic: word insertion, deletion, replacementCodec LM (autoregressive infilling)arXiv PaperGitHub Code🤗 English
🤗 Mandarin
F5-TTSSemantic: local content replacement / infillingFlow-matching DiTarXiv PaperGitHub Code🤗 Weights
FluentSpeechSemantic: content editing, disfluency correctionDiffusion (context-aware denoiser)arXiv PaperGitHub Code📁 Weights
EdiTTSSemantic: content / pitch edits in synthesized speechScore-based diffusion (Grad-TTS)arXiv PaperGitHub Code📦 Base

Music Models

ModelEditing TypesModel ArchitecturePaperCodeModel
YingMusic-Singer-PlusSemantic: lyrics
Instance: singer timbre replacement
Flow-matching DiTarXiv PaperGitHub Code🤗 Weights
ACE-Step 1.5Semantic: style / local repainting
Instance: track extraction / addition (base variant)
LM + flow-matching DiTarXiv PaperGitHub Code🤗 Turbo
🤗 Base variant
Instruct-MusicGenInstance: stem addition, removal, extractionCodec LM (MusicGen) + adaptersarXiv PaperGitHub Code🤗 Public-data retraining
MusicGen-StemInstance: stem replacement / addition (bass, drums, other)Multi-stream codec LMarXiv PaperGitHub Code🤗 Weights
MelodyFlowSemantic: genre, mood, style
Instance: instrumentation
Flow-matching DiTarXiv PaperHuggingFace Code🤗 Weights
AP-AdapterSemantic: genre / style transfer
Instance: instrument replacement
Diffusion U-Net + audio-prompt adapterarXiv PaperGitHub Code📁 Adapter
🤗 Base
AnchorSteerSemantic: genre / style
Instance: instrument changes
Diffusion DiT + structural/concept adaptersarXiv PaperGitHub Code🤗 Concept weights
📁 Structure adapter
🤗 Base (access terms)

Audio Models

ModelEditing TypesModel ArchitecturePaperCodeModel
MMEditAcoustic: loudness
Instance: event addition, removal, replacement, reordering
ALM + diffusion MMDiTarXiv PaperGitHub Code🤗 Weights
SAO-InstructAcoustic: filtering, denoising, restoration
Semantic: pitch / rate
Instance: event manipulation
Diffusion DiT (Stable Audio Open)arXiv PaperGitHub Code🤗 Weights
SmartDJ-EditorAcoustic: volume, reverb, spectral coloration
Instance: event addition, removal, extraction, relocation
Diffusion Transformer (U-DiT)arXiv PaperGitHub Code🤗 Editor weights
AudioEditorInstance: event addition, deletion, replacementDiffusion U-Net (Auffusion)arXiv PaperGitHub Code🤗 Base
CoherentAVEditInstance: video-conditioned sound-event replacementFlow-matching Transformer (MMAudio)arXiv PaperGitHub Code🤗 Weights

🏗️ Foundation Models for Audio Editing

1. Early Neural Editing Models

Before the foundation-model era, early neural audio editing methods mainly explored task-specific generative models for local reconstruction and attribute control.

2. Token-based Codec Language Models

Token-based codec language models cast audio editing as conditional generation over discrete audio tokens. After continuous audio is converted into compact discrete token sequences, target regions are edited through autoregressive continuation, infilling, or selective regeneration conditioned on context, prompts, or task controls.

3. Diffusion and Flow-Matching Models

Diffusion and flow-matching models formulate audio editing as conditional transformation in continuous acoustic spaces, such as mel-spectrograms or audio latents. Instead of infilling discrete tokens, they modify audio through conditional denoising, latent inversion, or continuous flow transformation, making them suitable for high-fidelity reconstruction, region-level refinement, and fine-grained acoustic control in complex scenarios.

4. Audio Editing Interfaces

Instruction-conditioned and multimodal interfaces for audio editing provide high-level control for foundation-model-based audio editing. They allow users to specify editing intents through natural language instructions, task prompts, reference audio, temporal regions, or visual cues, which are shifted into target spans, task embeddings, event locations, speaker references, or preservation constraints.


🧪 Training-based Audio Editing

Training-based approaches refer to audio editing methods that learn editing behaviors from supervised pairs, pseudo-pairs, or instruction-based triplets before inference. These methods explicitly optimize editing objectives, condition following, and preservation constraints, enabling stable and controllable editing. We group existing works into three categories based on their supervision and conditioning mechanisms, and discuss their core methods and functional scopes.

Overview of training-based audio editing methods

Figure 2: Overview of training-based audio editing methods.

ParadigmDescriptionRepresentative Scope
Task-specific TrainingOptimizes models for predefined editing functions or domains.text-based speech editing, prosody correction, source separation, music stem separation
Reference- and Attribute-based TrainingSpecifies the editing direction through reference audio, style examples, or attribute labels.voice conversion, timbre transfer, emotion editing, mixing style transfer
Instruction-conditioned TrainingLearns from instruction-input-output triplets to follow natural-language editing requests.addition, deletion, replacement, inpainting, super-resolution, music remixing, expressive refinement

🪄 Training-free Audio Editing

Training-free approaches adapt pretrained audio generative models to editing without parameter updates. They operate by manipulating inference-time mechanisms, such as inversion, attention control, prompt or guidance adjustment, and mask-based constraints. We group existing methods into three common categories, which are often combined to improve localization, preservation, and controllability. Since token-based autoregressive models are less naturally suited to training-free editing, this section mainly focuses on non-autoregressive paradigms, especially diffusion-based foundation models.

Overview of training-free audio editing methods

Figure 3: Overview of training-free audio editing methods.

ParadigmDescriptionRepresentative Scope
Inversion-Based EditingMaps source audio back into the latent, noise, or trajectory space of a pretrained generative model, then edits it by modifying conditions or sampling trajectories.DDPM/DDIM inversion, latent inversion, flow-based inversion, speech or music reconstruction and editing
Attention-Controlled EditingGuides pretrained generative models by modifying or reusing internal attention patterns without parameter updates.cross-attention event localization, self-attention preservation, prompt-level manipulation
Mask- and Region-Guided EditingSpecifies where to edit and where to preserve the source audio in waveform, spectrogram, latent, or source-component spaces.localized editing, inpainting, restoration, source-level manipulation
Token-Level Editing with Codec ModelsManipulates discrete audio tokens through masking, infilling, continuation, or selective regeneration at inference time.speech infilling, localized resynthesis, codec-token editing

📦 Resources

📊 Available Datasets

Public datasets for audio editing and controllable audio generation, grouped by their primary audio domain.

This non-exhaustive list highlights datasets suited to audio editing or widely used in the community, with availability verified by the repository maintainers for every entry.

Paired indicates released source–target audio, mixture–stem correspondence, or explicitly matched control/technique takes (✅ / ❌); shared transcripts, audio–text alignment, or audio–MIDI alignment alone do not count. † marks an editing use that requires task construction or adaptation, rather than native editing supervision. Editing types follow our Acoustic / Instance / Semantic taxonomy.

Durations are approximate, without adding together alternate modalities or mixture stems. Text refers to transcripts, captions or instructions; label-only metadata are described in Annotation.

Speech

NamePaperDataset / CodeDurationPairedEditing TypesAnnotationModalities
VoiceBank+DEMAND (28-spk)Paper LinkDataShare Data≈10 h✅ Noisy/cleanAcousticTranscript; noise/SNR conditionsAudio, Text
LibriTTS-RarXiv PaperOpenSLR Restored
OpenSLR Original
≈585 h✅ Original/restoredAcoustic; Semantic†Transcript; speaker labels; model-restored audioAudio, Text
LibriSpeechPaper LinkOpenSLR Data≈1,000 h❌Semantic†; Instance†Transcript; speaker/chapter labelsAudio, Text
VCTK v0.92Dataset RecordDataShare Data≈44 h❌Instance†; Semantic†Transcript; speaker/accent labelsAudio, Text
AISHELL-3arXiv PaperOpenSLR Data≈85 h❌Semantic†; Instance†Mandarin transcript; phonetic transcription; speaker labelsAudio, Text
Hi-Fi TTSarXiv PaperOpenSLR Data≈292 h❌Semantic†; Instance†Transcript; speaker labelsAudio, Text
LJSpeech v1.1Dataset ReleaseDownload Data≈24 h❌Semantic†Transcript; normalized textAudio, Text
RAVDESS (speech)Paper LinkZenodo Data≈1.7 h❌Semantic†Label: emotion, intensity, speaker; fixed transcriptsAudio, Text, Video
CREMA-DPaper LinkGitHub Code
GitLab Mirror
≈5.3 h❌Semantic†Label: emotion/intensity; perceptual ratings; fixed transcriptsAudio, Text, Video

Music

NamePaperDataset / CodeDurationPairedEditing TypesAnnotationModalities
GTSingerarXiv PaperGitHub Code
HuggingFace Dataset
Google Drive Data
≈80.6 h singing
+16.2 h speech
✅ Controlled/parallel takesSemantic; Instance†Label: technique/style; aligned lyrics/phonemes; scoresAudio, Text, MusicXML
Slakh2100arXiv PaperGitHub Code
Zenodo Data
≈145 h✅ Mixture/stemsInstanceLabel: instrument; aligned MIDI; stem metadataAudio, MIDI
MUSDB18-HQarXiv PaperGitHub Code
Zenodo Data
≈10 h✅ Mixture/stemsInstanceLabel: vocals, drums, bass, otherAudio
MAESTRO v3arXiv PaperProject Page
Download Data
≈199 h❌Semantic†Aligned MIDI: pitch, timing, velocity, pedals; piece metadataAudio, MIDI
NSyntharXiv PaperProject Page≈340 h❌Instance†; Semantic†Label: instrument, pitch, velocity, timbral qualitiesAudio
Groove MIDI DatasetarXiv PaperProject Page
Download Data
≈13.6 h❌Semantic†Aligned MIDI; tempo/style labels; performance timing/velocityAudio, MIDI
MusicCapsarXiv PaperHuggingFace Metadata≈15.3 h❌Semantic†; Instance†Caption; musical aspect labelsAudio, Text
MTG-JamendoPublication RecordGitHub Code
Download Data
≈3,770 h❌Semantic†; Instance†Label: genre, instrument, mood/themeAudio
FMA (large)arXiv PaperGitHub Code
Download Data
≈888 h❌Semantic†Label: genre hierarchy; track/artist metadataAudio

Audio

NamePaperDataset / CodeDurationPairedEditing TypesAnnotationModalities
FUSSarXiv PaperGitHub Code
Zenodo Data
≈61 h mixtures✅ Mixture/sources; dry/reverberantInstance; AcousticSource/time metadata; mixing parameters; no event labelsAudio
AudioSetPaper LinkDataset Metadata≈5,790 h❌Instance†Label: sound-event ontology; clip-level multi-labelsAudio, Video (upstream)
AudioCaps v1Paper LinkGitHub Metadata≈143 h❌Instance†; Semantic†Caption: one or five descriptions per clipAudio, Text
Clotho v2.1arXiv PaperZenodo Data≈37 h
(5,929 labeled clips)
❌Instance†; Semantic†Caption: five per clip; Freesound keywordsAudio, Text
WavCapsarXiv PaperGitHub Code
HuggingFace Dataset
≈7,568 h❌Instance†; Semantic†LLM-assisted captions; source descriptions/metadataAudio, Text
FSD50KarXiv PaperZenodo Data≈108 h❌Instance†Label: 200 sound-event classes; clip-level multi-labelsAudio
ESC-50Paper LinkGitHub Code≈2.8 h❌Instance†Label: 50 environmental sound classesAudio
UrbanSound8KPaper LinkProject Page
Zenodo Data
≈8.8 h❌Instance†Label: 10 urban sound classes; salience; source timestampsAudio
VGGSoundarXiv PaperGitHub Metadata≈550 h❌Instance†Label: audio-visual event class; video timestampsAudio, Video (upstream)

Unified

These corpora combine speech, music, and general sounds.

NamePaperDataset / CodeDurationPairedEditing TypesAnnotationModalities
AudioEdit (Audio-Omni)arXiv PaperGitHub Code
HuggingFace Dataset
≈2,686 h
(966,794 task pairs)
✅ Source/edited targetInstanceInstruct: add, remove, extract, source transformationAudio, Text
Divide and Remaster v2arXiv PaperGitHub Code
Zenodo Data
≈81 h✅ Mixture/stemsInstanceTranscript; music genre; sound labels/timestampsAudio, Text
MUSANarXiv PaperOpenSLR Data≈109 h❌Acoustic†; Instance†Label: speech/music/noise; speech and music metadataAudio

🛠️ Data Tools

Open-source tools for constructing editing data and annotating existing recordings. Supported Task Type follows our Acoustic / Semantic / Instance taxonomy and indicates the editing supervision that each tool can help construct. Unified covers tools applicable across speech, music and general audio.

Tools for Data Generation

Synthesis, source separation, mixing and signal processing for constructing audio examples and source–target pairs.

Speech
ToolSupported Task TypeWhat It ConstructsControl LevelCodeModel
Qwen3-TTSSemantic; InstanceText-aligned utterances with instruction-controlled delivery or a reference speaker.UtteranceGitHub CodeHugging Face Base
Hugging Face CustomVoice
CosyVoice3Semantic; InstanceText-aligned speech with voice cloning and prompted language, emotion or delivery.Utterance; pronunciation unitsGitHub CodeHugging Face Model
MaskGCTSemantic; InstanceText-conditioned speech with a reference voice and configurable total duration.Utterance; total durationGitHub CodeHugging Face Model
Seed-VCInstanceVoice-converted recordings paired with their source speech for speaker/timbre replacement.Utterance / source recordingGitHub CodeHugging Face Model
AuKAcoustic; Semantic; InstanceInstruction-edited speech for content, delivery, voice, enhancement and target-speaker tasks.Utterance; text-specified word / phraseGitHub CodeHugging Face Model
Music
ToolSupported Task TypeWhat It ConstructsControl LevelCodeModel
MusicGenSemanticText- or melody-conditioned music clips and continuations for style/content-controlled examples.Clip; melody sequenceGitHub CodeHugging Face Melody
DemucsInstanceEstimated vocal, drum, bass and other stems for extraction, removal and remix pair construction.Stem / trackGitHub CodeModel Checkpoint
SpleeterInstanceEstimated 2-, 4- or 5-stem decompositions for source removal, extraction and remixing.Stem / trackGitHub CodeGitHub Checkpoints
FluidSynthSemantic; InstanceAudio rendered from MIDI and a SoundFont, aligned with notes, velocities and instrument assignments.Note; MIDI control event / trackGitHub Code
Audio
ToolSupported Task TypeWhat It ConstructsControl LevelCodeModel
AudioLDM 2InstanceText-conditioned sound clips to use as source assets in insertion or replacement examples.ClipGitHub CodeHugging Face Model
AudioSepInstanceText-selected source estimates from mixtures for extraction and removal pair construction.Described source / clipGitHub CodeHugging Face Checkpoints
ScaperAcoustic; InstanceSynthetic soundscapes with event labels, onset/offset times, SNRs and optional isolated event tracks.Event; start time / duration / SNRGitHub Code
SpatialScaperAcoustic; InstanceSpatialized soundscapes with event activity, source trajectories and room-response conditions.Event / trajectory / sceneGitHub Code
Unified
ToolSupported Task TypeWhat It ConstructsControl LevelCodeModel
SAM-AudioInstancePrompt-selected target and residual audio for extraction, removal and remix examples.Source; temporal-span promptsGitHub CodeHugging Face Model
Access request
AudiomentationsAcoustic; SemanticAugmented audio for clean/degraded and pitch/tempo contrast pairs using noise, gain, filtering and other transforms.Clip; selected segment via slicingGitHub Code
PedalboardAcousticEffect-processed audio for dry/wet or clean/degraded pairs using EQ, gain, compression, distortion and reverb.Clip / processing blockGitHub Code
PyroomacousticsAcoustic; InstanceRoom impulse responses and microphone mixtures from positioned sources, including dry/reverberant pairs.Scene / source positionGitHub Code

Tools for Data Annotation

Tools for extracting or creating content, attribute and temporal annotations from existing audio.

Speech
ToolSupported Task TypeWhat It AnnotatesAnnotation LevelCodeModel
Montreal Forced Aligner (MFA)SemanticWord and phone boundaries obtained by aligning speech with supplied transcripts and pronunciation dictionaries.Word / phonemeGitHub CodeModel Acoustic models
WhisperXSemanticASR transcripts with word timestamps from a language-specific alignment model.Utterance / wordGitHub CodeHugging Face ASR
Hugging Face EN aligner
Qwen3-ASR + ForcedAlignerSemanticTranscripts, language labels and text-unit timestamps using the released ASR and forced-alignment models.Utterance / wordGitHub CodeHugging Face ASR
Hugging Face Aligner
pyannote.audioInstanceSpeaker-labeled speech turns and overlapping-speaker activity.Speaker turn / segmentGitHub CodeHugging Face Community-1
Accept access terms
Silero VADInstanceSpeech/non-speech probabilities and detected speech start/end times.Frame / speech segmentGitHub CodeGitHub Weights
emotion2vec+SemanticSpeech-emotion labels and scores, with optional learned emotion representations.Utterance (labels); frame (features)GitHub CodeHugging Face Large
FunASR / SenseVoiceSmallSemantic; InstanceTranscripts, language and emotion tags, and audio-event tags such as laughter or applause.Utterance / VAD segmentGitHub CodeHugging Face SenseVoiceSmall
Praat / ParselmouthAcoustic; SemanticPitch, formants and intensity tracks; manually defined TextGrid points and intervals in Praat.Frame; word / phoneme / interval (manual)GitHub Praat
GitHub Parselmouth
Music
ToolSupported Task TypeWhat It AnnotatesAnnotation LevelCodeModel
RMVPESemanticVocal F0 trajectories from polyphonic music.FrameGitHub CodeGoogle Drive ROSVOT bundle
CREPESemanticMonophonic F0 estimates and confidence values.FrameGitHub CodeGitHub Weights
ROSVOTSemanticSinging-note pitches and onset/offset times, with word boundaries from its RWBD component.Note / wordGitHub CodeGoogle Drive Checkpoints
Basic PitchSemanticPolyphonic note events and pitch bends exported as MIDI; most effective on one instrument at a time.Note; frame-level pitch contourGitHub CodeGitHub Weights
All-In-One Music Structure AnalyzerSemanticTempo, beat/downbeat timestamps and labeled sections such as verse, chorus and bridge.Beat / downbeat / sectionGitHub CodeHugging Face Models
Music FlamingoSemantic; InstanceMusic captions and question–answer annotations about instrumentation, harmony, mood, structure and lyrics.Clip / full track (free-form text)GitHub CodeHugging Face Model
Audio
ToolSupported Task TypeWhat It AnnotatesAnnotation LevelCodeModel
PANNsInstanceSound-event class scores and frame-wise activity with the released decision-level detection models.Clip / frameGitHub CodeZenodo Models
HTS-ATInstanceSound-event tags and temporal class-activation estimates in localization mode.Clip / frameGitHub CodeGoogle Drive Models
YAMNetInstanceScores for 521 sound-event classes from overlapping audio windows.0.96 s window; 0.48 s hopGitHub CodeModel Checkpoint
Unified
ToolSupported Task TypeWhat It AnnotatesAnnotation LevelCodeModel
Qwen3-Omni CaptionerAcoustic; Semantic; InstanceDetailed audio captions covering speech, music, sound events and acoustic characteristics.Clip / recording (free-form text)GitHub CodeHugging Face Captioner
Audio Flamingo 3Semantic; InstancePrompted transcripts, captions, event descriptions and audio question–answer annotations.Clip / recording (free-form text)GitHub CodeHugging Face Model
Label StudioAcoustic; Semantic; InstanceHuman-authored clip labels, time-region labels and transcriptions using configurable audio templates.Clip / manually selected intervalGitHub Code

🧪 Benchmarks

Public evaluation resources for audio editing. Editing categories follow this survey's taxonomy: Acoustic / Semantic / Instance / Composite, where Composite combines different categories within one request. Evaluation method describes the scorer: Expert models, MLLM, or Hybrid (including agent-based evaluation).

MMAE

  • TL;DR: Tests instruction following and preservation across speech, music, sound, and their mixtures with 2,000 cases (≈8.0 h), six complexity levels, and 17,741 verification rubrics.
  • Paper: arXiv; Code: GitHub; Dataset: Hugging Face.
  • Audio modalities: Speech; Music; Audio.
  • Editing categories: Acoustic; Semantic; Instance; Composite.
  • Evaluation method: MLLM — Qwen3-Omni judges individual rubrics with majority voting, yielding Instruction Following Rate (IFR), Consistency Rate (CR), and Exact Match Rate (EMR).
  • Best reported results:
    • Single model: On the full benchmark in the original comparison, Step-Audio-EditX leads IFR (44.86%) and CR (58.88%), while Ming-UniAudio leads EMR (3.20%); Audio-Omni reaches 4.99% EMR on the separate 801-case, ≤10 s subset. The newer AuK baseline without Prompt Enhancer reports 7.58% EMR on the 1,003-case single-operation subset.
    • Agent / LLM-assisted system: The challenge agent baseline, combining an LLM router with DSP, SAM-Audio, and AuK, reports 7.45% EMR, 44.06% IFR, and 74.63% CR on all 2,000 cases. Separately, AuK-Flash with Prompt Enhancer reaches 13.85% EMR on MMAE-Speech only; these scopes are not directly comparable.

SpeechEditBench

  • TL;DR: Separates edit success from linguistic-content preservation in bilingual speech editing through 4,700 cases (≈9.4 h) covering seven atomic attributes and multi-attribute instructions.
  • Paper: arXiv; Code: GitHub; Dataset: Hugging Face, v1.1.
  • Audio modalities: Speech.
  • Editing categories: Acoustic; Semantic; Instance; Composite.
  • Evaluation method: Hybrid — ASR, speaker verification, acoustic/prosodic measurements, and a Gemini audio judge produce target success, content-preservation success, and joint success.
  • Best reported results:
    • Single model: By task, GPT-Realtime reaches 96.67% content, 68.67% style, and 47.00% paralinguistic joint success; Gemini-Live reaches 27.79% emotion and 11.00% compositional joint success. The newer AuK comparison reports 71.33% prosody joint success, but covers only five task types.

Ming-Freeform-Audio-Edit

  • TL;DR: Evaluates timestamp-free, instruction-guided speech changes with ≈3.3k instruction cases across Basic/Full lexical edits and five attribute-control tasks in Chinese and English.
  • Paper: arXiv; Code: GitHub; Dataset: Hugging Face; Project Page: Ming-UniAudio.
  • Audio modalities: Speech.
  • Editing categories: Acoustic; Semantic.
  • Evaluation method: Hybrid — Whisper/Paraformer and WavLM measure transcription and speaker preservation; signal measurements assess speed/volume control, and an audio-capable judge assesses emotion/dialect conversion.
  • Best reported results:
    • Single model: AuK reports 3.09% / 3.96% WER and 91.47% / 85.25% edit accuracy on Full Chinese / English, averaged over deletion, insertion, and substitution; these lead the checked lexical-editing comparisons.

Step-Audio-Edit-Benchmark

  • TL;DR: Evaluates expressive and iterative speech editing using 8 speakers and 8,800 text prompts for emotion, speaking style, and paralinguistics, with reference voices released and output duration dependent on synthesis.
  • Paper: arXiv; Code: GitHub; Dataset: Prompt texts · Reference audio; Project Page: Step-Audio-EditX.
  • Audio modalities: Speech.
  • Editing categories: Semantic.
  • Evaluation method: MLLM — Gemini-2.5-Pro measures emotion/style classification accuracy and rates paralinguistic realization on a 1–3 scale.
  • Best reported results:
    • Single model: Step-Audio-EditX reports 71.0% emotion accuracy and 66.2% style accuracy after three editing iterations, and 2.89/3 paralinguistic score after one iteration in the published native-input comparison.

LyricEditBench (INTERSPEECH 2026)

  • TL;DR: Tests melody-preserving lyric modification with 7,200 bilingual cases, each using a ≤15 s melody reference, across six lyric-editing scenarios and both self-timbre and cross-timbre settings.
  • Paper: arXiv; Code: GitHub; Dataset: Hugging Face; Project Page: YingMusic-Singer-Plus.
  • Audio modalities: Music.
  • Editing categories: Semantic; Instance; Composite (lyric changes together with singer-identity transfer in the cross-timbre setting).
  • Evaluation method: Expert models — singing ASR for Phoneme Error Rate (PER), WavLM for speaker similarity, RMVPE for F0 correlation, and VocalVerse2 for vocal quality, supplemented by human listening tests.
  • Best reported results:
    • Single model: YingMusic-Singer outperforms Vevo2 on lyric intelligibility, melody adherence, and vocal quality in the published comparison; for Chinese partial substitution, self-timbre, it achieves 2.14% PER and 0.9615 F0 correlation. Vevo2 retains an advantage in speaker similarity on this setting.

ZoME-Bench (ACM MM 2025)

  • TL;DR: Provides 1,100 music-editing cases (10 s each; ≈3.1 h counted per case) across instrument, genre, mood, rhythm, melody, and background changes, with captions and instructions supporting both prompt-based and instruction-based evaluation.
  • Paper: MEDIC; Code: MEDIC repository (implementation not released) · later evaluation code; Dataset: Hugging Face metadata (source audio retrieved separately from MusicCaps/YouTube); Project Page: MEDIC.
  • Audio modalities: Music.
  • Editing categories: Semantic; Instance.
  • Evaluation method: Expert models — audio–text alignment and perceptual/structural measures such as CLAP, LPAPS, and chroma similarity, supplemented by human ratings.
  • Best reported results:
    • Single model / non-agent editor: In the later AnchorSteer instrument-editing comparison, its conditioned variant achieves the highest CLAP (0.395) and GAP (0.279) among the compared methods, while its unconditioned variant preserves more structure (chroma similarity 0.470, versus 0.238 for the conditioned variant).

MelodiaEdit (AAAI 2026)

  • TL;DR: Evaluates instrument, genre, and mood changes while preserving musical structure through 2,015 editing pairs drawn from 180 released source clips (≈0.86 h unique audio), combining synthesized and real music.
  • Paper: AAAI proceedings; Code: GitHub (data release; evaluation implementation not released); Dataset: Audio and prompts; Project Page: Melodia.
  • Audio modalities: Music.
  • Editing categories: Semantic; Instance.
  • Evaluation method: Expert models — CLAP, LPAPS, chroma similarity, FAD, and combined adherence/preservation scores, supplemented by human listening tests.
  • Best reported results:
    • Single model: In the published comparison, Melodia has the highest CLAP (0.39) and lowest LPAPS (3.11) on MelodiaEdit; MusicMagus instead leads chroma similarity (0.73) and FAD (0.57), illustrating the alignment–preservation trade-off.

AvED-Bench (WACV 2026)

  • TL;DR: Tests synchronized sound-event and visual-entity replacement using 110 audio-video clips (10 s each; ≈18.3 min) curated from VGGSound with source/target descriptions.
  • Paper: arXiv; Code: GitHub; Dataset: Benchmark CSV (source clips retrieved separately from VGGSound/YouTube); Project Page: AvED.
  • Audio modalities: Audio (with video input).
  • Editing categories: Instance; editing the video alongside the sound does not by itself constitute Composite audio editing.
  • Evaluation method: Expert models — embedding-based audio–text/audio–video alignment and perceptual preservation metrics, supplemented by human judgments.
  • Best reported results:
    • Single model / non-agent system: The later CoherentAVEdit comparison reports the strongest listening-test results for VACE → CoherentAVEdit: 3.7/5 audio–text fidelity, 3.8/5 audio–visual alignment, and 3.9/5 structure preservation on a six-video subjective subset; this is a sequential, non-agent pipeline, not a single joint model or a full-set aggregate ranking.

AVE-Compass

  • TL;DR: Diagnoses instruction following and preservation in free-form audio-video editing through 145 source videos (up to 10 s), 196 instructions, and 2,688 checklist items covering 28 editing operations.
  • Paper: arXiv; Code: GitHub; Dataset: Hugging Face; Project Page: AVE-Compass.
  • Audio modalities: Speech; Music; Audio (with video input).
  • Editing categories: Acoustic; Semantic; Instance; Composite (when the audio request itself crosses categories).
  • Evaluation method: Hybrid — checklist-based MLLM judging combines with expert models for audio quality, audio–video/lip synchronization, and visual preservation.
  • Best reported results:
    • Single model: On the official leaderboard, Wan2.7 leads overall Editing Intent at 42.4/100 (audio: 24.8); LTX2 has the highest audio-only Editing Intent among the listed single models (26.4/100).
    • Agent: AVE-Agent (Wan) leads with 59.8/100 overall Editing Intent and 50.2/100 audio Editing Intent on the same leaderboard.

📏 Evaluation Metrics

Metrics are grouped by the four evaluation dimensions used in this survey. ↑ / ↓ indicate higher / lower is better. Reference / Inputs lists the information needed alongside the edited output.

Instruction Adherence

MetricWhat it measuresAudio ModalitiesReference / InputsPaper / StandardCode / Model
WER / CER ↓ASR transcription errors relative to the requested words or characters.SpeechTarget transcript; ASR transcript of the edited output.arXivJiWER
ASR
Emotion classification accuracy ↑Agreement between the predicted emotion and the requested emotion label.SpeechTarget emotion label; an emotion classifier with a matching label set.arXivemotion2vec
CLAP audio–text similarity ↑Cosine similarity between output audio and the desired audio description.Music; AudioCaption describing the desired result; a specified CLAP checkpoint.arXivCLAP
Event Occurrence Score (EOS) ↑Minimum event-level CLAP similarity after text-guided source separation; checks coverage of requested events.AudioDesired event descriptions; event decomposition and separated event tracks.Paper

Preservation and Locality

For local edits, compare the regions or sources that should remain unchanged.

MetricWhat it measuresAudio ModalitiesReference / InputsPaper / StandardCode / Model
Speaker embedding cosine similarity ↑Retention of speaker identity in the edited speech.SpeechSource speaker audio; the same speaker-verification encoder for both recordings.arXivHF Model
Multi-resolution STFT distance ↓Spectral convergence and log-magnitude differences across several time–frequency resolutions.Speech; Music; AudioAligned source/output audio from the non-edited regions.arXivauraloss
CLAP audio–audio similarity ↑Semantic similarity between source and edited audio embeddings; a broad preservation proxy.Music; AudioSource audio; matching non-target regions or stems for local comparison.arXivCLAP
LPAPS distance ↓Perceptual distance between audio representations from a pretrained feature network.Music; AudioSource audio; aligned non-edited regions for local comparison.arXivLPAPS

Temporal and Structural Consistency

MetricWhat it measuresAudio ModalitiesReference / InputsPaper / StandardCode / Model
Boundary error ↓Mean or median absolute timing error of predicted speech-segment boundaries.SpeechManual boundary annotations and predicted boundaries, in the same time unit.Paper
Word-level Dynamic Time Warping (WDTW) ↓Length-normalized DTW distance over matched word segments in source and edited speech.SpeechSource and edited speech; both transcripts and word-level forced alignments.arXiv
Melody accuracy ↑Frame-wise agreement of the dominant pitch class between reference and edited music.MusicReference melody/audio; aligned pitch-class sequences.arXivEditGen
F0 Pearson correlation ↑Correlation between reference and output vocal-pitch contours.Music (vocals)Reference vocal audio; aligned F0 contours extracted with the same model.arXivPitch extractor
Pitch extractor
Chroma similarity / Chroma DTW similarity ↑Pitch-class distribution similarity, or frame-wise similarity after DTW alignment.MusicSource/reference music; consistently extracted chromagrams.arXivMuseCPEval
Beat F1 ↑Precision–recall balance of matching beat timestamps within a 70 ms tolerance.MusicReference and output beat timestamps.arXivMuseCPEval
Dynamics correlation ↑Frame-wise Pearson correlation of reference and output loudness trajectories.MusicReference dynamics/audio; aligned loudness trajectories.arXivEditGen
Structural pairwise F-measure / ARI ↑Agreement of musical section assignments, with ARI correcting for chance agreement.MusicSource/reference and output segmentations in a shared time frame.arXivMuseCPEval
Event Sequence Score (ESS) ↑Kendall-style rank agreement between the described and detected event order.AudioDesired event ordering; onset estimates from separated event tracks.Paper

Audio Quality and Naturalness

MetricWhat it measuresAudio ModalitiesReference / InputsPaper / StandardCode / Model
MOS / CMOS ↑Human ratings of output quality or comparative quality against another recording.Speech; Music; AudioListeners and a task-specific rating protocol; comparison audio for CMOS.StandardSpeech listening tests
Speech listening tests
MOSNet predicted MOS ↑Automatic prediction of speech naturalness ratings, developed for voice conversion.SpeechEdited speech.arXivMOSNet
UTMOSv2 predicted MOS ↑Predicted naturalness MOS, developed for high-quality synthetic speech.SpeechEdited speech.arXivGitHub
HF Model
SpeechJudge-GRM (pairwise)Paired naturalness ratings and preference, accompanied by a generated explanation.SpeechTarget transcript and two candidate speech recordings for the same text.arXivGitHub
HF Model
DNSMOS P.835 ↑Predicted speech-signal, background-noise and overall quality scores.SpeechEdited speech.arXivDNSMOS
NISQA ↑Predicted overall speech quality and degradation dimensions; NISQA-TTS targets synthetic-speech naturalness.SpeechEdited speech; the appropriate NISQA checkpoint.arXivNISQA
PAM ↑No-reference audio quality from an audio–language model using contrasting positive and negative quality prompts.Speech; Music; AudioEdited audio; fixed quality prompts and the PAM implementation's MS-CLAP backbone.arXivGitHub
PESQ ↑Reference-based perceptual speech quality after degradation or restoration.SpeechCorresponding clean target speech; 8 kHz narrowband or 16 kHz wideband mode.StandardPESQ
STOI ↑Estimated intelligibility of degraded or enhanced speech.SpeechTime-aligned clean target speech.Paperpystoi
SI-SDR ↑Target-signal reconstruction fidelity after compensating for a global scale difference.Speech; Music; AudioTime-aligned target waveform or isolated target source.arXivTorchMetrics
NOMAD distance ↓Perceptual speech degradation measured in a learned embedding space.SpeechClean speech references; matching linguistic content is not required.arXivNOMAD
SpeechBERTScore ↑Reference-aware speech quality proxy using greedy matching of self-supervised speech features.SpeechNatural reference speech; a fixed encoder, layer and precision/recall/F1 variant.arXivGitHub
Fréchet Audio Distance (FAD) ↓Distance between output and reference audio-embedding distributions.Music; AudioReference audio collection; the same embedding backbone and preprocessing.arXivFADtk

Multi-dimensional Evaluators

Reusable models and toolkits for multi-dimensional assessment of editing results and audio aesthetics.

EvaluatorAudio ModalitiesEvaluation DimensionsReference / InputsPaperCode / Model
AuditEval (SSL / LLM)AudioQuality, editing relevance and faithfulness to the source.Source and edited audio; original and target descriptions.arXivGitHub
ModelScope
MuseCPEvalMusicHarmony, rhythm, structure and melody preservation, with additional timbre metrics in the toolkit.Source and edited music; selected musical attributes to preserve.arXivGitHub
MMAE rubric evaluator (Qwen3-Omni)Speech; Music; AudioInstruction Following Rate (IFR), Consistency Rate (CR) and Exact Match Rate (EMR).Source/output audio, editing instructions and sample-specific MMAE rubrics.arXivGitHub
Audiobox AestheticsSpeech; Music; AudioContent Enjoyment (CE), Content Usefulness (CU), Production Complexity (PC) and Production Quality (PQ).Edited audio only.arXivGitHub
HF Model
SongEval scoring modelMusic (songs)Overall coherence, memorability, vocal breathing/phrasing naturalness, structural clarity and overall musicality.Full-length song audio with vocals and accompaniment.arXivGitHub
Weights
MuseCriticMusic (songs)Coherence, musicality, memorability, structural clarity and vocal naturalness; returns scores and a natural-language critique.Full-length song audio; the released aesthetic rubric.arXivGitHub
HF Model

🔮 Challenges and Future Directions

Foundation-model-based audio editing still faces several system-level challenges:

  1. Complex editing.
    Real-world audio entangles semantic events, speaker identity, acoustic attributes, background ambience, rhythm, spatial cues, and reverberation. Future systems should support precise source/event localization, attribute-level modification, and reliable preservation of non-target content across speech, music, and general audio.
  2. Robustness in open-domain settings.
    Editing models should remain reliable under noise, reverberation, overlapping sources, long-form context, and ambiguous instructions. Better instruction grounding, long-context modeling, iterative refinement, and self-verification are important directions toward robust real-world editing.
  3. Faithful and editing-specific evaluation.
    Existing evaluation often mixes generation quality with editing quality. Future benchmarks should explicitly annotate edit targets, operations, preservation regions, and relevant control signals, allowing edit success and non-target preservation to be evaluated separately.
  4. Safety, copyright, and misuse prevention.
    Modern editing systems can realistically modify speech content, speaker identity, emotion, environmental sounds, and music. Practical deployment therefore requires complementary mechanisms for provenance, watermarking, manipulated-audio detection, and responsible data licensing.

Citation

If You find this survey or repository useful, please cite our paper:

@article{pan2026audio,
  title={Audio Editing in the Era of Foundation Models: A Survey},
  author={Pan, Changhao and Fan, Yifei and Zhuo, Fan and Chen, Yifu and Guo, Wenxiang and Zhang, Yu and Li, Ruiqi and Zhu, Zhiyuan and Yang, Rui and Ji, Shengpeng and others},
  journal={arXiv preprint arXiv:2606.23139},
  year={2026}
}

Contributing

This repo is meant to keep growing. If an audio editing model, dataset, or benchmark is missing, please feel free to open an issue or a pull request.


📄 License

Unless otherwise noted below, original content created for this repository is licensed under the MIT License.

The survey paper and content reproduced or adapted from it, including assets/taxonomy_overview.png, assets/train-based.png, and assets/train-free.png, remain under CC BY-NC-SA 4.0. The MIT license does not relicense these materials.

Linked third-party papers, code, models, model weights, datasets, and tools are governed by their respective licenses.

aacl-2026
audio-editing
awesome-list
foundation-models
survey

MM-Speech/AudioEditSurvey

[AACL-IJCNLP] A survey of foundation-model-based audio editing across speech, music, and general audio, covering task taxonomies, training-based and training-free methods, datasets, and evaluation.

32

11 commits

updated Sep 29, 2026

See the code

README

Awesome Audio Editing

Audio Editing in the Era of Foundation Models: A Survey

AACL-IJCNLP 2026

Changhao Pan1,*, Yifei Fan1,*, Fan Zhuo1,*, Yifu Chen1, Wenxiang Guo1,
Yu Zhang2, Ruiqi Li2, Zhiyuan Zhu1, Rui Yang1, Shengpeng Ji3,
Chenyuhao Wen1, Jiayang Xu1, Ke Lei1, Xiaoda Yang1, Jingyu Lu1, Zhou Zhao1,†

1Zhejiang University  ·  2ByteDance  ·  3Hunyuan Team, Tencent

*Equal contribution  ·  †Corresponding author

arXiv AACL-IJCNLP 2026 Project Page GitHub stars

🌐 Languages

English · 简体中文 · 한국어

🚀 Quick Start

This repository is the official repository for Audio Editing in the Era of Foundation Models: A Survey, Which is accepted by AACL-IJCNLP 2026.

  • We establish a unified taxonomy of acoustic, semantic, and instance editing across speech, music, and general audio, clarifying what each task changes and what it should preserve to support consistent comparisons across editing goals.
  • We review mainstream audio editing techniques through foundation-model architectures and learning paradigms, with an emphasis on their core mechanisms and suitability for different editing scenarios.
  • (Updated Recently) We curate publicly available audio editing models, and summarize their supported task categories and key strengths to help readers identify suitable models.
  • (Updated Recently) We organize publicly available datasets, data construction tools, evaluation benchmarks, and metrics for audio editing, summarizing the audio domains, editing categories, and evaluation dimensions they cover.

🔥What's new

  • 📦 [2026/09] This repository has moved to MM-Speech/AudioEditSurvey for better management.
  • 🏆 [2026/09] Our paper has been accepted to the AACL-IJCNLP 2026!
  • 🎉 [2026/06] We have officially released this survey repository for Audio Editing Models, with the preprint available on arXiv.

Contents

  1. Introduction
  2. Scope
  3. Overall
  4. Foundation Models for Audio Editing
  5. Training-based Audio Editing
  6. Training-free Audio Editing
  7. Resources
  8. Challenges and Future Directions
  9. Citation
  10. Contributing

📌 Introduction

This is the official repository for Audio Editing in the Era of Foundation Models: A Survey, accepted to AACL-IJCNLP 2026. It is maintained by MM-Speech and collects papers and resources for foundation-model-based audio editing.

Abstract
Audio editing aims to modify a given synthetic or real-world audio signal to meet users' specific needs. As a promising yet challenging direction in AIGC, it has attracted increasing attention in recent years. With the rapid progress of text-to-audio and text-to-speech generation, powerful audio generation models have become the primary foundation for modern audio editing systems. In this survey, we provide a comprehensive review of foundation-model-based audio editing. We first define the scope of audio editing from a unified perspective and present a detailed taxonomy of existing editing tasks. We then summarize the major foundation-model paradigms for audio editing, and review representative approaches from both training-based and training-free perspectives. In addition, we systematically discuss related resources, including datasets, data construction tools, and evaluation protocols. Finally, we identify open challenges in this field and outline promising directions for future research.


🎯 Scope

In this survey, we focus on works that make direct contributions to audio editing in the era of foundation models. To ensure a precise and focused discussion, we adopt two main inclusion criteria: (1) the task should center on audio editing, which we define as modifying the acoustic attributes, instances, or content of an existing audio recording, without transformations so substantial that they amount to generating an entirely new audio sample; (2) the method should rely on mainstream audio foundation model paradigms. Accordingly, we do not cover works primarily focused on audio generation, nor do we provide an extensive discussion of signal-processing-based audio editing methods. In addition, to maintain a focused scope, spatial audio (multi-channel formats such as binaural stereo and first-order Ambisonics (FOA)) and related editing techniques are beyond the main scope of this survey.


🧭 Overall

🗂️ Taxonomy Overview

Taxonomy of Audio Editing Tasks

Figure 1: Taxonomy of audio editing tasks.

🧩 Taxonomy Details

CategoryDefinitionRepresentative Editing Goals
Acoustic EditingModifies low-level perceptual attributes while preserving the overall structure and source characteristics of the original audio.restoration, reverberation editing, loudness/mixing control, equalization, spectral texture editing
Semantic EditingModifies high-level interpretable information conveyed by audio while maintaining task-irrelevant properties.linguistic editing, expressive editing, stylistic editing
Instance EditingManipulates identifiable audio entities while preserving the remaining scene and source relationships.replacement, deletion/extraction, insertion, overlay/remixing

📚 Representative Audio Editing Methods

Representative editors with publicly released implementations and model weights, subject to each project’s license. Editing types follow our taxonomy; Unified groups editors supporting multiple audio domains, with their supported domains listed in the table. Base links the pretrained backbone used by an editor, while Adapter links its additional learned weights.

Unified Models

ModelAudio DomainEditing TypesModel ArchitecturePaperCodeModel
Audio-OmniSpeech; Music; AudioInstance: addition, removal, extraction, source transformationMLLM + rectified-flow DiTarXiv PaperGitHub Code🤗 Weights
AudioMorphixSpeech; Music; AudioSemantic: pitch / time stretching
Instance: addition, removal, replacement, time shifting
Diffusion U-Net (Tango / AudioLDM)arXiv PaperHuggingFace Code🤗 Base (Tango 2)
🤗 Base (AudioLDM)
AuK / AuK-FlashSpeech; MusicAcoustic: restoration, loudness
Semantic: words, lyrics, expression
Instance: timbre, source extraction
MLLM + rectified-flow DiTarXiv PaperGitHub Code🤗 AuK
🤗 Flash
Vevo2Speech; MusicSemantic: content, lyrics, prosody, style
Instance: voice / singer conversion
Codec LM + flow-matching decoderarXiv PaperGitHub Code🤗 Weights
DirectAudioEditMusic; AudioInstance: text-guided event replacement / addition / removalDiffusion U-Net (Tango 2 / AudioLDM2)arXiv PaperGitHub Code🤗 Base (Tango 2)
🤗 Base (audio)
DDPM Inversion (ZETA)Music; AudioSemantic: musical style
Instance: instrument / sound-event changes
Diffusion U-Net (AudioLDM2)arXiv PaperGitHub Code🤗 Base (audio)
🤗 Base (music)

Speech Models

ModelEditing TypesModel ArchitecturePaperCodeModel
Ming-UniAudio-EditAcoustic: denoising, loudness
Semantic: content, prosody, emotion, dialect
Continuous-token LM + diffusion headarXiv PaperGitHub Code🤗 Weights
Step-Audio-EditXSemantic: emotion, speaking style, paralinguistics, pronunciationCodec LM + flow-matching decoderarXiv PaperGitHub Code🤗 Weights
CosyEditSemantic: word insertion, deletion, replacementCodec LM + flow-matching decoderarXiv PaperGitHub Code🤗 Weights
VoiceCraft-XSemantic: multilingual content editingCodec LM (autoregressive infilling)arXiv PaperGitHub Code🤗 Weights
VoiceCraftSemantic: word insertion, deletion, replacementCodec LM (autoregressive infilling)arXiv PaperGitHub Code🤗 Weights
SSR-SpeechSemantic: word insertion, deletion, replacementCodec LM (autoregressive infilling)arXiv PaperGitHub Code🤗 English
🤗 Mandarin
F5-TTSSemantic: local content replacement / infillingFlow-matching DiTarXiv PaperGitHub Code🤗 Weights
FluentSpeechSemantic: content editing, disfluency correctionDiffusion (context-aware denoiser)arXiv PaperGitHub Code📁 Weights
EdiTTSSemantic: content / pitch edits in synthesized speechScore-based diffusion (Grad-TTS)arXiv PaperGitHub Code📦 Base

Music Models

ModelEditing TypesModel ArchitecturePaperCodeModel
YingMusic-Singer-PlusSemantic: lyrics
Instance: singer timbre replacement
Flow-matching DiTarXiv PaperGitHub Code🤗 Weights
ACE-Step 1.5Semantic: style / local repainting
Instance: track extraction / addition (base variant)
LM + flow-matching DiTarXiv PaperGitHub Code🤗 Turbo
🤗 Base variant
Instruct-MusicGenInstance: stem addition, removal, extractionCodec LM (MusicGen) + adaptersarXiv PaperGitHub Code🤗 Public-data retraining
MusicGen-StemInstance: stem replacement / addition (bass, drums, other)Multi-stream codec LMarXiv PaperGitHub Code🤗 Weights
MelodyFlowSemantic: genre, mood, style
Instance: instrumentation
Flow-matching DiTarXiv PaperHuggingFace Code🤗 Weights
AP-AdapterSemantic: genre / style transfer
Instance: instrument replacement
Diffusion U-Net + audio-prompt adapterarXiv PaperGitHub Code📁 Adapter
🤗 Base
AnchorSteerSemantic: genre / style
Instance: instrument changes
Diffusion DiT + structural/concept adaptersarXiv PaperGitHub Code🤗 Concept weights
📁 Structure adapter
🤗 Base (access terms)

Audio Models

ModelEditing TypesModel ArchitecturePaperCodeModel
MMEditAcoustic: loudness
Instance: event addition, removal, replacement, reordering
ALM + diffusion MMDiTarXiv PaperGitHub Code🤗 Weights
SAO-InstructAcoustic: filtering, denoising, restoration
Semantic: pitch / rate
Instance: event manipulation
Diffusion DiT (Stable Audio Open)arXiv PaperGitHub Code🤗 Weights
SmartDJ-EditorAcoustic: volume, reverb, spectral coloration
Instance: event addition, removal, extraction, relocation
Diffusion Transformer (U-DiT)arXiv PaperGitHub Code🤗 Editor weights
AudioEditorInstance: event addition, deletion, replacementDiffusion U-Net (Auffusion)arXiv PaperGitHub Code🤗 Base
CoherentAVEditInstance: video-conditioned sound-event replacementFlow-matching Transformer (MMAudio)arXiv PaperGitHub Code🤗 Weights

🏗️ Foundation Models for Audio Editing

1. Early Neural Editing Models

Before the foundation-model era, early neural audio editing methods mainly explored task-specific generative models for local reconstruction and attribute control.

2. Token-based Codec Language Models

Token-based codec language models cast audio editing as conditional generation over discrete audio tokens. After continuous audio is converted into compact discrete token sequences, target regions are edited through autoregressive continuation, infilling, or selective regeneration conditioned on context, prompts, or task controls.

3. Diffusion and Flow-Matching Models

Diffusion and flow-matching models formulate audio editing as conditional transformation in continuous acoustic spaces, such as mel-spectrograms or audio latents. Instead of infilling discrete tokens, they modify audio through conditional denoising, latent inversion, or continuous flow transformation, making them suitable for high-fidelity reconstruction, region-level refinement, and fine-grained acoustic control in complex scenarios.

4. Audio Editing Interfaces

Instruction-conditioned and multimodal interfaces for audio editing provide high-level control for foundation-model-based audio editing. They allow users to specify editing intents through natural language instructions, task prompts, reference audio, temporal regions, or visual cues, which are shifted into target spans, task embeddings, event locations, speaker references, or preservation constraints.


🧪 Training-based Audio Editing

Training-based approaches refer to audio editing methods that learn editing behaviors from supervised pairs, pseudo-pairs, or instruction-based triplets before inference. These methods explicitly optimize editing objectives, condition following, and preservation constraints, enabling stable and controllable editing. We group existing works into three categories based on their supervision and conditioning mechanisms, and discuss their core methods and functional scopes.

Overview of training-based audio editing methods

Figure 2: Overview of training-based audio editing methods.

ParadigmDescriptionRepresentative Scope
Task-specific TrainingOptimizes models for predefined editing functions or domains.text-based speech editing, prosody correction, source separation, music stem separation
Reference- and Attribute-based TrainingSpecifies the editing direction through reference audio, style examples, or attribute labels.voice conversion, timbre transfer, emotion editing, mixing style transfer
Instruction-conditioned TrainingLearns from instruction-input-output triplets to follow natural-language editing requests.addition, deletion, replacement, inpainting, super-resolution, music remixing, expressive refinement

🪄 Training-free Audio Editing

Training-free approaches adapt pretrained audio generative models to editing without parameter updates. They operate by manipulating inference-time mechanisms, such as inversion, attention control, prompt or guidance adjustment, and mask-based constraints. We group existing methods into three common categories, which are often combined to improve localization, preservation, and controllability. Since token-based autoregressive models are less naturally suited to training-free editing, this section mainly focuses on non-autoregressive paradigms, especially diffusion-based foundation models.

Overview of training-free audio editing methods

Figure 3: Overview of training-free audio editing methods.

ParadigmDescriptionRepresentative Scope
Inversion-Based EditingMaps source audio back into the latent, noise, or trajectory space of a pretrained generative model, then edits it by modifying conditions or sampling trajectories.DDPM/DDIM inversion, latent inversion, flow-based inversion, speech or music reconstruction and editing
Attention-Controlled EditingGuides pretrained generative models by modifying or reusing internal attention patterns without parameter updates.cross-attention event localization, self-attention preservation, prompt-level manipulation
Mask- and Region-Guided EditingSpecifies where to edit and where to preserve the source audio in waveform, spectrogram, latent, or source-component spaces.localized editing, inpainting, restoration, source-level manipulation
Token-Level Editing with Codec ModelsManipulates discrete audio tokens through masking, infilling, continuation, or selective regeneration at inference time.speech infilling, localized resynthesis, codec-token editing

📦 Resources

📊 Available Datasets

Public datasets for audio editing and controllable audio generation, grouped by their primary audio domain.

This non-exhaustive list highlights datasets suited to audio editing or widely used in the community, with availability verified by the repository maintainers for every entry.

Paired indicates released source–target audio, mixture–stem correspondence, or explicitly matched control/technique takes (✅ / ❌); shared transcripts, audio–text alignment, or audio–MIDI alignment alone do not count. † marks an editing use that requires task construction or adaptation, rather than native editing supervision. Editing types follow our Acoustic / Instance / Semantic taxonomy.

Durations are approximate, without adding together alternate modalities or mixture stems. Text refers to transcripts, captions or instructions; label-only metadata are described in Annotation.

Speech

NamePaperDataset / CodeDurationPairedEditing TypesAnnotationModalities
VoiceBank+DEMAND (28-spk)Paper LinkDataShare Data≈10 h✅ Noisy/cleanAcousticTranscript; noise/SNR conditionsAudio, Text
LibriTTS-RarXiv PaperOpenSLR Restored
OpenSLR Original
≈585 h✅ Original/restoredAcoustic; Semantic†Transcript; speaker labels; model-restored audioAudio, Text
LibriSpeechPaper LinkOpenSLR Data≈1,000 h❌Semantic†; Instance†Transcript; speaker/chapter labelsAudio, Text
VCTK v0.92Dataset RecordDataShare Data≈44 h❌Instance†; Semantic†Transcript; speaker/accent labelsAudio, Text
AISHELL-3arXiv PaperOpenSLR Data≈85 h❌Semantic†; Instance†Mandarin transcript; phonetic transcription; speaker labelsAudio, Text
Hi-Fi TTSarXiv PaperOpenSLR Data≈292 h❌Semantic†; Instance†Transcript; speaker labelsAudio, Text
LJSpeech v1.1Dataset ReleaseDownload Data≈24 h❌Semantic†Transcript; normalized textAudio, Text
RAVDESS (speech)Paper LinkZenodo Data≈1.7 h❌Semantic†Label: emotion, intensity, speaker; fixed transcriptsAudio, Text, Video
CREMA-DPaper LinkGitHub Code
GitLab Mirror
≈5.3 h❌Semantic†Label: emotion/intensity; perceptual ratings; fixed transcriptsAudio, Text, Video

Music

NamePaperDataset / CodeDurationPairedEditing TypesAnnotationModalities
GTSingerarXiv PaperGitHub Code
HuggingFace Dataset
Google Drive Data
≈80.6 h singing
+16.2 h speech
✅ Controlled/parallel takesSemantic; Instance†Label: technique/style; aligned lyrics/phonemes; scoresAudio, Text, MusicXML
Slakh2100arXiv PaperGitHub Code
Zenodo Data
≈145 h✅ Mixture/stemsInstanceLabel: instrument; aligned MIDI; stem metadataAudio, MIDI
MUSDB18-HQarXiv PaperGitHub Code
Zenodo Data
≈10 h✅ Mixture/stemsInstanceLabel: vocals, drums, bass, otherAudio
MAESTRO v3arXiv PaperProject Page
Download Data
≈199 h❌Semantic†Aligned MIDI: pitch, timing, velocity, pedals; piece metadataAudio, MIDI
NSyntharXiv PaperProject Page≈340 h❌Instance†; Semantic†Label: instrument, pitch, velocity, timbral qualitiesAudio
Groove MIDI DatasetarXiv PaperProject Page
Download Data
≈13.6 h❌Semantic†Aligned MIDI; tempo/style labels; performance timing/velocityAudio, MIDI
MusicCapsarXiv PaperHuggingFace Metadata≈15.3 h❌Semantic†; Instance†Caption; musical aspect labelsAudio, Text
MTG-JamendoPublication RecordGitHub Code
Download Data
≈3,770 h❌Semantic†; Instance†Label: genre, instrument, mood/themeAudio
FMA (large)arXiv PaperGitHub Code
Download Data
≈888 h❌Semantic†Label: genre hierarchy; track/artist metadataAudio

Audio

NamePaperDataset / CodeDurationPairedEditing TypesAnnotationModalities
FUSSarXiv PaperGitHub Code
Zenodo Data
≈61 h mixtures✅ Mixture/sources; dry/reverberantInstance; AcousticSource/time metadata; mixing parameters; no event labelsAudio
AudioSetPaper LinkDataset Metadata≈5,790 h❌Instance†Label: sound-event ontology; clip-level multi-labelsAudio, Video (upstream)
AudioCaps v1Paper LinkGitHub Metadata≈143 h❌Instance†; Semantic†Caption: one or five descriptions per clipAudio, Text
Clotho v2.1arXiv PaperZenodo Data≈37 h
(5,929 labeled clips)
❌Instance†; Semantic†Caption: five per clip; Freesound keywordsAudio, Text
WavCapsarXiv PaperGitHub Code
HuggingFace Dataset
≈7,568 h❌Instance†; Semantic†LLM-assisted captions; source descriptions/metadataAudio, Text
FSD50KarXiv PaperZenodo Data≈108 h❌Instance†Label: 200 sound-event classes; clip-level multi-labelsAudio
ESC-50Paper LinkGitHub Code≈2.8 h❌Instance†Label: 50 environmental sound classesAudio
UrbanSound8KPaper LinkProject Page
Zenodo Data
≈8.8 h❌Instance†Label: 10 urban sound classes; salience; source timestampsAudio
VGGSoundarXiv PaperGitHub Metadata≈550 h❌Instance†Label: audio-visual event class; video timestampsAudio, Video (upstream)

Unified

These corpora combine speech, music, and general sounds.

NamePaperDataset / CodeDurationPairedEditing TypesAnnotationModalities
AudioEdit (Audio-Omni)arXiv PaperGitHub Code
HuggingFace Dataset
≈2,686 h
(966,794 task pairs)
✅ Source/edited targetInstanceInstruct: add, remove, extract, source transformationAudio, Text
Divide and Remaster v2arXiv PaperGitHub Code
Zenodo Data
≈81 h✅ Mixture/stemsInstanceTranscript; music genre; sound labels/timestampsAudio, Text
MUSANarXiv PaperOpenSLR Data≈109 h❌Acoustic†; Instance†Label: speech/music/noise; speech and music metadataAudio

🛠️ Data Tools

Open-source tools for constructing editing data and annotating existing recordings. Supported Task Type follows our Acoustic / Semantic / Instance taxonomy and indicates the editing supervision that each tool can help construct. Unified covers tools applicable across speech, music and general audio.

Tools for Data Generation

Synthesis, source separation, mixing and signal processing for constructing audio examples and source–target pairs.

Speech
ToolSupported Task TypeWhat It ConstructsControl LevelCodeModel
Qwen3-TTSSemantic; InstanceText-aligned utterances with instruction-controlled delivery or a reference speaker.UtteranceGitHub CodeHugging Face Base
Hugging Face CustomVoice
CosyVoice3Semantic; InstanceText-aligned speech with voice cloning and prompted language, emotion or delivery.Utterance; pronunciation unitsGitHub CodeHugging Face Model
MaskGCTSemantic; InstanceText-conditioned speech with a reference voice and configurable total duration.Utterance; total durationGitHub CodeHugging Face Model
Seed-VCInstanceVoice-converted recordings paired with their source speech for speaker/timbre replacement.Utterance / source recordingGitHub CodeHugging Face Model
AuKAcoustic; Semantic; InstanceInstruction-edited speech for content, delivery, voice, enhancement and target-speaker tasks.Utterance; text-specified word / phraseGitHub CodeHugging Face Model
Music
ToolSupported Task TypeWhat It ConstructsControl LevelCodeModel
MusicGenSemanticText- or melody-conditioned music clips and continuations for style/content-controlled examples.Clip; melody sequenceGitHub CodeHugging Face Melody
DemucsInstanceEstimated vocal, drum, bass and other stems for extraction, removal and remix pair construction.Stem / trackGitHub CodeModel Checkpoint
SpleeterInstanceEstimated 2-, 4- or 5-stem decompositions for source removal, extraction and remixing.Stem / trackGitHub CodeGitHub Checkpoints
FluidSynthSemantic; InstanceAudio rendered from MIDI and a SoundFont, aligned with notes, velocities and instrument assignments.Note; MIDI control event / trackGitHub Code
Audio
ToolSupported Task TypeWhat It ConstructsControl LevelCodeModel
AudioLDM 2InstanceText-conditioned sound clips to use as source assets in insertion or replacement examples.ClipGitHub CodeHugging Face Model
AudioSepInstanceText-selected source estimates from mixtures for extraction and removal pair construction.Described source / clipGitHub CodeHugging Face Checkpoints
ScaperAcoustic; InstanceSynthetic soundscapes with event labels, onset/offset times, SNRs and optional isolated event tracks.Event; start time / duration / SNRGitHub Code
SpatialScaperAcoustic; InstanceSpatialized soundscapes with event activity, source trajectories and room-response conditions.Event / trajectory / sceneGitHub Code
Unified
ToolSupported Task TypeWhat It ConstructsControl LevelCodeModel
SAM-AudioInstancePrompt-selected target and residual audio for extraction, removal and remix examples.Source; temporal-span promptsGitHub CodeHugging Face Model
Access request
AudiomentationsAcoustic; SemanticAugmented audio for clean/degraded and pitch/tempo contrast pairs using noise, gain, filtering and other transforms.Clip; selected segment via slicingGitHub Code
PedalboardAcousticEffect-processed audio for dry/wet or clean/degraded pairs using EQ, gain, compression, distortion and reverb.Clip / processing blockGitHub Code
PyroomacousticsAcoustic; InstanceRoom impulse responses and microphone mixtures from positioned sources, including dry/reverberant pairs.Scene / source positionGitHub Code

Tools for Data Annotation

Tools for extracting or creating content, attribute and temporal annotations from existing audio.

Speech
ToolSupported Task TypeWhat It AnnotatesAnnotation LevelCodeModel
Montreal Forced Aligner (MFA)SemanticWord and phone boundaries obtained by aligning speech with supplied transcripts and pronunciation dictionaries.Word / phonemeGitHub CodeModel Acoustic models
WhisperXSemanticASR transcripts with word timestamps from a language-specific alignment model.Utterance / wordGitHub CodeHugging Face ASR
Hugging Face EN aligner
Qwen3-ASR + ForcedAlignerSemanticTranscripts, language labels and text-unit timestamps using the released ASR and forced-alignment models.Utterance / wordGitHub CodeHugging Face ASR
Hugging Face Aligner
pyannote.audioInstanceSpeaker-labeled speech turns and overlapping-speaker activity.Speaker turn / segmentGitHub CodeHugging Face Community-1
Accept access terms
Silero VADInstanceSpeech/non-speech probabilities and detected speech start/end times.Frame / speech segmentGitHub CodeGitHub Weights
emotion2vec+SemanticSpeech-emotion labels and scores, with optional learned emotion representations.Utterance (labels); frame (features)GitHub CodeHugging Face Large
FunASR / SenseVoiceSmallSemantic; InstanceTranscripts, language and emotion tags, and audio-event tags such as laughter or applause.Utterance / VAD segmentGitHub CodeHugging Face SenseVoiceSmall
Praat / ParselmouthAcoustic; SemanticPitch, formants and intensity tracks; manually defined TextGrid points and intervals in Praat.Frame; word / phoneme / interval (manual)GitHub Praat
GitHub Parselmouth
Music
ToolSupported Task TypeWhat It AnnotatesAnnotation LevelCodeModel
RMVPESemanticVocal F0 trajectories from polyphonic music.FrameGitHub CodeGoogle Drive ROSVOT bundle
CREPESemanticMonophonic F0 estimates and confidence values.FrameGitHub CodeGitHub Weights
ROSVOTSemanticSinging-note pitches and onset/offset times, with word boundaries from its RWBD component.Note / wordGitHub CodeGoogle Drive Checkpoints
Basic PitchSemanticPolyphonic note events and pitch bends exported as MIDI; most effective on one instrument at a time.Note; frame-level pitch contourGitHub CodeGitHub Weights
All-In-One Music Structure AnalyzerSemanticTempo, beat/downbeat timestamps and labeled sections such as verse, chorus and bridge.Beat / downbeat / sectionGitHub CodeHugging Face Models
Music FlamingoSemantic; InstanceMusic captions and question–answer annotations about instrumentation, harmony, mood, structure and lyrics.Clip / full track (free-form text)GitHub CodeHugging Face Model
Audio
ToolSupported Task TypeWhat It AnnotatesAnnotation LevelCodeModel
PANNsInstanceSound-event class scores and frame-wise activity with the released decision-level detection models.Clip / frameGitHub CodeZenodo Models
HTS-ATInstanceSound-event tags and temporal class-activation estimates in localization mode.Clip / frameGitHub CodeGoogle Drive Models
YAMNetInstanceScores for 521 sound-event classes from overlapping audio windows.0.96 s window; 0.48 s hopGitHub CodeModel Checkpoint
Unified
ToolSupported Task TypeWhat It AnnotatesAnnotation LevelCodeModel
Qwen3-Omni CaptionerAcoustic; Semantic; InstanceDetailed audio captions covering speech, music, sound events and acoustic characteristics.Clip / recording (free-form text)GitHub CodeHugging Face Captioner
Audio Flamingo 3Semantic; InstancePrompted transcripts, captions, event descriptions and audio question–answer annotations.Clip / recording (free-form text)GitHub CodeHugging Face Model
Label StudioAcoustic; Semantic; InstanceHuman-authored clip labels, time-region labels and transcriptions using configurable audio templates.Clip / manually selected intervalGitHub Code

🧪 Benchmarks

Public evaluation resources for audio editing. Editing categories follow this survey's taxonomy: Acoustic / Semantic / Instance / Composite, where Composite combines different categories within one request. Evaluation method describes the scorer: Expert models, MLLM, or Hybrid (including agent-based evaluation).

MMAE

  • TL;DR: Tests instruction following and preservation across speech, music, sound, and their mixtures with 2,000 cases (≈8.0 h), six complexity levels, and 17,741 verification rubrics.
  • Paper: arXiv; Code: GitHub; Dataset: Hugging Face.
  • Audio modalities: Speech; Music; Audio.
  • Editing categories: Acoustic; Semantic; Instance; Composite.
  • Evaluation method: MLLM — Qwen3-Omni judges individual rubrics with majority voting, yielding Instruction Following Rate (IFR), Consistency Rate (CR), and Exact Match Rate (EMR).
  • Best reported results:
    • Single model: On the full benchmark in the original comparison, Step-Audio-EditX leads IFR (44.86%) and CR (58.88%), while Ming-UniAudio leads EMR (3.20%); Audio-Omni reaches 4.99% EMR on the separate 801-case, ≤10 s subset. The newer AuK baseline without Prompt Enhancer reports 7.58% EMR on the 1,003-case single-operation subset.
    • Agent / LLM-assisted system: The challenge agent baseline, combining an LLM router with DSP, SAM-Audio, and AuK, reports 7.45% EMR, 44.06% IFR, and 74.63% CR on all 2,000 cases. Separately, AuK-Flash with Prompt Enhancer reaches 13.85% EMR on MMAE-Speech only; these scopes are not directly comparable.

SpeechEditBench

  • TL;DR: Separates edit success from linguistic-content preservation in bilingual speech editing through 4,700 cases (≈9.4 h) covering seven atomic attributes and multi-attribute instructions.
  • Paper: arXiv; Code: GitHub; Dataset: Hugging Face, v1.1.
  • Audio modalities: Speech.
  • Editing categories: Acoustic; Semantic; Instance; Composite.
  • Evaluation method: Hybrid — ASR, speaker verification, acoustic/prosodic measurements, and a Gemini audio judge produce target success, content-preservation success, and joint success.
  • Best reported results:
    • Single model: By task, GPT-Realtime reaches 96.67% content, 68.67% style, and 47.00% paralinguistic joint success; Gemini-Live reaches 27.79% emotion and 11.00% compositional joint success. The newer AuK comparison reports 71.33% prosody joint success, but covers only five task types.

Ming-Freeform-Audio-Edit

  • TL;DR: Evaluates timestamp-free, instruction-guided speech changes with ≈3.3k instruction cases across Basic/Full lexical edits and five attribute-control tasks in Chinese and English.
  • Paper: arXiv; Code: GitHub; Dataset: Hugging Face; Project Page: Ming-UniAudio.
  • Audio modalities: Speech.
  • Editing categories: Acoustic; Semantic.
  • Evaluation method: Hybrid — Whisper/Paraformer and WavLM measure transcription and speaker preservation; signal measurements assess speed/volume control, and an audio-capable judge assesses emotion/dialect conversion.
  • Best reported results:
    • Single model: AuK reports 3.09% / 3.96% WER and 91.47% / 85.25% edit accuracy on Full Chinese / English, averaged over deletion, insertion, and substitution; these lead the checked lexical-editing comparisons.

Step-Audio-Edit-Benchmark

  • TL;DR: Evaluates expressive and iterative speech editing using 8 speakers and 8,800 text prompts for emotion, speaking style, and paralinguistics, with reference voices released and output duration dependent on synthesis.
  • Paper: arXiv; Code: GitHub; Dataset: Prompt texts · Reference audio; Project Page: Step-Audio-EditX.
  • Audio modalities: Speech.
  • Editing categories: Semantic.
  • Evaluation method: MLLM — Gemini-2.5-Pro measures emotion/style classification accuracy and rates paralinguistic realization on a 1–3 scale.
  • Best reported results:
    • Single model: Step-Audio-EditX reports 71.0% emotion accuracy and 66.2% style accuracy after three editing iterations, and 2.89/3 paralinguistic score after one iteration in the published native-input comparison.

LyricEditBench (INTERSPEECH 2026)

  • TL;DR: Tests melody-preserving lyric modification with 7,200 bilingual cases, each using a ≤15 s melody reference, across six lyric-editing scenarios and both self-timbre and cross-timbre settings.
  • Paper: arXiv; Code: GitHub; Dataset: Hugging Face; Project Page: YingMusic-Singer-Plus.
  • Audio modalities: Music.
  • Editing categories: Semantic; Instance; Composite (lyric changes together with singer-identity transfer in the cross-timbre setting).
  • Evaluation method: Expert models — singing ASR for Phoneme Error Rate (PER), WavLM for speaker similarity, RMVPE for F0 correlation, and VocalVerse2 for vocal quality, supplemented by human listening tests.
  • Best reported results:
    • Single model: YingMusic-Singer outperforms Vevo2 on lyric intelligibility, melody adherence, and vocal quality in the published comparison; for Chinese partial substitution, self-timbre, it achieves 2.14% PER and 0.9615 F0 correlation. Vevo2 retains an advantage in speaker similarity on this setting.

ZoME-Bench (ACM MM 2025)

  • TL;DR: Provides 1,100 music-editing cases (10 s each; ≈3.1 h counted per case) across instrument, genre, mood, rhythm, melody, and background changes, with captions and instructions supporting both prompt-based and instruction-based evaluation.
  • Paper: MEDIC; Code: MEDIC repository (implementation not released) · later evaluation code; Dataset: Hugging Face metadata (source audio retrieved separately from MusicCaps/YouTube); Project Page: MEDIC.
  • Audio modalities: Music.
  • Editing categories: Semantic; Instance.
  • Evaluation method: Expert models — audio–text alignment and perceptual/structural measures such as CLAP, LPAPS, and chroma similarity, supplemented by human ratings.
  • Best reported results:
    • Single model / non-agent editor: In the later AnchorSteer instrument-editing comparison, its conditioned variant achieves the highest CLAP (0.395) and GAP (0.279) among the compared methods, while its unconditioned variant preserves more structure (chroma similarity 0.470, versus 0.238 for the conditioned variant).

MelodiaEdit (AAAI 2026)

  • TL;DR: Evaluates instrument, genre, and mood changes while preserving musical structure through 2,015 editing pairs drawn from 180 released source clips (≈0.86 h unique audio), combining synthesized and real music.
  • Paper: AAAI proceedings; Code: GitHub (data release; evaluation implementation not released); Dataset: Audio and prompts; Project Page: Melodia.
  • Audio modalities: Music.
  • Editing categories: Semantic; Instance.
  • Evaluation method: Expert models — CLAP, LPAPS, chroma similarity, FAD, and combined adherence/preservation scores, supplemented by human listening tests.
  • Best reported results:
    • Single model: In the published comparison, Melodia has the highest CLAP (0.39) and lowest LPAPS (3.11) on MelodiaEdit; MusicMagus instead leads chroma similarity (0.73) and FAD (0.57), illustrating the alignment–preservation trade-off.

AvED-Bench (WACV 2026)

  • TL;DR: Tests synchronized sound-event and visual-entity replacement using 110 audio-video clips (10 s each; ≈18.3 min) curated from VGGSound with source/target descriptions.
  • Paper: arXiv; Code: GitHub; Dataset: Benchmark CSV (source clips retrieved separately from VGGSound/YouTube); Project Page: AvED.
  • Audio modalities: Audio (with video input).
  • Editing categories: Instance; editing the video alongside the sound does not by itself constitute Composite audio editing.
  • Evaluation method: Expert models — embedding-based audio–text/audio–video alignment and perceptual preservation metrics, supplemented by human judgments.
  • Best reported results:
    • Single model / non-agent system: The later CoherentAVEdit comparison reports the strongest listening-test results for VACE → CoherentAVEdit: 3.7/5 audio–text fidelity, 3.8/5 audio–visual alignment, and 3.9/5 structure preservation on a six-video subjective subset; this is a sequential, non-agent pipeline, not a single joint model or a full-set aggregate ranking.

AVE-Compass

  • TL;DR: Diagnoses instruction following and preservation in free-form audio-video editing through 145 source videos (up to 10 s), 196 instructions, and 2,688 checklist items covering 28 editing operations.
  • Paper: arXiv; Code: GitHub; Dataset: Hugging Face; Project Page: AVE-Compass.
  • Audio modalities: Speech; Music; Audio (with video input).
  • Editing categories: Acoustic; Semantic; Instance; Composite (when the audio request itself crosses categories).
  • Evaluation method: Hybrid — checklist-based MLLM judging combines with expert models for audio quality, audio–video/lip synchronization, and visual preservation.
  • Best reported results:
    • Single model: On the official leaderboard, Wan2.7 leads overall Editing Intent at 42.4/100 (audio: 24.8); LTX2 has the highest audio-only Editing Intent among the listed single models (26.4/100).
    • Agent: AVE-Agent (Wan) leads with 59.8/100 overall Editing Intent and 50.2/100 audio Editing Intent on the same leaderboard.

📏 Evaluation Metrics

Metrics are grouped by the four evaluation dimensions used in this survey. ↑ / ↓ indicate higher / lower is better. Reference / Inputs lists the information needed alongside the edited output.

Instruction Adherence

MetricWhat it measuresAudio ModalitiesReference / InputsPaper / StandardCode / Model
WER / CER ↓ASR transcription errors relative to the requested words or characters.SpeechTarget transcript; ASR transcript of the edited output.arXivJiWER
ASR
Emotion classification accuracy ↑Agreement between the predicted emotion and the requested emotion label.SpeechTarget emotion label; an emotion classifier with a matching label set.arXivemotion2vec
CLAP audio–text similarity ↑Cosine similarity between output audio and the desired audio description.Music; AudioCaption describing the desired result; a specified CLAP checkpoint.arXivCLAP
Event Occurrence Score (EOS) ↑Minimum event-level CLAP similarity after text-guided source separation; checks coverage of requested events.AudioDesired event descriptions; event decomposition and separated event tracks.Paper

Preservation and Locality

For local edits, compare the regions or sources that should remain unchanged.

MetricWhat it measuresAudio ModalitiesReference / InputsPaper / StandardCode / Model
Speaker embedding cosine similarity ↑Retention of speaker identity in the edited speech.SpeechSource speaker audio; the same speaker-verification encoder for both recordings.arXivHF Model
Multi-resolution STFT distance ↓Spectral convergence and log-magnitude differences across several time–frequency resolutions.Speech; Music; AudioAligned source/output audio from the non-edited regions.arXivauraloss
CLAP audio–audio similarity ↑Semantic similarity between source and edited audio embeddings; a broad preservation proxy.Music; AudioSource audio; matching non-target regions or stems for local comparison.arXivCLAP
LPAPS distance ↓Perceptual distance between audio representations from a pretrained feature network.Music; AudioSource audio; aligned non-edited regions for local comparison.arXivLPAPS

Temporal and Structural Consistency

MetricWhat it measuresAudio ModalitiesReference / InputsPaper / StandardCode / Model
Boundary error ↓Mean or median absolute timing error of predicted speech-segment boundaries.SpeechManual boundary annotations and predicted boundaries, in the same time unit.Paper
Word-level Dynamic Time Warping (WDTW) ↓Length-normalized DTW distance over matched word segments in source and edited speech.SpeechSource and edited speech; both transcripts and word-level forced alignments.arXiv
Melody accuracy ↑Frame-wise agreement of the dominant pitch class between reference and edited music.MusicReference melody/audio; aligned pitch-class sequences.arXivEditGen
F0 Pearson correlation ↑Correlation between reference and output vocal-pitch contours.Music (vocals)Reference vocal audio; aligned F0 contours extracted with the same model.arXivPitch extractor
Pitch extractor
Chroma similarity / Chroma DTW similarity ↑Pitch-class distribution similarity, or frame-wise similarity after DTW alignment.MusicSource/reference music; consistently extracted chromagrams.arXivMuseCPEval
Beat F1 ↑Precision–recall balance of matching beat timestamps within a 70 ms tolerance.MusicReference and output beat timestamps.arXivMuseCPEval
Dynamics correlation ↑Frame-wise Pearson correlation of reference and output loudness trajectories.MusicReference dynamics/audio; aligned loudness trajectories.arXivEditGen
Structural pairwise F-measure / ARI ↑Agreement of musical section assignments, with ARI correcting for chance agreement.MusicSource/reference and output segmentations in a shared time frame.arXivMuseCPEval
Event Sequence Score (ESS) ↑Kendall-style rank agreement between the described and detected event order.AudioDesired event ordering; onset estimates from separated event tracks.Paper

Audio Quality and Naturalness

MetricWhat it measuresAudio ModalitiesReference / InputsPaper / StandardCode / Model
MOS / CMOS ↑Human ratings of output quality or comparative quality against another recording.Speech; Music; AudioListeners and a task-specific rating protocol; comparison audio for CMOS.StandardSpeech listening tests
Speech listening tests
MOSNet predicted MOS ↑Automatic prediction of speech naturalness ratings, developed for voice conversion.SpeechEdited speech.arXivMOSNet
UTMOSv2 predicted MOS ↑Predicted naturalness MOS, developed for high-quality synthetic speech.SpeechEdited speech.arXivGitHub
HF Model
SpeechJudge-GRM (pairwise)Paired naturalness ratings and preference, accompanied by a generated explanation.SpeechTarget transcript and two candidate speech recordings for the same text.arXivGitHub
HF Model
DNSMOS P.835 ↑Predicted speech-signal, background-noise and overall quality scores.SpeechEdited speech.arXivDNSMOS
NISQA ↑Predicted overall speech quality and degradation dimensions; NISQA-TTS targets synthetic-speech naturalness.SpeechEdited speech; the appropriate NISQA checkpoint.arXivNISQA
PAM ↑No-reference audio quality from an audio–language model using contrasting positive and negative quality prompts.Speech; Music; AudioEdited audio; fixed quality prompts and the PAM implementation's MS-CLAP backbone.arXivGitHub
PESQ ↑Reference-based perceptual speech quality after degradation or restoration.SpeechCorresponding clean target speech; 8 kHz narrowband or 16 kHz wideband mode.StandardPESQ
STOI ↑Estimated intelligibility of degraded or enhanced speech.SpeechTime-aligned clean target speech.Paperpystoi
SI-SDR ↑Target-signal reconstruction fidelity after compensating for a global scale difference.Speech; Music; AudioTime-aligned target waveform or isolated target source.arXivTorchMetrics
NOMAD distance ↓Perceptual speech degradation measured in a learned embedding space.SpeechClean speech references; matching linguistic content is not required.arXivNOMAD
SpeechBERTScore ↑Reference-aware speech quality proxy using greedy matching of self-supervised speech features.SpeechNatural reference speech; a fixed encoder, layer and precision/recall/F1 variant.arXivGitHub
Fréchet Audio Distance (FAD) ↓Distance between output and reference audio-embedding distributions.Music; AudioReference audio collection; the same embedding backbone and preprocessing.arXivFADtk

Multi-dimensional Evaluators

Reusable models and toolkits for multi-dimensional assessment of editing results and audio aesthetics.

EvaluatorAudio ModalitiesEvaluation DimensionsReference / InputsPaperCode / Model
AuditEval (SSL / LLM)AudioQuality, editing relevance and faithfulness to the source.Source and edited audio; original and target descriptions.arXivGitHub
ModelScope
MuseCPEvalMusicHarmony, rhythm, structure and melody preservation, with additional timbre metrics in the toolkit.Source and edited music; selected musical attributes to preserve.arXivGitHub
MMAE rubric evaluator (Qwen3-Omni)Speech; Music; AudioInstruction Following Rate (IFR), Consistency Rate (CR) and Exact Match Rate (EMR).Source/output audio, editing instructions and sample-specific MMAE rubrics.arXivGitHub
Audiobox AestheticsSpeech; Music; AudioContent Enjoyment (CE), Content Usefulness (CU), Production Complexity (PC) and Production Quality (PQ).Edited audio only.arXivGitHub
HF Model
SongEval scoring modelMusic (songs)Overall coherence, memorability, vocal breathing/phrasing naturalness, structural clarity and overall musicality.Full-length song audio with vocals and accompaniment.arXivGitHub
Weights
MuseCriticMusic (songs)Coherence, musicality, memorability, structural clarity and vocal naturalness; returns scores and a natural-language critique.Full-length song audio; the released aesthetic rubric.arXivGitHub
HF Model

🔮 Challenges and Future Directions

Foundation-model-based audio editing still faces several system-level challenges:

  1. Complex editing.
    Real-world audio entangles semantic events, speaker identity, acoustic attributes, background ambience, rhythm, spatial cues, and reverberation. Future systems should support precise source/event localization, attribute-level modification, and reliable preservation of non-target content across speech, music, and general audio.
  2. Robustness in open-domain settings.
    Editing models should remain reliable under noise, reverberation, overlapping sources, long-form context, and ambiguous instructions. Better instruction grounding, long-context modeling, iterative refinement, and self-verification are important directions toward robust real-world editing.
  3. Faithful and editing-specific evaluation.
    Existing evaluation often mixes generation quality with editing quality. Future benchmarks should explicitly annotate edit targets, operations, preservation regions, and relevant control signals, allowing edit success and non-target preservation to be evaluated separately.
  4. Safety, copyright, and misuse prevention.
    Modern editing systems can realistically modify speech content, speaker identity, emotion, environmental sounds, and music. Practical deployment therefore requires complementary mechanisms for provenance, watermarking, manipulated-audio detection, and responsible data licensing.

Citation

If You find this survey or repository useful, please cite our paper:

@article{pan2026audio,
  title={Audio Editing in the Era of Foundation Models: A Survey},
  author={Pan, Changhao and Fan, Yifei and Zhuo, Fan and Chen, Yifu and Guo, Wenxiang and Zhang, Yu and Li, Ruiqi and Zhu, Zhiyuan and Yang, Rui and Ji, Shengpeng and others},
  journal={arXiv preprint arXiv:2606.23139},
  year={2026}
}

Contributing

This repo is meant to keep growing. If an audio editing model, dataset, or benchmark is missing, please feel free to open an issue or a pull request.


📄 License

Unless otherwise noted below, original content created for this repository is licensed under the MIT License.

The survey paper and content reproduced or adapted from it, including assets/taxonomy_overview.png, assets/train-based.png, and assets/train-free.png, remain under CC BY-NC-SA 4.0. The MIT license does not relicense these materials.

Linked third-party papers, code, models, model weights, datasets, and tools are governed by their respective licenses.

aacl-2026
audio-editing
awesome-list
foundation-models
survey