A curated list of models, benchmarks, tools and guides for audio editing
46
39 commits
updated Sep 19, 2026
A curated list of models, benchmarks, tools and guides for audio editing
Welcome to PR if you want to add some resources.
| Date | Title | Relevant Resources |
|---|---|---|
| 2027 | Audio-Editing-Challenge (ICASSP 2027) | GitHub/Challenge Website |
| Date | Title | Relevant Resources |
|---|---|---|
| 2026-06 | MMAE: A Massive Multitask Audio Editing Benchmark | arXiv/Github/HuggingFace |
| Date | Title | Relevant Resources |
|---|---|---|
| 2026-06 | SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing | arXiv/Github/HuggingFace |
| 2025-11 | Ming-Freeform-Audio-Edit | arXiv/Github/HuggingFace |
| 2025-11 | Step-Audio-Edit-Benchmark | arXiv/Github |
| 2025-09 | ISSE: An Instruction-Guided Speech Style Editing Dataset and Benchmark | arXiv/HuggingFace/Demo |
| Date | Title | Relevant Resources |
|---|---|---|
| 2026-03 | LyricEditBench: The First Benchmark for Melody-Preserving Lyric Modification Evaluation | arXiv/Github/HuggingFace |
| 2025-12 | Evaluating Music Context Preservation: A Multi-facet Framework for Music Editing Systems | arXiv/Github |
| Date | Title | Relevant Resources |
|---|---|---|
| 2026-09 | AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing | arXiv/Github/HuggingFace/ModelScope/Demo |
| 2026-08 | FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation | arXiv/Github/HuggingFace/ModelScope/Demo |
| 2026-08 | FireRedTTS3: Unified Speech Generation and Editing with Semantically Enriched Speech Representations | arXiv/Github/HuggingFace/Demo |
| 2026-08 | VoiceDesigner: Text-to-Voice Generation and Editing via Unified Diffusion Modeling and Data Augmentation | arXiv/Demo |
| 2026-08 | Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing | arXiv |
| 2026-08 | dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model | arXiv/Demo |
| 2026-06 | UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling | arXiv/Demo |
| 2026-05 | CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS | arXiv/Demo |
| 2026-01 | CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models | arXiv/Github/HuggingFace/Demo |
| 2025-12 | MiMo-Audio: Audio Language Models are Few-Shot Learners | arXiv/Github/HuggingFace |
| 2025-11 | Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation | arXiv/Github/HuggingFace/Demo |
| 2025-11 | Step-Audio-EditX Technical Report | arXiv/Github/HuggingFace/Demo |
| 2025-11 | VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing | arXiv/Github |
| 2024-09 | SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis | arXiv/Github/HuggingFace |
| 2024-07 | Speech Editing – a Summary | arXiv |
| 2024-05 | InstructSpeech: Following Speech Editing Instructions via Large Language Models | Paper/Demo |
| 2024-03 | VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild | arXiv/Github/Demo |
| 2023-05 | FluentSpeech: Stutter-Oriented Automatic Speech Editing with Context-Aware Diffusion Models | arXiv/Github |
| Date | Title | Relevant Resources |
|---|---|---|
| 2026-09 | One-Stage Multi-Task Instruction-Guided 3D Spatial Audio Editing | arXiv/Github/Demo |
| 2026-06 | Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption | arXiv/Github/Demo |
| 2026-06 | DirectAudioEdit: Inversion-Free Text-Guided Audio Editing via Diffusion Prediction Contrast | arXiv/Demo |
| 2026-05 | UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion | arXiv/Github/Demo |
| 2026-05 | SpongeBob: Sync-Aware Harmonious Audio-Visual Generative Editing | arXiv/Github/Demo |
| 2026-04 | Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing | arXiv/Github/HuggingFace/Demo |
| 2026-04 | VoiceToInstrument: Convert voice to instrumental tracks using AI | HomePage/Blog |
| 2026-02 | Bagpiper: Solving Open-Ended Audio Tasks via Rich Captions | arXiv/Demo |
| 2026-02 | AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing | arXiv/Demo |
| 2025-12 | MMEDIT: A Unified Framework for Multi-Type Audio Editing via Audio Language Model | arXiv/Github/HuggingFace/Demo |
| 2025-10 | SAO-Instruct: Free-form Audio Editing using Natural Language Instructions | arXiv/Github/Demo |
| 2025-10 | UALM: Unified Audio Language Model for Understanding, Generation and Reasoning | arXiv |
| 2025-09 | Guiding Audio Editing with Audio Language Model | arXiv/Github |
| 2025-09 | Recomposer: Event-roll-guided generative audio editing | arXiv |
| 2025-06 | ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing | arXiv/Github/Demo |
| 2025-05 | AudioMorphix: Training-free audio editing with diffusion probabilistic models | arXiv/Github/Demo |
| 2024-09 | AudioEditor: A Training-Free Diffusion-Based Audio Editing Framework | arXiv/Github/Demo |
| 2024-06 | Prompt-guided Precise Audio Editing with Diffusion Models | arXiv |
| 2024-03 | WavCraft: Audio Editing and Generation with Large Language Models | arXiv/Github/Demo |
| 2024-02 | Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion | arXiv/Github/Demo |
| 2023-04 | AUDIT: Audio Editing by Following Instructions with Latent Diffusion Models | arXiv/Demo |
| Date | Title | Relevant Resources |
|---|---|---|
| 2026-09 | AURA: Unified Multimodal Framework for Conversational Music Editing | arXiv/Github/HuggingFace/Demo |
| 2026-08 | CLASVS: Continuous-Latent Autoregression for Melody-Preserving Lyric Editing in Singing Voice Synthesis | arXiv/Demo |
| 2026-08 | P-MUSE: Prompt-MIDI-Optional Model for Unified Instrumental Music Synthesis and Editing | arXiv/Github/Demo |
| 2026-07 | RIME: Enabling Large-Scale Agentic Music Post-Production | arXiv |
| 2026-02 | ACE-Step 1.5: Pushing the Boundaries of Open-Source Music Generation | arXiv/Demo |
| 2025-11 | Melodia: Training-Free Music Editing Guided by Attention Probing in Diffusion Models | arXiv |
| 2025-11 | MusRec: Zero-Shot Text-to-Music Editing via Rectified Flow and Diffusion Transformers | arXiv |
| 2025-06 | ACE-Step: A Step Towards Music Generation Foundation Model | arXiv/Github/HuggingFace |
| 2025-04 | SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing | arXiv/Github/Demo |
| 2024-12 | SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor | arXiv/Demo |
| 2024-07 | High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching (MelodyFlow) | arXiv/Demo/HuggingFace |
| 2024-05 | Instruct-MusicGen: Unlocking Text-to-Music Editing for Music Language Models via Instruction Tuning | arXiv/Github/Demo |
| 2024-05 | DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation | arXiv/Demo |
| 2024-02 | MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models | arXiv/Github/Demo |
| 2024-01 | DITTO: Diffusion Inference-Time T-Optimization for Music Generation | arXiv/Demo |
| 2023-10 | Loop Copilot: Conducting AI Ensembles for Music Generation and Iterative Editing | arXiv/Github |
A curated list of models, benchmarks, tools and guides for audio editing
46
39 commits
updated Sep 19, 2026
A curated list of models, benchmarks, tools and guides for audio editing
Welcome to PR if you want to add some resources.
| Date | Title | Relevant Resources |
|---|---|---|
| 2027 | Audio-Editing-Challenge (ICASSP 2027) | GitHub/Challenge Website |
| Date | Title | Relevant Resources |
|---|---|---|
| 2026-06 | MMAE: A Massive Multitask Audio Editing Benchmark | arXiv/Github/HuggingFace |
| Date | Title | Relevant Resources |
|---|---|---|
| 2026-06 | SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing | arXiv/Github/HuggingFace |
| 2025-11 | Ming-Freeform-Audio-Edit | arXiv/Github/HuggingFace |
| 2025-11 | Step-Audio-Edit-Benchmark | arXiv/Github |
| 2025-09 | ISSE: An Instruction-Guided Speech Style Editing Dataset and Benchmark | arXiv/HuggingFace/Demo |
| Date | Title | Relevant Resources |
|---|---|---|
| 2026-03 | LyricEditBench: The First Benchmark for Melody-Preserving Lyric Modification Evaluation | arXiv/Github/HuggingFace |
| 2025-12 | Evaluating Music Context Preservation: A Multi-facet Framework for Music Editing Systems | arXiv/Github |
| Date | Title | Relevant Resources |
|---|---|---|
| 2026-09 | AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing | arXiv/Github/HuggingFace/ModelScope/Demo |
| 2026-08 | FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation | arXiv/Github/HuggingFace/ModelScope/Demo |
| 2026-08 | FireRedTTS3: Unified Speech Generation and Editing with Semantically Enriched Speech Representations | arXiv/Github/HuggingFace/Demo |
| 2026-08 | VoiceDesigner: Text-to-Voice Generation and Editing via Unified Diffusion Modeling and Data Augmentation | arXiv/Demo |
| 2026-08 | Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing | arXiv |
| 2026-08 | dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model | arXiv/Demo |
| 2026-06 | UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling | arXiv/Demo |
| 2026-05 | CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS | arXiv/Demo |
| 2026-01 | CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models | arXiv/Github/HuggingFace/Demo |
| 2025-12 | MiMo-Audio: Audio Language Models are Few-Shot Learners | arXiv/Github/HuggingFace |
| 2025-11 | Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation | arXiv/Github/HuggingFace/Demo |
| 2025-11 | Step-Audio-EditX Technical Report | arXiv/Github/HuggingFace/Demo |
| 2025-11 | VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing | arXiv/Github |
| 2024-09 | SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis | arXiv/Github/HuggingFace |
| 2024-07 | Speech Editing – a Summary | arXiv |
| 2024-05 | InstructSpeech: Following Speech Editing Instructions via Large Language Models | Paper/Demo |
| 2024-03 | VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild | arXiv/Github/Demo |
| 2023-05 | FluentSpeech: Stutter-Oriented Automatic Speech Editing with Context-Aware Diffusion Models | arXiv/Github |
| Date | Title | Relevant Resources |
|---|---|---|
| 2026-09 | One-Stage Multi-Task Instruction-Guided 3D Spatial Audio Editing | arXiv/Github/Demo |
| 2026-06 | Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption | arXiv/Github/Demo |
| 2026-06 | DirectAudioEdit: Inversion-Free Text-Guided Audio Editing via Diffusion Prediction Contrast | arXiv/Demo |
| 2026-05 | UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion | arXiv/Github/Demo |
| 2026-05 | SpongeBob: Sync-Aware Harmonious Audio-Visual Generative Editing | arXiv/Github/Demo |
| 2026-04 | Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing | arXiv/Github/HuggingFace/Demo |
| 2026-04 | VoiceToInstrument: Convert voice to instrumental tracks using AI | HomePage/Blog |
| 2026-02 | Bagpiper: Solving Open-Ended Audio Tasks via Rich Captions | arXiv/Demo |
| 2026-02 | AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing | arXiv/Demo |
| 2025-12 | MMEDIT: A Unified Framework for Multi-Type Audio Editing via Audio Language Model | arXiv/Github/HuggingFace/Demo |
| 2025-10 | SAO-Instruct: Free-form Audio Editing using Natural Language Instructions | arXiv/Github/Demo |
| 2025-10 | UALM: Unified Audio Language Model for Understanding, Generation and Reasoning | arXiv |
| 2025-09 | Guiding Audio Editing with Audio Language Model | arXiv/Github |
| 2025-09 | Recomposer: Event-roll-guided generative audio editing | arXiv |
| 2025-06 | ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing | arXiv/Github/Demo |
| 2025-05 | AudioMorphix: Training-free audio editing with diffusion probabilistic models | arXiv/Github/Demo |
| 2024-09 | AudioEditor: A Training-Free Diffusion-Based Audio Editing Framework | arXiv/Github/Demo |
| 2024-06 | Prompt-guided Precise Audio Editing with Diffusion Models | arXiv |
| 2024-03 | WavCraft: Audio Editing and Generation with Large Language Models | arXiv/Github/Demo |
| 2024-02 | Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion | arXiv/Github/Demo |
| 2023-04 | AUDIT: Audio Editing by Following Instructions with Latent Diffusion Models | arXiv/Demo |
| Date | Title | Relevant Resources |
|---|---|---|
| 2026-09 | AURA: Unified Multimodal Framework for Conversational Music Editing | arXiv/Github/HuggingFace/Demo |
| 2026-08 | CLASVS: Continuous-Latent Autoregression for Melody-Preserving Lyric Editing in Singing Voice Synthesis | arXiv/Demo |
| 2026-08 | P-MUSE: Prompt-MIDI-Optional Model for Unified Instrumental Music Synthesis and Editing | arXiv/Github/Demo |
| 2026-07 | RIME: Enabling Large-Scale Agentic Music Post-Production | arXiv |
| 2026-02 | ACE-Step 1.5: Pushing the Boundaries of Open-Source Music Generation | arXiv/Demo |
| 2025-11 | Melodia: Training-Free Music Editing Guided by Attention Probing in Diffusion Models | arXiv |
| 2025-11 | MusRec: Zero-Shot Text-to-Music Editing via Rectified Flow and Diffusion Transformers | arXiv |
| 2025-06 | ACE-Step: A Step Towards Music Generation Foundation Model | arXiv/Github/HuggingFace |
| 2025-04 | SteerMusic: Enhanced Musical Consistency for Zero-shot Text-guided and Personalized Music Editing | arXiv/Github/Demo |
| 2024-12 | SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor | arXiv/Demo |
| 2024-07 | High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching (MelodyFlow) | arXiv/Demo/HuggingFace |
| 2024-05 | Instruct-MusicGen: Unlocking Text-to-Music Editing for Music Language Models via Instruction Tuning | arXiv/Github/Demo |
| 2024-05 | DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation | arXiv/Demo |
| 2024-02 | MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models | arXiv/Github/Demo |
| 2024-01 | DITTO: Diffusion Inference-Time T-Optimization for Music Generation | arXiv/Demo |
| 2023-10 | Loop Copilot: Conducting AI Ensembles for Music Generation and Iterative Editing | arXiv/Github |