Locality-aware continual unlearning (LACU), a framework that stabilizes sequential concept removal
Efficiently estimate which training data groups influenced diffusion model outputs
Theoretically analyzes the MaskGIT sampler, poviding a choose-then-sample (CTS) formulation
Training flow-map models in RAE latent with consistency mid-training for trajectory-aware initialization
A framework for Identify which training examples influenced specific concepts within the diffusion model
CMT reduces the training cost of diffusion-based flow map models by up to 90% while reaching SOTA performance
Improved object-centric diffusion learning with registers and contrastive alignment
An improved mechanism for applying classifier-free guidance in discrete diffusion
Learning conditional, unconditional, and matching-aware discriminator with adaptive weighting mechanism (cSAN)
Leveraging discrete diffusion models as priors for inverse problems
Propose tensor-decomposition-based PEFT method, showing its effectiveness on T-to-I generation tasks
Theoretical analysis of limitation of current discrete diffusion and a method for effectively capturing element-wise dependency
Classifier-Free Guidance inside the Attraction Basin May Cause Memorization
A general method to find an optimal sampling schedule for inference in discrete diffusion
A method efficiently leverages online human feedback to fine-tune Stable Diffusion for various range of tasks
An enhanced multimodal representation using weighted point clouds and its theoretical benefits
A 64x64 pre-trained diffusion model is all you need for 1-step high-resolution SOTA generation
Unified framework enables diverse samplers and 1-step generation SOTAs
Applications:
[SoundGen]
Enhancing GAN with metrizable discriminators
Applications:
[Vocoder]
Fast, Efficient, Training-Free, and Controllable diffusion-based generation method
Achieving blind inversion using DDPM
Applications:
[DeReverb]
[SpeechEnhance]
<div class="tile">
<h3>MCA</h3>
<a href=""><img src="./assets/mca.png"></a>
<h5>
[EMNLP]
<a href="https://arxiv.org/abs/2510.15543">[arXiv]</a>
[code]
</h5>
<p>MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval</p>
<div class="tile_venue">EMNLP26</div>
</div>
<div class="tile">
<h3>MLLMCLIP</h3>
<a href=""><img src="./assets/mllmclip.png"></a>
<h5>
[EMNLP]
[arXiv]
[code]
</h5>
<p>MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations</p>
<div class="tile_venue">EMNLP26</div>
</div>
<div class="tile">
<h3>Syn-Omni</h3>
<a href=""><img src="./assets/synomni.png"></a>
<h5>
[EMNLP]
[arXiv]
[code]
</h5>
<p>Syn-Omni: Structured Specialization and Progressive Collaboration for Omnimodal Embeddings</p>
<div class="tile_venue">EMNLP26</div>
</div>
<div class="tile">
<h3>DynaVieW</h3>
<a href=""><img src="./assets/dynaview.png"></a>
<h5>
<a href="https://icml.cc/virtual/2026/poster/65080">[ICML]</a>
<a href="https://arxiv.org/abs/2607.04112">[arXiv]</a>
<a href="https://github.com/Silin159/DynaVieW">[code]</a>
</h5>
<p>DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics</p>
<div class="tile_venue">ICML26</div>
</div>
<div class="tile">
<h3>LRPO</h3>
<a href=""><img src="./assets/lrpo.png"></a>
<h5>
<a href="https://icml.cc/virtual/2026/poster/66718">[ICML]</a>
<a href="https://arxiv.org/abs/2605.25360">[arXiv]</a>
<a href="https://github.com/Guochry/LRPO">[code]</a>
</h5>
<p>Learning to Route Languages for Multilingual Policy Optimization</p>
<div class="tile_venue">ICML26</div>
</div>
<div class="tile">
<h3>DeepResonance</h3>
<a href=""><img src="./assets/deepresonance.png"></a>
<h5>
<a href="https://aclanthology.org/2025.emnlp-main.653/">[EMNLP]</a>
<a href="https://arxiv.org/abs/2502.12623">[arXiv]</a>
<a href="https://github.com/sony/DeepResonance">[code]</a>
</h5>
<p>DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning</p>
<div class="tile_venue">EMNLP25</div>
</div>
<div class="tile">
<h3>CARE</h3>
<a href=""><img src="./assets/care.png"></a>
<h5>
<a href="https://aclanthology.org/2025.emnlp-main.1669/">[EMNLP]</a>
<a href="https://arxiv.org/abs/2504.05154">[arXiv]</a>
<a href="https://github.com/Guochry/CARE">[data]</a>
</h5>
<p>CARE: Assessing the Impact of Multilingual Human Preference Learning on Cultural Awareness</p>
<div class="tile_venue">EMNLP25</div>
</div>
<div class="tile">
<h3>BiAug</h3>
<a href=""><img src="./assets/biaug.png"></a>
<h5>
[MRR@ICCV25]
<a href="https://arxiv.org/abs/2310.01330">[arXiv]</a>
</h5>
<p>Towards reporting bias in visual-language datasets: bimodal augmentation by decoupling object-attribute association</p>
<div class="tile_venue">ICCV25 MRR Workshop</div>
</div>
<div class="tile">
<h3>GLOV</h3>
<a href=""><img src="./assets/glov.png"></a>
<h5>
[TMLR]
<a href="https://arxiv.org/abs/2410.06154">[arXiv]</a>
</h5>
<p>GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models</p>
<div class="tile_venue">TMLR</div>
</div>
<div class="tile">
<h3>Music-to-MVD</h3>
<a href=""><img src="./assets/mvd.png"></a>
<h5>
<a href="https://aclanthology.org/2025.repl4nlp-1.4.pdf">[RepL4NLP@NAACL25]</a>
<a href="https://arxiv.org/abs/2503.11190">[arXiv]</a>
</h5>
<p>Cross-Modal Learning for Music-to-Music-Video Description Generation</p>
<div class="tile_venue">NAACL25 RepL4NLP Workshop</div>
</div>
<div class="tile">
<h3>VinaBench</h3>
<a href=""><img src="./assets/vinabench.png"></a>
<h5>
[CVPR]
<a href="https://arxiv.org/abs/2503.20871">[arXiv]</a>
<a href="https://silin159.github.io/Vina-Bench/">[data]</a>
</h5>
<p>VinaBench: Benchmark for Faithful and Consistent Visual Narratives</p>
<div class="tile_venue">CVPR25</div>
</div>
<div class="tile">
<h3>OpenMU</h3>
<a href=""><img src="./assets/openmu.png"></a>
<h5>
<a href="https://arxiv.org/abs/2410.15573">[arXiv]</a>
<a href="https://huggingface.co/datasets/Sony/OpenMU-Bench">[data]</a>
<a href="https://mzhaojp22.github.io/open_music_understanding/">[demo]</a>
</h5>
<p>OpenMU: Your Swiss Army Knife for Music Understanding</p>
<div class="tile_venue">ISMIR2024 Late Breaking Demos</div>
</div>
<div class="tile">
<h3>DiffuCOMET</h3>
<a href="https://arxiv.org/abs/2402.17011"><img src="./assets/diffcomet.png"></a>
<h5>
<a href="https://aclanthology.org/2024.acl-long.264/">[ACL]</a>
<a href="https://arxiv.org/abs/2402.17011">[arXiv]</a>
<a href="https://github.com/Silin159/DiffuCOMET">[code]</a>
</h5>
<p>DiffuCOMET: Contextual Commonsense Knowledge Diffusion</p>
<div class="tile_venue">ACL24</div>
</div>
<div class="tile">
<h3>CyCLIPs/CyCLAPs</h3>
<a href="https://arxiv.org/abs/2310.13267"><img src="./assets/cyclips.png"></a>
<h5>
<a href="https://aclanthology.org/2024.findings-acl.293/">[ACL]</a>
<a href="https://arxiv.org/abs/2310.13267">[arXiv]</a>
</h5>
<p>On the Language Encoder of Contrastive Cross-modal Models</p>
<div class="tile_venue">ACL24</div>
</div>
<div class="tile">
<h3>DIIR</h3>
<a href="https://arxiv.org/abs/2403.15737"><img src="./assets/diir.png"></a>
<h5>
<a href="https://aclanthology.org/2024.findings-acl.782/">[ACL]</a>
<a href="https://arxiv.org/abs/2403.15737">[arXiv]</a>
<a href="https://github.com/zhouhanxie/DIIR">[code]</a>
</h5>
<p>Few-shot Dialogue Strategy Learning for Motivational Interviewing via Inductive Reasoning</p>
<div class="tile_venue">ACL24</div>
</div>
<div class="tile">
<h3>PeaCok</h3>
<a href="https://arxiv.org/abs/2305.02364"><img src="./assets/personas.png"></a>
<h5>
<a href="https://aclanthology.org/2023.acl-long.362/">[ACL]</a>
<a href="https://arxiv.org/abs/2305.02364">[arXiv]</a>
<a href="https://github.com/Silin159/PeaCoK">[code]</a>
</h5>
<p>PeaCoK: Persona Commonsense Knowledge for Consistent and Engaging Narratives<br>(Outstanding Paper Award)</p>
<div class="tile_venue">ACL23</div>
</div>
<div class="tile">
<h3>ComFact</h3>
<a href="https://aclanthology.org/2022.findings-emnlp.120/"><img src="./assets/comfact.png"></a>
<h5>
<a href="https://aclanthology.org/2022.findings-emnlp.120/">[EMNLP]</a>
<a href="https://arxiv.org/abs/2210.12678">[arXiv]</a>
<a href="https://github.com/epfl-nlp/ComFact">[code]</a>
</h5>
<p>ComFact: A Benchmark for Linking Contextual Commonsense Knowledge</p>
<div class="tile_venue">EMNLP22 Findings</div>
</div>
<div class="tile" style="background-color: white;"></div>
<div class="tile" style="background-color: white;"></div>
Automatic music mixing using a generative model of effect embeddings
Automatic Music Sample Identification with Multi-Track Contrastive Learning
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
Large-Scale Training Data Attribution for Music Generative Models via Unlearning
Reductive, exclusionary, normalising: The limits of generative AI music
Can Large Language Models Predict Audio Effects Parameters from Natural Language?
Inference-Time Optimisation for Vocal Effects Style Transfer using DiffVox
SOTA Fx representation: Extract instrument-wise audio effects representations from music mixtures
Inference Time Optimization for Music Mastering Style Transfer
Reverse Engineering of Music Mixing Graphs with Differentiable Processors and Iterative Pruning
Supervised contrastive learning from weakly-labeled audio segments for musical version matching
Music Foundation Model as Generic Booster for Music Downstream Tasks
DiffVox: A Differentiable Model for Capturing and Analysing Professional Effects Distributions
VRVQ: Variable Bitrate Residual Vector Quantization for Audio Compression
Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer
Searching For Music Mixing Graphs: A Pruning Approach
Improving Unsupervised Clean-to-Rendered Guitar Tone Transformation Using GANs and Integrated Unaligned Clean Data
MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models
Timbre-Trap: A Low-Resource Framework for Instrument-Agnostic Music Transcription
VRDMG: Vocal Restoration via Diffusion Posterior Sampling with Multiple Guidance
Automatic Piano Transcription with Hierarchical Frequency-Time Transformer
An Attention-based Approach To Hierarchical Multi-label Music Instrument Classification
Unsupervised Vocal Dereverberation with Diffusion-based Generative Models
Music Mixing Style Transfer: A Contrastive Learning Approach to Disentangle Audio Effects
DiffRoll: Diffusion-based Generative Music Transcription with Unsupervised Pretraining Capability
Hierarchical Diffusion Models for Singing Voice Neural Vocoder
Distortion Audio Effects: Learning How to Recover the Clean Signal
Automatic Music Mixing with Deep Learning and Out-of-Domain Data
Odoriko: A Shape-Aware Multimodal Diffusion Framework for Human Motion
Spectral Prior for Reducing Exposure Bias in Diffusion Models
Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance
PAVAS: a framework for generating physically plausible audio from video, by integrating physics estimation
Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models
VIRTUE: Visual-Interactive Text-Image Universal Embedder
A diffusion-based post-processor for perceptually improving speech enhancement and separation outputs
CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation
TITAN-Guide: Taming Inference-Time AligNment for Guided Text-to-Video Diffusion Models
MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation
SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation
Mining Your Own Secrets: Diffusion Classifier Scores for Continual Personalization of Text-to-Image Diffusion Models
GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping
SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
Zero- and Few-shot Sound Event Localization and Detection
STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events
Extending Audio Masked Autoencoders Toward Audio Restoration
Diffiner: A Versatile Diffusion-based Generative Refiner for Speech Enhancement
CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled Videos
Semantic Acoustic Imaging for Sound Event Localization and Detection from Spatial Audio and Audiovisual Scenes
CSS
60.4%
JavaScript
39.4%
Locality-aware continual unlearning (LACU), a framework that stabilizes sequential concept removal
Efficiently estimate which training data groups influenced diffusion model outputs
Theoretically analyzes the MaskGIT sampler, poviding a choose-then-sample (CTS) formulation
Training flow-map models in RAE latent with consistency mid-training for trajectory-aware initialization
A framework for Identify which training examples influenced specific concepts within the diffusion model
CMT reduces the training cost of diffusion-based flow map models by up to 90% while reaching SOTA performance
Improved object-centric diffusion learning with registers and contrastive alignment
An improved mechanism for applying classifier-free guidance in discrete diffusion
Learning conditional, unconditional, and matching-aware discriminator with adaptive weighting mechanism (cSAN)
Leveraging discrete diffusion models as priors for inverse problems
Propose tensor-decomposition-based PEFT method, showing its effectiveness on T-to-I generation tasks
Theoretical analysis of limitation of current discrete diffusion and a method for effectively capturing element-wise dependency
Classifier-Free Guidance inside the Attraction Basin May Cause Memorization
A general method to find an optimal sampling schedule for inference in discrete diffusion
A method efficiently leverages online human feedback to fine-tune Stable Diffusion for various range of tasks
An enhanced multimodal representation using weighted point clouds and its theoretical benefits
A 64x64 pre-trained diffusion model is all you need for 1-step high-resolution SOTA generation
Unified framework enables diverse samplers and 1-step generation SOTAs
Applications:
[SoundGen]
Enhancing GAN with metrizable discriminators
Applications:
[Vocoder]
Fast, Efficient, Training-Free, and Controllable diffusion-based generation method
Achieving blind inversion using DDPM
Applications:
[DeReverb]
[SpeechEnhance]
<div class="tile">
<h3>MCA</h3>
<a href=""><img src="./assets/mca.png"></a>
<h5>
[EMNLP]
<a href="https://arxiv.org/abs/2510.15543">[arXiv]</a>
[code]
</h5>
<p>MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval</p>
<div class="tile_venue">EMNLP26</div>
</div>
<div class="tile">
<h3>MLLMCLIP</h3>
<a href=""><img src="./assets/mllmclip.png"></a>
<h5>
[EMNLP]
[arXiv]
[code]
</h5>
<p>MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations</p>
<div class="tile_venue">EMNLP26</div>
</div>
<div class="tile">
<h3>Syn-Omni</h3>
<a href=""><img src="./assets/synomni.png"></a>
<h5>
[EMNLP]
[arXiv]
[code]
</h5>
<p>Syn-Omni: Structured Specialization and Progressive Collaboration for Omnimodal Embeddings</p>
<div class="tile_venue">EMNLP26</div>
</div>
<div class="tile">
<h3>DynaVieW</h3>
<a href=""><img src="./assets/dynaview.png"></a>
<h5>
<a href="https://icml.cc/virtual/2026/poster/65080">[ICML]</a>
<a href="https://arxiv.org/abs/2607.04112">[arXiv]</a>
<a href="https://github.com/Silin159/DynaVieW">[code]</a>
</h5>
<p>DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics</p>
<div class="tile_venue">ICML26</div>
</div>
<div class="tile">
<h3>LRPO</h3>
<a href=""><img src="./assets/lrpo.png"></a>
<h5>
<a href="https://icml.cc/virtual/2026/poster/66718">[ICML]</a>
<a href="https://arxiv.org/abs/2605.25360">[arXiv]</a>
<a href="https://github.com/Guochry/LRPO">[code]</a>
</h5>
<p>Learning to Route Languages for Multilingual Policy Optimization</p>
<div class="tile_venue">ICML26</div>
</div>
<div class="tile">
<h3>DeepResonance</h3>
<a href=""><img src="./assets/deepresonance.png"></a>
<h5>
<a href="https://aclanthology.org/2025.emnlp-main.653/">[EMNLP]</a>
<a href="https://arxiv.org/abs/2502.12623">[arXiv]</a>
<a href="https://github.com/sony/DeepResonance">[code]</a>
</h5>
<p>DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning</p>
<div class="tile_venue">EMNLP25</div>
</div>
<div class="tile">
<h3>CARE</h3>
<a href=""><img src="./assets/care.png"></a>
<h5>
<a href="https://aclanthology.org/2025.emnlp-main.1669/">[EMNLP]</a>
<a href="https://arxiv.org/abs/2504.05154">[arXiv]</a>
<a href="https://github.com/Guochry/CARE">[data]</a>
</h5>
<p>CARE: Assessing the Impact of Multilingual Human Preference Learning on Cultural Awareness</p>
<div class="tile_venue">EMNLP25</div>
</div>
<div class="tile">
<h3>BiAug</h3>
<a href=""><img src="./assets/biaug.png"></a>
<h5>
[MRR@ICCV25]
<a href="https://arxiv.org/abs/2310.01330">[arXiv]</a>
</h5>
<p>Towards reporting bias in visual-language datasets: bimodal augmentation by decoupling object-attribute association</p>
<div class="tile_venue">ICCV25 MRR Workshop</div>
</div>
<div class="tile">
<h3>GLOV</h3>
<a href=""><img src="./assets/glov.png"></a>
<h5>
[TMLR]
<a href="https://arxiv.org/abs/2410.06154">[arXiv]</a>
</h5>
<p>GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models</p>
<div class="tile_venue">TMLR</div>
</div>
<div class="tile">
<h3>Music-to-MVD</h3>
<a href=""><img src="./assets/mvd.png"></a>
<h5>
<a href="https://aclanthology.org/2025.repl4nlp-1.4.pdf">[RepL4NLP@NAACL25]</a>
<a href="https://arxiv.org/abs/2503.11190">[arXiv]</a>
</h5>
<p>Cross-Modal Learning for Music-to-Music-Video Description Generation</p>
<div class="tile_venue">NAACL25 RepL4NLP Workshop</div>
</div>
<div class="tile">
<h3>VinaBench</h3>
<a href=""><img src="./assets/vinabench.png"></a>
<h5>
[CVPR]
<a href="https://arxiv.org/abs/2503.20871">[arXiv]</a>
<a href="https://silin159.github.io/Vina-Bench/">[data]</a>
</h5>
<p>VinaBench: Benchmark for Faithful and Consistent Visual Narratives</p>
<div class="tile_venue">CVPR25</div>
</div>
<div class="tile">
<h3>OpenMU</h3>
<a href=""><img src="./assets/openmu.png"></a>
<h5>
<a href="https://arxiv.org/abs/2410.15573">[arXiv]</a>
<a href="https://huggingface.co/datasets/Sony/OpenMU-Bench">[data]</a>
<a href="https://mzhaojp22.github.io/open_music_understanding/">[demo]</a>
</h5>
<p>OpenMU: Your Swiss Army Knife for Music Understanding</p>
<div class="tile_venue">ISMIR2024 Late Breaking Demos</div>
</div>
<div class="tile">
<h3>DiffuCOMET</h3>
<a href="https://arxiv.org/abs/2402.17011"><img src="./assets/diffcomet.png"></a>
<h5>
<a href="https://aclanthology.org/2024.acl-long.264/">[ACL]</a>
<a href="https://arxiv.org/abs/2402.17011">[arXiv]</a>
<a href="https://github.com/Silin159/DiffuCOMET">[code]</a>
</h5>
<p>DiffuCOMET: Contextual Commonsense Knowledge Diffusion</p>
<div class="tile_venue">ACL24</div>
</div>
<div class="tile">
<h3>CyCLIPs/CyCLAPs</h3>
<a href="https://arxiv.org/abs/2310.13267"><img src="./assets/cyclips.png"></a>
<h5>
<a href="https://aclanthology.org/2024.findings-acl.293/">[ACL]</a>
<a href="https://arxiv.org/abs/2310.13267">[arXiv]</a>
</h5>
<p>On the Language Encoder of Contrastive Cross-modal Models</p>
<div class="tile_venue">ACL24</div>
</div>
<div class="tile">
<h3>DIIR</h3>
<a href="https://arxiv.org/abs/2403.15737"><img src="./assets/diir.png"></a>
<h5>
<a href="https://aclanthology.org/2024.findings-acl.782/">[ACL]</a>
<a href="https://arxiv.org/abs/2403.15737">[arXiv]</a>
<a href="https://github.com/zhouhanxie/DIIR">[code]</a>
</h5>
<p>Few-shot Dialogue Strategy Learning for Motivational Interviewing via Inductive Reasoning</p>
<div class="tile_venue">ACL24</div>
</div>
<div class="tile">
<h3>PeaCok</h3>
<a href="https://arxiv.org/abs/2305.02364"><img src="./assets/personas.png"></a>
<h5>
<a href="https://aclanthology.org/2023.acl-long.362/">[ACL]</a>
<a href="https://arxiv.org/abs/2305.02364">[arXiv]</a>
<a href="https://github.com/Silin159/PeaCoK">[code]</a>
</h5>
<p>PeaCoK: Persona Commonsense Knowledge for Consistent and Engaging Narratives<br>(Outstanding Paper Award)</p>
<div class="tile_venue">ACL23</div>
</div>
<div class="tile">
<h3>ComFact</h3>
<a href="https://aclanthology.org/2022.findings-emnlp.120/"><img src="./assets/comfact.png"></a>
<h5>
<a href="https://aclanthology.org/2022.findings-emnlp.120/">[EMNLP]</a>
<a href="https://arxiv.org/abs/2210.12678">[arXiv]</a>
<a href="https://github.com/epfl-nlp/ComFact">[code]</a>
</h5>
<p>ComFact: A Benchmark for Linking Contextual Commonsense Knowledge</p>
<div class="tile_venue">EMNLP22 Findings</div>
</div>
<div class="tile" style="background-color: white;"></div>
<div class="tile" style="background-color: white;"></div>
Automatic music mixing using a generative model of effect embeddings
Automatic Music Sample Identification with Multi-Track Contrastive Learning
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
Large-Scale Training Data Attribution for Music Generative Models via Unlearning
Reductive, exclusionary, normalising: The limits of generative AI music
Can Large Language Models Predict Audio Effects Parameters from Natural Language?
Inference-Time Optimisation for Vocal Effects Style Transfer using DiffVox
SOTA Fx representation: Extract instrument-wise audio effects representations from music mixtures
Inference Time Optimization for Music Mastering Style Transfer
Reverse Engineering of Music Mixing Graphs with Differentiable Processors and Iterative Pruning
Supervised contrastive learning from weakly-labeled audio segments for musical version matching
Music Foundation Model as Generic Booster for Music Downstream Tasks
DiffVox: A Differentiable Model for Capturing and Analysing Professional Effects Distributions
VRVQ: Variable Bitrate Residual Vector Quantization for Audio Compression
Latent Diffusion Bridges for Unsupervised Musical Audio Timbre Transfer
Searching For Music Mixing Graphs: A Pruning Approach
Improving Unsupervised Clean-to-Rendered Guitar Tone Transformation Using GANs and Integrated Unaligned Clean Data
MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models
Timbre-Trap: A Low-Resource Framework for Instrument-Agnostic Music Transcription
VRDMG: Vocal Restoration via Diffusion Posterior Sampling with Multiple Guidance
Automatic Piano Transcription with Hierarchical Frequency-Time Transformer
An Attention-based Approach To Hierarchical Multi-label Music Instrument Classification
Unsupervised Vocal Dereverberation with Diffusion-based Generative Models
Music Mixing Style Transfer: A Contrastive Learning Approach to Disentangle Audio Effects
DiffRoll: Diffusion-based Generative Music Transcription with Unsupervised Pretraining Capability
Hierarchical Diffusion Models for Singing Voice Neural Vocoder
Distortion Audio Effects: Learning How to Recover the Clean Signal
Automatic Music Mixing with Deep Learning and Out-of-Domain Data
Odoriko: A Shape-Aware Multimodal Diffusion Framework for Human Motion
Spectral Prior for Reducing Exposure Bias in Diffusion Models
Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance
PAVAS: a framework for generating physically plausible audio from video, by integrating physics estimation
Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models
VIRTUE: Visual-Interactive Text-Image Universal Embedder
A diffusion-based post-processor for perceptually improving speech enhancement and separation outputs
CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation
TITAN-Guide: Taming Inference-Time AligNment for Guided Text-to-Video Diffusion Models
MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation
SoundCTM: Unifying Score-based and Consistency Models for Full-band Text-to-Sound Generation
Mining Your Own Secrets: Diffusion Classifier Scores for Continual Personalization of Text-to-Image Diffusion Models
GenWarp: Single Image to Novel Views with Semantic-Preserving Generative Warping
SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
BigVSAN: Enhancing GAN-based Neural Vocoders with Slicing Adversarial Network
Zero- and Few-shot Sound Event Localization and Detection
STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound Events
Extending Audio Masked Autoencoders Toward Audio Restoration
Diffiner: A Versatile Diffusion-based Generative Refiner for Speech Enhancement
CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled Videos
Semantic Acoustic Imaging for Sound Event Localization and Detection from Spatial Audio and Audiovisual Scenes
CSS
60.4%
JavaScript
39.4%