DongKeon/Awesome-Speaker-Diarization

Some comprehensive papers about speaker diarization

372

12 commits

updated Mar 24, 2026

See the code

README

🎀 Awesome Speaker Diarization Awesome

Papers Code Notion DB Updated

πŸ“„ Paper Β· πŸ’» Code Β· πŸ“ Review Β· 🎬 Video/Demo Β· πŸ“Š Slides Β· πŸ”— Other


🧠 Core Methods

EEND TS-VAD Clustering Embedding Self-Supervised

πŸ”Œ Extensions

Online Multi-Channel Sep/TSE

🌐 Cross-Modal

ASR Vision NLP/LLM Emotion

πŸ”— Related

VAD/OSD/SCD Speaker Rec Personal VAD Spoofing TTS Child-Adult

πŸ“¦ Resources

Dataset Tools Reviews Measurement Scoring

Challenge


πŸ“– Overview β€” 1 paper

2020 (1 paper)
  • DIHARD Keynote Session: The yellow brick road of diarization, challenges and other neural paths πŸ“Š 🎬

πŸ“ Reviews β€” 3 papers

πŸ”₯ 2023-2024 (3 papers)
  • "Overview of Speaker Modeling and Its Applications: From the Lens of Deep Speaker Representation Learning," in Submitted to IEEE/ACM TASLP, 2024. πŸ“„
  • β€œA review of speaker diarization: Recent advances with deep learning”, in Computer Speech & Language, Volume 72, 2023. (USC) πŸ“„
  • "An Experimental Review of Speaker Diarization methods with application to Two-Speaker Conversational Telephone Speech recordings", in Computer Speech & Language, 2023. πŸ“„

πŸ“š EEND (End-to-End Neural Diarization)-based β€” 50 papers

πŸ”₯ 2025 (7 papers)
  • "Mamba-based Segmentation Model for Speaker Diarization," Proc. ICASSP, 2025. (NTT) πŸ“„ πŸ’»
  • "Pushing the Limits of End-to-End Diarization," in Proc. Interspeech, 2025. πŸ“„
  • VBx-EEND-VC: "VBx for End-to-End Neural and Clustering-based Diarization," in arXiv:2510.19572, 2025. (BUT) πŸ“„
  • "Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling," in arXiv:2506.05593, 2025. (OSU) πŸ“„
  • DLF-EEND: "Dynamic Layer Fusion for End-to-End Speaker Diarization," in Proc. Interspeech, 2025. πŸ“„
  • "End-to-End Diarization utilizing Attractor Deep Clustering," in Proc. Interspeech, 2025. (JHU, OSU) πŸ“„
  • "Pretraining Multi-Speaker Identification for Neural Speaker Diarization," in Proc. Interspeech, 2025. (NTT) πŸ“„
2024 (9 papers)
  • "NTT speaker diarization system for CHiME-7: multi-domain, multi-microphone End-to-end and vector clustering diarization," in Proc. ICASSP, 2024. (NTT) πŸ“„
  • AED-EEND-EE: "Attention-based Encoder-Decoder End-to-End Neural Diarization with Embedding Enhancer," in IEEE/ACM TASLP, 2024. (SJTU) πŸ“„ πŸ“
  • "DiaPer: End-to-End Neural Diarization with Perceiver-Based Attractors," in IEEE/ACM TASLP, 2024. (BUT) πŸ“„ πŸ’» πŸ“
  • "EEND-DEMUX: End-to-End Neural Speaker Diarization via Demultiplexed Speaker Embeddings," in Submitted to IEEE SPL, 2024. (SNU) πŸ“„ πŸ“
  • "EEND-M2F: Masked-attention mask transformers for speaker diarization," in Proc. Interspeech, 2024. (Fano Labs) πŸ“„ πŸ“„ πŸ“
  • EEND-NAA (2): "End-to-End Neural Speaker Diarization with Non-Autoregressive Attractors", in IEEE/ACM TASLP, 2024. (JHU) πŸ“„ πŸ“
  • "From Modular to End-to-End Speaker Diarization," Ph.D. thesis, 2024. (BUT) πŸ“„
  • "On the calibration of powerset speaker diarization models," in Proc. Interspeech, 2024. (IRIT) πŸ“„ πŸ“„ πŸ’» πŸ“
  • Local-global EEND: "Speakers Unembedded: Embedding-free Approach to Long-form Neural Diarization," in Proc. Interspeech, 2024. (Amazon) πŸ“„ πŸ“
2023 (11 papers)
  • "Improving Transformer-based End-to-End Speaker Diarization by Assigning Auxiliary Losses to Attention Heads", in Proc. ICASSP, 2023. (HU) πŸ“„
  • EEND-NA: β€œNeural Diarization with Non-Autoregressive Intermediate Attractors”, in Proc. ICASSP, 2023. (LINE) πŸ“„
  • "TOLD: A Novel Two-Stage Overlap-Aware Framework for Speaker Diarization", in Proc. ICASSP, 2023. (Alibaba) πŸ“„ πŸ’»
  • "Improving End-to-End Neural Diarization Using Conversational Summary Representations", in Proc. Interspeech, 2023. (Fano Labs) πŸ“„
  • AED-EEND: β€œAttention-based Encoder-Decoder Network for End-to-End Neural Speaker Diarization with Target Speaker Attractor”, in Proc. Interspeech, 2023. (SJTU) πŸ“„ πŸ“
  • "Self-Distillation into Self-Attention Heads for Improving Transformer-based End-to-End Neural Speaker Diarization", in Proc. Interspeech, 2023. (HU) πŸ“„
  • "Powerset Multi-class Cross Entropy Loss for Neural Speaker Diarization", in Proc. Interspeech, 2023. (Pyannote) πŸ“„ πŸ’»
  • "End-to-End Neural Speaker Diarization with Absolute Speaker Loss", in Proc. Interspeech, 2023. (Pyannote) πŸ“„
  • "Blueprint Separable Subsampling and Aggregate Feature Conformer-Based End-to-End Neural Diarization", in Electronics, 2023. πŸ“„
  • EEND-TA: "Transformer Attractors for Robust and Efficient End-to-End Neural Diarization," in Proc. ASRU, 2023. (Fano Labs) πŸ“„
  • "Robust End-to-End Diarization with Domain Adaptive Training and Multi-Task Learning," in Proc. ASRU, 2023. (Fano Labs) πŸ“„
2022 (10 papers)
  • EEND-EDA (2): β€œEncoder-Decoder Based Attractor Calculation for End-to-End Neural Diarization”, in IEEE/ACM TASLP, 2022. (Hitachi) πŸ“„ πŸ“ πŸ’»
  • "DIVE: End-to-end Speech Diarization via Iterative Speaker Embedding", in Proc. ICASSP, 2022. (Google) πŸ“„
  • RX-EEND: β€œAuxiliary Loss of Transformer with Residual Connection for End-to-End Speaker Diarization”, in Proc. ICASSP, 2022. (GIST) πŸ“„ πŸ“
  • "End-to-end speaker diarization with transformer", in Proc. arXiv, 2022. πŸ“„
  • EEND-VC-iGMM: "Tight integration of neural and clustering-based diarization through deep unfolding of infinite Gaussian mixture model", in Proc. ICASSP, 2022. (NTT) πŸ“„
  • EDA-RC: "Robust End-to-end Speaker Diarization with Generic Neural Clustering", in Proc. Interspeech, 2022. (SJTU) πŸ“„
  • EEND-NAA: "End-to-End Neural Speaker Diarization with an Iterative Refinement of Non-Autoregressive Attention-based Attractors", in Proc. Interspeech, 2022. (JHU) πŸ“„ πŸ“
  • Graph-PIT: "Utterance-by-utterance overlap-aware neural diarization with Graph-PIT", in Proc. Interspeech, 2022. (NTT) πŸ“„ πŸ’»
  • "Efficient Transformers for End-to-End Neural Speaker Diarization", in Proc. IberSPEECH, 2022. πŸ“„
  • EEND-EDA-SpkAtt: "Towards End-to-end Speaker Diarization in the Wild", in arXiv:2211.01299v1, 2022. πŸ“„
2021 (7 papers)
  • CB-EEND: "End-to-end Neural Diarization: From Transformer to Conformer", in Proc. Interspeech, 2021. (Amazon) πŸ“„ πŸ“
  • TDCN-SA: "End-to-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings", in Proc. ICASSP, 2021. (Google) πŸ“„ πŸ“
  • "End-to-End Speaker Diarization Conditioned on Speech Activity and Overlap Detection", in Proc. IEEE SLT, 2021. (Hitachi) πŸ“„
  • EEND-VC (1): "Integrating end-to-end neural and clustering-based diarization: Getting the best of both worlds", in Proc. ICASSP, 2021. (NTT) πŸ“„ πŸ“ πŸ’»
  • EEND-VC (2): "Advances in integration of end-to-end neural and clustering-based diarization for real conversational speech", in Proc. Interspeech, 2021. (NTT) πŸ“„ πŸ“ πŸ’»
  • "Robust End-to-End Speaker Diarization with Conformer and Additive Margin Penalty," in Proc. Interspeech, 2021. (Fano Labs) πŸ“„
  • EEND-GLA: "Towards Neural Diarization for Unlimited Numbers of Speakers Using Global and Local Attractors", in Proc. ASRU, 2021. (Hitachi) πŸ“„ πŸ“
2020 (3 papers)
  • SA-EEND (2): β€œEnd-to-End Neural Diarization: Reformulating Speaker Diarization as Simple Multi-label Classification”, in arXiv:2003.02966, 2020. (Hitachi) πŸ“„ πŸ“
  • SC-EEND: "Neural Speaker Diarization with Speaker-Wise Chain Rule", in arXiv:2006.01796, 2020. (Hitachi) πŸ“„ πŸ“
  • EEND-EDA (1): β€œEnd-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors”, in Proc. Interspeech, 2020. (Hitachi) πŸ“„ πŸ“ πŸ’»
2023 (1 paper)
  • EEND-IAAE: "End-to-end neural speaker diarization with an iterative adaptive attractor estimation," in Neural Networks, Elsevier. πŸ“„ πŸ’»
2019 (2 papers)
  • BLSTM-EEND: "End-to-End Neural Speaker Diarization with Permutation-Free Objectives", in Proc. Interspeech, 2019. (Hitachi) πŸ“„
  • SA-EEND (1): β€œEnd-to-End Neural Speaker Diarization with Self-attention”, in Proc. ASRU, 2019. (Hitachi) πŸ“„ πŸ’» πŸ’» πŸ“

πŸ”₯ 2024 (2 papers)
  • "Do End-to-End Neural Diarization Attractors Need to Encode Speaker Characteristic Information?," in Proc. Odyssey, 2024. πŸ“„
  • "Leveraging Speaker Embeddings in End-to-End Neural Diarization for Two-Speaker Scenarios," in Proc. Odyssey, 2024. πŸ“„

πŸ“Œ Post-Processing β€” 3 papers

πŸ”₯ 2021-2024 (3 papers)
  • "DiaCorrect: Error Correction Back-end For Speaker Diarization," in Proc. ICASSP, 2024. (BUT) πŸ“„ πŸ’»
  • EENDasP: "End-to-End Speaker Diarization as Post-Processing", in Proc. ICASSP, 2021. (Hitachi) πŸ“„ πŸ“ πŸ’»
  • Dover-Lap: "DOVER-Lap: A Method for Combining Overlap-aware Diarization Outputs", in Proc. IEEE SLT, 2021. (JHU) πŸ“„ πŸ“ πŸ’»

🎯 Using Target Speaker Embedding β€” 17 papers

πŸ”₯ 2025 (4 papers)
  • "Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining," in Proc. ICASSP, 2025. πŸ“„
  • MIMO-TSVAD: "Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization," in IEEE/ACM TASLP, 2025. (DKU) πŸ“„
  • "Mitigating Non-Target Speaker Bias in Guided Speaker Embedding," in Proc. Interspeech, 2025. (NTT) πŸ“„
  • "Diarization-Guided Multi-Speaker Embeddings," in Proc. Interspeech, 2025. (Pyannote) πŸ“„
2024 (3 papers)
  • NSD-MS2S: "Neural Speaker Diarization Using Memory-Aware Multi-Speaker Embedding with Sequence-to-Sequence Architecture, " in Proc. ICASSP, 2024. (USTC) πŸ“„ πŸ’»
  • PET-TSVAD: "Profile-Error-Tolerant Target-Speaker Voice Activity Detection," in Proc. ICASSP, 2024. (Microsoft) πŸ“„
  • Flow-TSVAD: "Target-Speaker Voice Activity Detection via Latent Flow Matching," in arXiv:2409.04859, 2024. (DKU) πŸ“„
2023 (4 papers)
  • EDA-TS-VAD: β€œTarget Speaker Voice Activity Detection with Transformers and Its Integration with End-to-End Neural Diarization”, in Proc. ICASSP, 2023. (Microsoft) πŸ“„
  • Seq2Seq-TS-VAD: β€œTarget-Speaker Voice Activity Detection via Sequence-to-Sequence Prediction”, in Proc. ICASSP, 2023. (DKU) πŸ“„ πŸ“
  • QM-TS-VAD: "Unsupervised Adaptation with Quality-Aware Masking to Improve Target-Speaker Voice Activity Detection for Speaker Diarization", in Proc. Interspeech, 2023. (USTC) πŸ“„
  • "ANSD-MA-MSE: Adaptive Neural Speaker Diarization Using Memory-Aware Multi-Speaker Embedding," in IEEE/ACM TASLP, 2023. (USTC) πŸ“„ πŸ’»
2022 (3 papers)
  • SEND (2): "Speaker Embedding-aware Neural Diarization: an Efficient Framework for Overlapping Speech Diarization in Meeting Scenarios," in arXiv:2203.09767, 2022 (Alibaba) πŸ“„
  • MTEAD: "Multi-target Filter and Detector for Unknown-number Speaker Diarization", in IEEE SPL, 2022. πŸ“„
  • SOND: "Speaker Overlap-aware Neural Diarization for Multi-party Meeting Analysis", in Proc. EMNLP, 2022. (Alibaba) πŸ“„ πŸ’»
2020-2021 (3 papers)
  • SEND (1): "Speaker Embedding-aware Neural Diarization for Flexible Number of Speakers with Textual Information," in arXiv:2111.13694, 2021. (Alibaba) πŸ“„
  • TS-VAD: "Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario", in Proc. Interspeech, 2020. πŸ“„ πŸ’» πŸ“Š
  • β€œThe STC system for the CHiME-6 challenge,” in CHiME Workshop, 2020. πŸ“„

🎯 Target Speech Diarization β€” 1 paper

πŸ”₯ 2024 (1 paper)
  • PTSD: "Prompt-driven Target Speech Diarization," in Proc. ICASSP, 2024. (NUS) πŸ“„

πŸ”€ With Separation or Target Speaker Extraction β€” 12 papers

πŸ”₯ 2025 (3 papers)
  • "Robust Target Speaker Diarization and Separation via Augmented Speaker Embedding Sampling," in Proc. Interspeech, 2025. πŸ“„
  • S2SND: "Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation," in IEEE/ACM TASLP, 2025. (DKU) πŸ“„
  • "Exploring Speaker Diarization with Mixture of Experts," in arXiv:2506.14750, 2025. (USTC) πŸ“„
2024 (7 papers)
  • "TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddings", in IEEE/ACM TASLP, 2024. πŸ“„
  • "Continuous Target Speech Extraction: Enhancing Personalized Diarization and Extraction on Complex Recordings," in arXiv:2401.15993, 2024. (Tencent) πŸ“„ 🎬
  • "PixIT: Joint Training of Speaker Diarization and Speech Separation from Real-world Multi-speaker Recordings," in Proc. Odyssey, 2024. πŸ“„ πŸ’»
  • MC-EEND: "Multi-channel Conversational Speaker Separation via Neural Diarization," in IEEE/ACM TASLP, 2024. (OSU) πŸ“„
  • "USED: Universal Speaker Extraction and Diarization," in submitted to IEEE/ACM TASLP, 2024. (CUHK) πŸ“„ 🎬 πŸ”— πŸ“
  • "Neural Blind Source Separation and Diarization for Distant Speech Recognition," in Proc. Interspeech, 2024. (AIST) πŸ“„
  • "TalTech-IRIT-LIS Speaker and Language Diarization Systems for DISPLACE 2024," in Proc. Interspeech, 2024. (Pyannote) πŸ“„
2021-2022 (2 papers)
  • EEND-SS: "Joint End-to-End Neural Speaker Diarization and Speech Separation for Flexible Number of Speakers”, in Proc. SLT, 2022. (CMU) πŸ“„ πŸ“
  • "Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis," in Proc. SLT, 2021. (JHU) πŸ“„ πŸ”— πŸ“

πŸ“‘ Multi-Channel β€” 14 papers

πŸ”₯ 2025 (5 papers)
  • "Multi-channel Speaker Counting for EEND-VC-based Speaker Diarization on Multi-domain Conversation," in Proc. ICASSP, 2025. (NTT) πŸ“„ πŸ“
  • "Spatially Aware Self-Supervised Models for Multi-Channel Neural Speaker Diarization," in arXiv:2510.14551, 2025. (BUT) πŸ“„ πŸ’»
  • "Multi-Channel Sequence-to-Sequence Neural Diarization for The MISP 2025 Challenge," in arXiv:2505.16387, 2025. πŸ“„
  • "Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings," in Proc. ICASSP, 2025. πŸ“„
  • "Spatio-Spectral Diarization of Meetings by Combining TDOA-based Segmentation and Speaker Embedding-based Clustering," in Proc. Interspeech, 2025. πŸ“„
2024 (5 papers)
  • "UniX-Encoder: A Universal X-Channel Speech Encoder for Ad-Hoc Microphone Array Speech Processing," in arXiv:2310.16367, 2024. (JHU, Tencent) πŸ“„
  • "Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech Detection," in IEEE/ACM TASLP, 2024. πŸ“„
  • "A Spatial Long-Term Iterative Mask Estimation Approach for Multi-Channel Speaker Diarization and Speech Recognition," in Proc. ICASSP, 2024. (USTC) πŸ“„
  • MC-EEND: "Multi-channel Conversational Speaker Separation via Neural Diarization," in IEEE/ACM TASLP, 2024. (OSU) πŸ“„
  • "ASoBO: Attentive Beamformer Selection for Distant Speaker Diarization in Meetings," in Proc. Interspeech, 2024. (LIUM) πŸ“„
2022-2023 (4 papers)
  • "Mutual Learning of Single- and Multi-Channel End-to-End Neural Diarization," in Proc. IEEE SLT, 2023. (Hitachi) πŸ“„
  • "Semi-supervised multi-channel speaker diarization with cross-channel attention", in Proc. ASRU, 2023. (USTC) πŸ“„
  • "Multi-Channel End-to-End Neural Diarization with Distributed Microphones", in Proc. ICASSP, 2022. (Hitachi) πŸ“„
  • "Multi-Channel Speaker Diarization Using Spatial Features for Meetings", in Proc. ICASSP, 2022. (Tencent) πŸ“„

⚑ Online β€” 18 papers

πŸ”₯ 2024-2025 (8 papers)
  • SCDiar: "A Streaming Diarization System based on Speaker Change Detection and Speech Recognition," in Proc. ICASSP, 2025. πŸ“„
  • Streaming Sortformer: "Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering," in Proc. Interspeech, 2025. (NVIDIA) πŸ“„
  • OTS-VAD: "Online Neural Speaker Diarization With Target Speaker Tracking," in IEEE/ACM TASLP, 2024. (DKU) πŸ“„
  • FS-EEND: "Frame-wise streaming end-to-end speaker diarization with non-autoregressive self-attention-based attractors," in Proc. ICASSP, 2024. (Hangzhou) πŸ“„ πŸ’»
  • "Online speaker diarization of meetings guided by speech separation," in Proc. ICASSP, 2024. (LTCI) πŸ“„ πŸ’»
  • "Interrelate Training and Clustering for Online Speaker Diarization," in IEEE/ACM TASLP, 2024. πŸ“„
  • O-EENC-SD: "Efficient Online End-to-End Neural Clustering for Speaker Diarization," in Proc. ICASSP, 2025. πŸ“„
  • LS-EEND: "Long-Form Streaming End-to-End Neural Diarization with Online Attractor Extraction," in IEEE/ACM TASLP, 2025. (Westlake) πŸ“„ πŸ’»
2023 (2 papers)
  • "Absolute decision corrupts absolutely: conservative online speaker diarisation", in Proc. ICASSP, 2023. (Naver) πŸ“„
  • "A Reinforcement Learning Framework for Online Speaker Diarization", in Under Review. NeruIPS, 2023. (CU) πŸ“„
2022 (3 papers)
  • "Low-Latency Online Speaker Diarization with Graph-Based Label Generation", in Proc. Odyssey, 2022. (DKU) πŸ“„
  • EEND-GLA: "Online Neural Diarization of Unlimited Numbers of Speakers Using Global and Local Attractors", in IEEE/ACM TASLP, 2022. (Hitachi) πŸ“„
  • Online TS-VAD: "Online Target Speaker Voice Activity Detection for Speaker Diarization", in Proc. Interspeech, 2022. (DKU) πŸ“„
2021 (4 papers)
  • "Online End-to-End Neural Diarization with Speaker-Tracing Buffer", in Proc. IEEE SLT, 2021. (Hitachi) πŸ“„
  • BW-EDA-EEND: "BW-EDA-EEND: Streaming End-to-End Neural Speaker Diarization for a Variable Number of Speakers", in Proc. Interspeech, 2021. (Amazon) πŸ“„
  • FS-EEND: "Online Streaming End-to-End Neural Diarization Handling Overlapping Speech and Flexible Numbers of Speakers", in Proc. Interspeech, 2021. (Hitachi) πŸ“„ πŸ“
  • Diart: "Overlap-aware low-latency online speaker diarization based on end-to-end local segmentation", in Proc. ASRU, 2021. πŸ“„ πŸ’»
2020 (1 paper)
  • "Supervised online diarization with sample mean loss for multi-domain data", in Proc. ICASSP, 2020 πŸ“„ πŸ’»

πŸ”— Clustering-based β€” 21 papers

πŸ”₯ 2025 (3 papers)
  • E-SHARC: "End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization," in IEEE/ACM TASLP, 2025. (IISC) πŸ“„
  • Pyannote Community-1: "pyannote.audio 4.0 with community-1 open-source diarization model," 2025. πŸ”— πŸ’»
  • "Speaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation Algorithm," in Proc. Interspeech, 2025. πŸ“„
2024 (6 papers)
  • "Overlap-aware End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization," in submitted to IEEE/ACM TASLP, 2024. πŸ“„
  • "Apollo's Unheard Voices: Graph Attention Networks for Speaker Diarization and Clustering for Fearless Steps Apollo Collection," in Proc. ICASSP, 2024. (UTD) πŸ“„
  • "Multi-View Speaker Embedding Learning for Enhanced Stability and Discriminability," in Proc. ICASSP, 2024. (Tsinghua) πŸ“„
  • "Towards Unsupervised Speaker Diarization System for Multilingual Telephone Calls Using Pre-trained Whisper Model and Mixture of Sparse Autoencoders," in arXiv:2407.01963, 2024. πŸ“„
  • "Investigating Confidence Estimation Measures for Speaker Diarization," in Proc. Interspeech, 2024. πŸ“„
  • "Once more Diarization: Improving meeting transcription systems through segment-level speaker reassignment," in Proc. Interspeech, 2024. (PU) πŸ“„ πŸ“„ πŸ’»
2023 (5 papers)
  • SCALE: "Spectral Clustering-aware Learning of Embeddings for Speaker Diarisation", in Proc. ICASSP, 2023. (CAM) πŸ“„
  • SHARC: "Supervised Hierarchical Clustering using Graph Neural Networks for Speaker Diarization", in Proc. ICASSP, 2023. (IISC) πŸ“„
  • CDGCN: "Community Detection Graph Convolutional Network for Overlap-Aware Speaker Diarization," in Proc. ICASSP, 2023. (XMU) πŸ“„
  • "Pyannote.Audio 2.1: Speaker Diarization Pipeline: Principle, Benchmark and Recipe", in Proc. Interspeech, 2023. (CNRS) πŸ“„
  • GADEC: "Graph attention-based deep embedded clustering for speaker diarization,", in Speech Communication, 2023. (NJUPT) πŸ“„
2020-2022 (4 papers)
  • UMAP-Leiden: "Reformulating Speaker Diarization as Community Detection With Emphasis On Topological Structure", in Proc. ICASSP, 2022. (Alibaba) πŸ“„
  • Pyannote 2.0: "End-to-end speaker segmentation for overlap-aware resegmentation", in Proc. Interspeech, 2021. (CNRS) πŸ“„ πŸ’» 🎬
  • Pyannote: "pyannote.audio: neural building blocks for speaker diarization", in Proc. ICASSP, 2020. (CNRS) πŸ“„ πŸ’» 🎬
  • Resegmentation with VB: β€œOverlap-Aware Diarization: Resegmentation Using Neural End-to-End Overlapped Speech Detection”, in Proc. ICASSP, 2020. πŸ“„
2018 (1 paper)
2019 (2 papers)
  • DNC: "Discriminative Neural Clustering for Speaker Diarisation", in Proc. IEEE SLT, 2019. πŸ“„ πŸ’» πŸ“
  • NME-SC: β€œAuto-Tuning Spectral Clustering for Speaker Diarization Using Normalized Maximum Eigengap”, IEEE SPL, 2019. πŸ“„ πŸ’»

πŸ“ Variational Bayes and HMM β€” 24 papers

πŸ”₯ 2023-2024 (3 papers)
  • DVBx: "Discriminative Training of VBx Diarization", in Proc. ICASSP, 2024. (BUT) πŸ“„ πŸ’»
  • MS-VBx: "Multi-Stream Extension of Variational Bayesian HMM Clustering (MS-VBx) for Combined End-to-End and Vector Clustering-based Diarization", in Proc. Interspeech, 2023. (NTT) πŸ“„
  • "Generalized domain adaptation framework for parametric back-end in speaker recognition", in arXiv:2305.15567, 2023. πŸ“„
2021-2022 (4 papers)
  • "Bayesian HMM clustering of x-vector sequences (VBx) in speaker diarization: theory, implementation and analysis on standard tasks", in Computer Speech & Language, 2022. (BUT) πŸ“„
  • DCA-PLDA "A Speaker Verification Backend with Robust Performance across Conditions”, in Computer & Language, 2022. πŸ“„ πŸ’»
  • "Analysis of the but Diarization System for Voxconverse Challenge", in Proc. ICASSP, 2021. (BUT) πŸ“„ πŸ’»
  • "Discriminatively trained probabilistic linear discriminant analysis for speaker verification", in Proc. ICASSP, 2021. πŸ“„
2019-2020 (4 papers)
  • "Optimizing Bayesian Hmm Based X-Vector Clustering for the Second Dihard Speech Diarization Challenge", in Proc. ICASSP, 2020. (BUT) πŸ“„
  • β€œAnalysis of Speaker Diarization Based on Bayesian HMM With Eigenvoice Priors”, IEEE/ACM TASLP, 2019. (BUT) πŸ“„
  • "BUT System Description for DIHARD Speech Diarization Challenge 2019", in arXiv:1910.08847, 2019. (BUT) πŸ“„
  • "Bayesian HMM Based x-Vector Clustering for Speaker Diarization", in Proc. Interspeech, 2019. (BUT) πŸ“„
2018 (5 papers)
  • "Speaker Diarization based on Bayesian HMM with Eigenvoice Priors", in Proc. Odyssey, 2018. (BUT) πŸ“„
  • "VB-HMM Speaker Diarization with Enhanced and Refined Segment Representation", in Proc. Odyssey, 2018. (Tsinghua) πŸ“„
  • "Diarization is hard: some experiences and lessons learned for the JHU team in the inaugural DIHARD challenge", in Proc. Interspeech, 2018. πŸ“„
  • "The speaker partitioning problem", in Proc. Odyssey, 2018. πŸ“„
  • "Estimation of the Number of Speakers with Variational Bayesian PLDA in the DIHARD Diarization Challenge", in Proc. Interspeech, 2018. πŸ“„
2015-2017 (3 papers)
  • "Domain Adaptation of PLDA Models in Broadcast Diarization by Means of Unsupervised Speaker Clustering, in Proc. Interspeech, 2017. πŸ“„
  • "Iterative PLDA Adaptation for Speaker Diarization", in Proc. Interspeech, 2016. πŸ“„
  • "Diarization resegmentation in the factor analysis subspace", in Proc. ICASSP, 2015. πŸ“„
2011-2014 (3 papers)
  • "Speaker diarization with plda i-vector scoring and unsupervised calibration", in Proc. IEEE SLT, 2014. πŸ“„
  • "Unsupervised Methods for Speaker Diarization: An Integrated and Iterative Approach", IEEE/ACM TASLP, 2013. πŸ“„
  • "Analysis of i-vector length normalization in speaker recognition systems", in Proc. Interspeech, 2011. πŸ“„
2005-2008 (2 papers)
  • "Bayesian analysis of speaker diarization with eigenvoice priors", in CRIM, Montreal, Technical Report, 2008. πŸ“„
  • "Variational Bayesian methods for audio indexing", in Proc. ICMI-MLMI, 2005. πŸ“„

🧩 Embedding (With Clustering) β€” 14 papers

πŸ”₯ 2024 (4 papers)
  • "Geodesic interpolation of frame-wise speaker embeddings for the diarization of meeting scenarios", in Proc. ICASSP, 2024. (PU) πŸ“„ πŸ“
  • "Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization," in Proc. Odyssey, 2024. (IDLab) πŸ“„
  • "Efficient Speaker Embedding Extraction Using a Twofold Sliding Window Algorithm for Speaker Diarization," in Proc. Interspeech, 2024. (HU) πŸ“„
  • "Variable Segment Length and Domain-Adapted Feature Optimization for Speaker Diarization," in Proc. Interspeech, 2024. (XMU) πŸ“„ πŸ’»
2023 (5 papers)
  • "In Search of Strong Embedding Extractors For Speaker Diarization", in Proc. ICASSP, 2023. (Naver) πŸ“„ πŸ“
  • DR-DESA: "Advancing the dimensionality reduction of speaker embeddings for speaker diarisation: disentangling noise and informing speech activity", in Proc. ICASSP, 2023. (Naver) πŸ“„ πŸ“
  • HEE: "High-resolution embedding extractor for speaker diarisation", in Proc. ICASSP, 2023. (Naver) πŸ“„ πŸ“
  • "Frame-wise and overlap-robust speaker embeddings for meeting diarization", in Proc. ICASSP, 2023. (PU) πŸ“„ πŸ“
  • "A Teacher-Student approach for extracting informative speaker embeddings from speech mixtures", in Proc. Interspeech, 2023. (PU) πŸ“„
2022 (3 papers)
  • GAT+AA: "Multi-scale speaker embedding-based graph attention networks for speaker diarisation", in Proc. ICASSP, 2022. (Naver) πŸ“„
  • MSDD: "Multi-scale Speaker Diarization with Dynamic Scale Weighting", in Proc. Interspeech, 2022. (NVIDIA) πŸ“„ πŸ’» πŸ”—
  • PRISM: "PRISM: Pre-trained Indeterminate Speaker Representation Model for Speaker Diarization and Speaker Verification", in Proc. Interspeech, 2022. (Alibaba) πŸ“„
2021 (2 papers)
  • "Multi-Scale Speaker Diarization With Neural Affinity Score Fusion", in Proc. ICASSP, 2021. (USC) πŸ“„
  • AA+DR+NS: "Adapting Speaker Embeddings for Speaker Diarisation", in Proc. Interspeech, 2021. (Naver) πŸ“„ πŸ“

πŸͺͺ With Speaker Identification β€” 1 paper

πŸ”₯ 2024 (1 paper)
  • "Uncertainty Quantification in Machine Learning for Joint Speaker Diarization and Identification, in Submitted to IEEE/ACM TASLP, 2024. πŸ“„

πŸ”Š Speaker Recognition & Verification β€” 7 papers

πŸ”₯ 2024 (3 papers)
  • "Rethinking Session Variability: Leveraging Session Embeddings for Session Robustness in Speaker Verification," in Proc. ICASSP, 2024. (Naver) πŸ“„
  • "Leveraging In-the-Wild Data for Effective Self-Supervised Pretraining in Speaker Recognition," in Proc. ICASSP, 2024. (CUHK) πŸ“„
  • "Disentangled Representation Learning for Environment-agnostic Speaker Recognition," in Proc. Interspeech, 2024. (KAIST) πŸ“„ πŸ“„ πŸ’»
2023 (3 papers)
  • "Build a SRE Challenge System: Lessons from VoxSRC 2022 and CNSRC 2022," in Proc. Interspeech, 2023. (SJTU) πŸ“„
  • RecXi "Disentangling Voice and Content with Self-Supervision for Speaker Recognition," in Proc. NeurIPS, 2023. (A*STAR) πŸ“„
  • "ECAPA2: A Hybrid Neural Network Architecture and Training Strategy for Robust Speaker Embeddings," in Proc. ASRU, 2023. (IDLab) πŸ“„ πŸ”— πŸ“
2021 (1 paper)
  • "Xi-Vector Embedding for Speaker Recognition," in IEEE, SPL. (A*STAR) πŸ“„ πŸ“

πŸ“Š Scoring β€” 3 papers

πŸ”₯ 2019-2023 (3 papers)
  • β€œSimilarity Measurement of Segment-Level Speaker Embeddings in Speaker Diarization”, IEEE/ACM TASLP, 2023. (DKU) πŸ“„
  • "Self-Attentive Similarity Measurement Strategies in Speaker Diarization", in Proc. Interspeech, 2020. (DKU) πŸ“„
  • LSTM scoring: "LSTM based Similarity Measurement with Spectral Clustering for Speaker Diarization", in Proc. Interspeech, 2019. (DKU) πŸ“„

πŸ—£οΈ With ASR β€” 27 papers

πŸ”₯ 2025-2026 (8 papers)
  • SE-DiCoW: "Self-Enrolled Diarization-Conditioned Whisper," in arXiv:2601.19194, 2026. (BUT) πŸ“„
  • TagSpeech: "End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding," in arXiv:2601.06896, 2026. πŸ“„
  • DiCoW: "Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition," in Proc. ICASSP, 2025. (BUT) πŸ“„ πŸ’»
  • "Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition," in arXiv:2510.03723, 2025. (BUT) πŸ“„
  • "Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models," in arXiv:2506.05796, 2025. (DKU) πŸ“„
  • "Language Modelling for Speaker Diarization in Telephonic Interviews," in arXiv:2501.17893, 2025. πŸ“„
  • SC-SOT: "Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition," in Proc. Interspeech, 2025. πŸ“„
    • "Target Speaker ASR with Whisper," in Submitted to ICASSP, 2025. (BUT) πŸ“„ πŸ’»
2024 (12 papers)
  • "Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach,", in Proc. ICASSP, 2024. (NVIDIA) πŸ“„
  • WEEND: "Towards Word-Level End-to-End Neural Speaker Diarization with Auxiliary Network," in arXiv:2309.08489, 2024. (Google) πŸ“„ πŸ”—
  • "One model to rule them all ? Towards End-to-End Joint Speaker Diarization and Speech Recognition", in Proc. ICASSP, 2024. (CMU) πŸ“„
  • "Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization," in arXiv:2309.16482, 2024. (PU) πŸ“„
  • β€œJoint Inference of Speaker Diarization and ASR with Multi-Stage Information Sharing," in Proc. ICASSP, 2024. (DKU) πŸ“„
  • "Multitask Speech Recognition and Speaker Change Detection for Unknown Number of Speakers" in Proc. ICASSP, 2024. (Idiap) πŸ“„
  • "A Spatial Long-Term Iterative Mask Estimation Approach for Multi-Channel Speaker Diarization and Speech Recognition," in Proc. ICASSP, 2024. (USTC) πŸ“„
  • "On the Success and Limitations of Auxiliary Network Based Word-Level End-to-End Neural Speaker Diarization," in Proc. Interspeech, 2024. (Google) πŸ“„
  • Sortformer: "Seamless Integration of Speaker Diarization and ASR by Bridging Timestamps and Tokens," in Proc. ICML, 2025. (NVIDIA) πŸ“„ πŸ“
    • "Speaker Mask Transformer for Multi-talker Overlapped Speech Recognition," in arXiv:2312.10959, 2024. (NICT) πŸ“„
    • "On Speaker Attribution with SURT," in Proc. Odyssey, 2024. (JHU) πŸ“„
    • "Improving Speaker Assignment in Speaker-Attributed ASR for Real Meeting Applications," in Proc. Odyssey, 2024. (CNRS) πŸ“„
2023 (5 papers)
  • "Unified Modeling of Multi-Talker Overlapped Speech Recognition and Diarization with a Sidecar Separator", in Proc. Interspeech, 2023. (CUHK) πŸ“„
  • "Multi-resolution Approach to Identification of Spoken Languages and to Improve Overall Language Diarization System using Whisper Model", in Proc. Interspeech, 2023.
  • "Speaker Diarization for ASR Output with T-vectors: A Sequence Classification Approach", in Proc. Interspeech, 2023. πŸ“„
  • "Lexical Speaker Error Correction: Leveraging Language Models for Speaker Diarization Error Correction", in Proc. Interspeech, 2023. (Amazon) πŸ“„
    • "SA-Paraformer: Non-autoregressive End-to-End Speaker-Attributed ASR," in Proc. ASRU, 2023. (Alibaba) πŸ“„
2022 (2 papers)
  • "Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed ASR," in Proc. ICASSP, 2022. πŸ“„
  • "Tandem Multitask Training of Speaker Diarisation and Speech Recognition for Meeting Transcription", in Proc. Interspeech, 2022. πŸ“„

πŸ’¬ With NLP / LLM / Language β€” 12 papers

πŸ”₯ 2024-2025 (7 papers)
  • SpeakerLM: "End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models," in arXiv:2508.06372, 2025. πŸ“„
  • "Interactive Real-Time Speaker Diarization Correction with Human Feedback," in arXiv:2509.18377, 2025. πŸ“„
  • "DiariST: Streaming Speech Translation with Speaker Diarization," in Proc. ICASSP, 2024. (Microsoft) πŸ“„ πŸ’»
  • JPCP: "Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation," in arXiv:2309.10456, 2024. (Alibaba) πŸ“„
  • "DiarizationLM: Speaker Diarization Post-Processing with Large Language Models," in Proc. Interspeech, 2024. (Google) πŸ“„ πŸ“„ πŸ’» πŸ“
  • "LLM-based speaker diarization correction: A generalizable approach," in Submitted to IEEE/ACM TASLP, 2024. πŸ“„
  • "AG-LSEC: Audio Grounded Lexical Speaker Error Correction," in Proc. Interspeech, 2024. (Amazon) πŸ“„
2023 (3 papers)
  • "Exploring Speaker-Related Information in Spoken Language Understanding for Better Speaker Diarization", in Proc. ACL, 2023. (Alibaba) πŸ“„
  • MMSCD, "Encoder-decoder multimodal speaker change detection", in Proc. Interspeech, 2023. (Naver) πŸ“„
  • "Aligning Speakers: Evaluating and Visualizing Text-based Diarization Using Efficient Multiple Sequence Alignment,", in Proc. ICTAI, 2023. πŸ“„

🌐 Language Diarization β€” 2 papers

πŸ”₯ 2023 (2 papers)
  • "End-to-End Spoken Language Diarization with Wav2vec Embeddings", in Proc. Interspeech, 2023. πŸ“„ πŸ’»
  • "Multi-resolution Approach to Identification of Spoken Languages and To Improve Overall Language Diarization System Using Whisper Model," in Proc. Interspeech, 2023. πŸ“„

πŸ‘οΈ With Vision β€” 21 papers

πŸ”₯ 2025-2026 (4 papers)
  • CineSRD: "Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarization," in arXiv:2603.16966, 2026. πŸ“„
  • "Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization," in Proc. ACL, 2025. πŸ“„
  • "Cross-Attention and Self-Attention for Audio-visual Speaker Diarization," in arXiv:2506.02621, 2025. πŸ“„
  • "Count Your Speakers! Multitask Learning for Multimodal Speaker Diarization," in Proc. Interspeech, 2025. πŸ“„
2024 (6 papers)
  • "Speaker Diarization of Scripted Audiovisual Content," in arXiv:2308.02160, 2024. (Amazon) πŸ“„
  • "AFL-Net: Integrating Audio, Facial, and Lip Modalities with Cross-Attention for Robust Speaker Diarization in the Wild," in Proc. ICASSP, 2024. (Tencent) πŸ“„ 🎬
  • "Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation," in Proc. AAAI, 2024. (Tencent) πŸ“„
  • "3D-Speaker-Toolkit: An Open Source Toolkit for Multi-modal Speaker Verification and Diarization," in arXiv:2403.19971, 2024. (Alibaba) πŸ“„ πŸ’»
  • "Target Speech Diarization with Multimodal Prompts," in Submitted to IEEE/ACM TASLP, 2024. (NUS) πŸ“„
  • MFV-KSD: "Multi-Stage Face-Voice Association Learning with Keynote Speaker Diarization," in Submitted to ACM MM, 2024. πŸ“„ πŸ’»
2023 (5 papers)
  • "Audio-Visual Speaker Diarization in the Framework of Multi-User Human-Robot Interaction", in Proc. ICASSP, 2023. πŸ“„
  • STHG: "Spatial-Temporal Heterogeneous Graph Learning for Advanced Audio-Visual Diarization, in Proc. CVPR, 2023. (Intel) πŸ“„
  • "Uncertainty-Guided End-to-End Audio-Visual Speaker Diarization for Far-Field Recordings," in Proc. ACM MM, 2023. πŸ“„
  • "Joint Training or Not: An Exploration of Pre-trained Speech Models in Audio-Visual Speaker Diarization," in Springer Computer Science proceedings, 2023. πŸ“„
  • EEND-EDA++: "Late Audio-Visual Fusion for In-The-Wild Speaker Diarization," in arXiv:2211.01299v2, 2023. πŸ“„
2022 (3 papers)
  • AVA-AVD (AVR-Net): "AVA-AVD: Audio-Visual Speaker Diarization in the Wild", in Proc. ACM MM, 2022. πŸ“„ πŸ’» 🎬
  • "End-to-End Audio-Visual Neural Speaker Diarization", in Proc. Interspeech, 2022. (USTC) πŸ“„ πŸ’» πŸ“
  • DyViSE: "DyViSE: Dynamic Vision-Guided Speaker Embedding for Audio-Visual Speaker Diarization", in Proc. MMSP, 2022. (THU) πŸ“„ πŸ’»
2024 (1 paper)
  • "Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization," in Submitted to IEEE/ACM TASLP. (DKU) πŸ“„
2019-2020 (2 papers)
  • "Self-supervised learning for audio-visual speaker diarization", in Proc. ICASSP, 2020. (Tencent) πŸ“„ πŸ”—
  • "Who said that?: Audio-visual speaker diarisation of real-world meetings", in Proc. Interspeech, 2019. (Naver) πŸ“„

πŸ”₯ 2024 (1 paper)
  • "Spoof Diarization: "What Spoofed When" in Partially Spoofed Audio," in Proc. Interspeech, 2024. (IITK) πŸ“„

πŸ“Œ Speaker Anonymization β€” 1 paper

πŸ”₯ 2024 (1 paper)
  • "A Benchmark for Multi-speaker Anonymization," in Submitted to IEEE/ACM TASLP, 2024. (SIT) πŸ“„ πŸ’»

πŸ“Œ Singing Diarization β€” 1 paper

πŸ”₯ 2024 (1 paper)
  • "Song Data Cleansing for End-to-End Neural Singer Diarization Using Neural Analysis and Synthesis Framework," in Proc. Interspeech, 2024. (LY) πŸ“„

😊 With Emotion β€” 3 papers

πŸ”₯ 2023-2024 (3 papers)
  • "ED-TTS: Multi-scale Emotion Modeling using Cross-domain Emotion Diarization for Emotional Speech Synthesis, in Proc. ICASSP, 2024. πŸ“„
  • "Speech Emotion Diarization: Which Emotion Appears When?," in Proc. ASRU, 2023. (Zaion) πŸ“„
  • "EmoDiarize: Speaker Diarization and Emotion Identification from Speech Signals using Convolutional Neural Networks," in arxiv:2310.12851, 2023. πŸ“„

πŸŽ™οΈ Personal VAD β€” 2 papers

πŸ”₯ 2020-2023 (2 papers)
  • "SVVAD: Personal Voice Activity Detection for Speaker Verification", in Proc. Interspeech, 2023. πŸ“„
  • "Personal VAD: Speaker-Conditioned Voice Activity Detection", in Proc. Odyssey, 2020. (Google) πŸ“„

πŸ“ˆ VAD & OSD & SCD β€” 10 papers

πŸ”₯ 2024 (3 papers)
  • "USM-SCD: Multilingual Speaker Change Detection Based on Large Pretrained Foundation Models," in Proc. ICASSP, 2024. (Google) πŸ“„
  • "Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech Detection," in IEEE/ACM TASLP, 2024. πŸ“„
  • "Speaker Change Detection with Weighted-sum Knowledge Distillation based on Self-supervised Pre-trained Models," in Proc. Interspeech, 2024. πŸ“„
2023 (4 papers)
  • "Multitask Detection of Speaker Changes, Overlapping Speech and Voice Activity Using wav2vec 2.0," in Proc. ICASSP, 2023. πŸ“„ πŸ’»
  • "Semantic VAD: Low-Latency Voice Activity Detection for Speech Interaction," in Proc. Interspeech, 2023. πŸ“„
  • "Joint speech and overlap detection: a benchmark over multiple audio setup and speech domains," in arxiv:2307.13012, 2023. πŸ“„
  • "Advancing the study of Large-Scale Learning in Overlapped Speech Detection," in arXiv:2308.05987, 2023. πŸ“„
2022 (3 papers)
  • "Overlapped Speech Detection in Broadcast Streams Using X-vectors," in Proc. Interspeech, 2022. πŸ“„
  • "Overlapped speech and gender detection with WavLM pre-trained features," in Proc. Interspeech, 2022. πŸ“„
  • "Microphone Array Channel Combination Algorithms for Overlapped Speech Detection," in Proc. Interspeech, 2022. πŸ“„

πŸ“Š Dataset β€” 19 papers

πŸ”₯ 2024-2025 (6 papers)
  • "Conversations in the wild: Data collection, automatic generation and evaluation," in Computer Speech & Language, 2025. πŸ“„
  • M3SD: "Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset," in arXiv:2506.14427, 2025. πŸ“„
  • "VoxBlink: X-Large Speaker Verification Dataset on Camera", in Proc. ICASSP, 2024. πŸ“„ πŸ”—
  • "NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant Meeting Transcription," in arXiv:2401.08887, 2024. (MS) πŸ“„
  • "A Comparative Analysis of Speaker Diarization Models: Creating a Dataset for German Dialectal Speech," in Proc. ACL, 2024. πŸ“„
  • "ALLIES: A Speech Corpus for Segmentation, Speaker Diarization, Speech Recognition and Speaker Change Detection," in Proc. ACL, 2024. (LIUM) πŸ“„
2020-2022 (5 papers)
  • Ego4D: " Around the World in 3,000 Hours of Egocentric Video," in Proc. CVPR, 2022. (Meta) πŸ“„ πŸ’» πŸ”—
  • AliMeeting: "Summary On The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge," in Proc. ICASSP, 2022. (Alibaba) πŸ“„ πŸ”— πŸ’»
  • Voxconverse: "Spot the conversation: speaker diarisation in the wild", in Proc. Interspeech, 2020. (VGG, Naver) πŸ“„ πŸ’» πŸ”—
  • MSDWild: Multi-modal Speaker Diarization Dataset in the Wild, in Proc. Interspeech, 2020. πŸ“„ πŸ”—
  • "LibriMix: An Open-Source Dataset for Generalizable Speech Separation," in arXiv:2005.11262, 2020. πŸ“„ πŸ’»

πŸ“Š Simulated Dataset β€” 8 papers

πŸ”₯ 2023-2024 (4 papers)
  • "Enhancing low-latency speaker diarization with spatial dictionary learning," in Proc. ICASSP, 2024. (NTU) πŸ“„ πŸ“Š
  • "Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling," in Proc. ICASSP, 2024. (OSU) πŸ“„
  • "Multi-Speaker and Wide-Band Simulated Conversations as Training Data for End-to-End Neural Diarization", in Proc. ICASSP, 2023. (BUT) πŸ“„ πŸ’» πŸ“
  • "Property-Aware Multi-Speaker Data Simulation: A Probabilistic Modelling Technique for Synthetic Data Generation," in CHiME-7 Workshop, 2023. (NVIDIA) πŸ“„
2022 (3 papers)
  • "From simulated mixtures to simulated conversations as training data for end-to-end neural diarization" , in Proc. Interspeech, 2022. (BUT) πŸ“„ πŸ’» πŸ“
  • Markov selection: "Improving the naturalness of simulated conversations for end-to-end neural diarization", in Proc. Odyssey, 2022. (Hitachi) πŸ“„
  • EEND-EDA-SpkAtt: "Towards End-to-end Speaker Diarization in the Wild", in arXiv:2211.01299v1, 2022. πŸ“„
2019 (1 paper)
  • Concat-and-sum: "End-to-end neuarl speaker diarization with permuation-free objectives", in Proc. Interspeech, 2019. πŸ“„

πŸ› οΈ Tools β€” 1 paper

πŸ”₯ 2024 (1 paper)
  • "Gryannote open-source speaker diarization labeling tool," in Proc. Interspeech (Show and Tell), 2024. (IRIT) πŸ“„ πŸ’»

πŸ”„ Self-Supervised β€” 5 papers

πŸ”₯ 2025 (3 papers)
  • DiariZen: "Leveraging Self-Supervised Learning for Speaker Diarization," in Proc. ICASSP," 2025. (BUT) πŸ“„ πŸ“„ πŸ’» πŸ“
  • "Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models," in arXiv:2506.18623, 2025. (BUT) πŸ“„ πŸ’»
  • "Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization," in arXiv:2505.24111, 2025. πŸ“„
2022 (2 papers)
  • β€œSelf-supervised Speaker Diarization”, in Proc. Interspeech, 2022. πŸ“„
  • CSDA: "Continual Self-Supervised Domain Adaptation for End-to-End Speaker Diarization", in Proc. IEEE SLT, 2022. (CNRS) πŸ“„ πŸ’»

πŸ”ƒ Semi-Supervised β€” 1 paper

πŸ”₯ 2017 (1 paper)
  • "Active Learning Based Constrained Clustering For Speaker Diarization", in IEEE/ACM TASLP, 2017. (UT) πŸ“„

πŸ“ Measurement β€” 3 papers

πŸ”₯ 2022-2025 (3 papers)
  • SDBench: β€œA Comprehensive Benchmark Suite for Speaker Diarization,” in Proc. Interspeech, 2025. πŸ“„
  • β€œBenchmarking Diarization Models,” in arXiv:2509.26177, 2025. πŸ“„
  • BER: β€œBalanced Error Rate For Speaker Diarization”, in Proc. arXiv:2211.04304, 2022 πŸ“„ πŸ’»

πŸ‘Ά Child-Adult β€” 2 papers

πŸ”₯ 2023-2024 (2 papers)
  • "Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions," in Proc. Interspeech, 2024. (USC) πŸ“„
  • "Robust Self Supervised Speech Embeddings for Child-Adult Classification in Interactions involving Children with Autism," in Proc. Interspeech, 2023. πŸ“„

πŸ† Challenge


πŸ”Š VoxSRC (VoxCeleb Speaker Recognition Challenge)

πŸ“Œ VoxSRC-20 Track4 β€” 3 papers

Unknown (3 papers)

πŸ“Œ VoxSRC-21 Track4 β€” 3 papers

Unknown (3 papers)

πŸ“Œ VoxSRC-22 Track4 β€” 3 papers

Unknown (3 papers)

πŸ“Œ VoxSRC-23 Track4 β€” 5 papers

Unknown (5 papers)

πŸ“‘ M2MeT (Multi-channel Multi-party Meeting Transcription Grand Challenge)

πŸ“Œ 2022 M2MeT β€” 2 papers

Unknown (2 papers)

πŸ“Œ MISP (Multimodal Information Based Speech Processing)

πŸ“Œ 2022 MISP Track1 β€” 3 papers

Unknown (3 papers)

πŸ“Œ DIHARD

πŸ“Œ 2020 DIHARD III

πŸ“Œ Track1 β€” 3 papers

Unknown (3 papers)

πŸ“Œ Track2 β€” 3 papers

Unknown (3 papers)

πŸ“Œ Etc.


πŸ† The DISPLACE Challenge 2023 β€” 2 papers

πŸ”₯ 2023 (2 papers)
  • "The DISPLACE Challenge 2023 - DIarization of SPeaker and LAnguage in Conversational Environments," in Proc. Interspeech, 2023. πŸ“„ πŸ”—
  • "The SpeeD--ZevoTech submission at DISPLACE 2023," in Proc. Interspeech, 2023. πŸ“„

πŸ† MERLIon CCS Challenge 2023 β€” 1 paper

πŸ”₯ 2023 (1 paper)
  • "MERLIon CCS Challenge: A English-Mandarin code-switching child-directed speech corpus for language identification and diarization," in Proc. Interspeech, 2023. πŸ“„ πŸ”—

πŸ“Œ CHiME-6


πŸ† ICMC-ASR Grand Challenge (ICASSP2024) β€” 2 papers

πŸ”₯ 2023-2024 (2 papers)
  • "ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge," 2023. πŸ“„
  • "The NUS-HLT System for ICASSP2024 ICMC-ASR Grand Challenge," in Technical Report, 2023. πŸ“„

πŸ“Œ The Second DISPLACE β€” 1 paper

πŸ”₯ 2024 (1 paper)
  • "The Second DISPLACE Challenge : DIarization of SPeaker and LAnguage in Conversational Environments," in Proc. Interspeech, 2024. πŸ“„

πŸ“Œ CHiME-8 β€” 1 paper

πŸ”₯ 2024 (1 paper)
  • "The CHiME-8 DASR Challenge for Generalizable and Array Agnostic Distant Automatic Speech Recognition and Diarization," 2024. πŸ“„

πŸ“Œ MISP 2025 (Interspeech 2025) β€” 2 papers

πŸ”₯ 2025 (2 papers)
  • "The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition," 2025. πŸ“„
  • "Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge," in Proc. Interspeech, 2025. πŸ“„

πŸ”— Other Awesome Lists

awesome
awesome-list
speaker-diarization

Contributors

DongKeon

11 commits

haerski

1 commits

DongKeon/Awesome-Speaker-Diarization

Some comprehensive papers about speaker diarization

372

12 commits

updated Mar 24, 2026

See the code

README

🎀 Awesome Speaker Diarization Awesome

Papers Code Notion DB Updated

πŸ“„ Paper Β· πŸ’» Code Β· πŸ“ Review Β· 🎬 Video/Demo Β· πŸ“Š Slides Β· πŸ”— Other


🧠 Core Methods

EEND TS-VAD Clustering Embedding Self-Supervised

πŸ”Œ Extensions

Online Multi-Channel Sep/TSE

🌐 Cross-Modal

ASR Vision NLP/LLM Emotion

πŸ”— Related

VAD/OSD/SCD Speaker Rec Personal VAD Spoofing TTS Child-Adult

πŸ“¦ Resources

Dataset Tools Reviews Measurement Scoring

Challenge


πŸ“– Overview β€” 1 paper

2020 (1 paper)
  • DIHARD Keynote Session: The yellow brick road of diarization, challenges and other neural paths πŸ“Š 🎬

πŸ“ Reviews β€” 3 papers

πŸ”₯ 2023-2024 (3 papers)
  • "Overview of Speaker Modeling and Its Applications: From the Lens of Deep Speaker Representation Learning," in Submitted to IEEE/ACM TASLP, 2024. πŸ“„
  • β€œA review of speaker diarization: Recent advances with deep learning”, in Computer Speech & Language, Volume 72, 2023. (USC) πŸ“„
  • "An Experimental Review of Speaker Diarization methods with application to Two-Speaker Conversational Telephone Speech recordings", in Computer Speech & Language, 2023. πŸ“„

πŸ“š EEND (End-to-End Neural Diarization)-based β€” 50 papers

πŸ”₯ 2025 (7 papers)
  • "Mamba-based Segmentation Model for Speaker Diarization," Proc. ICASSP, 2025. (NTT) πŸ“„ πŸ’»
  • "Pushing the Limits of End-to-End Diarization," in Proc. Interspeech, 2025. πŸ“„
  • VBx-EEND-VC: "VBx for End-to-End Neural and Clustering-based Diarization," in arXiv:2510.19572, 2025. (BUT) πŸ“„
  • "Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling," in arXiv:2506.05593, 2025. (OSU) πŸ“„
  • DLF-EEND: "Dynamic Layer Fusion for End-to-End Speaker Diarization," in Proc. Interspeech, 2025. πŸ“„
  • "End-to-End Diarization utilizing Attractor Deep Clustering," in Proc. Interspeech, 2025. (JHU, OSU) πŸ“„
  • "Pretraining Multi-Speaker Identification for Neural Speaker Diarization," in Proc. Interspeech, 2025. (NTT) πŸ“„
2024 (9 papers)
  • "NTT speaker diarization system for CHiME-7: multi-domain, multi-microphone End-to-end and vector clustering diarization," in Proc. ICASSP, 2024. (NTT) πŸ“„
  • AED-EEND-EE: "Attention-based Encoder-Decoder End-to-End Neural Diarization with Embedding Enhancer," in IEEE/ACM TASLP, 2024. (SJTU) πŸ“„ πŸ“
  • "DiaPer: End-to-End Neural Diarization with Perceiver-Based Attractors," in IEEE/ACM TASLP, 2024. (BUT) πŸ“„ πŸ’» πŸ“
  • "EEND-DEMUX: End-to-End Neural Speaker Diarization via Demultiplexed Speaker Embeddings," in Submitted to IEEE SPL, 2024. (SNU) πŸ“„ πŸ“
  • "EEND-M2F: Masked-attention mask transformers for speaker diarization," in Proc. Interspeech, 2024. (Fano Labs) πŸ“„ πŸ“„ πŸ“
  • EEND-NAA (2): "End-to-End Neural Speaker Diarization with Non-Autoregressive Attractors", in IEEE/ACM TASLP, 2024. (JHU) πŸ“„ πŸ“
  • "From Modular to End-to-End Speaker Diarization," Ph.D. thesis, 2024. (BUT) πŸ“„
  • "On the calibration of powerset speaker diarization models," in Proc. Interspeech, 2024. (IRIT) πŸ“„ πŸ“„ πŸ’» πŸ“
  • Local-global EEND: "Speakers Unembedded: Embedding-free Approach to Long-form Neural Diarization," in Proc. Interspeech, 2024. (Amazon) πŸ“„ πŸ“
2023 (11 papers)
  • "Improving Transformer-based End-to-End Speaker Diarization by Assigning Auxiliary Losses to Attention Heads", in Proc. ICASSP, 2023. (HU) πŸ“„
  • EEND-NA: β€œNeural Diarization with Non-Autoregressive Intermediate Attractors”, in Proc. ICASSP, 2023. (LINE) πŸ“„
  • "TOLD: A Novel Two-Stage Overlap-Aware Framework for Speaker Diarization", in Proc. ICASSP, 2023. (Alibaba) πŸ“„ πŸ’»
  • "Improving End-to-End Neural Diarization Using Conversational Summary Representations", in Proc. Interspeech, 2023. (Fano Labs) πŸ“„
  • AED-EEND: β€œAttention-based Encoder-Decoder Network for End-to-End Neural Speaker Diarization with Target Speaker Attractor”, in Proc. Interspeech, 2023. (SJTU) πŸ“„ πŸ“
  • "Self-Distillation into Self-Attention Heads for Improving Transformer-based End-to-End Neural Speaker Diarization", in Proc. Interspeech, 2023. (HU) πŸ“„
  • "Powerset Multi-class Cross Entropy Loss for Neural Speaker Diarization", in Proc. Interspeech, 2023. (Pyannote) πŸ“„ πŸ’»
  • "End-to-End Neural Speaker Diarization with Absolute Speaker Loss", in Proc. Interspeech, 2023. (Pyannote) πŸ“„
  • "Blueprint Separable Subsampling and Aggregate Feature Conformer-Based End-to-End Neural Diarization", in Electronics, 2023. πŸ“„
  • EEND-TA: "Transformer Attractors for Robust and Efficient End-to-End Neural Diarization," in Proc. ASRU, 2023. (Fano Labs) πŸ“„
  • "Robust End-to-End Diarization with Domain Adaptive Training and Multi-Task Learning," in Proc. ASRU, 2023. (Fano Labs) πŸ“„
2022 (10 papers)
  • EEND-EDA (2): β€œEncoder-Decoder Based Attractor Calculation for End-to-End Neural Diarization”, in IEEE/ACM TASLP, 2022. (Hitachi) πŸ“„ πŸ“ πŸ’»
  • "DIVE: End-to-end Speech Diarization via Iterative Speaker Embedding", in Proc. ICASSP, 2022. (Google) πŸ“„
  • RX-EEND: β€œAuxiliary Loss of Transformer with Residual Connection for End-to-End Speaker Diarization”, in Proc. ICASSP, 2022. (GIST) πŸ“„ πŸ“
  • "End-to-end speaker diarization with transformer", in Proc. arXiv, 2022. πŸ“„
  • EEND-VC-iGMM: "Tight integration of neural and clustering-based diarization through deep unfolding of infinite Gaussian mixture model", in Proc. ICASSP, 2022. (NTT) πŸ“„
  • EDA-RC: "Robust End-to-end Speaker Diarization with Generic Neural Clustering", in Proc. Interspeech, 2022. (SJTU) πŸ“„
  • EEND-NAA: "End-to-End Neural Speaker Diarization with an Iterative Refinement of Non-Autoregressive Attention-based Attractors", in Proc. Interspeech, 2022. (JHU) πŸ“„ πŸ“
  • Graph-PIT: "Utterance-by-utterance overlap-aware neural diarization with Graph-PIT", in Proc. Interspeech, 2022. (NTT) πŸ“„ πŸ’»
  • "Efficient Transformers for End-to-End Neural Speaker Diarization", in Proc. IberSPEECH, 2022. πŸ“„
  • EEND-EDA-SpkAtt: "Towards End-to-end Speaker Diarization in the Wild", in arXiv:2211.01299v1, 2022. πŸ“„
2021 (7 papers)
  • CB-EEND: "End-to-end Neural Diarization: From Transformer to Conformer", in Proc. Interspeech, 2021. (Amazon) πŸ“„ πŸ“
  • TDCN-SA: "End-to-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings", in Proc. ICASSP, 2021. (Google) πŸ“„ πŸ“
  • "End-to-End Speaker Diarization Conditioned on Speech Activity and Overlap Detection", in Proc. IEEE SLT, 2021. (Hitachi) πŸ“„
  • EEND-VC (1): "Integrating end-to-end neural and clustering-based diarization: Getting the best of both worlds", in Proc. ICASSP, 2021. (NTT) πŸ“„ πŸ“ πŸ’»
  • EEND-VC (2): "Advances in integration of end-to-end neural and clustering-based diarization for real conversational speech", in Proc. Interspeech, 2021. (NTT) πŸ“„ πŸ“ πŸ’»
  • "Robust End-to-End Speaker Diarization with Conformer and Additive Margin Penalty," in Proc. Interspeech, 2021. (Fano Labs) πŸ“„
  • EEND-GLA: "Towards Neural Diarization for Unlimited Numbers of Speakers Using Global and Local Attractors", in Proc. ASRU, 2021. (Hitachi) πŸ“„ πŸ“
2020 (3 papers)
  • SA-EEND (2): β€œEnd-to-End Neural Diarization: Reformulating Speaker Diarization as Simple Multi-label Classification”, in arXiv:2003.02966, 2020. (Hitachi) πŸ“„ πŸ“
  • SC-EEND: "Neural Speaker Diarization with Speaker-Wise Chain Rule", in arXiv:2006.01796, 2020. (Hitachi) πŸ“„ πŸ“
  • EEND-EDA (1): β€œEnd-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractors”, in Proc. Interspeech, 2020. (Hitachi) πŸ“„ πŸ“ πŸ’»
2023 (1 paper)
  • EEND-IAAE: "End-to-end neural speaker diarization with an iterative adaptive attractor estimation," in Neural Networks, Elsevier. πŸ“„ πŸ’»
2019 (2 papers)
  • BLSTM-EEND: "End-to-End Neural Speaker Diarization with Permutation-Free Objectives", in Proc. Interspeech, 2019. (Hitachi) πŸ“„
  • SA-EEND (1): β€œEnd-to-End Neural Speaker Diarization with Self-attention”, in Proc. ASRU, 2019. (Hitachi) πŸ“„ πŸ’» πŸ’» πŸ“

πŸ”₯ 2024 (2 papers)
  • "Do End-to-End Neural Diarization Attractors Need to Encode Speaker Characteristic Information?," in Proc. Odyssey, 2024. πŸ“„
  • "Leveraging Speaker Embeddings in End-to-End Neural Diarization for Two-Speaker Scenarios," in Proc. Odyssey, 2024. πŸ“„

πŸ“Œ Post-Processing β€” 3 papers

πŸ”₯ 2021-2024 (3 papers)
  • "DiaCorrect: Error Correction Back-end For Speaker Diarization," in Proc. ICASSP, 2024. (BUT) πŸ“„ πŸ’»
  • EENDasP: "End-to-End Speaker Diarization as Post-Processing", in Proc. ICASSP, 2021. (Hitachi) πŸ“„ πŸ“ πŸ’»
  • Dover-Lap: "DOVER-Lap: A Method for Combining Overlap-aware Diarization Outputs", in Proc. IEEE SLT, 2021. (JHU) πŸ“„ πŸ“ πŸ’»

🎯 Using Target Speaker Embedding β€” 17 papers

πŸ”₯ 2025 (4 papers)
  • "Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining," in Proc. ICASSP, 2025. πŸ“„
  • MIMO-TSVAD: "Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization," in IEEE/ACM TASLP, 2025. (DKU) πŸ“„
  • "Mitigating Non-Target Speaker Bias in Guided Speaker Embedding," in Proc. Interspeech, 2025. (NTT) πŸ“„
  • "Diarization-Guided Multi-Speaker Embeddings," in Proc. Interspeech, 2025. (Pyannote) πŸ“„
2024 (3 papers)
  • NSD-MS2S: "Neural Speaker Diarization Using Memory-Aware Multi-Speaker Embedding with Sequence-to-Sequence Architecture, " in Proc. ICASSP, 2024. (USTC) πŸ“„ πŸ’»
  • PET-TSVAD: "Profile-Error-Tolerant Target-Speaker Voice Activity Detection," in Proc. ICASSP, 2024. (Microsoft) πŸ“„
  • Flow-TSVAD: "Target-Speaker Voice Activity Detection via Latent Flow Matching," in arXiv:2409.04859, 2024. (DKU) πŸ“„
2023 (4 papers)
  • EDA-TS-VAD: β€œTarget Speaker Voice Activity Detection with Transformers and Its Integration with End-to-End Neural Diarization”, in Proc. ICASSP, 2023. (Microsoft) πŸ“„
  • Seq2Seq-TS-VAD: β€œTarget-Speaker Voice Activity Detection via Sequence-to-Sequence Prediction”, in Proc. ICASSP, 2023. (DKU) πŸ“„ πŸ“
  • QM-TS-VAD: "Unsupervised Adaptation with Quality-Aware Masking to Improve Target-Speaker Voice Activity Detection for Speaker Diarization", in Proc. Interspeech, 2023. (USTC) πŸ“„
  • "ANSD-MA-MSE: Adaptive Neural Speaker Diarization Using Memory-Aware Multi-Speaker Embedding," in IEEE/ACM TASLP, 2023. (USTC) πŸ“„ πŸ’»
2022 (3 papers)
  • SEND (2): "Speaker Embedding-aware Neural Diarization: an Efficient Framework for Overlapping Speech Diarization in Meeting Scenarios," in arXiv:2203.09767, 2022 (Alibaba) πŸ“„
  • MTEAD: "Multi-target Filter and Detector for Unknown-number Speaker Diarization", in IEEE SPL, 2022. πŸ“„
  • SOND: "Speaker Overlap-aware Neural Diarization for Multi-party Meeting Analysis", in Proc. EMNLP, 2022. (Alibaba) πŸ“„ πŸ’»
2020-2021 (3 papers)
  • SEND (1): "Speaker Embedding-aware Neural Diarization for Flexible Number of Speakers with Textual Information," in arXiv:2111.13694, 2021. (Alibaba) πŸ“„
  • TS-VAD: "Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario", in Proc. Interspeech, 2020. πŸ“„ πŸ’» πŸ“Š
  • β€œThe STC system for the CHiME-6 challenge,” in CHiME Workshop, 2020. πŸ“„

🎯 Target Speech Diarization β€” 1 paper

πŸ”₯ 2024 (1 paper)
  • PTSD: "Prompt-driven Target Speech Diarization," in Proc. ICASSP, 2024. (NUS) πŸ“„

πŸ”€ With Separation or Target Speaker Extraction β€” 12 papers

πŸ”₯ 2025 (3 papers)
  • "Robust Target Speaker Diarization and Separation via Augmented Speaker Embedding Sampling," in Proc. Interspeech, 2025. πŸ“„
  • S2SND: "Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation," in IEEE/ACM TASLP, 2025. (DKU) πŸ“„
  • "Exploring Speaker Diarization with Mixture of Experts," in arXiv:2506.14750, 2025. (USTC) πŸ“„
2024 (7 papers)
  • "TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddings", in IEEE/ACM TASLP, 2024. πŸ“„
  • "Continuous Target Speech Extraction: Enhancing Personalized Diarization and Extraction on Complex Recordings," in arXiv:2401.15993, 2024. (Tencent) πŸ“„ 🎬
  • "PixIT: Joint Training of Speaker Diarization and Speech Separation from Real-world Multi-speaker Recordings," in Proc. Odyssey, 2024. πŸ“„ πŸ’»
  • MC-EEND: "Multi-channel Conversational Speaker Separation via Neural Diarization," in IEEE/ACM TASLP, 2024. (OSU) πŸ“„
  • "USED: Universal Speaker Extraction and Diarization," in submitted to IEEE/ACM TASLP, 2024. (CUHK) πŸ“„ 🎬 πŸ”— πŸ“
  • "Neural Blind Source Separation and Diarization for Distant Speech Recognition," in Proc. Interspeech, 2024. (AIST) πŸ“„
  • "TalTech-IRIT-LIS Speaker and Language Diarization Systems for DISPLACE 2024," in Proc. Interspeech, 2024. (Pyannote) πŸ“„
2021-2022 (2 papers)
  • EEND-SS: "Joint End-to-End Neural Speaker Diarization and Speech Separation for Flexible Number of Speakers”, in Proc. SLT, 2022. (CMU) πŸ“„ πŸ“
  • "Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis," in Proc. SLT, 2021. (JHU) πŸ“„ πŸ”— πŸ“

πŸ“‘ Multi-Channel β€” 14 papers

πŸ”₯ 2025 (5 papers)
  • "Multi-channel Speaker Counting for EEND-VC-based Speaker Diarization on Multi-domain Conversation," in Proc. ICASSP, 2025. (NTT) πŸ“„ πŸ“
  • "Spatially Aware Self-Supervised Models for Multi-Channel Neural Speaker Diarization," in arXiv:2510.14551, 2025. (BUT) πŸ“„ πŸ’»
  • "Multi-Channel Sequence-to-Sequence Neural Diarization for The MISP 2025 Challenge," in arXiv:2505.16387, 2025. πŸ“„
  • "Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings," in Proc. ICASSP, 2025. πŸ“„
  • "Spatio-Spectral Diarization of Meetings by Combining TDOA-based Segmentation and Speaker Embedding-based Clustering," in Proc. Interspeech, 2025. πŸ“„
2024 (5 papers)
  • "UniX-Encoder: A Universal X-Channel Speech Encoder for Ad-Hoc Microphone Array Speech Processing," in arXiv:2310.16367, 2024. (JHU, Tencent) πŸ“„
  • "Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech Detection," in IEEE/ACM TASLP, 2024. πŸ“„
  • "A Spatial Long-Term Iterative Mask Estimation Approach for Multi-Channel Speaker Diarization and Speech Recognition," in Proc. ICASSP, 2024. (USTC) πŸ“„
  • MC-EEND: "Multi-channel Conversational Speaker Separation via Neural Diarization," in IEEE/ACM TASLP, 2024. (OSU) πŸ“„
  • "ASoBO: Attentive Beamformer Selection for Distant Speaker Diarization in Meetings," in Proc. Interspeech, 2024. (LIUM) πŸ“„
2022-2023 (4 papers)
  • "Mutual Learning of Single- and Multi-Channel End-to-End Neural Diarization," in Proc. IEEE SLT, 2023. (Hitachi) πŸ“„
  • "Semi-supervised multi-channel speaker diarization with cross-channel attention", in Proc. ASRU, 2023. (USTC) πŸ“„
  • "Multi-Channel End-to-End Neural Diarization with Distributed Microphones", in Proc. ICASSP, 2022. (Hitachi) πŸ“„
  • "Multi-Channel Speaker Diarization Using Spatial Features for Meetings", in Proc. ICASSP, 2022. (Tencent) πŸ“„

⚑ Online β€” 18 papers

πŸ”₯ 2024-2025 (8 papers)
  • SCDiar: "A Streaming Diarization System based on Speaker Change Detection and Speech Recognition," in Proc. ICASSP, 2025. πŸ“„
  • Streaming Sortformer: "Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering," in Proc. Interspeech, 2025. (NVIDIA) πŸ“„
  • OTS-VAD: "Online Neural Speaker Diarization With Target Speaker Tracking," in IEEE/ACM TASLP, 2024. (DKU) πŸ“„
  • FS-EEND: "Frame-wise streaming end-to-end speaker diarization with non-autoregressive self-attention-based attractors," in Proc. ICASSP, 2024. (Hangzhou) πŸ“„ πŸ’»
  • "Online speaker diarization of meetings guided by speech separation," in Proc. ICASSP, 2024. (LTCI) πŸ“„ πŸ’»
  • "Interrelate Training and Clustering for Online Speaker Diarization," in IEEE/ACM TASLP, 2024. πŸ“„
  • O-EENC-SD: "Efficient Online End-to-End Neural Clustering for Speaker Diarization," in Proc. ICASSP, 2025. πŸ“„
  • LS-EEND: "Long-Form Streaming End-to-End Neural Diarization with Online Attractor Extraction," in IEEE/ACM TASLP, 2025. (Westlake) πŸ“„ πŸ’»
2023 (2 papers)
  • "Absolute decision corrupts absolutely: conservative online speaker diarisation", in Proc. ICASSP, 2023. (Naver) πŸ“„
  • "A Reinforcement Learning Framework for Online Speaker Diarization", in Under Review. NeruIPS, 2023. (CU) πŸ“„
2022 (3 papers)
  • "Low-Latency Online Speaker Diarization with Graph-Based Label Generation", in Proc. Odyssey, 2022. (DKU) πŸ“„
  • EEND-GLA: "Online Neural Diarization of Unlimited Numbers of Speakers Using Global and Local Attractors", in IEEE/ACM TASLP, 2022. (Hitachi) πŸ“„
  • Online TS-VAD: "Online Target Speaker Voice Activity Detection for Speaker Diarization", in Proc. Interspeech, 2022. (DKU) πŸ“„
2021 (4 papers)
  • "Online End-to-End Neural Diarization with Speaker-Tracing Buffer", in Proc. IEEE SLT, 2021. (Hitachi) πŸ“„
  • BW-EDA-EEND: "BW-EDA-EEND: Streaming End-to-End Neural Speaker Diarization for a Variable Number of Speakers", in Proc. Interspeech, 2021. (Amazon) πŸ“„
  • FS-EEND: "Online Streaming End-to-End Neural Diarization Handling Overlapping Speech and Flexible Numbers of Speakers", in Proc. Interspeech, 2021. (Hitachi) πŸ“„ πŸ“
  • Diart: "Overlap-aware low-latency online speaker diarization based on end-to-end local segmentation", in Proc. ASRU, 2021. πŸ“„ πŸ’»
2020 (1 paper)
  • "Supervised online diarization with sample mean loss for multi-domain data", in Proc. ICASSP, 2020 πŸ“„ πŸ’»

πŸ”— Clustering-based β€” 21 papers

πŸ”₯ 2025 (3 papers)
  • E-SHARC: "End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization," in IEEE/ACM TASLP, 2025. (IISC) πŸ“„
  • Pyannote Community-1: "pyannote.audio 4.0 with community-1 open-source diarization model," 2025. πŸ”— πŸ’»
  • "Speaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation Algorithm," in Proc. Interspeech, 2025. πŸ“„
2024 (6 papers)
  • "Overlap-aware End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization," in submitted to IEEE/ACM TASLP, 2024. πŸ“„
  • "Apollo's Unheard Voices: Graph Attention Networks for Speaker Diarization and Clustering for Fearless Steps Apollo Collection," in Proc. ICASSP, 2024. (UTD) πŸ“„
  • "Multi-View Speaker Embedding Learning for Enhanced Stability and Discriminability," in Proc. ICASSP, 2024. (Tsinghua) πŸ“„
  • "Towards Unsupervised Speaker Diarization System for Multilingual Telephone Calls Using Pre-trained Whisper Model and Mixture of Sparse Autoencoders," in arXiv:2407.01963, 2024. πŸ“„
  • "Investigating Confidence Estimation Measures for Speaker Diarization," in Proc. Interspeech, 2024. πŸ“„
  • "Once more Diarization: Improving meeting transcription systems through segment-level speaker reassignment," in Proc. Interspeech, 2024. (PU) πŸ“„ πŸ“„ πŸ’»
2023 (5 papers)
  • SCALE: "Spectral Clustering-aware Learning of Embeddings for Speaker Diarisation", in Proc. ICASSP, 2023. (CAM) πŸ“„
  • SHARC: "Supervised Hierarchical Clustering using Graph Neural Networks for Speaker Diarization", in Proc. ICASSP, 2023. (IISC) πŸ“„
  • CDGCN: "Community Detection Graph Convolutional Network for Overlap-Aware Speaker Diarization," in Proc. ICASSP, 2023. (XMU) πŸ“„
  • "Pyannote.Audio 2.1: Speaker Diarization Pipeline: Principle, Benchmark and Recipe", in Proc. Interspeech, 2023. (CNRS) πŸ“„
  • GADEC: "Graph attention-based deep embedded clustering for speaker diarization,", in Speech Communication, 2023. (NJUPT) πŸ“„
2020-2022 (4 papers)
  • UMAP-Leiden: "Reformulating Speaker Diarization as Community Detection With Emphasis On Topological Structure", in Proc. ICASSP, 2022. (Alibaba) πŸ“„
  • Pyannote 2.0: "End-to-end speaker segmentation for overlap-aware resegmentation", in Proc. Interspeech, 2021. (CNRS) πŸ“„ πŸ’» 🎬
  • Pyannote: "pyannote.audio: neural building blocks for speaker diarization", in Proc. ICASSP, 2020. (CNRS) πŸ“„ πŸ’» 🎬
  • Resegmentation with VB: β€œOverlap-Aware Diarization: Resegmentation Using Neural End-to-End Overlapped Speech Detection”, in Proc. ICASSP, 2020. πŸ“„
2018 (1 paper)
2019 (2 papers)
  • DNC: "Discriminative Neural Clustering for Speaker Diarisation", in Proc. IEEE SLT, 2019. πŸ“„ πŸ’» πŸ“
  • NME-SC: β€œAuto-Tuning Spectral Clustering for Speaker Diarization Using Normalized Maximum Eigengap”, IEEE SPL, 2019. πŸ“„ πŸ’»

πŸ“ Variational Bayes and HMM β€” 24 papers

πŸ”₯ 2023-2024 (3 papers)
  • DVBx: "Discriminative Training of VBx Diarization", in Proc. ICASSP, 2024. (BUT) πŸ“„ πŸ’»
  • MS-VBx: "Multi-Stream Extension of Variational Bayesian HMM Clustering (MS-VBx) for Combined End-to-End and Vector Clustering-based Diarization", in Proc. Interspeech, 2023. (NTT) πŸ“„
  • "Generalized domain adaptation framework for parametric back-end in speaker recognition", in arXiv:2305.15567, 2023. πŸ“„
2021-2022 (4 papers)
  • "Bayesian HMM clustering of x-vector sequences (VBx) in speaker diarization: theory, implementation and analysis on standard tasks", in Computer Speech & Language, 2022. (BUT) πŸ“„
  • DCA-PLDA "A Speaker Verification Backend with Robust Performance across Conditions”, in Computer & Language, 2022. πŸ“„ πŸ’»
  • "Analysis of the but Diarization System for Voxconverse Challenge", in Proc. ICASSP, 2021. (BUT) πŸ“„ πŸ’»
  • "Discriminatively trained probabilistic linear discriminant analysis for speaker verification", in Proc. ICASSP, 2021. πŸ“„
2019-2020 (4 papers)
  • "Optimizing Bayesian Hmm Based X-Vector Clustering for the Second Dihard Speech Diarization Challenge", in Proc. ICASSP, 2020. (BUT) πŸ“„
  • β€œAnalysis of Speaker Diarization Based on Bayesian HMM With Eigenvoice Priors”, IEEE/ACM TASLP, 2019. (BUT) πŸ“„
  • "BUT System Description for DIHARD Speech Diarization Challenge 2019", in arXiv:1910.08847, 2019. (BUT) πŸ“„
  • "Bayesian HMM Based x-Vector Clustering for Speaker Diarization", in Proc. Interspeech, 2019. (BUT) πŸ“„
2018 (5 papers)
  • "Speaker Diarization based on Bayesian HMM with Eigenvoice Priors", in Proc. Odyssey, 2018. (BUT) πŸ“„
  • "VB-HMM Speaker Diarization with Enhanced and Refined Segment Representation", in Proc. Odyssey, 2018. (Tsinghua) πŸ“„
  • "Diarization is hard: some experiences and lessons learned for the JHU team in the inaugural DIHARD challenge", in Proc. Interspeech, 2018. πŸ“„
  • "The speaker partitioning problem", in Proc. Odyssey, 2018. πŸ“„
  • "Estimation of the Number of Speakers with Variational Bayesian PLDA in the DIHARD Diarization Challenge", in Proc. Interspeech, 2018. πŸ“„
2015-2017 (3 papers)
  • "Domain Adaptation of PLDA Models in Broadcast Diarization by Means of Unsupervised Speaker Clustering, in Proc. Interspeech, 2017. πŸ“„
  • "Iterative PLDA Adaptation for Speaker Diarization", in Proc. Interspeech, 2016. πŸ“„
  • "Diarization resegmentation in the factor analysis subspace", in Proc. ICASSP, 2015. πŸ“„
2011-2014 (3 papers)
  • "Speaker diarization with plda i-vector scoring and unsupervised calibration", in Proc. IEEE SLT, 2014. πŸ“„
  • "Unsupervised Methods for Speaker Diarization: An Integrated and Iterative Approach", IEEE/ACM TASLP, 2013. πŸ“„
  • "Analysis of i-vector length normalization in speaker recognition systems", in Proc. Interspeech, 2011. πŸ“„
2005-2008 (2 papers)
  • "Bayesian analysis of speaker diarization with eigenvoice priors", in CRIM, Montreal, Technical Report, 2008. πŸ“„
  • "Variational Bayesian methods for audio indexing", in Proc. ICMI-MLMI, 2005. πŸ“„

🧩 Embedding (With Clustering) β€” 14 papers

πŸ”₯ 2024 (4 papers)
  • "Geodesic interpolation of frame-wise speaker embeddings for the diarization of meeting scenarios", in Proc. ICASSP, 2024. (PU) πŸ“„ πŸ“
  • "Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization," in Proc. Odyssey, 2024. (IDLab) πŸ“„
  • "Efficient Speaker Embedding Extraction Using a Twofold Sliding Window Algorithm for Speaker Diarization," in Proc. Interspeech, 2024. (HU) πŸ“„
  • "Variable Segment Length and Domain-Adapted Feature Optimization for Speaker Diarization," in Proc. Interspeech, 2024. (XMU) πŸ“„ πŸ’»
2023 (5 papers)
  • "In Search of Strong Embedding Extractors For Speaker Diarization", in Proc. ICASSP, 2023. (Naver) πŸ“„ πŸ“
  • DR-DESA: "Advancing the dimensionality reduction of speaker embeddings for speaker diarisation: disentangling noise and informing speech activity", in Proc. ICASSP, 2023. (Naver) πŸ“„ πŸ“
  • HEE: "High-resolution embedding extractor for speaker diarisation", in Proc. ICASSP, 2023. (Naver) πŸ“„ πŸ“
  • "Frame-wise and overlap-robust speaker embeddings for meeting diarization", in Proc. ICASSP, 2023. (PU) πŸ“„ πŸ“
  • "A Teacher-Student approach for extracting informative speaker embeddings from speech mixtures", in Proc. Interspeech, 2023. (PU) πŸ“„
2022 (3 papers)
  • GAT+AA: "Multi-scale speaker embedding-based graph attention networks for speaker diarisation", in Proc. ICASSP, 2022. (Naver) πŸ“„
  • MSDD: "Multi-scale Speaker Diarization with Dynamic Scale Weighting", in Proc. Interspeech, 2022. (NVIDIA) πŸ“„ πŸ’» πŸ”—
  • PRISM: "PRISM: Pre-trained Indeterminate Speaker Representation Model for Speaker Diarization and Speaker Verification", in Proc. Interspeech, 2022. (Alibaba) πŸ“„
2021 (2 papers)
  • "Multi-Scale Speaker Diarization With Neural Affinity Score Fusion", in Proc. ICASSP, 2021. (USC) πŸ“„
  • AA+DR+NS: "Adapting Speaker Embeddings for Speaker Diarisation", in Proc. Interspeech, 2021. (Naver) πŸ“„ πŸ“

πŸͺͺ With Speaker Identification β€” 1 paper

πŸ”₯ 2024 (1 paper)
  • "Uncertainty Quantification in Machine Learning for Joint Speaker Diarization and Identification, in Submitted to IEEE/ACM TASLP, 2024. πŸ“„

πŸ”Š Speaker Recognition & Verification β€” 7 papers

πŸ”₯ 2024 (3 papers)
  • "Rethinking Session Variability: Leveraging Session Embeddings for Session Robustness in Speaker Verification," in Proc. ICASSP, 2024. (Naver) πŸ“„
  • "Leveraging In-the-Wild Data for Effective Self-Supervised Pretraining in Speaker Recognition," in Proc. ICASSP, 2024. (CUHK) πŸ“„
  • "Disentangled Representation Learning for Environment-agnostic Speaker Recognition," in Proc. Interspeech, 2024. (KAIST) πŸ“„ πŸ“„ πŸ’»
2023 (3 papers)
  • "Build a SRE Challenge System: Lessons from VoxSRC 2022 and CNSRC 2022," in Proc. Interspeech, 2023. (SJTU) πŸ“„
  • RecXi "Disentangling Voice and Content with Self-Supervision for Speaker Recognition," in Proc. NeurIPS, 2023. (A*STAR) πŸ“„
  • "ECAPA2: A Hybrid Neural Network Architecture and Training Strategy for Robust Speaker Embeddings," in Proc. ASRU, 2023. (IDLab) πŸ“„ πŸ”— πŸ“
2021 (1 paper)
  • "Xi-Vector Embedding for Speaker Recognition," in IEEE, SPL. (A*STAR) πŸ“„ πŸ“

πŸ“Š Scoring β€” 3 papers

πŸ”₯ 2019-2023 (3 papers)
  • β€œSimilarity Measurement of Segment-Level Speaker Embeddings in Speaker Diarization”, IEEE/ACM TASLP, 2023. (DKU) πŸ“„
  • "Self-Attentive Similarity Measurement Strategies in Speaker Diarization", in Proc. Interspeech, 2020. (DKU) πŸ“„
  • LSTM scoring: "LSTM based Similarity Measurement with Spectral Clustering for Speaker Diarization", in Proc. Interspeech, 2019. (DKU) πŸ“„

πŸ—£οΈ With ASR β€” 27 papers

πŸ”₯ 2025-2026 (8 papers)
  • SE-DiCoW: "Self-Enrolled Diarization-Conditioned Whisper," in arXiv:2601.19194, 2026. (BUT) πŸ“„
  • TagSpeech: "End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding," in arXiv:2601.06896, 2026. πŸ“„
  • DiCoW: "Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition," in Proc. ICASSP, 2025. (BUT) πŸ“„ πŸ’»
  • "Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition," in arXiv:2510.03723, 2025. (BUT) πŸ“„
  • "Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models," in arXiv:2506.05796, 2025. (DKU) πŸ“„
  • "Language Modelling for Speaker Diarization in Telephonic Interviews," in arXiv:2501.17893, 2025. πŸ“„
  • SC-SOT: "Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition," in Proc. Interspeech, 2025. πŸ“„
    • "Target Speaker ASR with Whisper," in Submitted to ICASSP, 2025. (BUT) πŸ“„ πŸ’»
2024 (12 papers)
  • "Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach,", in Proc. ICASSP, 2024. (NVIDIA) πŸ“„
  • WEEND: "Towards Word-Level End-to-End Neural Speaker Diarization with Auxiliary Network," in arXiv:2309.08489, 2024. (Google) πŸ“„ πŸ”—
  • "One model to rule them all ? Towards End-to-End Joint Speaker Diarization and Speech Recognition", in Proc. ICASSP, 2024. (CMU) πŸ“„
  • "Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization," in arXiv:2309.16482, 2024. (PU) πŸ“„
  • β€œJoint Inference of Speaker Diarization and ASR with Multi-Stage Information Sharing," in Proc. ICASSP, 2024. (DKU) πŸ“„
  • "Multitask Speech Recognition and Speaker Change Detection for Unknown Number of Speakers" in Proc. ICASSP, 2024. (Idiap) πŸ“„
  • "A Spatial Long-Term Iterative Mask Estimation Approach for Multi-Channel Speaker Diarization and Speech Recognition," in Proc. ICASSP, 2024. (USTC) πŸ“„
  • "On the Success and Limitations of Auxiliary Network Based Word-Level End-to-End Neural Speaker Diarization," in Proc. Interspeech, 2024. (Google) πŸ“„
  • Sortformer: "Seamless Integration of Speaker Diarization and ASR by Bridging Timestamps and Tokens," in Proc. ICML, 2025. (NVIDIA) πŸ“„ πŸ“
    • "Speaker Mask Transformer for Multi-talker Overlapped Speech Recognition," in arXiv:2312.10959, 2024. (NICT) πŸ“„
    • "On Speaker Attribution with SURT," in Proc. Odyssey, 2024. (JHU) πŸ“„
    • "Improving Speaker Assignment in Speaker-Attributed ASR for Real Meeting Applications," in Proc. Odyssey, 2024. (CNRS) πŸ“„
2023 (5 papers)
  • "Unified Modeling of Multi-Talker Overlapped Speech Recognition and Diarization with a Sidecar Separator", in Proc. Interspeech, 2023. (CUHK) πŸ“„
  • "Multi-resolution Approach to Identification of Spoken Languages and to Improve Overall Language Diarization System using Whisper Model", in Proc. Interspeech, 2023.
  • "Speaker Diarization for ASR Output with T-vectors: A Sequence Classification Approach", in Proc. Interspeech, 2023. πŸ“„
  • "Lexical Speaker Error Correction: Leveraging Language Models for Speaker Diarization Error Correction", in Proc. Interspeech, 2023. (Amazon) πŸ“„
    • "SA-Paraformer: Non-autoregressive End-to-End Speaker-Attributed ASR," in Proc. ASRU, 2023. (Alibaba) πŸ“„
2022 (2 papers)
  • "Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed ASR," in Proc. ICASSP, 2022. πŸ“„
  • "Tandem Multitask Training of Speaker Diarisation and Speech Recognition for Meeting Transcription", in Proc. Interspeech, 2022. πŸ“„

πŸ’¬ With NLP / LLM / Language β€” 12 papers

πŸ”₯ 2024-2025 (7 papers)
  • SpeakerLM: "End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models," in arXiv:2508.06372, 2025. πŸ“„
  • "Interactive Real-Time Speaker Diarization Correction with Human Feedback," in arXiv:2509.18377, 2025. πŸ“„
  • "DiariST: Streaming Speech Translation with Speaker Diarization," in Proc. ICASSP, 2024. (Microsoft) πŸ“„ πŸ’»
  • JPCP: "Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation," in arXiv:2309.10456, 2024. (Alibaba) πŸ“„
  • "DiarizationLM: Speaker Diarization Post-Processing with Large Language Models," in Proc. Interspeech, 2024. (Google) πŸ“„ πŸ“„ πŸ’» πŸ“
  • "LLM-based speaker diarization correction: A generalizable approach," in Submitted to IEEE/ACM TASLP, 2024. πŸ“„
  • "AG-LSEC: Audio Grounded Lexical Speaker Error Correction," in Proc. Interspeech, 2024. (Amazon) πŸ“„
2023 (3 papers)
  • "Exploring Speaker-Related Information in Spoken Language Understanding for Better Speaker Diarization", in Proc. ACL, 2023. (Alibaba) πŸ“„
  • MMSCD, "Encoder-decoder multimodal speaker change detection", in Proc. Interspeech, 2023. (Naver) πŸ“„
  • "Aligning Speakers: Evaluating and Visualizing Text-based Diarization Using Efficient Multiple Sequence Alignment,", in Proc. ICTAI, 2023. πŸ“„

🌐 Language Diarization β€” 2 papers

πŸ”₯ 2023 (2 papers)
  • "End-to-End Spoken Language Diarization with Wav2vec Embeddings", in Proc. Interspeech, 2023. πŸ“„ πŸ’»
  • "Multi-resolution Approach to Identification of Spoken Languages and To Improve Overall Language Diarization System Using Whisper Model," in Proc. Interspeech, 2023. πŸ“„

πŸ‘οΈ With Vision β€” 21 papers

πŸ”₯ 2025-2026 (4 papers)
  • CineSRD: "Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarization," in arXiv:2603.16966, 2026. πŸ“„
  • "Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization," in Proc. ACL, 2025. πŸ“„
  • "Cross-Attention and Self-Attention for Audio-visual Speaker Diarization," in arXiv:2506.02621, 2025. πŸ“„
  • "Count Your Speakers! Multitask Learning for Multimodal Speaker Diarization," in Proc. Interspeech, 2025. πŸ“„
2024 (6 papers)
  • "Speaker Diarization of Scripted Audiovisual Content," in arXiv:2308.02160, 2024. (Amazon) πŸ“„
  • "AFL-Net: Integrating Audio, Facial, and Lip Modalities with Cross-Attention for Robust Speaker Diarization in the Wild," in Proc. ICASSP, 2024. (Tencent) πŸ“„ 🎬
  • "Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation," in Proc. AAAI, 2024. (Tencent) πŸ“„
  • "3D-Speaker-Toolkit: An Open Source Toolkit for Multi-modal Speaker Verification and Diarization," in arXiv:2403.19971, 2024. (Alibaba) πŸ“„ πŸ’»
  • "Target Speech Diarization with Multimodal Prompts," in Submitted to IEEE/ACM TASLP, 2024. (NUS) πŸ“„
  • MFV-KSD: "Multi-Stage Face-Voice Association Learning with Keynote Speaker Diarization," in Submitted to ACM MM, 2024. πŸ“„ πŸ’»
2023 (5 papers)
  • "Audio-Visual Speaker Diarization in the Framework of Multi-User Human-Robot Interaction", in Proc. ICASSP, 2023. πŸ“„
  • STHG: "Spatial-Temporal Heterogeneous Graph Learning for Advanced Audio-Visual Diarization, in Proc. CVPR, 2023. (Intel) πŸ“„
  • "Uncertainty-Guided End-to-End Audio-Visual Speaker Diarization for Far-Field Recordings," in Proc. ACM MM, 2023. πŸ“„
  • "Joint Training or Not: An Exploration of Pre-trained Speech Models in Audio-Visual Speaker Diarization," in Springer Computer Science proceedings, 2023. πŸ“„
  • EEND-EDA++: "Late Audio-Visual Fusion for In-The-Wild Speaker Diarization," in arXiv:2211.01299v2, 2023. πŸ“„
2022 (3 papers)
  • AVA-AVD (AVR-Net): "AVA-AVD: Audio-Visual Speaker Diarization in the Wild", in Proc. ACM MM, 2022. πŸ“„ πŸ’» 🎬
  • "End-to-End Audio-Visual Neural Speaker Diarization", in Proc. Interspeech, 2022. (USTC) πŸ“„ πŸ’» πŸ“
  • DyViSE: "DyViSE: Dynamic Vision-Guided Speaker Embedding for Audio-Visual Speaker Diarization", in Proc. MMSP, 2022. (THU) πŸ“„ πŸ’»
2024 (1 paper)
  • "Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization," in Submitted to IEEE/ACM TASLP. (DKU) πŸ“„
2019-2020 (2 papers)
  • "Self-supervised learning for audio-visual speaker diarization", in Proc. ICASSP, 2020. (Tencent) πŸ“„ πŸ”—
  • "Who said that?: Audio-visual speaker diarisation of real-world meetings", in Proc. Interspeech, 2019. (Naver) πŸ“„

πŸ”₯ 2024 (1 paper)
  • "Spoof Diarization: "What Spoofed When" in Partially Spoofed Audio," in Proc. Interspeech, 2024. (IITK) πŸ“„

πŸ“Œ Speaker Anonymization β€” 1 paper

πŸ”₯ 2024 (1 paper)
  • "A Benchmark for Multi-speaker Anonymization," in Submitted to IEEE/ACM TASLP, 2024. (SIT) πŸ“„ πŸ’»

πŸ“Œ Singing Diarization β€” 1 paper

πŸ”₯ 2024 (1 paper)
  • "Song Data Cleansing for End-to-End Neural Singer Diarization Using Neural Analysis and Synthesis Framework," in Proc. Interspeech, 2024. (LY) πŸ“„

😊 With Emotion β€” 3 papers

πŸ”₯ 2023-2024 (3 papers)
  • "ED-TTS: Multi-scale Emotion Modeling using Cross-domain Emotion Diarization for Emotional Speech Synthesis, in Proc. ICASSP, 2024. πŸ“„
  • "Speech Emotion Diarization: Which Emotion Appears When?," in Proc. ASRU, 2023. (Zaion) πŸ“„
  • "EmoDiarize: Speaker Diarization and Emotion Identification from Speech Signals using Convolutional Neural Networks," in arxiv:2310.12851, 2023. πŸ“„

πŸŽ™οΈ Personal VAD β€” 2 papers

πŸ”₯ 2020-2023 (2 papers)
  • "SVVAD: Personal Voice Activity Detection for Speaker Verification", in Proc. Interspeech, 2023. πŸ“„
  • "Personal VAD: Speaker-Conditioned Voice Activity Detection", in Proc. Odyssey, 2020. (Google) πŸ“„

πŸ“ˆ VAD & OSD & SCD β€” 10 papers

πŸ”₯ 2024 (3 papers)
  • "USM-SCD: Multilingual Speaker Change Detection Based on Large Pretrained Foundation Models," in Proc. ICASSP, 2024. (Google) πŸ“„
  • "Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech Detection," in IEEE/ACM TASLP, 2024. πŸ“„
  • "Speaker Change Detection with Weighted-sum Knowledge Distillation based on Self-supervised Pre-trained Models," in Proc. Interspeech, 2024. πŸ“„
2023 (4 papers)
  • "Multitask Detection of Speaker Changes, Overlapping Speech and Voice Activity Using wav2vec 2.0," in Proc. ICASSP, 2023. πŸ“„ πŸ’»
  • "Semantic VAD: Low-Latency Voice Activity Detection for Speech Interaction," in Proc. Interspeech, 2023. πŸ“„
  • "Joint speech and overlap detection: a benchmark over multiple audio setup and speech domains," in arxiv:2307.13012, 2023. πŸ“„
  • "Advancing the study of Large-Scale Learning in Overlapped Speech Detection," in arXiv:2308.05987, 2023. πŸ“„
2022 (3 papers)
  • "Overlapped Speech Detection in Broadcast Streams Using X-vectors," in Proc. Interspeech, 2022. πŸ“„
  • "Overlapped speech and gender detection with WavLM pre-trained features," in Proc. Interspeech, 2022. πŸ“„
  • "Microphone Array Channel Combination Algorithms for Overlapped Speech Detection," in Proc. Interspeech, 2022. πŸ“„

πŸ“Š Dataset β€” 19 papers

πŸ”₯ 2024-2025 (6 papers)
  • "Conversations in the wild: Data collection, automatic generation and evaluation," in Computer Speech & Language, 2025. πŸ“„
  • M3SD: "Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset," in arXiv:2506.14427, 2025. πŸ“„
  • "VoxBlink: X-Large Speaker Verification Dataset on Camera", in Proc. ICASSP, 2024. πŸ“„ πŸ”—
  • "NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant Meeting Transcription," in arXiv:2401.08887, 2024. (MS) πŸ“„
  • "A Comparative Analysis of Speaker Diarization Models: Creating a Dataset for German Dialectal Speech," in Proc. ACL, 2024. πŸ“„
  • "ALLIES: A Speech Corpus for Segmentation, Speaker Diarization, Speech Recognition and Speaker Change Detection," in Proc. ACL, 2024. (LIUM) πŸ“„
2020-2022 (5 papers)
  • Ego4D: " Around the World in 3,000 Hours of Egocentric Video," in Proc. CVPR, 2022. (Meta) πŸ“„ πŸ’» πŸ”—
  • AliMeeting: "Summary On The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge," in Proc. ICASSP, 2022. (Alibaba) πŸ“„ πŸ”— πŸ’»
  • Voxconverse: "Spot the conversation: speaker diarisation in the wild", in Proc. Interspeech, 2020. (VGG, Naver) πŸ“„ πŸ’» πŸ”—
  • MSDWild: Multi-modal Speaker Diarization Dataset in the Wild, in Proc. Interspeech, 2020. πŸ“„ πŸ”—
  • "LibriMix: An Open-Source Dataset for Generalizable Speech Separation," in arXiv:2005.11262, 2020. πŸ“„ πŸ’»

πŸ“Š Simulated Dataset β€” 8 papers

πŸ”₯ 2023-2024 (4 papers)
  • "Enhancing low-latency speaker diarization with spatial dictionary learning," in Proc. ICASSP, 2024. (NTU) πŸ“„ πŸ“Š
  • "Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling," in Proc. ICASSP, 2024. (OSU) πŸ“„
  • "Multi-Speaker and Wide-Band Simulated Conversations as Training Data for End-to-End Neural Diarization", in Proc. ICASSP, 2023. (BUT) πŸ“„ πŸ’» πŸ“
  • "Property-Aware Multi-Speaker Data Simulation: A Probabilistic Modelling Technique for Synthetic Data Generation," in CHiME-7 Workshop, 2023. (NVIDIA) πŸ“„
2022 (3 papers)
  • "From simulated mixtures to simulated conversations as training data for end-to-end neural diarization" , in Proc. Interspeech, 2022. (BUT) πŸ“„ πŸ’» πŸ“
  • Markov selection: "Improving the naturalness of simulated conversations for end-to-end neural diarization", in Proc. Odyssey, 2022. (Hitachi) πŸ“„
  • EEND-EDA-SpkAtt: "Towards End-to-end Speaker Diarization in the Wild", in arXiv:2211.01299v1, 2022. πŸ“„
2019 (1 paper)
  • Concat-and-sum: "End-to-end neuarl speaker diarization with permuation-free objectives", in Proc. Interspeech, 2019. πŸ“„

πŸ› οΈ Tools β€” 1 paper

πŸ”₯ 2024 (1 paper)
  • "Gryannote open-source speaker diarization labeling tool," in Proc. Interspeech (Show and Tell), 2024. (IRIT) πŸ“„ πŸ’»

πŸ”„ Self-Supervised β€” 5 papers

πŸ”₯ 2025 (3 papers)
  • DiariZen: "Leveraging Self-Supervised Learning for Speaker Diarization," in Proc. ICASSP," 2025. (BUT) πŸ“„ πŸ“„ πŸ’» πŸ“
  • "Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models," in arXiv:2506.18623, 2025. (BUT) πŸ“„ πŸ’»
  • "Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization," in arXiv:2505.24111, 2025. πŸ“„
2022 (2 papers)
  • β€œSelf-supervised Speaker Diarization”, in Proc. Interspeech, 2022. πŸ“„
  • CSDA: "Continual Self-Supervised Domain Adaptation for End-to-End Speaker Diarization", in Proc. IEEE SLT, 2022. (CNRS) πŸ“„ πŸ’»

πŸ”ƒ Semi-Supervised β€” 1 paper

πŸ”₯ 2017 (1 paper)
  • "Active Learning Based Constrained Clustering For Speaker Diarization", in IEEE/ACM TASLP, 2017. (UT) πŸ“„

πŸ“ Measurement β€” 3 papers

πŸ”₯ 2022-2025 (3 papers)
  • SDBench: β€œA Comprehensive Benchmark Suite for Speaker Diarization,” in Proc. Interspeech, 2025. πŸ“„
  • β€œBenchmarking Diarization Models,” in arXiv:2509.26177, 2025. πŸ“„
  • BER: β€œBalanced Error Rate For Speaker Diarization”, in Proc. arXiv:2211.04304, 2022 πŸ“„ πŸ’»

πŸ‘Ά Child-Adult β€” 2 papers

πŸ”₯ 2023-2024 (2 papers)
  • "Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions," in Proc. Interspeech, 2024. (USC) πŸ“„
  • "Robust Self Supervised Speech Embeddings for Child-Adult Classification in Interactions involving Children with Autism," in Proc. Interspeech, 2023. πŸ“„

πŸ† Challenge


πŸ”Š VoxSRC (VoxCeleb Speaker Recognition Challenge)

πŸ“Œ VoxSRC-20 Track4 β€” 3 papers

Unknown (3 papers)

πŸ“Œ VoxSRC-21 Track4 β€” 3 papers

Unknown (3 papers)

πŸ“Œ VoxSRC-22 Track4 β€” 3 papers

Unknown (3 papers)

πŸ“Œ VoxSRC-23 Track4 β€” 5 papers

Unknown (5 papers)

πŸ“‘ M2MeT (Multi-channel Multi-party Meeting Transcription Grand Challenge)

πŸ“Œ 2022 M2MeT β€” 2 papers

Unknown (2 papers)

πŸ“Œ MISP (Multimodal Information Based Speech Processing)

πŸ“Œ 2022 MISP Track1 β€” 3 papers

Unknown (3 papers)

πŸ“Œ DIHARD

πŸ“Œ 2020 DIHARD III

πŸ“Œ Track1 β€” 3 papers

Unknown (3 papers)

πŸ“Œ Track2 β€” 3 papers

Unknown (3 papers)

πŸ“Œ Etc.


πŸ† The DISPLACE Challenge 2023 β€” 2 papers

πŸ”₯ 2023 (2 papers)
  • "The DISPLACE Challenge 2023 - DIarization of SPeaker and LAnguage in Conversational Environments," in Proc. Interspeech, 2023. πŸ“„ πŸ”—
  • "The SpeeD--ZevoTech submission at DISPLACE 2023," in Proc. Interspeech, 2023. πŸ“„

πŸ† MERLIon CCS Challenge 2023 β€” 1 paper

πŸ”₯ 2023 (1 paper)
  • "MERLIon CCS Challenge: A English-Mandarin code-switching child-directed speech corpus for language identification and diarization," in Proc. Interspeech, 2023. πŸ“„ πŸ”—

πŸ“Œ CHiME-6


πŸ† ICMC-ASR Grand Challenge (ICASSP2024) β€” 2 papers

πŸ”₯ 2023-2024 (2 papers)
  • "ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge," 2023. πŸ“„
  • "The NUS-HLT System for ICASSP2024 ICMC-ASR Grand Challenge," in Technical Report, 2023. πŸ“„

πŸ“Œ The Second DISPLACE β€” 1 paper

πŸ”₯ 2024 (1 paper)
  • "The Second DISPLACE Challenge : DIarization of SPeaker and LAnguage in Conversational Environments," in Proc. Interspeech, 2024. πŸ“„

πŸ“Œ CHiME-8 β€” 1 paper

πŸ”₯ 2024 (1 paper)
  • "The CHiME-8 DASR Challenge for Generalizable and Array Agnostic Distant Automatic Speech Recognition and Diarization," 2024. πŸ“„

πŸ“Œ MISP 2025 (Interspeech 2025) β€” 2 papers

πŸ”₯ 2025 (2 papers)
  • "The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition," 2025. πŸ“„
  • "Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge," in Proc. Interspeech, 2025. πŸ“„

πŸ”— Other Awesome Lists

awesome
awesome-list
speaker-diarization

Contributors

DongKeon

11 commits

haerski

1 commits