"Mamba-based Segmentation Model for Speaker Diarization," Proc. ICASSP, 2025. (NTT) ππ»
"Pushing the Limits of End-to-End Diarization," in Proc. Interspeech, 2025. π
VBx-EEND-VC: "VBx for End-to-End Neural and Clustering-based Diarization," in arXiv:2510.19572, 2025. (BUT) π
"Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling," in arXiv:2506.05593, 2025. (OSU) π
DLF-EEND: "Dynamic Layer Fusion for End-to-End Speaker Diarization," in Proc. Interspeech, 2025. π
"End-to-End Diarization utilizing Attractor Deep Clustering," in Proc. Interspeech, 2025. (JHU, OSU) π
"Pretraining Multi-Speaker Identification for Neural Speaker Diarization," in Proc. Interspeech, 2025. (NTT) π
2024 (9 papers)
"NTT speaker diarization system for CHiME-7: multi-domain, multi-microphone End-to-end and vector clustering diarization," in Proc. ICASSP, 2024. (NTT) π
AED-EEND-EE: "Attention-based Encoder-Decoder End-to-End Neural Diarization with Embedding Enhancer," in IEEE/ACM TASLP, 2024. (SJTU) ππ
"DiaPer: End-to-End Neural Diarization with Perceiver-Based Attractors," in IEEE/ACM TASLP, 2024. (BUT) ππ»π
"EEND-DEMUX: End-to-End Neural Speaker Diarization via Demultiplexed Speaker Embeddings," in Submitted to IEEE SPL, 2024. (SNU) ππ
"EEND-M2F: Masked-attention mask transformers for speaker diarization," in Proc. Interspeech, 2024. (Fano Labs) πππ
EEND-NAA (2): "End-to-End Neural Speaker Diarization with Non-Autoregressive Attractors", in IEEE/ACM TASLP, 2024. (JHU) ππ
"On the calibration of powerset speaker diarization models," in Proc. Interspeech, 2024. (IRIT) πππ»π
Local-global EEND: "Speakers Unembedded: Embedding-free Approach to Long-form Neural Diarization," in Proc. Interspeech, 2024. (Amazon) ππ
2023 (11 papers)
"Improving Transformer-based End-to-End Speaker Diarization by Assigning Auxiliary Losses to Attention Heads", in Proc. ICASSP, 2023. (HU) π
EEND-NA: βNeural Diarization with Non-Autoregressive Intermediate Attractorsβ, in Proc. ICASSP, 2023. (LINE) π
"TOLD: A Novel Two-Stage Overlap-Aware Framework for Speaker Diarization", in Proc. ICASSP, 2023. (Alibaba) ππ»
"Improving End-to-End Neural Diarization Using Conversational Summary Representations", in Proc. Interspeech, 2023. (Fano Labs) π
AED-EEND: βAttention-based Encoder-Decoder Network for End-to-End Neural Speaker Diarization with Target Speaker Attractorβ, in Proc. Interspeech, 2023. (SJTU) ππ
"Self-Distillation into Self-Attention Heads for Improving Transformer-based End-to-End Neural Speaker Diarization", in Proc. Interspeech, 2023. (HU) π
"Powerset Multi-class Cross Entropy Loss for Neural Speaker Diarization", in Proc. Interspeech, 2023. (Pyannote) ππ»
"End-to-End Neural Speaker Diarization with Absolute Speaker Loss", in Proc. Interspeech, 2023. (Pyannote) π
"Blueprint Separable Subsampling and Aggregate Feature Conformer-Based End-to-End Neural Diarization", in Electronics, 2023. π
EEND-TA: "Transformer Attractors for Robust and Efficient End-to-End Neural Diarization," in Proc. ASRU, 2023. (Fano Labs) π
"Robust End-to-End Diarization with Domain Adaptive Training and Multi-Task Learning," in Proc. ASRU, 2023. (Fano Labs) π
2022 (10 papers)
EEND-EDA (2): βEncoder-Decoder Based Attractor Calculation for End-to-End Neural Diarizationβ, in IEEE/ACM TASLP, 2022. (Hitachi) πππ»
"DIVE: End-to-end Speech Diarization via Iterative Speaker Embedding", in Proc. ICASSP, 2022. (Google) π
RX-EEND: βAuxiliary Loss of Transformer with Residual Connection for End-to-End Speaker Diarizationβ, in Proc. ICASSP, 2022. (GIST) ππ
"End-to-end speaker diarization with transformer", in Proc. arXiv, 2022. π
EEND-VC-iGMM: "Tight integration of neural and clustering-based diarization through deep unfolding of infinite Gaussian mixture model", in Proc. ICASSP, 2022. (NTT) π
EDA-RC: "Robust End-to-end Speaker Diarization with Generic Neural Clustering", in Proc. Interspeech, 2022. (SJTU) π
EEND-NAA: "End-to-End Neural Speaker Diarization with an Iterative Refinement of Non-Autoregressive Attention-based Attractors", in Proc. Interspeech, 2022. (JHU) ππ
Graph-PIT: "Utterance-by-utterance overlap-aware neural diarization with Graph-PIT", in Proc. Interspeech, 2022. (NTT) ππ»
"Efficient Transformers for End-to-End Neural Speaker Diarization", in Proc. IberSPEECH, 2022. π
EEND-EDA-SpkAtt: "Towards End-to-end Speaker Diarization in the Wild", in arXiv:2211.01299v1, 2022. π
2021 (7 papers)
CB-EEND: "End-to-end Neural Diarization: From Transformer to Conformer", in Proc. Interspeech, 2021. (Amazon) ππ
TDCN-SA: "End-to-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings", in Proc. ICASSP, 2021. (Google) ππ
"End-to-End Speaker Diarization Conditioned on Speech Activity and Overlap Detection", in Proc. IEEE SLT, 2021. (Hitachi) π
EEND-VC (1): "Integrating end-to-end neural and clustering-based diarization: Getting the best of both worlds", in Proc. ICASSP, 2021. (NTT) πππ»
EEND-VC (2): "Advances in integration of end-to-end neural and clustering-based diarization for real conversational speech", in Proc. Interspeech, 2021. (NTT) πππ»
"Robust End-to-End Speaker Diarization with Conformer and Additive Margin Penalty," in Proc. Interspeech, 2021. (Fano Labs) π
EEND-GLA: "Towards Neural Diarization for Unlimited Numbers of Speakers Using Global and Local Attractors", in Proc. ASRU, 2021. (Hitachi) ππ
2020 (3 papers)
SA-EEND (2): βEnd-to-End Neural Diarization: Reformulating Speaker Diarization as Simple Multi-label Classificationβ, in arXiv:2003.02966, 2020. (Hitachi) ππ
SC-EEND: "Neural Speaker Diarization with Speaker-Wise Chain Rule", in arXiv:2006.01796, 2020. (Hitachi) ππ
EEND-EDA (1): βEnd-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractorsβ, in Proc. Interspeech, 2020. (Hitachi) πππ»
2023 (1 paper)
EEND-IAAE: "End-to-end neural speaker diarization with an iterative adaptive attractor estimation," in Neural Networks, Elsevier. ππ»
2019 (2 papers)
BLSTM-EEND: "End-to-End Neural Speaker Diarization with Permutation-Free Objectives", in Proc. Interspeech, 2019. (Hitachi) π
SA-EEND (1): βEnd-to-End Neural Speaker Diarization with Self-attentionβ, in Proc. ASRU, 2019. (Hitachi) ππ»π»π
π Related Speaker information β 2 papers
π₯ 2024 (2 papers)
"Do End-to-End Neural Diarization Attractors Need to Encode Speaker Characteristic Information?," in Proc. Odyssey, 2024. π
"Leveraging Speaker Embeddings in End-to-End Neural Diarization for Two-Speaker Scenarios," in Proc. Odyssey, 2024. π
π Post-Processing β 3 papers
π₯ 2021-2024 (3 papers)
"DiaCorrect: Error Correction Back-end For Speaker Diarization," in Proc. ICASSP, 2024. (BUT) ππ»
EENDasP: "End-to-End Speaker Diarization as Post-Processing", in Proc. ICASSP, 2021. (Hitachi) πππ»
Dover-Lap: "DOVER-Lap: A Method for Combining Overlap-aware Diarization Outputs", in Proc. IEEE SLT, 2021. (JHU) πππ»
π― Using Target Speaker Embedding β 17 papers
π₯ 2025 (4 papers)
"Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining," in Proc. ICASSP, 2025. π
MIMO-TSVAD: "Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization," in IEEE/ACM TASLP, 2025. (DKU) π
"Mitigating Non-Target Speaker Bias in Guided Speaker Embedding," in Proc. Interspeech, 2025. (NTT) π
"Diarization-Guided Multi-Speaker Embeddings," in Proc. Interspeech, 2025. (Pyannote) π
2024 (3 papers)
NSD-MS2S: "Neural Speaker Diarization Using Memory-Aware Multi-Speaker Embedding with Sequence-to-Sequence Architecture, " in Proc. ICASSP, 2024. (USTC) ππ»
PET-TSVAD: "Profile-Error-Tolerant Target-Speaker Voice Activity Detection," in Proc. ICASSP, 2024. (Microsoft) π
Flow-TSVAD: "Target-Speaker Voice Activity Detection via Latent Flow Matching," in arXiv:2409.04859, 2024. (DKU) π
2023 (4 papers)
EDA-TS-VAD: βTarget Speaker Voice Activity Detection with Transformers and Its Integration with End-to-End Neural Diarizationβ, in Proc. ICASSP, 2023. (Microsoft) π
Seq2Seq-TS-VAD: βTarget-Speaker Voice Activity Detection via Sequence-to-Sequence Predictionβ, in Proc. ICASSP, 2023. (DKU) ππ
QM-TS-VAD: "Unsupervised Adaptation with Quality-Aware Masking to Improve Target-Speaker Voice Activity Detection for Speaker Diarization", in Proc. Interspeech, 2023. (USTC) π
"ANSD-MA-MSE: Adaptive Neural Speaker Diarization Using Memory-Aware Multi-Speaker Embedding," in IEEE/ACM TASLP, 2023. (USTC) ππ»
2022 (3 papers)
SEND (2): "Speaker Embedding-aware Neural Diarization: an Efficient Framework for Overlapping Speech Diarization in Meeting Scenarios," in arXiv:2203.09767, 2022 (Alibaba) π
MTEAD: "Multi-target Filter and Detector for Unknown-number Speaker Diarization", in IEEE SPL, 2022. π
SOND: "Speaker Overlap-aware Neural Diarization for Multi-party Meeting Analysis", in Proc. EMNLP, 2022. (Alibaba) ππ»
2020-2021 (3 papers)
SEND (1): "Speaker Embedding-aware Neural Diarization for Flexible Number of Speakers with Textual Information," in arXiv:2111.13694, 2021. (Alibaba) π
TS-VAD: "Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario", in Proc. Interspeech, 2020. ππ»π
βThe STC system for the CHiME-6 challenge,β in CHiME Workshop, 2020. π
π― Target Speech Diarization β 1 paper
π₯ 2024 (1 paper)
PTSD: "Prompt-driven Target Speech Diarization," in Proc. ICASSP, 2024. (NUS) π
π With Separation or Target Speaker Extraction β 12 papers
π₯ 2025 (3 papers)
"Robust Target Speaker Diarization and Separation via Augmented Speaker Embedding Sampling," in Proc. Interspeech, 2025. π
S2SND: "Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation," in IEEE/ACM TASLP, 2025. (DKU) π
"Exploring Speaker Diarization with Mixture of Experts," in arXiv:2506.14750, 2025. (USTC) π
2024 (7 papers)
"TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddings", in IEEE/ACM TASLP, 2024. π
"Continuous Target Speech Extraction: Enhancing Personalized Diarization and Extraction on Complex Recordings," in arXiv:2401.15993, 2024. (Tencent) ππ¬
"PixIT: Joint Training of Speaker Diarization and Speech Separation from Real-world Multi-speaker Recordings," in Proc. Odyssey, 2024. ππ»
MC-EEND: "Multi-channel Conversational Speaker Separation via Neural Diarization," in IEEE/ACM TASLP, 2024. (OSU) π
"USED: Universal Speaker Extraction and Diarization," in submitted to IEEE/ACM TASLP, 2024. (CUHK) ππ¬ππ
"Neural Blind Source Separation and Diarization for Distant Speech Recognition," in Proc. Interspeech, 2024. (AIST) π
"TalTech-IRIT-LIS Speaker and Language Diarization Systems for DISPLACE 2024," in Proc. Interspeech, 2024. (Pyannote) π
2021-2022 (2 papers)
EEND-SS: "Joint End-to-End Neural Speaker Diarization and Speech Separation for Flexible Number of Speakersβ, in Proc. SLT, 2022. (CMU) ππ
"Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis," in Proc. SLT, 2021. (JHU) πππ
π‘ Multi-Channel β 14 papers
π₯ 2025 (5 papers)
"Multi-channel Speaker Counting for EEND-VC-based Speaker Diarization on Multi-domain Conversation," in Proc. ICASSP, 2025. (NTT) ππ
"Spatially Aware Self-Supervised Models for Multi-Channel Neural Speaker Diarization," in arXiv:2510.14551, 2025. (BUT) ππ»
"Multi-Channel Sequence-to-Sequence Neural Diarization for The MISP 2025 Challenge," in arXiv:2505.16387, 2025. π
"Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings," in Proc. ICASSP, 2025. π
"Spatio-Spectral Diarization of Meetings by Combining TDOA-based Segmentation and Speaker Embedding-based Clustering," in Proc. Interspeech, 2025. π
2024 (5 papers)
"UniX-Encoder: A Universal X-Channel Speech Encoder for Ad-Hoc Microphone Array Speech Processing," in arXiv:2310.16367, 2024. (JHU, Tencent) π
"Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech Detection," in IEEE/ACM TASLP, 2024. π
"A Spatial Long-Term Iterative Mask Estimation Approach for Multi-Channel Speaker Diarization and Speech Recognition," in Proc. ICASSP, 2024. (USTC) π
MC-EEND: "Multi-channel Conversational Speaker Separation via Neural Diarization," in IEEE/ACM TASLP, 2024. (OSU) π
"ASoBO: Attentive Beamformer Selection for Distant Speaker Diarization in Meetings," in Proc. Interspeech, 2024. (LIUM) π
2022-2023 (4 papers)
"Mutual Learning of Single- and Multi-Channel End-to-End Neural Diarization," in Proc. IEEE SLT, 2023. (Hitachi) π
"Semi-supervised multi-channel speaker diarization with cross-channel attention", in Proc. ASRU, 2023. (USTC) π
"Multi-Channel End-to-End Neural Diarization with Distributed Microphones", in Proc. ICASSP, 2022. (Hitachi) π
"Multi-Channel Speaker Diarization Using Spatial Features for Meetings", in Proc. ICASSP, 2022. (Tencent) π
β‘ Online β 18 papers
π₯ 2024-2025 (8 papers)
SCDiar: "A Streaming Diarization System based on Speaker Change Detection and Speech Recognition," in Proc. ICASSP, 2025. π
Streaming Sortformer: "Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering," in Proc. Interspeech, 2025. (NVIDIA) π
OTS-VAD: "Online Neural Speaker Diarization With Target Speaker Tracking," in IEEE/ACM TASLP, 2024. (DKU) π
FS-EEND: "Frame-wise streaming end-to-end speaker diarization with non-autoregressive self-attention-based attractors," in Proc. ICASSP, 2024. (Hangzhou) ππ»
"Online speaker diarization of meetings guided by speech separation," in Proc. ICASSP, 2024. (LTCI) ππ»
"Interrelate Training and Clustering for Online Speaker Diarization," in IEEE/ACM TASLP, 2024. π
O-EENC-SD: "Efficient Online End-to-End Neural Clustering for Speaker Diarization," in Proc. ICASSP, 2025. π
LS-EEND: "Long-Form Streaming End-to-End Neural Diarization with Online Attractor Extraction," in IEEE/ACM TASLP, 2025. (Westlake) ππ»
2023 (2 papers)
"Absolute decision corrupts absolutely: conservative online speaker diarisation", in Proc. ICASSP, 2023. (Naver) π
"A Reinforcement Learning Framework for Online Speaker Diarization", in Under Review. NeruIPS, 2023. (CU) π
2022 (3 papers)
"Low-Latency Online Speaker Diarization with Graph-Based Label Generation", in Proc. Odyssey, 2022. (DKU) π
EEND-GLA: "Online Neural Diarization of Unlimited Numbers of Speakers Using Global and Local Attractors", in IEEE/ACM TASLP, 2022. (Hitachi) π
Online TS-VAD: "Online Target Speaker Voice Activity Detection for Speaker Diarization", in Proc. Interspeech, 2022. (DKU) π
2021 (4 papers)
"Online End-to-End Neural Diarization with Speaker-Tracing Buffer", in Proc. IEEE SLT, 2021. (Hitachi) π
BW-EDA-EEND: "BW-EDA-EEND: Streaming End-to-End Neural Speaker Diarization for a Variable Number of Speakers", in Proc. Interspeech, 2021. (Amazon) π
FS-EEND: "Online Streaming End-to-End Neural Diarization Handling Overlapping Speech and Flexible Numbers of Speakers", in Proc. Interspeech, 2021. (Hitachi) ππ
Diart: "Overlap-aware low-latency online speaker diarization based on end-to-end local segmentation", in Proc. ASRU, 2021. ππ»
2020 (1 paper)
"Supervised online diarization with sample mean loss for multi-domain data", in Proc. ICASSP, 2020 ππ»
π Clustering-based β 21 papers
π₯ 2025 (3 papers)
E-SHARC: "End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization," in IEEE/ACM TASLP, 2025. (IISC) π
"Speaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation Algorithm," in Proc. Interspeech, 2025. π
2024 (6 papers)
"Overlap-aware End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization," in submitted to IEEE/ACM TASLP, 2024. π
"Apollo's Unheard Voices: Graph Attention Networks for Speaker Diarization and Clustering for Fearless Steps Apollo Collection," in Proc. ICASSP, 2024. (UTD) π
"Multi-View Speaker Embedding Learning for Enhanced Stability and Discriminability," in Proc. ICASSP, 2024. (Tsinghua) π
"Towards Unsupervised Speaker Diarization System for Multilingual Telephone Calls Using Pre-trained Whisper Model and Mixture of Sparse Autoencoders," in arXiv:2407.01963, 2024. π
"Investigating Confidence Estimation Measures for Speaker Diarization," in Proc. Interspeech, 2024. π
"Once more Diarization: Improving meeting transcription systems through segment-level speaker reassignment," in Proc. Interspeech, 2024. (PU) πππ»
2023 (5 papers)
SCALE: "Spectral Clustering-aware Learning of Embeddings for Speaker Diarisation", in Proc. ICASSP, 2023. (CAM) π
SHARC: "Supervised Hierarchical Clustering using Graph Neural Networks for Speaker Diarization", in Proc. ICASSP, 2023. (IISC) π
CDGCN: "Community Detection Graph Convolutional Network for Overlap-Aware Speaker Diarization," in Proc. ICASSP, 2023. (XMU) π
"Pyannote.Audio 2.1: Speaker Diarization Pipeline: Principle, Benchmark and Recipe", in Proc. Interspeech, 2023. (CNRS) π
GADEC: "Graph attention-based deep embedded clustering for speaker diarization,", in Speech Communication, 2023. (NJUPT) π
2020-2022 (4 papers)
UMAP-Leiden: "Reformulating Speaker Diarization as Community Detection With Emphasis On Topological Structure", in Proc. ICASSP, 2022. (Alibaba) π
Pyannote 2.0: "End-to-end speaker segmentation for overlap-aware resegmentation", in Proc. Interspeech, 2021. (CNRS) ππ»π¬
Pyannote: "pyannote.audio: neural building blocks for speaker diarization", in Proc. ICASSP, 2020. (CNRS) ππ»π¬
Resegmentation with VB: βOverlap-Aware Diarization: Resegmentation Using Neural End-to-End Overlapped Speech Detectionβ, in Proc. ICASSP, 2020. π
DNC: "Discriminative Neural Clustering for Speaker Diarisation", in Proc. IEEE SLT, 2019. ππ»π
NME-SC: βAuto-Tuning Spectral Clustering for Speaker Diarization Using Normalized Maximum Eigengapβ, IEEE SPL, 2019. ππ»
π Variational Bayes and HMM β 24 papers
π₯ 2023-2024 (3 papers)
DVBx: "Discriminative Training of VBx Diarization", in Proc. ICASSP, 2024. (BUT) ππ»
MS-VBx: "Multi-Stream Extension of Variational Bayesian HMM Clustering (MS-VBx) for Combined End-to-End and Vector Clustering-based Diarization", in Proc. Interspeech, 2023. (NTT) π
"Generalized domain adaptation framework for parametric back-end in speaker recognition", in arXiv:2305.15567, 2023. π
2021-2022 (4 papers)
"Bayesian HMM clustering of x-vector sequences (VBx) in speaker diarization: theory, implementation and analysis on standard tasks", in Computer Speech & Language, 2022. (BUT) π
DCA-PLDA "A Speaker Verification Backend with Robust Performance across Conditionsβ, in Computer & Language, 2022. ππ»
"Analysis of the but Diarization System for Voxconverse Challenge", in Proc. ICASSP, 2021. (BUT) ππ»
"Discriminatively trained probabilistic linear discriminant analysis for speaker verification", in Proc. ICASSP, 2021. π
2019-2020 (4 papers)
"Optimizing Bayesian Hmm Based X-Vector Clustering for the Second Dihard Speech Diarization Challenge", in Proc. ICASSP, 2020. (BUT) π
βAnalysis of Speaker Diarization Based on Bayesian HMM With Eigenvoice Priorsβ, IEEE/ACM TASLP, 2019. (BUT) π
"BUT System Description for DIHARD Speech Diarization Challenge 2019", in arXiv:1910.08847, 2019. (BUT) π
"Bayesian HMM Based x-Vector Clustering for Speaker Diarization", in Proc. Interspeech, 2019. (BUT) π
2018 (5 papers)
"Speaker Diarization based on Bayesian HMM with Eigenvoice Priors", in Proc. Odyssey, 2018. (BUT) π
"VB-HMM Speaker Diarization with Enhanced and Refined Segment Representation", in Proc. Odyssey, 2018. (Tsinghua) π
"Diarization is hard: some experiences and lessons learned for the JHU team in the inaugural DIHARD challenge", in Proc. Interspeech, 2018. π
"The speaker partitioning problem", in Proc. Odyssey, 2018. π
"Estimation of the Number of Speakers with Variational Bayesian PLDA in the DIHARD Diarization Challenge", in Proc. Interspeech, 2018. π
2015-2017 (3 papers)
"Domain Adaptation of PLDA Models in Broadcast Diarization by Means of Unsupervised Speaker Clustering, in Proc. Interspeech, 2017. π
"Iterative PLDA Adaptation for Speaker Diarization", in Proc. Interspeech, 2016. π
"Diarization resegmentation in the factor analysis subspace", in Proc. ICASSP, 2015. π
2011-2014 (3 papers)
"Speaker diarization with plda i-vector scoring and unsupervised calibration", in Proc. IEEE SLT, 2014. π
"Unsupervised Methods for Speaker Diarization: An Integrated and Iterative Approach", IEEE/ACM TASLP, 2013. π
"Analysis of i-vector length normalization in speaker recognition systems", in Proc. Interspeech, 2011. π
2005-2008 (2 papers)
"Bayesian analysis of speaker diarization with eigenvoice priors", in CRIM, Montreal, Technical Report, 2008. π
"Variational Bayesian methods for audio indexing", in Proc. ICMI-MLMI, 2005. π
"Geodesic interpolation of frame-wise speaker embeddings for the diarization of meeting scenarios", in Proc. ICASSP, 2024. (PU) ππ
"Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization," in Proc. Odyssey, 2024. (IDLab) π
"Efficient Speaker Embedding Extraction Using a Twofold Sliding Window Algorithm for Speaker Diarization," in Proc. Interspeech, 2024. (HU) π
"Variable Segment Length and Domain-Adapted Feature Optimization for Speaker Diarization," in Proc. Interspeech, 2024. (XMU) ππ»
2023 (5 papers)
"In Search of Strong Embedding Extractors For Speaker Diarization", in Proc. ICASSP, 2023. (Naver) ππ
DR-DESA: "Advancing the dimensionality reduction of speaker embeddings for speaker diarisation: disentangling noise and informing speech activity", in Proc. ICASSP, 2023. (Naver) ππ
HEE: "High-resolution embedding extractor for speaker diarisation", in Proc. ICASSP, 2023. (Naver) ππ
"Frame-wise and overlap-robust speaker embeddings for meeting diarization", in Proc. ICASSP, 2023. (PU) ππ
"A Teacher-Student approach for extracting informative speaker embeddings from speech mixtures", in Proc. Interspeech, 2023. (PU) π
2022 (3 papers)
GAT+AA: "Multi-scale speaker embedding-based graph attention networks for speaker diarisation", in Proc. ICASSP, 2022. (Naver) π
MSDD: "Multi-scale Speaker Diarization with Dynamic Scale Weighting", in Proc. Interspeech, 2022. (NVIDIA) ππ»π
PRISM: "PRISM: Pre-trained Indeterminate Speaker Representation Model for Speaker Diarization and Speaker Verification", in Proc. Interspeech, 2022. (Alibaba) π
2021 (2 papers)
"Multi-Scale Speaker Diarization With Neural Affinity Score Fusion", in Proc. ICASSP, 2021. (USC) π
AA+DR+NS: "Adapting Speaker Embeddings for Speaker Diarisation", in Proc. Interspeech, 2021. (Naver) ππ
πͺͺ With Speaker Identification β 1 paper
π₯ 2024 (1 paper)
"Uncertainty Quantification in Machine Learning for Joint Speaker Diarization and Identification, in Submitted to IEEE/ACM TASLP, 2024. π
"Rethinking Session Variability: Leveraging Session Embeddings for Session Robustness in Speaker Verification," in Proc. ICASSP, 2024. (Naver) π
"Leveraging In-the-Wild Data for Effective Self-Supervised Pretraining in Speaker Recognition," in Proc. ICASSP, 2024. (CUHK) π
"Disentangled Representation Learning for Environment-agnostic Speaker Recognition," in Proc. Interspeech, 2024. (KAIST) πππ»
2023 (3 papers)
"Build a SRE Challenge System: Lessons from VoxSRC 2022 and CNSRC 2022," in Proc. Interspeech, 2023. (SJTU) π
RecXi "Disentangling Voice and Content with Self-Supervision for Speaker Recognition," in Proc. NeurIPS, 2023. (A*STAR) π
"ECAPA2: A Hybrid Neural Network Architecture and Training Strategy for Robust Speaker Embeddings," in Proc. ASRU, 2023. (IDLab) πππ
2021 (1 paper)
"Xi-Vector Embedding for Speaker Recognition," in IEEE, SPL. (A*STAR) ππ
π Scoring β 3 papers
π₯ 2019-2023 (3 papers)
βSimilarity Measurement of Segment-Level Speaker Embeddings in Speaker Diarizationβ, IEEE/ACM TASLP, 2023. (DKU) π
"Self-Attentive Similarity Measurement Strategies in Speaker Diarization", in Proc. Interspeech, 2020. (DKU) π
LSTM scoring: "LSTM based Similarity Measurement with Spectral Clustering for Speaker Diarization", in Proc. Interspeech, 2019. (DKU) π
π£οΈ With ASR β 27 papers
π₯ 2025-2026 (8 papers)
SE-DiCoW: "Self-Enrolled Diarization-Conditioned Whisper," in arXiv:2601.19194, 2026. (BUT) π
TagSpeech: "End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding," in arXiv:2601.06896, 2026. π
DiCoW: "Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition," in Proc. ICASSP, 2025. (BUT) ππ»
"Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition," in arXiv:2510.03723, 2025. (BUT) π
"Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models," in arXiv:2506.05796, 2025. (DKU) π
"Language Modelling for Speaker Diarization in Telephonic Interviews," in arXiv:2501.17893, 2025. π
SC-SOT: "Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition," in Proc. Interspeech, 2025. π
"Target Speaker ASR with Whisper," in Submitted to ICASSP, 2025. (BUT) ππ»
2024 (12 papers)
"Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach,", in Proc. ICASSP, 2024. (NVIDIA) π
WEEND: "Towards Word-Level End-to-End Neural Speaker Diarization with Auxiliary Network," in arXiv:2309.08489, 2024. (Google) ππ
"One model to rule them all ? Towards End-to-End Joint Speaker Diarization and Speech Recognition", in Proc. ICASSP, 2024. (CMU) π
"Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization," in arXiv:2309.16482, 2024. (PU) π
βJoint Inference of Speaker Diarization and ASR with Multi-Stage Information Sharing," in Proc. ICASSP, 2024. (DKU) π
"Multitask Speech Recognition and Speaker Change Detection for Unknown Number of Speakers" in Proc. ICASSP, 2024. (Idiap) π
"A Spatial Long-Term Iterative Mask Estimation Approach for Multi-Channel Speaker Diarization and Speech Recognition," in Proc. ICASSP, 2024. (USTC) π
"On the Success and Limitations of Auxiliary Network Based Word-Level End-to-End Neural Speaker Diarization," in Proc. Interspeech, 2024. (Google) π
Sortformer: "Seamless Integration of Speaker Diarization and ASR by Bridging Timestamps and Tokens," in Proc. ICML, 2025. (NVIDIA) ππ
"Speaker Mask Transformer for Multi-talker Overlapped Speech Recognition," in arXiv:2312.10959, 2024. (NICT) π
"On Speaker Attribution with SURT," in Proc. Odyssey, 2024. (JHU) π
"Improving Speaker Assignment in Speaker-Attributed ASR for Real Meeting Applications," in Proc. Odyssey, 2024. (CNRS) π
2023 (5 papers)
"Unified Modeling of Multi-Talker Overlapped Speech Recognition and Diarization with a Sidecar Separator", in Proc. Interspeech, 2023. (CUHK) π
"Multi-resolution Approach to Identification of Spoken Languages and to Improve Overall Language Diarization System using Whisper Model", in Proc. Interspeech, 2023.
"Speaker Diarization for ASR Output with T-vectors: A Sequence Classification Approach", in Proc. Interspeech, 2023. π
"Lexical Speaker Error Correction: Leveraging Language Models for Speaker Diarization Error Correction", in Proc. Interspeech, 2023. (Amazon) π
"SA-Paraformer: Non-autoregressive End-to-End Speaker-Attributed ASR," in Proc. ASRU, 2023. (Alibaba) π
2022 (2 papers)
"Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed ASR," in Proc. ICASSP, 2022. π
"Tandem Multitask Training of Speaker Diarisation and Speech Recognition for Meeting Transcription", in Proc. Interspeech, 2022. π
π¬ With NLP / LLM / Language β 12 papers
π₯ 2024-2025 (7 papers)
SpeakerLM: "End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models," in arXiv:2508.06372, 2025. π
"Interactive Real-Time Speaker Diarization Correction with Human Feedback," in arXiv:2509.18377, 2025. π
"DiariST: Streaming Speech Translation with Speaker Diarization," in Proc. ICASSP, 2024. (Microsoft) ππ»
JPCP: "Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation," in arXiv:2309.10456, 2024. (Alibaba) π
"DiarizationLM: Speaker Diarization Post-Processing with Large Language Models," in Proc. Interspeech, 2024. (Google) πππ»π
"LLM-based speaker diarization correction: A generalizable approach," in Submitted to IEEE/ACM TASLP, 2024. π
"AG-LSEC: Audio Grounded Lexical Speaker Error Correction," in Proc. Interspeech, 2024. (Amazon) π
2023 (3 papers)
"Exploring Speaker-Related Information in Spoken Language Understanding for Better Speaker Diarization", in Proc. ACL, 2023. (Alibaba) π
MMSCD, "Encoder-decoder multimodal speaker change detection", in Proc. Interspeech, 2023. (Naver) π
"Aligning Speakers: Evaluating and Visualizing Text-based Diarization Using Efficient Multiple Sequence Alignment,", in Proc. ICTAI, 2023. π
π Language Diarization β 2 papers
π₯ 2023 (2 papers)
"End-to-End Spoken Language Diarization with Wav2vec Embeddings", in Proc. Interspeech, 2023. ππ»
"Multi-resolution Approach to Identification of Spoken Languages and To Improve Overall Language Diarization System Using Whisper Model," in Proc. Interspeech, 2023. π
ποΈ With Vision β 21 papers
π₯ 2025-2026 (4 papers)
CineSRD: "Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarization," in arXiv:2603.16966, 2026. π
"Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization," in Proc. ACL, 2025. π
"Cross-Attention and Self-Attention for Audio-visual Speaker Diarization," in arXiv:2506.02621, 2025. π
"Count Your Speakers! Multitask Learning for Multimodal Speaker Diarization," in Proc. Interspeech, 2025. π
2024 (6 papers)
"Speaker Diarization of Scripted Audiovisual Content," in arXiv:2308.02160, 2024. (Amazon) π
"AFL-Net: Integrating Audio, Facial, and Lip Modalities with Cross-Attention for Robust Speaker Diarization in the Wild," in Proc. ICASSP, 2024. (Tencent) ππ¬
"Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation," in Proc. AAAI, 2024. (Tencent) π
"3D-Speaker-Toolkit: An Open Source Toolkit for Multi-modal Speaker Verification and Diarization," in arXiv:2403.19971, 2024. (Alibaba) ππ»
"Target Speech Diarization with Multimodal Prompts," in Submitted to IEEE/ACM TASLP, 2024. (NUS) π
MFV-KSD: "Multi-Stage Face-Voice Association Learning with Keynote Speaker Diarization," in Submitted to ACM MM, 2024. ππ»
2023 (5 papers)
"Audio-Visual Speaker Diarization in the Framework of Multi-User Human-Robot Interaction", in Proc. ICASSP, 2023. π
STHG: "Spatial-Temporal Heterogeneous Graph Learning for Advanced Audio-Visual Diarization, in Proc. CVPR, 2023. (Intel) π
"Uncertainty-Guided End-to-End Audio-Visual Speaker Diarization for Far-Field Recordings," in Proc. ACM MM, 2023. π
"Joint Training or Not: An Exploration of Pre-trained Speech Models in Audio-Visual Speaker Diarization," in Springer Computer Science proceedings, 2023. π
EEND-EDA++: "Late Audio-Visual Fusion for In-The-Wild Speaker Diarization," in arXiv:2211.01299v2, 2023. π
2022 (3 papers)
AVA-AVD (AVR-Net): "AVA-AVD: Audio-Visual Speaker Diarization in the Wild", in Proc. ACM MM, 2022. ππ»π¬
"End-to-End Audio-Visual Neural Speaker Diarization", in Proc. Interspeech, 2022. (USTC) ππ»π
DyViSE: "DyViSE: Dynamic Vision-Guided Speaker Embedding for Audio-Visual Speaker Diarization", in Proc. MMSP, 2022. (THU) ππ»
2024 (1 paper)
"Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization," in Submitted to IEEE/ACM TASLP. (DKU) π
2019-2020 (2 papers)
"Self-supervised learning for audio-visual speaker diarization", in Proc. ICASSP, 2020. (Tencent) ππ
"Who said that?: Audio-visual speaker diarisation of real-world meetings", in Proc. Interspeech, 2019. (Naver) π
π‘οΈ Related Spoofing β 1 paper
π₯ 2024 (1 paper)
"Spoof Diarization: "What Spoofed When" in Partially Spoofed Audio," in Proc. Interspeech, 2024. (IITK) π
π Related TTS
π Speaker Anonymization β 1 paper
π₯ 2024 (1 paper)
"A Benchmark for Multi-speaker Anonymization," in Submitted to IEEE/ACM TASLP, 2024. (SIT) ππ»
π Singing Diarization β 1 paper
π₯ 2024 (1 paper)
"Song Data Cleansing for End-to-End Neural Singer Diarization Using Neural Analysis and Synthesis Framework," in Proc. Interspeech, 2024. (LY) π
π With Emotion β 3 papers
π₯ 2023-2024 (3 papers)
"ED-TTS: Multi-scale Emotion Modeling using Cross-domain Emotion Diarization for Emotional Speech Synthesis, in Proc. ICASSP, 2024. π
"Speech Emotion Diarization: Which Emotion Appears When?," in Proc. ASRU, 2023. (Zaion) π
"EmoDiarize: Speaker Diarization and Emotion Identification from Speech Signals using Convolutional Neural Networks," in arxiv:2310.12851, 2023. π
ποΈ Personal VAD β 2 papers
π₯ 2020-2023 (2 papers)
"SVVAD: Personal Voice Activity Detection for Speaker Verification", in Proc. Interspeech, 2023. π
"Personal VAD: Speaker-Conditioned Voice Activity Detection", in Proc. Odyssey, 2020. (Google) π
π VAD & OSD & SCD β 10 papers
π₯ 2024 (3 papers)
"USM-SCD: Multilingual Speaker Change Detection Based on Large Pretrained Foundation Models," in Proc. ICASSP, 2024. (Google) π
"Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech Detection," in IEEE/ACM TASLP, 2024. π
"Speaker Change Detection with Weighted-sum Knowledge Distillation based on Self-supervised Pre-trained Models," in Proc. Interspeech, 2024. π
2023 (4 papers)
"Multitask Detection of Speaker Changes, Overlapping Speech and Voice Activity Using wav2vec 2.0," in Proc. ICASSP, 2023. ππ»
"Semantic VAD: Low-Latency Voice Activity Detection for Speech Interaction," in Proc. Interspeech, 2023. π
"Joint speech and overlap detection: a benchmark over multiple audio setup and speech domains," in arxiv:2307.13012, 2023. π
"Advancing the study of Large-Scale Learning in Overlapped Speech Detection," in arXiv:2308.05987, 2023. π
2022 (3 papers)
"Overlapped Speech Detection in Broadcast Streams Using X-vectors," in Proc. Interspeech, 2022. π
"Overlapped speech and gender detection with WavLM pre-trained features," in Proc. Interspeech, 2022. π
"Microphone Array Channel Combination Algorithms for Overlapped Speech Detection," in Proc. Interspeech, 2022. π
π Dataset β 19 papers
π₯ 2024-2025 (6 papers)
"Conversations in the wild: Data collection, automatic generation and evaluation," in Computer Speech & Language, 2025. π
M3SD: "Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset," in arXiv:2506.14427, 2025. π
"VoxBlink: X-Large Speaker Verification Dataset on Camera", in Proc. ICASSP, 2024. ππ
"NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant Meeting Transcription," in arXiv:2401.08887, 2024. (MS) π
"A Comparative Analysis of Speaker Diarization Models: Creating a Dataset for German Dialectal Speech," in Proc. ACL, 2024. π
"ALLIES: A Speech Corpus for Segmentation, Speaker Diarization, Speech Recognition and Speaker Change Detection," in Proc. ACL, 2024. (LIUM) π
2020-2022 (5 papers)
Ego4D: " Around the World in 3,000 Hours of Egocentric Video," in Proc. CVPR, 2022. (Meta) ππ»π
AliMeeting: "Summary On The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge," in Proc. ICASSP, 2022. (Alibaba) πππ»
Voxconverse: "Spot the conversation: speaker diarisation in the wild", in Proc. Interspeech, 2020. (VGG, Naver) ππ»π
MSDWild: Multi-modal Speaker Diarization Dataset in the Wild, in Proc. Interspeech, 2020. ππ
"LibriMix: An Open-Source Dataset for Generalizable Speech Separation," in arXiv:2005.11262, 2020. ππ»
π Simulated Dataset β 8 papers
π₯ 2023-2024 (4 papers)
"Enhancing low-latency speaker diarization with spatial dictionary learning," in Proc. ICASSP, 2024. (NTU) ππ
"Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling," in Proc. ICASSP, 2024. (OSU) π
"Multi-Speaker and Wide-Band Simulated Conversations as Training Data for End-to-End Neural Diarization", in Proc. ICASSP, 2023. (BUT) ππ»π
"Property-Aware Multi-Speaker Data Simulation: A Probabilistic Modelling Technique for Synthetic Data Generation," in CHiME-7 Workshop, 2023. (NVIDIA) π
2022 (3 papers)
"From simulated mixtures to simulated conversations as training data for end-to-end neural diarization" , in Proc. Interspeech, 2022. (BUT) ππ»π
Markov selection: "Improving the naturalness of simulated conversations for end-to-end neural diarization", in Proc. Odyssey, 2022. (Hitachi) π
EEND-EDA-SpkAtt: "Towards End-to-end Speaker Diarization in the Wild", in arXiv:2211.01299v1, 2022. π
2019 (1 paper)
Concat-and-sum: "End-to-end neuarl speaker diarization with permuation-free objectives", in Proc. Interspeech, 2019. π
π οΈ Tools β 1 paper
π₯ 2024 (1 paper)
"Gryannote open-source speaker diarization labeling tool," in Proc. Interspeech (Show and Tell), 2024. (IRIT) ππ»
π Self-Supervised β 5 papers
π₯ 2025 (3 papers)
DiariZen: "Leveraging Self-Supervised Learning for Speaker Diarization," in Proc. ICASSP," 2025. (BUT) πππ»π
"Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models," in arXiv:2506.18623, 2025. (BUT) ππ»
"Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization," in arXiv:2505.24111, 2025. π
2022 (2 papers)
βSelf-supervised Speaker Diarizationβ, in Proc. Interspeech, 2022. π
CSDA: "Continual Self-Supervised Domain Adaptation for End-to-End Speaker Diarization", in Proc. IEEE SLT, 2022. (CNRS) ππ»
π Semi-Supervised β 1 paper
π₯ 2017 (1 paper)
"Active Learning Based Constrained Clustering For Speaker Diarization", in IEEE/ACM TASLP, 2017. (UT) π
π Measurement β 3 papers
π₯ 2022-2025 (3 papers)
SDBench: βA Comprehensive Benchmark Suite for Speaker Diarization,β in Proc. Interspeech, 2025. π
βBenchmarking Diarization Models,β in arXiv:2509.26177, 2025. π
BER: βBalanced Error Rate For Speaker Diarizationβ, in Proc. arXiv:2211.04304, 2022 ππ»
πΆ Child-Adult β 2 papers
π₯ 2023-2024 (2 papers)
"Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions," in Proc. Interspeech, 2024. (USC) π
"Robust Self Supervised Speech Embeddings for Child-Adult Classification in Interactions involving Children with Autism," in Proc. Interspeech, 2023. π
"The DISPLACE Challenge 2023 - DIarization of SPeaker and LAnguage in Conversational Environments," in Proc. Interspeech, 2023. ππ
"The SpeeD--ZevoTech submission at DISPLACE 2023," in Proc. Interspeech, 2023. π
π MERLIon CCS Challenge 2023 β 1 paper
π₯ 2023 (1 paper)
"MERLIon CCS Challenge: A English-Mandarin code-switching child-directed speech corpus for language identification and diarization," in Proc. Interspeech, 2023. ππ
π CHiME-6
π ICMC-ASR Grand Challenge (ICASSP2024) β 2 papers
"Mamba-based Segmentation Model for Speaker Diarization," Proc. ICASSP, 2025. (NTT) ππ»
"Pushing the Limits of End-to-End Diarization," in Proc. Interspeech, 2025. π
VBx-EEND-VC: "VBx for End-to-End Neural and Clustering-based Diarization," in arXiv:2510.19572, 2025. (BUT) π
"Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling," in arXiv:2506.05593, 2025. (OSU) π
DLF-EEND: "Dynamic Layer Fusion for End-to-End Speaker Diarization," in Proc. Interspeech, 2025. π
"End-to-End Diarization utilizing Attractor Deep Clustering," in Proc. Interspeech, 2025. (JHU, OSU) π
"Pretraining Multi-Speaker Identification for Neural Speaker Diarization," in Proc. Interspeech, 2025. (NTT) π
2024 (9 papers)
"NTT speaker diarization system for CHiME-7: multi-domain, multi-microphone End-to-end and vector clustering diarization," in Proc. ICASSP, 2024. (NTT) π
AED-EEND-EE: "Attention-based Encoder-Decoder End-to-End Neural Diarization with Embedding Enhancer," in IEEE/ACM TASLP, 2024. (SJTU) ππ
"DiaPer: End-to-End Neural Diarization with Perceiver-Based Attractors," in IEEE/ACM TASLP, 2024. (BUT) ππ»π
"EEND-DEMUX: End-to-End Neural Speaker Diarization via Demultiplexed Speaker Embeddings," in Submitted to IEEE SPL, 2024. (SNU) ππ
"EEND-M2F: Masked-attention mask transformers for speaker diarization," in Proc. Interspeech, 2024. (Fano Labs) πππ
EEND-NAA (2): "End-to-End Neural Speaker Diarization with Non-Autoregressive Attractors", in IEEE/ACM TASLP, 2024. (JHU) ππ
"On the calibration of powerset speaker diarization models," in Proc. Interspeech, 2024. (IRIT) πππ»π
Local-global EEND: "Speakers Unembedded: Embedding-free Approach to Long-form Neural Diarization," in Proc. Interspeech, 2024. (Amazon) ππ
2023 (11 papers)
"Improving Transformer-based End-to-End Speaker Diarization by Assigning Auxiliary Losses to Attention Heads", in Proc. ICASSP, 2023. (HU) π
EEND-NA: βNeural Diarization with Non-Autoregressive Intermediate Attractorsβ, in Proc. ICASSP, 2023. (LINE) π
"TOLD: A Novel Two-Stage Overlap-Aware Framework for Speaker Diarization", in Proc. ICASSP, 2023. (Alibaba) ππ»
"Improving End-to-End Neural Diarization Using Conversational Summary Representations", in Proc. Interspeech, 2023. (Fano Labs) π
AED-EEND: βAttention-based Encoder-Decoder Network for End-to-End Neural Speaker Diarization with Target Speaker Attractorβ, in Proc. Interspeech, 2023. (SJTU) ππ
"Self-Distillation into Self-Attention Heads for Improving Transformer-based End-to-End Neural Speaker Diarization", in Proc. Interspeech, 2023. (HU) π
"Powerset Multi-class Cross Entropy Loss for Neural Speaker Diarization", in Proc. Interspeech, 2023. (Pyannote) ππ»
"End-to-End Neural Speaker Diarization with Absolute Speaker Loss", in Proc. Interspeech, 2023. (Pyannote) π
"Blueprint Separable Subsampling and Aggregate Feature Conformer-Based End-to-End Neural Diarization", in Electronics, 2023. π
EEND-TA: "Transformer Attractors for Robust and Efficient End-to-End Neural Diarization," in Proc. ASRU, 2023. (Fano Labs) π
"Robust End-to-End Diarization with Domain Adaptive Training and Multi-Task Learning," in Proc. ASRU, 2023. (Fano Labs) π
2022 (10 papers)
EEND-EDA (2): βEncoder-Decoder Based Attractor Calculation for End-to-End Neural Diarizationβ, in IEEE/ACM TASLP, 2022. (Hitachi) πππ»
"DIVE: End-to-end Speech Diarization via Iterative Speaker Embedding", in Proc. ICASSP, 2022. (Google) π
RX-EEND: βAuxiliary Loss of Transformer with Residual Connection for End-to-End Speaker Diarizationβ, in Proc. ICASSP, 2022. (GIST) ππ
"End-to-end speaker diarization with transformer", in Proc. arXiv, 2022. π
EEND-VC-iGMM: "Tight integration of neural and clustering-based diarization through deep unfolding of infinite Gaussian mixture model", in Proc. ICASSP, 2022. (NTT) π
EDA-RC: "Robust End-to-end Speaker Diarization with Generic Neural Clustering", in Proc. Interspeech, 2022. (SJTU) π
EEND-NAA: "End-to-End Neural Speaker Diarization with an Iterative Refinement of Non-Autoregressive Attention-based Attractors", in Proc. Interspeech, 2022. (JHU) ππ
Graph-PIT: "Utterance-by-utterance overlap-aware neural diarization with Graph-PIT", in Proc. Interspeech, 2022. (NTT) ππ»
"Efficient Transformers for End-to-End Neural Speaker Diarization", in Proc. IberSPEECH, 2022. π
EEND-EDA-SpkAtt: "Towards End-to-end Speaker Diarization in the Wild", in arXiv:2211.01299v1, 2022. π
2021 (7 papers)
CB-EEND: "End-to-end Neural Diarization: From Transformer to Conformer", in Proc. Interspeech, 2021. (Amazon) ππ
TDCN-SA: "End-to-End Diarization for Variable Number of Speakers with Local-Global Networks and Discriminative Speaker Embeddings", in Proc. ICASSP, 2021. (Google) ππ
"End-to-End Speaker Diarization Conditioned on Speech Activity and Overlap Detection", in Proc. IEEE SLT, 2021. (Hitachi) π
EEND-VC (1): "Integrating end-to-end neural and clustering-based diarization: Getting the best of both worlds", in Proc. ICASSP, 2021. (NTT) πππ»
EEND-VC (2): "Advances in integration of end-to-end neural and clustering-based diarization for real conversational speech", in Proc. Interspeech, 2021. (NTT) πππ»
"Robust End-to-End Speaker Diarization with Conformer and Additive Margin Penalty," in Proc. Interspeech, 2021. (Fano Labs) π
EEND-GLA: "Towards Neural Diarization for Unlimited Numbers of Speakers Using Global and Local Attractors", in Proc. ASRU, 2021. (Hitachi) ππ
2020 (3 papers)
SA-EEND (2): βEnd-to-End Neural Diarization: Reformulating Speaker Diarization as Simple Multi-label Classificationβ, in arXiv:2003.02966, 2020. (Hitachi) ππ
SC-EEND: "Neural Speaker Diarization with Speaker-Wise Chain Rule", in arXiv:2006.01796, 2020. (Hitachi) ππ
EEND-EDA (1): βEnd-to-End Speaker Diarization for an Unknown Number of Speakers with Encoder-Decoder Based Attractorsβ, in Proc. Interspeech, 2020. (Hitachi) πππ»
2023 (1 paper)
EEND-IAAE: "End-to-end neural speaker diarization with an iterative adaptive attractor estimation," in Neural Networks, Elsevier. ππ»
2019 (2 papers)
BLSTM-EEND: "End-to-End Neural Speaker Diarization with Permutation-Free Objectives", in Proc. Interspeech, 2019. (Hitachi) π
SA-EEND (1): βEnd-to-End Neural Speaker Diarization with Self-attentionβ, in Proc. ASRU, 2019. (Hitachi) ππ»π»π
π Related Speaker information β 2 papers
π₯ 2024 (2 papers)
"Do End-to-End Neural Diarization Attractors Need to Encode Speaker Characteristic Information?," in Proc. Odyssey, 2024. π
"Leveraging Speaker Embeddings in End-to-End Neural Diarization for Two-Speaker Scenarios," in Proc. Odyssey, 2024. π
π Post-Processing β 3 papers
π₯ 2021-2024 (3 papers)
"DiaCorrect: Error Correction Back-end For Speaker Diarization," in Proc. ICASSP, 2024. (BUT) ππ»
EENDasP: "End-to-End Speaker Diarization as Post-Processing", in Proc. ICASSP, 2021. (Hitachi) πππ»
Dover-Lap: "DOVER-Lap: A Method for Combining Overlap-aware Diarization Outputs", in Proc. IEEE SLT, 2021. (JHU) πππ»
π― Using Target Speaker Embedding β 17 papers
π₯ 2025 (4 papers)
"Noise-Robust Target-Speaker Voice Activity Detection Through Self-Supervised Pretraining," in Proc. ICASSP, 2025. π
MIMO-TSVAD: "Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization," in IEEE/ACM TASLP, 2025. (DKU) π
"Mitigating Non-Target Speaker Bias in Guided Speaker Embedding," in Proc. Interspeech, 2025. (NTT) π
"Diarization-Guided Multi-Speaker Embeddings," in Proc. Interspeech, 2025. (Pyannote) π
2024 (3 papers)
NSD-MS2S: "Neural Speaker Diarization Using Memory-Aware Multi-Speaker Embedding with Sequence-to-Sequence Architecture, " in Proc. ICASSP, 2024. (USTC) ππ»
PET-TSVAD: "Profile-Error-Tolerant Target-Speaker Voice Activity Detection," in Proc. ICASSP, 2024. (Microsoft) π
Flow-TSVAD: "Target-Speaker Voice Activity Detection via Latent Flow Matching," in arXiv:2409.04859, 2024. (DKU) π
2023 (4 papers)
EDA-TS-VAD: βTarget Speaker Voice Activity Detection with Transformers and Its Integration with End-to-End Neural Diarizationβ, in Proc. ICASSP, 2023. (Microsoft) π
Seq2Seq-TS-VAD: βTarget-Speaker Voice Activity Detection via Sequence-to-Sequence Predictionβ, in Proc. ICASSP, 2023. (DKU) ππ
QM-TS-VAD: "Unsupervised Adaptation with Quality-Aware Masking to Improve Target-Speaker Voice Activity Detection for Speaker Diarization", in Proc. Interspeech, 2023. (USTC) π
"ANSD-MA-MSE: Adaptive Neural Speaker Diarization Using Memory-Aware Multi-Speaker Embedding," in IEEE/ACM TASLP, 2023. (USTC) ππ»
2022 (3 papers)
SEND (2): "Speaker Embedding-aware Neural Diarization: an Efficient Framework for Overlapping Speech Diarization in Meeting Scenarios," in arXiv:2203.09767, 2022 (Alibaba) π
MTEAD: "Multi-target Filter and Detector for Unknown-number Speaker Diarization", in IEEE SPL, 2022. π
SOND: "Speaker Overlap-aware Neural Diarization for Multi-party Meeting Analysis", in Proc. EMNLP, 2022. (Alibaba) ππ»
2020-2021 (3 papers)
SEND (1): "Speaker Embedding-aware Neural Diarization for Flexible Number of Speakers with Textual Information," in arXiv:2111.13694, 2021. (Alibaba) π
TS-VAD: "Target-Speaker Voice Activity Detection: a Novel Approach for Multi-Speaker Diarization in a Dinner Party Scenario", in Proc. Interspeech, 2020. ππ»π
βThe STC system for the CHiME-6 challenge,β in CHiME Workshop, 2020. π
π― Target Speech Diarization β 1 paper
π₯ 2024 (1 paper)
PTSD: "Prompt-driven Target Speech Diarization," in Proc. ICASSP, 2024. (NUS) π
π With Separation or Target Speaker Extraction β 12 papers
π₯ 2025 (3 papers)
"Robust Target Speaker Diarization and Separation via Augmented Speaker Embedding Sampling," in Proc. Interspeech, 2025. π
S2SND: "Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection and Representation," in IEEE/ACM TASLP, 2025. (DKU) π
"Exploring Speaker Diarization with Mixture of Experts," in arXiv:2506.14750, 2025. (USTC) π
2024 (7 papers)
"TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddings", in IEEE/ACM TASLP, 2024. π
"Continuous Target Speech Extraction: Enhancing Personalized Diarization and Extraction on Complex Recordings," in arXiv:2401.15993, 2024. (Tencent) ππ¬
"PixIT: Joint Training of Speaker Diarization and Speech Separation from Real-world Multi-speaker Recordings," in Proc. Odyssey, 2024. ππ»
MC-EEND: "Multi-channel Conversational Speaker Separation via Neural Diarization," in IEEE/ACM TASLP, 2024. (OSU) π
"USED: Universal Speaker Extraction and Diarization," in submitted to IEEE/ACM TASLP, 2024. (CUHK) ππ¬ππ
"Neural Blind Source Separation and Diarization for Distant Speech Recognition," in Proc. Interspeech, 2024. (AIST) π
"TalTech-IRIT-LIS Speaker and Language Diarization Systems for DISPLACE 2024," in Proc. Interspeech, 2024. (Pyannote) π
2021-2022 (2 papers)
EEND-SS: "Joint End-to-End Neural Speaker Diarization and Speech Separation for Flexible Number of Speakersβ, in Proc. SLT, 2022. (CMU) ππ
"Integration of speech separation, diarization, and recognition for multi-speaker meetings: System description, comparison, and analysis," in Proc. SLT, 2021. (JHU) πππ
π‘ Multi-Channel β 14 papers
π₯ 2025 (5 papers)
"Multi-channel Speaker Counting for EEND-VC-based Speaker Diarization on Multi-domain Conversation," in Proc. ICASSP, 2025. (NTT) ππ
"Spatially Aware Self-Supervised Models for Multi-Channel Neural Speaker Diarization," in arXiv:2510.14551, 2025. (BUT) ππ»
"Multi-Channel Sequence-to-Sequence Neural Diarization for The MISP 2025 Challenge," in arXiv:2505.16387, 2025. π
"Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings," in Proc. ICASSP, 2025. π
"Spatio-Spectral Diarization of Meetings by Combining TDOA-based Segmentation and Speaker Embedding-based Clustering," in Proc. Interspeech, 2025. π
2024 (5 papers)
"UniX-Encoder: A Universal X-Channel Speech Encoder for Ad-Hoc Microphone Array Speech Processing," in arXiv:2310.16367, 2024. (JHU, Tencent) π
"Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech Detection," in IEEE/ACM TASLP, 2024. π
"A Spatial Long-Term Iterative Mask Estimation Approach for Multi-Channel Speaker Diarization and Speech Recognition," in Proc. ICASSP, 2024. (USTC) π
MC-EEND: "Multi-channel Conversational Speaker Separation via Neural Diarization," in IEEE/ACM TASLP, 2024. (OSU) π
"ASoBO: Attentive Beamformer Selection for Distant Speaker Diarization in Meetings," in Proc. Interspeech, 2024. (LIUM) π
2022-2023 (4 papers)
"Mutual Learning of Single- and Multi-Channel End-to-End Neural Diarization," in Proc. IEEE SLT, 2023. (Hitachi) π
"Semi-supervised multi-channel speaker diarization with cross-channel attention", in Proc. ASRU, 2023. (USTC) π
"Multi-Channel End-to-End Neural Diarization with Distributed Microphones", in Proc. ICASSP, 2022. (Hitachi) π
"Multi-Channel Speaker Diarization Using Spatial Features for Meetings", in Proc. ICASSP, 2022. (Tencent) π
β‘ Online β 18 papers
π₯ 2024-2025 (8 papers)
SCDiar: "A Streaming Diarization System based on Speaker Change Detection and Speech Recognition," in Proc. ICASSP, 2025. π
Streaming Sortformer: "Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering," in Proc. Interspeech, 2025. (NVIDIA) π
OTS-VAD: "Online Neural Speaker Diarization With Target Speaker Tracking," in IEEE/ACM TASLP, 2024. (DKU) π
FS-EEND: "Frame-wise streaming end-to-end speaker diarization with non-autoregressive self-attention-based attractors," in Proc. ICASSP, 2024. (Hangzhou) ππ»
"Online speaker diarization of meetings guided by speech separation," in Proc. ICASSP, 2024. (LTCI) ππ»
"Interrelate Training and Clustering for Online Speaker Diarization," in IEEE/ACM TASLP, 2024. π
O-EENC-SD: "Efficient Online End-to-End Neural Clustering for Speaker Diarization," in Proc. ICASSP, 2025. π
LS-EEND: "Long-Form Streaming End-to-End Neural Diarization with Online Attractor Extraction," in IEEE/ACM TASLP, 2025. (Westlake) ππ»
2023 (2 papers)
"Absolute decision corrupts absolutely: conservative online speaker diarisation", in Proc. ICASSP, 2023. (Naver) π
"A Reinforcement Learning Framework for Online Speaker Diarization", in Under Review. NeruIPS, 2023. (CU) π
2022 (3 papers)
"Low-Latency Online Speaker Diarization with Graph-Based Label Generation", in Proc. Odyssey, 2022. (DKU) π
EEND-GLA: "Online Neural Diarization of Unlimited Numbers of Speakers Using Global and Local Attractors", in IEEE/ACM TASLP, 2022. (Hitachi) π
Online TS-VAD: "Online Target Speaker Voice Activity Detection for Speaker Diarization", in Proc. Interspeech, 2022. (DKU) π
2021 (4 papers)
"Online End-to-End Neural Diarization with Speaker-Tracing Buffer", in Proc. IEEE SLT, 2021. (Hitachi) π
BW-EDA-EEND: "BW-EDA-EEND: Streaming End-to-End Neural Speaker Diarization for a Variable Number of Speakers", in Proc. Interspeech, 2021. (Amazon) π
FS-EEND: "Online Streaming End-to-End Neural Diarization Handling Overlapping Speech and Flexible Numbers of Speakers", in Proc. Interspeech, 2021. (Hitachi) ππ
Diart: "Overlap-aware low-latency online speaker diarization based on end-to-end local segmentation", in Proc. ASRU, 2021. ππ»
2020 (1 paper)
"Supervised online diarization with sample mean loss for multi-domain data", in Proc. ICASSP, 2020 ππ»
π Clustering-based β 21 papers
π₯ 2025 (3 papers)
E-SHARC: "End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization," in IEEE/ACM TASLP, 2025. (IISC) π
"Speaker Diarization with Overlapping Community Detection Using Graph Attention Networks and Label Propagation Algorithm," in Proc. Interspeech, 2025. π
2024 (6 papers)
"Overlap-aware End-to-End Supervised Hierarchical Graph Clustering for Speaker Diarization," in submitted to IEEE/ACM TASLP, 2024. π
"Apollo's Unheard Voices: Graph Attention Networks for Speaker Diarization and Clustering for Fearless Steps Apollo Collection," in Proc. ICASSP, 2024. (UTD) π
"Multi-View Speaker Embedding Learning for Enhanced Stability and Discriminability," in Proc. ICASSP, 2024. (Tsinghua) π
"Towards Unsupervised Speaker Diarization System for Multilingual Telephone Calls Using Pre-trained Whisper Model and Mixture of Sparse Autoencoders," in arXiv:2407.01963, 2024. π
"Investigating Confidence Estimation Measures for Speaker Diarization," in Proc. Interspeech, 2024. π
"Once more Diarization: Improving meeting transcription systems through segment-level speaker reassignment," in Proc. Interspeech, 2024. (PU) πππ»
2023 (5 papers)
SCALE: "Spectral Clustering-aware Learning of Embeddings for Speaker Diarisation", in Proc. ICASSP, 2023. (CAM) π
SHARC: "Supervised Hierarchical Clustering using Graph Neural Networks for Speaker Diarization", in Proc. ICASSP, 2023. (IISC) π
CDGCN: "Community Detection Graph Convolutional Network for Overlap-Aware Speaker Diarization," in Proc. ICASSP, 2023. (XMU) π
"Pyannote.Audio 2.1: Speaker Diarization Pipeline: Principle, Benchmark and Recipe", in Proc. Interspeech, 2023. (CNRS) π
GADEC: "Graph attention-based deep embedded clustering for speaker diarization,", in Speech Communication, 2023. (NJUPT) π
2020-2022 (4 papers)
UMAP-Leiden: "Reformulating Speaker Diarization as Community Detection With Emphasis On Topological Structure", in Proc. ICASSP, 2022. (Alibaba) π
Pyannote 2.0: "End-to-end speaker segmentation for overlap-aware resegmentation", in Proc. Interspeech, 2021. (CNRS) ππ»π¬
Pyannote: "pyannote.audio: neural building blocks for speaker diarization", in Proc. ICASSP, 2020. (CNRS) ππ»π¬
Resegmentation with VB: βOverlap-Aware Diarization: Resegmentation Using Neural End-to-End Overlapped Speech Detectionβ, in Proc. ICASSP, 2020. π
DNC: "Discriminative Neural Clustering for Speaker Diarisation", in Proc. IEEE SLT, 2019. ππ»π
NME-SC: βAuto-Tuning Spectral Clustering for Speaker Diarization Using Normalized Maximum Eigengapβ, IEEE SPL, 2019. ππ»
π Variational Bayes and HMM β 24 papers
π₯ 2023-2024 (3 papers)
DVBx: "Discriminative Training of VBx Diarization", in Proc. ICASSP, 2024. (BUT) ππ»
MS-VBx: "Multi-Stream Extension of Variational Bayesian HMM Clustering (MS-VBx) for Combined End-to-End and Vector Clustering-based Diarization", in Proc. Interspeech, 2023. (NTT) π
"Generalized domain adaptation framework for parametric back-end in speaker recognition", in arXiv:2305.15567, 2023. π
2021-2022 (4 papers)
"Bayesian HMM clustering of x-vector sequences (VBx) in speaker diarization: theory, implementation and analysis on standard tasks", in Computer Speech & Language, 2022. (BUT) π
DCA-PLDA "A Speaker Verification Backend with Robust Performance across Conditionsβ, in Computer & Language, 2022. ππ»
"Analysis of the but Diarization System for Voxconverse Challenge", in Proc. ICASSP, 2021. (BUT) ππ»
"Discriminatively trained probabilistic linear discriminant analysis for speaker verification", in Proc. ICASSP, 2021. π
2019-2020 (4 papers)
"Optimizing Bayesian Hmm Based X-Vector Clustering for the Second Dihard Speech Diarization Challenge", in Proc. ICASSP, 2020. (BUT) π
βAnalysis of Speaker Diarization Based on Bayesian HMM With Eigenvoice Priorsβ, IEEE/ACM TASLP, 2019. (BUT) π
"BUT System Description for DIHARD Speech Diarization Challenge 2019", in arXiv:1910.08847, 2019. (BUT) π
"Bayesian HMM Based x-Vector Clustering for Speaker Diarization", in Proc. Interspeech, 2019. (BUT) π
2018 (5 papers)
"Speaker Diarization based on Bayesian HMM with Eigenvoice Priors", in Proc. Odyssey, 2018. (BUT) π
"VB-HMM Speaker Diarization with Enhanced and Refined Segment Representation", in Proc. Odyssey, 2018. (Tsinghua) π
"Diarization is hard: some experiences and lessons learned for the JHU team in the inaugural DIHARD challenge", in Proc. Interspeech, 2018. π
"The speaker partitioning problem", in Proc. Odyssey, 2018. π
"Estimation of the Number of Speakers with Variational Bayesian PLDA in the DIHARD Diarization Challenge", in Proc. Interspeech, 2018. π
2015-2017 (3 papers)
"Domain Adaptation of PLDA Models in Broadcast Diarization by Means of Unsupervised Speaker Clustering, in Proc. Interspeech, 2017. π
"Iterative PLDA Adaptation for Speaker Diarization", in Proc. Interspeech, 2016. π
"Diarization resegmentation in the factor analysis subspace", in Proc. ICASSP, 2015. π
2011-2014 (3 papers)
"Speaker diarization with plda i-vector scoring and unsupervised calibration", in Proc. IEEE SLT, 2014. π
"Unsupervised Methods for Speaker Diarization: An Integrated and Iterative Approach", IEEE/ACM TASLP, 2013. π
"Analysis of i-vector length normalization in speaker recognition systems", in Proc. Interspeech, 2011. π
2005-2008 (2 papers)
"Bayesian analysis of speaker diarization with eigenvoice priors", in CRIM, Montreal, Technical Report, 2008. π
"Variational Bayesian methods for audio indexing", in Proc. ICMI-MLMI, 2005. π
"Geodesic interpolation of frame-wise speaker embeddings for the diarization of meeting scenarios", in Proc. ICASSP, 2024. (PU) ππ
"Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization," in Proc. Odyssey, 2024. (IDLab) π
"Efficient Speaker Embedding Extraction Using a Twofold Sliding Window Algorithm for Speaker Diarization," in Proc. Interspeech, 2024. (HU) π
"Variable Segment Length and Domain-Adapted Feature Optimization for Speaker Diarization," in Proc. Interspeech, 2024. (XMU) ππ»
2023 (5 papers)
"In Search of Strong Embedding Extractors For Speaker Diarization", in Proc. ICASSP, 2023. (Naver) ππ
DR-DESA: "Advancing the dimensionality reduction of speaker embeddings for speaker diarisation: disentangling noise and informing speech activity", in Proc. ICASSP, 2023. (Naver) ππ
HEE: "High-resolution embedding extractor for speaker diarisation", in Proc. ICASSP, 2023. (Naver) ππ
"Frame-wise and overlap-robust speaker embeddings for meeting diarization", in Proc. ICASSP, 2023. (PU) ππ
"A Teacher-Student approach for extracting informative speaker embeddings from speech mixtures", in Proc. Interspeech, 2023. (PU) π
2022 (3 papers)
GAT+AA: "Multi-scale speaker embedding-based graph attention networks for speaker diarisation", in Proc. ICASSP, 2022. (Naver) π
MSDD: "Multi-scale Speaker Diarization with Dynamic Scale Weighting", in Proc. Interspeech, 2022. (NVIDIA) ππ»π
PRISM: "PRISM: Pre-trained Indeterminate Speaker Representation Model for Speaker Diarization and Speaker Verification", in Proc. Interspeech, 2022. (Alibaba) π
2021 (2 papers)
"Multi-Scale Speaker Diarization With Neural Affinity Score Fusion", in Proc. ICASSP, 2021. (USC) π
AA+DR+NS: "Adapting Speaker Embeddings for Speaker Diarisation", in Proc. Interspeech, 2021. (Naver) ππ
πͺͺ With Speaker Identification β 1 paper
π₯ 2024 (1 paper)
"Uncertainty Quantification in Machine Learning for Joint Speaker Diarization and Identification, in Submitted to IEEE/ACM TASLP, 2024. π
"Rethinking Session Variability: Leveraging Session Embeddings for Session Robustness in Speaker Verification," in Proc. ICASSP, 2024. (Naver) π
"Leveraging In-the-Wild Data for Effective Self-Supervised Pretraining in Speaker Recognition," in Proc. ICASSP, 2024. (CUHK) π
"Disentangled Representation Learning for Environment-agnostic Speaker Recognition," in Proc. Interspeech, 2024. (KAIST) πππ»
2023 (3 papers)
"Build a SRE Challenge System: Lessons from VoxSRC 2022 and CNSRC 2022," in Proc. Interspeech, 2023. (SJTU) π
RecXi "Disentangling Voice and Content with Self-Supervision for Speaker Recognition," in Proc. NeurIPS, 2023. (A*STAR) π
"ECAPA2: A Hybrid Neural Network Architecture and Training Strategy for Robust Speaker Embeddings," in Proc. ASRU, 2023. (IDLab) πππ
2021 (1 paper)
"Xi-Vector Embedding for Speaker Recognition," in IEEE, SPL. (A*STAR) ππ
π Scoring β 3 papers
π₯ 2019-2023 (3 papers)
βSimilarity Measurement of Segment-Level Speaker Embeddings in Speaker Diarizationβ, IEEE/ACM TASLP, 2023. (DKU) π
"Self-Attentive Similarity Measurement Strategies in Speaker Diarization", in Proc. Interspeech, 2020. (DKU) π
LSTM scoring: "LSTM based Similarity Measurement with Spectral Clustering for Speaker Diarization", in Proc. Interspeech, 2019. (DKU) π
π£οΈ With ASR β 27 papers
π₯ 2025-2026 (8 papers)
SE-DiCoW: "Self-Enrolled Diarization-Conditioned Whisper," in arXiv:2601.19194, 2026. (BUT) π
TagSpeech: "End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding," in arXiv:2601.06896, 2026. π
DiCoW: "Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition," in Proc. ICASSP, 2025. (BUT) ππ»
"Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition," in arXiv:2510.03723, 2025. (BUT) π
"Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models," in arXiv:2506.05796, 2025. (DKU) π
"Language Modelling for Speaker Diarization in Telephonic Interviews," in arXiv:2501.17893, 2025. π
SC-SOT: "Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition," in Proc. Interspeech, 2025. π
"Target Speaker ASR with Whisper," in Submitted to ICASSP, 2025. (BUT) ππ»
2024 (12 papers)
"Enhancing Speaker Diarization with Large Language Models: A Contextual Beam Search Approach,", in Proc. ICASSP, 2024. (NVIDIA) π
WEEND: "Towards Word-Level End-to-End Neural Speaker Diarization with Auxiliary Network," in arXiv:2309.08489, 2024. (Google) ππ
"One model to rule them all ? Towards End-to-End Joint Speaker Diarization and Speech Recognition", in Proc. ICASSP, 2024. (CMU) π
"Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization," in arXiv:2309.16482, 2024. (PU) π
βJoint Inference of Speaker Diarization and ASR with Multi-Stage Information Sharing," in Proc. ICASSP, 2024. (DKU) π
"Multitask Speech Recognition and Speaker Change Detection for Unknown Number of Speakers" in Proc. ICASSP, 2024. (Idiap) π
"A Spatial Long-Term Iterative Mask Estimation Approach for Multi-Channel Speaker Diarization and Speech Recognition," in Proc. ICASSP, 2024. (USTC) π
"On the Success and Limitations of Auxiliary Network Based Word-Level End-to-End Neural Speaker Diarization," in Proc. Interspeech, 2024. (Google) π
Sortformer: "Seamless Integration of Speaker Diarization and ASR by Bridging Timestamps and Tokens," in Proc. ICML, 2025. (NVIDIA) ππ
"Speaker Mask Transformer for Multi-talker Overlapped Speech Recognition," in arXiv:2312.10959, 2024. (NICT) π
"On Speaker Attribution with SURT," in Proc. Odyssey, 2024. (JHU) π
"Improving Speaker Assignment in Speaker-Attributed ASR for Real Meeting Applications," in Proc. Odyssey, 2024. (CNRS) π
2023 (5 papers)
"Unified Modeling of Multi-Talker Overlapped Speech Recognition and Diarization with a Sidecar Separator", in Proc. Interspeech, 2023. (CUHK) π
"Multi-resolution Approach to Identification of Spoken Languages and to Improve Overall Language Diarization System using Whisper Model", in Proc. Interspeech, 2023.
"Speaker Diarization for ASR Output with T-vectors: A Sequence Classification Approach", in Proc. Interspeech, 2023. π
"Lexical Speaker Error Correction: Leveraging Language Models for Speaker Diarization Error Correction", in Proc. Interspeech, 2023. (Amazon) π
"SA-Paraformer: Non-autoregressive End-to-End Speaker-Attributed ASR," in Proc. ASRU, 2023. (Alibaba) π
2022 (2 papers)
"Transcribe-to-Diarize: Neural Speaker Diarization for Unlimited Number of Speakers using End-to-End Speaker-Attributed ASR," in Proc. ICASSP, 2022. π
"Tandem Multitask Training of Speaker Diarisation and Speech Recognition for Meeting Transcription", in Proc. Interspeech, 2022. π
π¬ With NLP / LLM / Language β 12 papers
π₯ 2024-2025 (7 papers)
SpeakerLM: "End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models," in arXiv:2508.06372, 2025. π
"Interactive Real-Time Speaker Diarization Correction with Human Feedback," in arXiv:2509.18377, 2025. π
"DiariST: Streaming Speech Translation with Speaker Diarization," in Proc. ICASSP, 2024. (Microsoft) ππ»
JPCP: "Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation," in arXiv:2309.10456, 2024. (Alibaba) π
"DiarizationLM: Speaker Diarization Post-Processing with Large Language Models," in Proc. Interspeech, 2024. (Google) πππ»π
"LLM-based speaker diarization correction: A generalizable approach," in Submitted to IEEE/ACM TASLP, 2024. π
"AG-LSEC: Audio Grounded Lexical Speaker Error Correction," in Proc. Interspeech, 2024. (Amazon) π
2023 (3 papers)
"Exploring Speaker-Related Information in Spoken Language Understanding for Better Speaker Diarization", in Proc. ACL, 2023. (Alibaba) π
MMSCD, "Encoder-decoder multimodal speaker change detection", in Proc. Interspeech, 2023. (Naver) π
"Aligning Speakers: Evaluating and Visualizing Text-based Diarization Using Efficient Multiple Sequence Alignment,", in Proc. ICTAI, 2023. π
π Language Diarization β 2 papers
π₯ 2023 (2 papers)
"End-to-End Spoken Language Diarization with Wav2vec Embeddings", in Proc. Interspeech, 2023. ππ»
"Multi-resolution Approach to Identification of Spoken Languages and To Improve Overall Language Diarization System Using Whisper Model," in Proc. Interspeech, 2023. π
ποΈ With Vision β 21 papers
π₯ 2025-2026 (4 papers)
CineSRD: "Leveraging Visual, Acoustic, and Linguistic Cues for Open-World Visual Media Speaker Diarization," in arXiv:2603.16966, 2026. π
"Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization," in Proc. ACL, 2025. π
"Cross-Attention and Self-Attention for Audio-visual Speaker Diarization," in arXiv:2506.02621, 2025. π
"Count Your Speakers! Multitask Learning for Multimodal Speaker Diarization," in Proc. Interspeech, 2025. π
2024 (6 papers)
"Speaker Diarization of Scripted Audiovisual Content," in arXiv:2308.02160, 2024. (Amazon) π
"AFL-Net: Integrating Audio, Facial, and Lip Modalities with Cross-Attention for Robust Speaker Diarization in the Wild," in Proc. ICASSP, 2024. (Tencent) ππ¬
"Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation," in Proc. AAAI, 2024. (Tencent) π
"3D-Speaker-Toolkit: An Open Source Toolkit for Multi-modal Speaker Verification and Diarization," in arXiv:2403.19971, 2024. (Alibaba) ππ»
"Target Speech Diarization with Multimodal Prompts," in Submitted to IEEE/ACM TASLP, 2024. (NUS) π
MFV-KSD: "Multi-Stage Face-Voice Association Learning with Keynote Speaker Diarization," in Submitted to ACM MM, 2024. ππ»
2023 (5 papers)
"Audio-Visual Speaker Diarization in the Framework of Multi-User Human-Robot Interaction", in Proc. ICASSP, 2023. π
STHG: "Spatial-Temporal Heterogeneous Graph Learning for Advanced Audio-Visual Diarization, in Proc. CVPR, 2023. (Intel) π
"Uncertainty-Guided End-to-End Audio-Visual Speaker Diarization for Far-Field Recordings," in Proc. ACM MM, 2023. π
"Joint Training or Not: An Exploration of Pre-trained Speech Models in Audio-Visual Speaker Diarization," in Springer Computer Science proceedings, 2023. π
EEND-EDA++: "Late Audio-Visual Fusion for In-The-Wild Speaker Diarization," in arXiv:2211.01299v2, 2023. π
2022 (3 papers)
AVA-AVD (AVR-Net): "AVA-AVD: Audio-Visual Speaker Diarization in the Wild", in Proc. ACM MM, 2022. ππ»π¬
"End-to-End Audio-Visual Neural Speaker Diarization", in Proc. Interspeech, 2022. (USTC) ππ»π
DyViSE: "DyViSE: Dynamic Vision-Guided Speaker Embedding for Audio-Visual Speaker Diarization", in Proc. MMSP, 2022. (THU) ππ»
2024 (1 paper)
"Multi-Input Multi-Output Target-Speaker Voice Activity Detection For Unified, Flexible, and Robust Audio-Visual Speaker Diarization," in Submitted to IEEE/ACM TASLP. (DKU) π
2019-2020 (2 papers)
"Self-supervised learning for audio-visual speaker diarization", in Proc. ICASSP, 2020. (Tencent) ππ
"Who said that?: Audio-visual speaker diarisation of real-world meetings", in Proc. Interspeech, 2019. (Naver) π
π‘οΈ Related Spoofing β 1 paper
π₯ 2024 (1 paper)
"Spoof Diarization: "What Spoofed When" in Partially Spoofed Audio," in Proc. Interspeech, 2024. (IITK) π
π Related TTS
π Speaker Anonymization β 1 paper
π₯ 2024 (1 paper)
"A Benchmark for Multi-speaker Anonymization," in Submitted to IEEE/ACM TASLP, 2024. (SIT) ππ»
π Singing Diarization β 1 paper
π₯ 2024 (1 paper)
"Song Data Cleansing for End-to-End Neural Singer Diarization Using Neural Analysis and Synthesis Framework," in Proc. Interspeech, 2024. (LY) π
π With Emotion β 3 papers
π₯ 2023-2024 (3 papers)
"ED-TTS: Multi-scale Emotion Modeling using Cross-domain Emotion Diarization for Emotional Speech Synthesis, in Proc. ICASSP, 2024. π
"Speech Emotion Diarization: Which Emotion Appears When?," in Proc. ASRU, 2023. (Zaion) π
"EmoDiarize: Speaker Diarization and Emotion Identification from Speech Signals using Convolutional Neural Networks," in arxiv:2310.12851, 2023. π
ποΈ Personal VAD β 2 papers
π₯ 2020-2023 (2 papers)
"SVVAD: Personal Voice Activity Detection for Speaker Verification", in Proc. Interspeech, 2023. π
"Personal VAD: Speaker-Conditioned Voice Activity Detection", in Proc. Odyssey, 2020. (Google) π
π VAD & OSD & SCD β 10 papers
π₯ 2024 (3 papers)
"USM-SCD: Multilingual Speaker Change Detection Based on Large Pretrained Foundation Models," in Proc. ICASSP, 2024. (Google) π
"Channel-Combination Algorithms for Robust Distant Voice Activity and Overlapped Speech Detection," in IEEE/ACM TASLP, 2024. π
"Speaker Change Detection with Weighted-sum Knowledge Distillation based on Self-supervised Pre-trained Models," in Proc. Interspeech, 2024. π
2023 (4 papers)
"Multitask Detection of Speaker Changes, Overlapping Speech and Voice Activity Using wav2vec 2.0," in Proc. ICASSP, 2023. ππ»
"Semantic VAD: Low-Latency Voice Activity Detection for Speech Interaction," in Proc. Interspeech, 2023. π
"Joint speech and overlap detection: a benchmark over multiple audio setup and speech domains," in arxiv:2307.13012, 2023. π
"Advancing the study of Large-Scale Learning in Overlapped Speech Detection," in arXiv:2308.05987, 2023. π
2022 (3 papers)
"Overlapped Speech Detection in Broadcast Streams Using X-vectors," in Proc. Interspeech, 2022. π
"Overlapped speech and gender detection with WavLM pre-trained features," in Proc. Interspeech, 2022. π
"Microphone Array Channel Combination Algorithms for Overlapped Speech Detection," in Proc. Interspeech, 2022. π
π Dataset β 19 papers
π₯ 2024-2025 (6 papers)
"Conversations in the wild: Data collection, automatic generation and evaluation," in Computer Speech & Language, 2025. π
M3SD: "Multi-modal, Multi-scenario and Multi-language Speaker Diarization Dataset," in arXiv:2506.14427, 2025. π
"VoxBlink: X-Large Speaker Verification Dataset on Camera", in Proc. ICASSP, 2024. ππ
"NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant Meeting Transcription," in arXiv:2401.08887, 2024. (MS) π
"A Comparative Analysis of Speaker Diarization Models: Creating a Dataset for German Dialectal Speech," in Proc. ACL, 2024. π
"ALLIES: A Speech Corpus for Segmentation, Speaker Diarization, Speech Recognition and Speaker Change Detection," in Proc. ACL, 2024. (LIUM) π
2020-2022 (5 papers)
Ego4D: " Around the World in 3,000 Hours of Egocentric Video," in Proc. CVPR, 2022. (Meta) ππ»π
AliMeeting: "Summary On The ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge," in Proc. ICASSP, 2022. (Alibaba) πππ»
Voxconverse: "Spot the conversation: speaker diarisation in the wild", in Proc. Interspeech, 2020. (VGG, Naver) ππ»π
MSDWild: Multi-modal Speaker Diarization Dataset in the Wild, in Proc. Interspeech, 2020. ππ
"LibriMix: An Open-Source Dataset for Generalizable Speech Separation," in arXiv:2005.11262, 2020. ππ»
π Simulated Dataset β 8 papers
π₯ 2023-2024 (4 papers)
"Enhancing low-latency speaker diarization with spatial dictionary learning," in Proc. ICASSP, 2024. (NTU) ππ
"Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling," in Proc. ICASSP, 2024. (OSU) π
"Multi-Speaker and Wide-Band Simulated Conversations as Training Data for End-to-End Neural Diarization", in Proc. ICASSP, 2023. (BUT) ππ»π
"Property-Aware Multi-Speaker Data Simulation: A Probabilistic Modelling Technique for Synthetic Data Generation," in CHiME-7 Workshop, 2023. (NVIDIA) π
2022 (3 papers)
"From simulated mixtures to simulated conversations as training data for end-to-end neural diarization" , in Proc. Interspeech, 2022. (BUT) ππ»π
Markov selection: "Improving the naturalness of simulated conversations for end-to-end neural diarization", in Proc. Odyssey, 2022. (Hitachi) π
EEND-EDA-SpkAtt: "Towards End-to-end Speaker Diarization in the Wild", in arXiv:2211.01299v1, 2022. π
2019 (1 paper)
Concat-and-sum: "End-to-end neuarl speaker diarization with permuation-free objectives", in Proc. Interspeech, 2019. π
π οΈ Tools β 1 paper
π₯ 2024 (1 paper)
"Gryannote open-source speaker diarization labeling tool," in Proc. Interspeech (Show and Tell), 2024. (IRIT) ππ»
π Self-Supervised β 5 papers
π₯ 2025 (3 papers)
DiariZen: "Leveraging Self-Supervised Learning for Speaker Diarization," in Proc. ICASSP," 2025. (BUT) πππ»π
"Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models," in arXiv:2506.18623, 2025. (BUT) ππ»
"Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization," in arXiv:2505.24111, 2025. π
2022 (2 papers)
βSelf-supervised Speaker Diarizationβ, in Proc. Interspeech, 2022. π
CSDA: "Continual Self-Supervised Domain Adaptation for End-to-End Speaker Diarization", in Proc. IEEE SLT, 2022. (CNRS) ππ»
π Semi-Supervised β 1 paper
π₯ 2017 (1 paper)
"Active Learning Based Constrained Clustering For Speaker Diarization", in IEEE/ACM TASLP, 2017. (UT) π
π Measurement β 3 papers
π₯ 2022-2025 (3 papers)
SDBench: βA Comprehensive Benchmark Suite for Speaker Diarization,β in Proc. Interspeech, 2025. π
βBenchmarking Diarization Models,β in arXiv:2509.26177, 2025. π
BER: βBalanced Error Rate For Speaker Diarizationβ, in Proc. arXiv:2211.04304, 2022 ππ»
πΆ Child-Adult β 2 papers
π₯ 2023-2024 (2 papers)
"Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions," in Proc. Interspeech, 2024. (USC) π
"Robust Self Supervised Speech Embeddings for Child-Adult Classification in Interactions involving Children with Autism," in Proc. Interspeech, 2023. π
"The DISPLACE Challenge 2023 - DIarization of SPeaker and LAnguage in Conversational Environments," in Proc. Interspeech, 2023. ππ
"The SpeeD--ZevoTech submission at DISPLACE 2023," in Proc. Interspeech, 2023. π
π MERLIon CCS Challenge 2023 β 1 paper
π₯ 2023 (1 paper)
"MERLIon CCS Challenge: A English-Mandarin code-switching child-directed speech corpus for language identification and diarization," in Proc. Interspeech, 2023. ππ
π CHiME-6
π ICMC-ASR Grand Challenge (ICASSP2024) β 2 papers