Wespeaker implementations for speaker recognition and verification: U3-xi,Uncertainty aware AAM Softmax, Score normalization, calibration, and unofficial RecXi (NeurIPS 2023) source code.
Python
26
349 commits
updated Sep 19, 2026
💡 WeSpeaker implementations for uncertainty-aware speaker recognition and verification, covering uncertainty estimation, speaker embedding learning, scoring, score normalization, and calibration.
Keywords: RecXi, NeurIPS 2023, speaker recognition, speaker verification, speaker embedding, WeSpeaker implementation, source code, ECAPA-TDNN, uncertainty estimation, uncertainty-aware speaker modeling.
Core idea · Getting started · U³-xi · Robust speaker modeling · Unified back-end · RecXi · Citation

Editable draw.io source · Vector SVG
The goal is to estimate higher uncertainty when speech provides less reliable speaker evidence, such as under noise or reverberation. Frame-level precision estimates guide Bayesian pooling so reliable frames contribute more to the utterance representation. The model outputs both a speaker embedding and its associated uncertainty.
Conceptual illustration; the waveforms and distributions are schematic, not experimental measurements.
This project builds on WeSpeaker. Refer to the original WeSpeaker README for environment setup and the VoxCeleb recipe for data preparation. Use the configurations and scripts in this repository for the uncertainty-aware experiments described below.
Run recipe scripts from examples/voxceleb/v2 after configuring the dataset paths, GPU IDs, and experiment directory for your environment.
Junjie Li, Kong Aik Lee
The Hong Kong Polytechnic University
Email: junjie98.li@connect.polyu.hk
U³-xi incorporates uncertainty into speaker embedding learning and verification. The implementation provides:
Training and evaluation recipes are available in examples/voxceleb/v2. The main implementation entry points are listed below.
The table reports equal error rate (EER, %) on the cleaned VoxCeleb1-O, VoxCeleb1-E, and VoxCeleb1-H trial lists. Lower values are better. LM denotes large-margin fine-tuning, AS-Norm denotes adaptive score normalization, and QMF denotes quality measure function calibration. ✓ and × indicate whether a component is enabled; rows with a blank model name use the model listed above; blank FLOPs values for named models were not reported.
| Model | Parameters | FLOPs | LM | AS-Norm | QMF | Vox1-O-clean | Vox1-E-clean | Vox1-H-clean |
|---|---|---|---|---|---|---|---|---|
| ECAPA_TDNN_GLOB_c512-ASTP-emb192 | 6.19M | 1.04G | × | × | × | 1.069 | 1.209 | 2.310 |
| × | ✓ | × | 0.957 | 1.128 | 2.105 | |||
| ✓ | × | × | 0.878 | 1.072 | 2.007 | |||
| ✓ | ✓ | × | 0.782 | 1.005 | 1.824 | |||
| ECAPA_TDNN_GLOB_c1024-ASTP-emb192 | 14.65M | 2.65G | × | × | × | 0.856 | 1.072 | 2.059 |
| × | ✓ | × | 0.808 | 0.990 | 1.874 | |||
| ✓ | × | × | 0.798 | 0.993 | 1.883 | |||
| ✓ | ✓ | × | 0.728 | 0.929 | 1.721 | |||
| ✓ | ✓ | ✓ | 0.707 | 0.894 | 1.615 | |||
| ResNet34-TSTP-emb256 | 6.63M | 4.55G | × | × | × | 0.867 | 1.049 | 1.959 |
| × | ✓ | × | 0.787 | 0.964 | 1.726 | |||
| × | ✓ | ✓ | 0.718 | 0.911 | 1.606 | |||
| ✓ | × | × | 0.797 | 0.937 | 1.695 | |||
| ✓ | ✓ | × | 0.723 | 0.867 | 1.532 | |||
| ✓ | ✓ | ✓ | 0.659 | 0.821 | 1.437 | |||
| XI_VEC_ECAPA_TDNN_c512 | 5.9M | 1.04G | × | × | × | 0.995 | 1.130 | 2.169 |
| × | ✓ | × | 0.883 | 1.056 | 1.976 | |||
| ✓ | × | × | 0.909 | 1.000 | 1.855 | |||
| ✓ | ✓ | × | 0.787 | 0.930 | 1.693 | |||
| ECAPA_TDNN_c512_u_cube_xi(ours) | 6.7M | 1.20G | × | × | × | 0.782 | 1.016 | 1.888 |
| ResNet34_u_cube_xi(ours) | 7.9M | × | × | × | 0.867 | 0.868 | 1.641 | |
| ReDimNet-B2(ours) | 5.5M | × | × | × | 0.606 | 0.779 | 1.494 | |
| ✓ | × | × | 0.489 | 0.698 | 1.311 | |||
| ✓ | ✓ | × | 0.399 | 0.638 | 1.170 |
These checkpoints were retrained, so their results may differ slightly from those reported in the paper.
This work addresses uncertainty estimation and calibration under domain shifts through two methods:
Implementation scope: This release includes Inter- and Intra-Speaker-Aware Uncertainty Softmax only. UCDA is described in the paper but is not included in this repository.
Training note: train.py supports alpha scheduling for USphereFace2 inter-intra.
The table reports EER (%) and minimum detection cost function (minDCF) values for in-domain evaluation on VoxCeleb1 and cross-domain evaluation on CNCeleb. Lower values are better. RI denotes relative improvement over the corresponding benchmark, and links in the loss column point to model checkpoints.
| Model | # Param. | Loss | Uncertainty-aware cosine score | In-domain | Cross-domain | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Vox1-O | Vox1-E | Vox1-H | RI (%) | CNCeleb | RI (%) | ||||||||
| EER | minDCF | EER | minDCF | EER | minDCF | EER | minDCF | ||||||
| ECAPA512 | 6.19 M | AAM-Softmax | No | 1.069 | 0.122 | 1.209 | 0.136 | 2.310 | 0.226 | Benchmark | 15.314 | 0.633 | Benchmark |
| 6.69 M | UAAM-Softmax | No | 0.856 | 0.109 | 1.064 | 0.121 | 1.982 | 0.195 | 13.57 | 13.706 | 0.608 | 7.23 | |
| Yes | 0.782 | 0.100 | 1.016 | 0.115 | 1.888 | 0.187 | 18.64 | 10.271 | 1.000 | -12.52 | |||
| 6.69 | UAAM-Softmax inter-intra | No | 0.936 | 0.102 | 1.050 | 0.122 | 1.978 | 0.195 | 13.40 | 13.974 | 0.581 | 8.48 | |
| Yes | 0.840 | 0.086 | 0.965 | 0.110 | 1.833 | 0.189 | 21.22 | 10.781 | 0.835 | -1.16 | |||
| 6.19 M | AM-Softmax | No | 1.005 | 0.107 | 1.206 | 0.133 | 2.254 | 0.221 | Benchmark | 14.162 | 0.611 | Benchmark | |
| 6.69 M | UAM-Softmax inter-intra | No | 0.888 | 0.099 | 1.076 | 0.119 | 1.973 | 0.186 | 11.46 | 12.436 | 0.553 | 10.84 | |
| Yes | 0.808 | 0.084 | 0.991 | 0.109 | 1.794 | 0.178 | 19.46 | 9.411 | 1.000 | -15.03 | |||
| 6.19 M | SphereFace2 | No | 0.963 | 0.108 | 1.121 | 0.125 | 1.967 | 0.199 | Benchmark | 12.582 | 0.573 | Benchmark | |
| 6.69 M | USphereFace2 inter-intra | No | 0.856 | 0.104 | 1.035 | 0.119 | 1.918 | 0.196 | 5.21 | 12.265 | 0.550 | 3.27 | |
| Yes | 0.739 | 0.102 | 0.965 | 0.108 | 1.771 | 0.178 | 12.81 | 10.560 | 0.624 | 3.59 | |||
| ResNet34 | 6.63 M | AAM-Softmax | No | 0.867 | 0.091 | 1.049 | 0.121 | 1.960 | 0.192 | Benchmark | 11.090 | 0.488 | Benchmark |
| 7.92 M | UAAM-Softmax | No | 0.888 | 0.085 | 0.900 | 0.099 | 1.712 | 0.175 | 9.68 | 11.732 | 0.513 | -5.46 | |
| Yes | 0.867 | 0.078 | 0.868 | 0.095 | 1.641 | 0.172 | 13.29 | 10.082 | 0.541 | -0.89 | |||
| 7.92 M | UAAM-Softmax inter-intra | No | 0.904 | 0.070 | 0.933 | 0.098 | 1.658 | 0.165 | 13.06 | 12.116 | 0.505 | -6.37 | |
| Yes | 0.813 | 0.075 | 0.847 | 0.091 | 1.532 | 0.167 | 17.12 | 9.631 | 0.539 | 1.35 | |||
| 7.92 M | USphereFace2 | No | 1.483 | 0.148 | 1.451 | 0.156 | 2.112 | 0.206 | -36.00 | 11.441 | 0.512 | -4.04 | |
| Yes | 1.340 | 0.156 | 1.357 | 0.150 | 1.986 | 0.193 | -30.19 | 10.949 | 0.499 | -0.49 | |||
| ReDimNet-B2 | 4.89 M | AAM-Softmax | No | 0.782 | 0.064 | 0.907 | 0.097 | 1.667 | 0.162 | Benchmark | 12.385 | 0.552 | Benchmark |
| 5.46 M | UAAM-Softmax | No | 0.649 | 0.073 | 0.801 | 0.089 | 1.532 | 0.153 | 6.09 | 13.464 | 0.552 | -4.36 | |
| Yes | 0.606 | 0.065 | 0.779 | 0.091 | 1.494 | 0.157 | 9.12 | 9.479 | 1.000 | -28.85 | |||
| 5.46 M | UAAM-Softmax inter-intra | No | 0.686 | 0.070 | 0.802 | 0.090 | 1.536 | 0.151 | 6.06 | 12.132 | 0.516 | 4.28 | |
| Yes | 0.627 | 0.064 | 0.758 | 0.088 | 1.434 | 0.153 | 10.84 | 8.607 | 0.838 | -10.65 | |||
| 5.46 M | USphereFace2 inter-intra | No | 0.622 | 0.052 | 0.776 | 0.085 | 1.440 | 0.146 | 14.92 | 12.081 | 0.515 | 4.58 | |
| Yes | 0.622 | 0.051 | 0.774 | 0.084 | 1.433 | 0.145 | 15.56 | 11.899 | 0.506 | 6.13 | |||
This work carries embedding uncertainty through the speaker verification back-end, from similarity scoring to normalization and calibration. Each utterance is represented by a speaker embedding and its covariance estimate.
The unified pipeline includes three components:
The VoxCeleb recipe includes uncertainty extraction, scoring, normalization, and calibration. The corresponding scripts are listed in pipeline order below.
Uncertainty extraction
Uncertainty-aware cosine scoring
Uncertainty-aware AS-Norm (UAS-Norm)
Uncertainty-aware calibration (UQMF)
The table reports EER (%) and minDCF on Vox1-O, Vox1-E, and Vox1-H. Lower values are better. Bold and underlined values denote the best and second-best results, respectively, within each architecture group. RI is the average relative improvement over the corresponding architecture-specific benchmark across the six VoxCeleb metrics. The suffixes -o and -u denote the original and uncertainty-aware variants, respectively; ①–③ identify the UAS-Norm components evaluated in the paper.
| Row | Model | # Param. | Cosine score | AS-Norm | QMFs | Vox1-O | Vox1-E | Vox1-H | RI (%) | |||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| EER | minDCF | EER | minDCF | EER | minDCF | |||||||
| 1 | ECAPA† | 6.19 M | scos-o | 1.069 | 0.122 | 1.209 | 0.136 | 2.310 | 0.226 | Benchmark | ||
| 2 | ECAPA+xi | 6.69 M | scos-o | 0.936 | 0.100 | 1.054 | 0.123 | 1.978 | 0.195 | 13.49 | ||
| 3 | scos-o | sAS-o | 0.840 | 0.111 | 0.979 | 0.115 | 1.809 | 0.173 | 18.34 | |||
| 4 | scos-o | sAS-o | sQMF-o | 0.766 | 0.106 | 0.932 | 0.108 | 1.693 | 0.167 | 22.96 | ||
| 5 | scos-u | 0.840 | 0.086 | 0.965 | 0.110 | 1.833 | 0.189 | 21.21 | ||||
| 6 | scos-u | sAS-u (①) | 0.761 | 0.101 | 0.913 | 0.104 | 1.677 | 0.167 | 24.59 | |||
| 7 | scos-u | sAS-u (①+②) | 0.761 | 0.100 | 0.912 | 0.104 | 1.675 | 0.166 | 24.83 | |||
| 8 | scos-u | sAS-u (①+②+③) | 0.750 | 0.098 | 0.906 | 0.100 | 1.663 | 0.167 | 25.86 | |||
| 9 | scos-u | sAS-u | sQMF-u | 0.750 | 0.095 | 0.892 | 0.090 | 1.629 | 0.166 | 28.01 | ||
| 10 | ResNet† | 6.63 M | scos-o | 0.867 | 0.091 | 1.049 | 0.121 | 1.960 | 0.192 | Benchmark | ||
| 11 | ResNet+xi | 7.92 M | scos-o | 0.904 | 0.070 | 0.933 | 0.098 | 1.658 | 0.165 | 13.06 | ||
| 12 | scos-o | sAS-o | 0.888 | 0.079 | 0.922 | 0.096 | 1.618 | 0.163 | 12.68 | |||
| 13 | scos-o | sAS-o | sQMF-o | 0.782 | 0.065 | 0.842 | 0.090 | 1.489 | 0.151 | 21.52 | ||
| 14 | scos-u | 0.813 | 0.075 | 0.847 | 0.091 | 1.532 | 0.167 | 17.12 | ||||
| 15 | scos-u | sAS-u | 0.771 | 0.052 | 0.805 | 0.086 | 1.400 | 0.145 | 26.53 | |||
| 16 | scos-u | sAS-u | sQMF-u | 0.745 | 0.053 | 0.793 | 0.086 | 1.386 | 0.145 | 27.15 | ||
† Results obtained directly from pretrained WeSpeaker models.
If you use this repository in your research, please cite the papers relevant to the methods you use. BibTeX entries are provided below; select an IEEE bibliography style in your manuscript if required.
@inproceedings{li2026xiplus,
author = {Junjie Li and Kong Aik Lee and Duc-Tuan Truong and Tianchi Liu and Man-Wai Mak},
title = {{XI+}: Uncertainty Supervision for Robust Speaker Embedding},
booktitle = {2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
year = {2026},
pages = {18847--18851},
doi = {10.1109/ICASSP55912.2026.11463369}
}
@article{li2026u3xi,
author = {Junjie Li and Kong Aik Lee},
title = {{U3}-xi: Pushing the Boundaries of Speaker Recognition by Incorporating Uncertainty},
journal = {arXiv preprint arXiv:2601.15719},
year = {2026},
doi = {10.48550/arXiv.2601.15719},
url = {https://arxiv.org/abs/2601.15719}
}
@article{li2026robust,
author = {Junjie Li and Yang Xiao and Kong Aik Lee},
title = {Towards Robust Uncertainty-Aware Speaker Modeling},
journal = {arXiv preprint arXiv:2607.04937},
year = {2026},
doi = {10.48550/arXiv.2607.04937},
url = {https://arxiv.org/abs/2607.04937}
}
@article{li2026unified,
author = {Junjie Li and Kong Aik Lee},
title = {A Unified Uncertainty-Aware Back-End for Speaker Verification: Scoring, Normalization, and Calibration},
journal = {arXiv preprint arXiv:2609.01221},
year = {2026},
doi = {10.48550/arXiv.2609.01221},
url = {https://arxiv.org/abs/2609.01221}
}
If you use the RecXi implementation, please also cite the original RecXi paper:
@inproceedings{liu2023disentangling,
author = {Tianchi Liu and Kong Aik Lee and Qiongqiong Wang and Haizhou Li},
title = {Disentangling Voice and Content with Self-Supervision for Speaker Recognition},
booktitle = {Advances in Neural Information Processing Systems},
volume = {36},
year = {2023},
url = {https://proceedings.neurips.cc/paper_files/paper/2023/hash/9d276b0a087efdd2404f3295b26c24c1-Abstract-Conference.html}
}
Please also acknowledge the underlying toolkit using the WeSpeaker citations.
Python
86.4%
C++
7.9%
Shell
4.1%
CMake
1.0%
Wespeaker implementations for speaker recognition and verification: U3-xi,Uncertainty aware AAM Softmax, Score normalization, calibration, and unofficial RecXi (NeurIPS 2023) source code.
Python
26
349 commits
updated Sep 19, 2026
💡 WeSpeaker implementations for uncertainty-aware speaker recognition and verification, covering uncertainty estimation, speaker embedding learning, scoring, score normalization, and calibration.
Keywords: RecXi, NeurIPS 2023, speaker recognition, speaker verification, speaker embedding, WeSpeaker implementation, source code, ECAPA-TDNN, uncertainty estimation, uncertainty-aware speaker modeling.
Core idea · Getting started · U³-xi · Robust speaker modeling · Unified back-end · RecXi · Citation

Editable draw.io source · Vector SVG
The goal is to estimate higher uncertainty when speech provides less reliable speaker evidence, such as under noise or reverberation. Frame-level precision estimates guide Bayesian pooling so reliable frames contribute more to the utterance representation. The model outputs both a speaker embedding and its associated uncertainty.
Conceptual illustration; the waveforms and distributions are schematic, not experimental measurements.
This project builds on WeSpeaker. Refer to the original WeSpeaker README for environment setup and the VoxCeleb recipe for data preparation. Use the configurations and scripts in this repository for the uncertainty-aware experiments described below.
Run recipe scripts from examples/voxceleb/v2 after configuring the dataset paths, GPU IDs, and experiment directory for your environment.
Junjie Li, Kong Aik Lee
The Hong Kong Polytechnic University
Email: junjie98.li@connect.polyu.hk
U³-xi incorporates uncertainty into speaker embedding learning and verification. The implementation provides:
Training and evaluation recipes are available in examples/voxceleb/v2. The main implementation entry points are listed below.
The table reports equal error rate (EER, %) on the cleaned VoxCeleb1-O, VoxCeleb1-E, and VoxCeleb1-H trial lists. Lower values are better. LM denotes large-margin fine-tuning, AS-Norm denotes adaptive score normalization, and QMF denotes quality measure function calibration. ✓ and × indicate whether a component is enabled; rows with a blank model name use the model listed above; blank FLOPs values for named models were not reported.
| Model | Parameters | FLOPs | LM | AS-Norm | QMF | Vox1-O-clean | Vox1-E-clean | Vox1-H-clean |
|---|---|---|---|---|---|---|---|---|
| ECAPA_TDNN_GLOB_c512-ASTP-emb192 | 6.19M | 1.04G | × | × | × | 1.069 | 1.209 | 2.310 |
| × | ✓ | × | 0.957 | 1.128 | 2.105 | |||
| ✓ | × | × | 0.878 | 1.072 | 2.007 | |||
| ✓ | ✓ | × | 0.782 | 1.005 | 1.824 | |||
| ECAPA_TDNN_GLOB_c1024-ASTP-emb192 | 14.65M | 2.65G | × | × | × | 0.856 | 1.072 | 2.059 |
| × | ✓ | × | 0.808 | 0.990 | 1.874 | |||
| ✓ | × | × | 0.798 | 0.993 | 1.883 | |||
| ✓ | ✓ | × | 0.728 | 0.929 | 1.721 | |||
| ✓ | ✓ | ✓ | 0.707 | 0.894 | 1.615 | |||
| ResNet34-TSTP-emb256 | 6.63M | 4.55G | × | × | × | 0.867 | 1.049 | 1.959 |
| × | ✓ | × | 0.787 | 0.964 | 1.726 | |||
| × | ✓ | ✓ | 0.718 | 0.911 | 1.606 | |||
| ✓ | × | × | 0.797 | 0.937 | 1.695 | |||
| ✓ | ✓ | × | 0.723 | 0.867 | 1.532 | |||
| ✓ | ✓ | ✓ | 0.659 | 0.821 | 1.437 | |||
| XI_VEC_ECAPA_TDNN_c512 | 5.9M | 1.04G | × | × | × | 0.995 | 1.130 | 2.169 |
| × | ✓ | × | 0.883 | 1.056 | 1.976 | |||
| ✓ | × | × | 0.909 | 1.000 | 1.855 | |||
| ✓ | ✓ | × | 0.787 | 0.930 | 1.693 | |||
| ECAPA_TDNN_c512_u_cube_xi(ours) | 6.7M | 1.20G | × | × | × | 0.782 | 1.016 | 1.888 |
| ResNet34_u_cube_xi(ours) | 7.9M | × | × | × | 0.867 | 0.868 | 1.641 | |
| ReDimNet-B2(ours) | 5.5M | × | × | × | 0.606 | 0.779 | 1.494 | |
| ✓ | × | × | 0.489 | 0.698 | 1.311 | |||
| ✓ | ✓ | × | 0.399 | 0.638 | 1.170 |
These checkpoints were retrained, so their results may differ slightly from those reported in the paper.
This work addresses uncertainty estimation and calibration under domain shifts through two methods:
Implementation scope: This release includes Inter- and Intra-Speaker-Aware Uncertainty Softmax only. UCDA is described in the paper but is not included in this repository.
Training note: train.py supports alpha scheduling for USphereFace2 inter-intra.
The table reports EER (%) and minimum detection cost function (minDCF) values for in-domain evaluation on VoxCeleb1 and cross-domain evaluation on CNCeleb. Lower values are better. RI denotes relative improvement over the corresponding benchmark, and links in the loss column point to model checkpoints.
| Model | # Param. | Loss | Uncertainty-aware cosine score | In-domain | Cross-domain | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Vox1-O | Vox1-E | Vox1-H | RI (%) | CNCeleb | RI (%) | ||||||||
| EER | minDCF | EER | minDCF | EER | minDCF | EER | minDCF | ||||||
| ECAPA512 | 6.19 M | AAM-Softmax | No | 1.069 | 0.122 | 1.209 | 0.136 | 2.310 | 0.226 | Benchmark | 15.314 | 0.633 | Benchmark |
| 6.69 M | UAAM-Softmax | No | 0.856 | 0.109 | 1.064 | 0.121 | 1.982 | 0.195 | 13.57 | 13.706 | 0.608 | 7.23 | |
| Yes | 0.782 | 0.100 | 1.016 | 0.115 | 1.888 | 0.187 | 18.64 | 10.271 | 1.000 | -12.52 | |||
| 6.69 | UAAM-Softmax inter-intra | No | 0.936 | 0.102 | 1.050 | 0.122 | 1.978 | 0.195 | 13.40 | 13.974 | 0.581 | 8.48 | |
| Yes | 0.840 | 0.086 | 0.965 | 0.110 | 1.833 | 0.189 | 21.22 | 10.781 | 0.835 | -1.16 | |||
| 6.19 M | AM-Softmax | No | 1.005 | 0.107 | 1.206 | 0.133 | 2.254 | 0.221 | Benchmark | 14.162 | 0.611 | Benchmark | |
| 6.69 M | UAM-Softmax inter-intra | No | 0.888 | 0.099 | 1.076 | 0.119 | 1.973 | 0.186 | 11.46 | 12.436 | 0.553 | 10.84 | |
| Yes | 0.808 | 0.084 | 0.991 | 0.109 | 1.794 | 0.178 | 19.46 | 9.411 | 1.000 | -15.03 | |||
| 6.19 M | SphereFace2 | No | 0.963 | 0.108 | 1.121 | 0.125 | 1.967 | 0.199 | Benchmark | 12.582 | 0.573 | Benchmark | |
| 6.69 M | USphereFace2 inter-intra | No | 0.856 | 0.104 | 1.035 | 0.119 | 1.918 | 0.196 | 5.21 | 12.265 | 0.550 | 3.27 | |
| Yes | 0.739 | 0.102 | 0.965 | 0.108 | 1.771 | 0.178 | 12.81 | 10.560 | 0.624 | 3.59 | |||
| ResNet34 | 6.63 M | AAM-Softmax | No | 0.867 | 0.091 | 1.049 | 0.121 | 1.960 | 0.192 | Benchmark | 11.090 | 0.488 | Benchmark |
| 7.92 M | UAAM-Softmax | No | 0.888 | 0.085 | 0.900 | 0.099 | 1.712 | 0.175 | 9.68 | 11.732 | 0.513 | -5.46 | |
| Yes | 0.867 | 0.078 | 0.868 | 0.095 | 1.641 | 0.172 | 13.29 | 10.082 | 0.541 | -0.89 | |||
| 7.92 M | UAAM-Softmax inter-intra | No | 0.904 | 0.070 | 0.933 | 0.098 | 1.658 | 0.165 | 13.06 | 12.116 | 0.505 | -6.37 | |
| Yes | 0.813 | 0.075 | 0.847 | 0.091 | 1.532 | 0.167 | 17.12 | 9.631 | 0.539 | 1.35 | |||
| 7.92 M | USphereFace2 | No | 1.483 | 0.148 | 1.451 | 0.156 | 2.112 | 0.206 | -36.00 | 11.441 | 0.512 | -4.04 | |
| Yes | 1.340 | 0.156 | 1.357 | 0.150 | 1.986 | 0.193 | -30.19 | 10.949 | 0.499 | -0.49 | |||
| ReDimNet-B2 | 4.89 M | AAM-Softmax | No | 0.782 | 0.064 | 0.907 | 0.097 | 1.667 | 0.162 | Benchmark | 12.385 | 0.552 | Benchmark |
| 5.46 M | UAAM-Softmax | No | 0.649 | 0.073 | 0.801 | 0.089 | 1.532 | 0.153 | 6.09 | 13.464 | 0.552 | -4.36 | |
| Yes | 0.606 | 0.065 | 0.779 | 0.091 | 1.494 | 0.157 | 9.12 | 9.479 | 1.000 | -28.85 | |||
| 5.46 M | UAAM-Softmax inter-intra | No | 0.686 | 0.070 | 0.802 | 0.090 | 1.536 | 0.151 | 6.06 | 12.132 | 0.516 | 4.28 | |
| Yes | 0.627 | 0.064 | 0.758 | 0.088 | 1.434 | 0.153 | 10.84 | 8.607 | 0.838 | -10.65 | |||
| 5.46 M | USphereFace2 inter-intra | No | 0.622 | 0.052 | 0.776 | 0.085 | 1.440 | 0.146 | 14.92 | 12.081 | 0.515 | 4.58 | |
| Yes | 0.622 | 0.051 | 0.774 | 0.084 | 1.433 | 0.145 | 15.56 | 11.899 | 0.506 | 6.13 | |||
This work carries embedding uncertainty through the speaker verification back-end, from similarity scoring to normalization and calibration. Each utterance is represented by a speaker embedding and its covariance estimate.
The unified pipeline includes three components:
The VoxCeleb recipe includes uncertainty extraction, scoring, normalization, and calibration. The corresponding scripts are listed in pipeline order below.
Uncertainty extraction
Uncertainty-aware cosine scoring
Uncertainty-aware AS-Norm (UAS-Norm)
Uncertainty-aware calibration (UQMF)
The table reports EER (%) and minDCF on Vox1-O, Vox1-E, and Vox1-H. Lower values are better. Bold and underlined values denote the best and second-best results, respectively, within each architecture group. RI is the average relative improvement over the corresponding architecture-specific benchmark across the six VoxCeleb metrics. The suffixes -o and -u denote the original and uncertainty-aware variants, respectively; ①–③ identify the UAS-Norm components evaluated in the paper.
| Row | Model | # Param. | Cosine score | AS-Norm | QMFs | Vox1-O | Vox1-E | Vox1-H | RI (%) | |||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| EER | minDCF | EER | minDCF | EER | minDCF | |||||||
| 1 | ECAPA† | 6.19 M | scos-o | 1.069 | 0.122 | 1.209 | 0.136 | 2.310 | 0.226 | Benchmark | ||
| 2 | ECAPA+xi | 6.69 M | scos-o | 0.936 | 0.100 | 1.054 | 0.123 | 1.978 | 0.195 | 13.49 | ||
| 3 | scos-o | sAS-o | 0.840 | 0.111 | 0.979 | 0.115 | 1.809 | 0.173 | 18.34 | |||
| 4 | scos-o | sAS-o | sQMF-o | 0.766 | 0.106 | 0.932 | 0.108 | 1.693 | 0.167 | 22.96 | ||
| 5 | scos-u | 0.840 | 0.086 | 0.965 | 0.110 | 1.833 | 0.189 | 21.21 | ||||
| 6 | scos-u | sAS-u (①) | 0.761 | 0.101 | 0.913 | 0.104 | 1.677 | 0.167 | 24.59 | |||
| 7 | scos-u | sAS-u (①+②) | 0.761 | 0.100 | 0.912 | 0.104 | 1.675 | 0.166 | 24.83 | |||
| 8 | scos-u | sAS-u (①+②+③) | 0.750 | 0.098 | 0.906 | 0.100 | 1.663 | 0.167 | 25.86 | |||
| 9 | scos-u | sAS-u | sQMF-u | 0.750 | 0.095 | 0.892 | 0.090 | 1.629 | 0.166 | 28.01 | ||
| 10 | ResNet† | 6.63 M | scos-o | 0.867 | 0.091 | 1.049 | 0.121 | 1.960 | 0.192 | Benchmark | ||
| 11 | ResNet+xi | 7.92 M | scos-o | 0.904 | 0.070 | 0.933 | 0.098 | 1.658 | 0.165 | 13.06 | ||
| 12 | scos-o | sAS-o | 0.888 | 0.079 | 0.922 | 0.096 | 1.618 | 0.163 | 12.68 | |||
| 13 | scos-o | sAS-o | sQMF-o | 0.782 | 0.065 | 0.842 | 0.090 | 1.489 | 0.151 | 21.52 | ||
| 14 | scos-u | 0.813 | 0.075 | 0.847 | 0.091 | 1.532 | 0.167 | 17.12 | ||||
| 15 | scos-u | sAS-u | 0.771 | 0.052 | 0.805 | 0.086 | 1.400 | 0.145 | 26.53 | |||
| 16 | scos-u | sAS-u | sQMF-u | 0.745 | 0.053 | 0.793 | 0.086 | 1.386 | 0.145 | 27.15 | ||
† Results obtained directly from pretrained WeSpeaker models.
If you use this repository in your research, please cite the papers relevant to the methods you use. BibTeX entries are provided below; select an IEEE bibliography style in your manuscript if required.
@inproceedings{li2026xiplus,
author = {Junjie Li and Kong Aik Lee and Duc-Tuan Truong and Tianchi Liu and Man-Wai Mak},
title = {{XI+}: Uncertainty Supervision for Robust Speaker Embedding},
booktitle = {2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
year = {2026},
pages = {18847--18851},
doi = {10.1109/ICASSP55912.2026.11463369}
}
@article{li2026u3xi,
author = {Junjie Li and Kong Aik Lee},
title = {{U3}-xi: Pushing the Boundaries of Speaker Recognition by Incorporating Uncertainty},
journal = {arXiv preprint arXiv:2601.15719},
year = {2026},
doi = {10.48550/arXiv.2601.15719},
url = {https://arxiv.org/abs/2601.15719}
}
@article{li2026robust,
author = {Junjie Li and Yang Xiao and Kong Aik Lee},
title = {Towards Robust Uncertainty-Aware Speaker Modeling},
journal = {arXiv preprint arXiv:2607.04937},
year = {2026},
doi = {10.48550/arXiv.2607.04937},
url = {https://arxiv.org/abs/2607.04937}
}
@article{li2026unified,
author = {Junjie Li and Kong Aik Lee},
title = {A Unified Uncertainty-Aware Back-End for Speaker Verification: Scoring, Normalization, and Calibration},
journal = {arXiv preprint arXiv:2609.01221},
year = {2026},
doi = {10.48550/arXiv.2609.01221},
url = {https://arxiv.org/abs/2609.01221}
}
If you use the RecXi implementation, please also cite the original RecXi paper:
@inproceedings{liu2023disentangling,
author = {Tianchi Liu and Kong Aik Lee and Qiongqiong Wang and Haizhou Li},
title = {Disentangling Voice and Content with Self-Supervision for Speaker Recognition},
booktitle = {Advances in Neural Information Processing Systems},
volume = {36},
year = {2023},
url = {https://proceedings.neurips.cc/paper_files/paper/2023/hash/9d276b0a087efdd2404f3295b26c24c1-Abstract-Conference.html}
}
Please also acknowledge the underlying toolkit using the WeSpeaker citations.
Python
86.4%
C++
7.9%
Shell
4.1%
CMake
1.0%