mrjunjieli/wespeaker_u_cube

Wespeaker implementations for speaker recognition and verification: U3-xi,Uncertainty aware AAM Softmax, Score normalization, calibration, and unofficial RecXi (NeurIPS 2023) source code.

Python

26

349 commits

updated Sep 19, 2026

See the code

README

WeSpeaker U³-xi

💡 WeSpeaker implementations for uncertainty-aware speaker recognition and verification, covering uncertainty estimation, speaker embedding learning, scoring, score normalization, and calibration.

Keywords: RecXi, NeurIPS 2023, speaker recognition, speaker verification, speaker embedding, WeSpeaker implementation, source code, ECAPA-TDNN, uncertainty estimation, uncertainty-aware speaker modeling.

Core idea · Getting started · U³-xi · Robust speaker modeling · Unified back-end · RecXi · Citation

📰 News

  • [Sep 2026] 🎉 Our paper "$\mathcal{U}^3$-xi: Pushing the Boundaries of Speaker Recognition by Incorporating Uncertainty" has been accepted by IEEE Transactions on Audio, Speech, and Language Processing (TASLP)!

Core Idea

Uncertainty-aware speaker embedding: less reliable speech has higher estimated uncertainty, and precision-weighted pooling reduces its influence on the speaker representation.

Editable draw.io source · Vector SVG

The goal is to estimate higher uncertainty when speech provides less reliable speaker evidence, such as under noise or reverberation. Frame-level precision estimates guide Bayesian pooling so reliable frames contribute more to the utterance representation. The model outputs both a speaker embedding and its associated uncertainty.

Conceptual illustration; the waveforms and distributions are schematic, not experimental measurements.

📚 Papers and Implementations

Getting Started

This project builds on WeSpeaker. Refer to the original WeSpeaker README for environment setup and the VoxCeleb recipe for data preparation. Use the configurations and scripts in this repository for the uncertainty-aware experiments described below.

Run recipe scripts from examples/voxceleb/v2 after configuring the dataset paths, GPU IDs, and experiment directory for your environment.

1. $\mathcal{U}^3$-xi: Pushing the Boundaries of Speaker Recognition by Incorporating Uncertainty

Junjie Li, Kong Aik Lee
The Hong Kong Polytechnic University

arXiv:2601.15719 Accepted by IEEE TASLP Models on Hugging Face Repository visitors GitHub stars

Email: junjie98.li@connect.polyu.hk

Key Highlights

U³-xi incorporates uncertainty into speaker embedding learning and verification. The implementation provides:

  • Multi-view uncertainty estimation: a multi-view self-attention (MVA) module estimates the uncertainty of frame-level representations.
  • Global uncertainty supervision: predicted uncertainty adjusts the softmax scale during training.
  • Uncertainty-aware cosine scoring: embedding uncertainty is incorporated into verification scores.

Experiments

Training and evaluation recipes are available in examples/voxceleb/v2. The main implementation entry points are listed below.

Results

The table reports equal error rate (EER, %) on the cleaned VoxCeleb1-O, VoxCeleb1-E, and VoxCeleb1-H trial lists. Lower values are better. LM denotes large-margin fine-tuning, AS-Norm denotes adaptive score normalization, and QMF denotes quality measure function calibration. ✓ and × indicate whether a component is enabled; rows with a blank model name use the model listed above; blank FLOPs values for named models were not reported.

ModelParametersFLOPsLMAS-NormQMFVox1-O-cleanVox1-E-cleanVox1-H-clean
ECAPA_TDNN_GLOB_c512-ASTP-emb1926.19M1.04G×××1.0691.2092.310
×✓×0.9571.1282.105
✓××0.8781.0722.007
✓✓×0.7821.0051.824
ECAPA_TDNN_GLOB_c1024-ASTP-emb19214.65M2.65G×××0.8561.0722.059
×✓×0.8080.9901.874
✓××0.7980.9931.883
✓✓×0.7280.9291.721
✓✓✓0.7070.8941.615
ResNet34-TSTP-emb2566.63M4.55G×××0.8671.0491.959
×✓×0.7870.9641.726
×✓✓0.7180.9111.606
✓××0.7970.9371.695
✓✓×0.7230.8671.532
✓✓✓0.6590.8211.437
XI_VEC_ECAPA_TDNN_c5125.9M1.04G×××0.9951.1302.169
×✓×0.8831.0561.976
✓××0.9091.0001.855
✓✓×0.7870.9301.693
ECAPA_TDNN_c512_u_cube_xi(ours)6.7M1.20G×××0.7821.0161.888
ResNet34_u_cube_xi(ours)7.9M×××0.8670.8681.641
ReDimNet-B2(ours)5.5M×××0.6060.7791.494
✓××0.4890.6981.311
✓✓×0.3990.6381.170

Pretrained Models

These checkpoints were retrained, so their results may differ slightly from those reported in the paper.

2. Towards Robust Uncertainty-Aware Speaker Modeling

arXiv:2607.04937

This work addresses uncertainty estimation and calibration under domain shifts through two methods:

  • Inter- and Intra-Speaker-Aware Uncertainty Softmax: incorporates inter-speaker separability and intra-speaker variability into uncertainty learning.
  • Uncertainty-Calibrated Domain Adaptation (UCDA): addresses uncertainty miscalibration caused by domain mismatch.

Implementation scope: This release includes Inter- and Intra-Speaker-Aware Uncertainty Softmax only. UCDA is described in the paper but is not included in this repository.

Implementations

Training note: train.py supports alpha scheduling for USphereFace2 inter-intra.

Results

The table reports EER (%) and minimum detection cost function (minDCF) values for in-domain evaluation on VoxCeleb1 and cross-domain evaluation on CNCeleb. Lower values are better. RI denotes relative improvement over the corresponding benchmark, and links in the loss column point to model checkpoints.

Model# Param.LossUncertainty-aware cosine scoreIn-domainCross-domain
Vox1-OVox1-EVox1-HRI (%)CNCelebRI (%)
EERminDCFEERminDCFEERminDCFEERminDCF
ECAPA5126.19 MAAM-SoftmaxNo1.0690.1221.2090.1362.3100.226Benchmark15.3140.633Benchmark
6.69 MUAAM-SoftmaxNo0.8560.1091.0640.1211.9820.19513.5713.7060.6087.23
Yes0.7820.1001.0160.1151.8880.18718.6410.2711.000-12.52
6.69UAAM-Softmax inter-intraNo0.9360.1021.0500.1221.9780.19513.4013.9740.5818.48
Yes0.8400.0860.9650.1101.8330.18921.2210.7810.835-1.16
6.19 MAM-SoftmaxNo1.0050.1071.2060.1332.2540.221Benchmark14.1620.611Benchmark
6.69 MUAM-Softmax inter-intraNo0.8880.0991.0760.1191.9730.18611.4612.4360.55310.84
Yes0.8080.0840.9910.1091.7940.17819.469.4111.000-15.03
6.19 MSphereFace2No0.9630.1081.1210.1251.9670.199Benchmark12.5820.573Benchmark
6.69 MUSphereFace2 inter-intraNo0.8560.1041.0350.1191.9180.1965.2112.2650.5503.27
Yes0.7390.1020.9650.1081.7710.17812.8110.5600.6243.59
ResNet346.63 MAAM-SoftmaxNo0.8670.0911.0490.1211.9600.192Benchmark11.0900.488Benchmark
7.92 MUAAM-SoftmaxNo0.8880.0850.9000.0991.7120.1759.6811.7320.513-5.46
Yes0.8670.0780.8680.0951.6410.17213.2910.0820.541-0.89
7.92 MUAAM-Softmax inter-intraNo0.9040.0700.9330.0981.6580.16513.0612.1160.505-6.37
Yes0.8130.0750.8470.0911.5320.16717.129.6310.5391.35
7.92 MUSphereFace2No1.4830.1481.4510.1562.1120.206-36.0011.4410.512-4.04
Yes1.3400.1561.3570.1501.9860.193-30.1910.9490.499-0.49
ReDimNet-B24.89 MAAM-SoftmaxNo0.7820.0640.9070.0971.6670.162Benchmark12.3850.552Benchmark
5.46 MUAAM-SoftmaxNo0.6490.0730.8010.0891.5320.1536.0913.4640.552-4.36
Yes0.6060.0650.7790.0911.4940.1579.129.4791.000-28.85
5.46 MUAAM-Softmax inter-intraNo0.6860.0700.8020.0901.5360.1516.0612.1320.5164.28
Yes0.6270.0640.7580.0881.4340.15310.848.6070.838-10.65
5.46 MUSphereFace2 inter-intraNo0.6220.0520.7760.0851.4400.14614.9212.0810.5154.58
Yes0.6220.0510.7740.0841.4330.14515.5611.8990.5066.13

3. A Unified Uncertainty-Aware Back-End for Speaker Verification: Scoring, Normalization, and Calibration

arXiv:2609.01221

This work carries embedding uncertainty through the speaker verification back-end, from similarity scoring to normalization and calibration. Each utterance is represented by a speaker embedding and its covariance estimate.

The unified pipeline includes three components:

  • Uncertainty-aware cosine scoring: adjusts similarity scoring using embedding uncertainty.
  • Uncertainty-aware AS-Norm (UAS-Norm): incorporates uncertainty into score normalization.
  • Uncertainty-aware quality measure function calibration (UQMF): uses uncertainty-related features for score calibration.

Experiments

The VoxCeleb recipe includes uncertainty extraction, scoring, normalization, and calibration. The corresponding scripts are listed in pipeline order below.

Results

The table reports EER (%) and minDCF on Vox1-O, Vox1-E, and Vox1-H. Lower values are better. Bold and underlined values denote the best and second-best results, respectively, within each architecture group. RI is the average relative improvement over the corresponding architecture-specific benchmark across the six VoxCeleb metrics. The suffixes -o and -u denote the original and uncertainty-aware variants, respectively; ①–③ identify the UAS-Norm components evaluated in the paper.

RowModel# Param.Cosine scoreAS-NormQMFsVox1-OVox1-EVox1-HRI (%)
EERminDCFEERminDCFEERminDCF
1ECAPA†6.19 Mscos-o1.0690.1221.2090.1362.3100.226Benchmark
2ECAPA+xi6.69 Mscos-o0.9360.1001.0540.1231.9780.19513.49
3scos-osAS-o0.8400.1110.9790.1151.8090.17318.34
4scos-osAS-osQMF-o0.7660.1060.9320.1081.6930.16722.96
5scos-u0.8400.0860.9650.1101.8330.18921.21
6scos-usAS-u (①)0.7610.1010.9130.1041.6770.16724.59
7scos-usAS-u (①+②)0.7610.1000.9120.1041.6750.16624.83
8scos-usAS-u (①+②+③)0.7500.0980.9060.1001.6630.16725.86
9scos-usAS-usQMF-u0.7500.0950.8920.0901.6290.16628.01
10ResNet†6.63 Mscos-o0.8670.0911.0490.1211.9600.192Benchmark
11ResNet+xi7.92 Mscos-o0.9040.0700.9330.0981.6580.16513.06
12scos-osAS-o0.8880.0790.9220.0961.6180.16312.68
13scos-osAS-osQMF-o0.7820.0650.8420.0901.4890.15121.52
14scos-u0.8130.0750.8470.0911.5320.16717.12
15scos-usAS-u0.7710.0520.8050.0861.4000.14526.53
16scos-usAS-usQMF-u0.7450.0530.7930.0861.3860.14527.15

† Results obtained directly from pretrained WeSpeaker models.

Citation

If you use this repository in your research, please cite the papers relevant to the methods you use. BibTeX entries are provided below; select an IEEE bibliography style in your manuscript if required.

@inproceedings{li2026xiplus,
  author    = {Junjie Li and Kong Aik Lee and Duc-Tuan Truong and Tianchi Liu and Man-Wai Mak},
  title     = {{XI+}: Uncertainty Supervision for Robust Speaker Embedding},
  booktitle = {2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
  year      = {2026},
  pages     = {18847--18851},
  doi       = {10.1109/ICASSP55912.2026.11463369}
}

@article{li2026u3xi,
  author  = {Junjie Li and Kong Aik Lee},
  title   = {{U3}-xi: Pushing the Boundaries of Speaker Recognition by Incorporating Uncertainty},
  journal = {arXiv preprint arXiv:2601.15719},
  year    = {2026},
  doi     = {10.48550/arXiv.2601.15719},
  url     = {https://arxiv.org/abs/2601.15719}
}

@article{li2026robust,
  author  = {Junjie Li and Yang Xiao and Kong Aik Lee},
  title   = {Towards Robust Uncertainty-Aware Speaker Modeling},
  journal = {arXiv preprint arXiv:2607.04937},
  year    = {2026},
  doi     = {10.48550/arXiv.2607.04937},
  url     = {https://arxiv.org/abs/2607.04937}
}

@article{li2026unified,
  author  = {Junjie Li and Kong Aik Lee},
  title   = {A Unified Uncertainty-Aware Back-End for Speaker Verification: Scoring, Normalization, and Calibration},
  journal = {arXiv preprint arXiv:2609.01221},
  year    = {2026},
  doi     = {10.48550/arXiv.2609.01221},
  url     = {https://arxiv.org/abs/2609.01221}
}

If you use the RecXi implementation, please also cite the original RecXi paper:

@inproceedings{liu2023disentangling,
  author    = {Tianchi Liu and Kong Aik Lee and Qiongqiong Wang and Haizhou Li},
  title     = {Disentangling Voice and Content with Self-Supervision for Speaker Recognition},
  booktitle = {Advances in Neural Information Processing Systems},
  volume    = {36},
  year      = {2023},
  url       = {https://proceedings.neurips.cc/paper_files/paper/2023/hash/9d276b0a087efdd2404f3295b26c24c1-Abstract-Conference.html}
}

Please also acknowledge the underlying toolkit using the WeSpeaker citations.

aam-softmax
disentangle
ecapa-tdnn
recxi
speaker-recognition
speaker-verification
u3-xi
uncertianty-aware
uncertianty-estimation
voxceleb
wespeaker

mrjunjieli/wespeaker_u_cube

Wespeaker implementations for speaker recognition and verification: U3-xi,Uncertainty aware AAM Softmax, Score normalization, calibration, and unofficial RecXi (NeurIPS 2023) source code.

Python

26

349 commits

updated Sep 19, 2026

See the code

README

WeSpeaker U³-xi

💡 WeSpeaker implementations for uncertainty-aware speaker recognition and verification, covering uncertainty estimation, speaker embedding learning, scoring, score normalization, and calibration.

Keywords: RecXi, NeurIPS 2023, speaker recognition, speaker verification, speaker embedding, WeSpeaker implementation, source code, ECAPA-TDNN, uncertainty estimation, uncertainty-aware speaker modeling.

Core idea · Getting started · U³-xi · Robust speaker modeling · Unified back-end · RecXi · Citation

📰 News

  • [Sep 2026] 🎉 Our paper "$\mathcal{U}^3$-xi: Pushing the Boundaries of Speaker Recognition by Incorporating Uncertainty" has been accepted by IEEE Transactions on Audio, Speech, and Language Processing (TASLP)!

Core Idea

Uncertainty-aware speaker embedding: less reliable speech has higher estimated uncertainty, and precision-weighted pooling reduces its influence on the speaker representation.

Editable draw.io source · Vector SVG

The goal is to estimate higher uncertainty when speech provides less reliable speaker evidence, such as under noise or reverberation. Frame-level precision estimates guide Bayesian pooling so reliable frames contribute more to the utterance representation. The model outputs both a speaker embedding and its associated uncertainty.

Conceptual illustration; the waveforms and distributions are schematic, not experimental measurements.

📚 Papers and Implementations

Getting Started

This project builds on WeSpeaker. Refer to the original WeSpeaker README for environment setup and the VoxCeleb recipe for data preparation. Use the configurations and scripts in this repository for the uncertainty-aware experiments described below.

Run recipe scripts from examples/voxceleb/v2 after configuring the dataset paths, GPU IDs, and experiment directory for your environment.

1. $\mathcal{U}^3$-xi: Pushing the Boundaries of Speaker Recognition by Incorporating Uncertainty

Junjie Li, Kong Aik Lee
The Hong Kong Polytechnic University

arXiv:2601.15719 Accepted by IEEE TASLP Models on Hugging Face Repository visitors GitHub stars

Email: junjie98.li@connect.polyu.hk

Key Highlights

U³-xi incorporates uncertainty into speaker embedding learning and verification. The implementation provides:

  • Multi-view uncertainty estimation: a multi-view self-attention (MVA) module estimates the uncertainty of frame-level representations.
  • Global uncertainty supervision: predicted uncertainty adjusts the softmax scale during training.
  • Uncertainty-aware cosine scoring: embedding uncertainty is incorporated into verification scores.

Experiments

Training and evaluation recipes are available in examples/voxceleb/v2. The main implementation entry points are listed below.

Results

The table reports equal error rate (EER, %) on the cleaned VoxCeleb1-O, VoxCeleb1-E, and VoxCeleb1-H trial lists. Lower values are better. LM denotes large-margin fine-tuning, AS-Norm denotes adaptive score normalization, and QMF denotes quality measure function calibration. ✓ and × indicate whether a component is enabled; rows with a blank model name use the model listed above; blank FLOPs values for named models were not reported.

ModelParametersFLOPsLMAS-NormQMFVox1-O-cleanVox1-E-cleanVox1-H-clean
ECAPA_TDNN_GLOB_c512-ASTP-emb1926.19M1.04G×××1.0691.2092.310
×✓×0.9571.1282.105
✓××0.8781.0722.007
✓✓×0.7821.0051.824
ECAPA_TDNN_GLOB_c1024-ASTP-emb19214.65M2.65G×××0.8561.0722.059
×✓×0.8080.9901.874
✓××0.7980.9931.883
✓✓×0.7280.9291.721
✓✓✓0.7070.8941.615
ResNet34-TSTP-emb2566.63M4.55G×××0.8671.0491.959
×✓×0.7870.9641.726
×✓✓0.7180.9111.606
✓××0.7970.9371.695
✓✓×0.7230.8671.532
✓✓✓0.6590.8211.437
XI_VEC_ECAPA_TDNN_c5125.9M1.04G×××0.9951.1302.169
×✓×0.8831.0561.976
✓××0.9091.0001.855
✓✓×0.7870.9301.693
ECAPA_TDNN_c512_u_cube_xi(ours)6.7M1.20G×××0.7821.0161.888
ResNet34_u_cube_xi(ours)7.9M×××0.8670.8681.641
ReDimNet-B2(ours)5.5M×××0.6060.7791.494
✓××0.4890.6981.311
✓✓×0.3990.6381.170

Pretrained Models

These checkpoints were retrained, so their results may differ slightly from those reported in the paper.

2. Towards Robust Uncertainty-Aware Speaker Modeling

arXiv:2607.04937

This work addresses uncertainty estimation and calibration under domain shifts through two methods:

  • Inter- and Intra-Speaker-Aware Uncertainty Softmax: incorporates inter-speaker separability and intra-speaker variability into uncertainty learning.
  • Uncertainty-Calibrated Domain Adaptation (UCDA): addresses uncertainty miscalibration caused by domain mismatch.

Implementation scope: This release includes Inter- and Intra-Speaker-Aware Uncertainty Softmax only. UCDA is described in the paper but is not included in this repository.

Implementations

Training note: train.py supports alpha scheduling for USphereFace2 inter-intra.

Results

The table reports EER (%) and minimum detection cost function (minDCF) values for in-domain evaluation on VoxCeleb1 and cross-domain evaluation on CNCeleb. Lower values are better. RI denotes relative improvement over the corresponding benchmark, and links in the loss column point to model checkpoints.

Model# Param.LossUncertainty-aware cosine scoreIn-domainCross-domain
Vox1-OVox1-EVox1-HRI (%)CNCelebRI (%)
EERminDCFEERminDCFEERminDCFEERminDCF
ECAPA5126.19 MAAM-SoftmaxNo1.0690.1221.2090.1362.3100.226Benchmark15.3140.633Benchmark
6.69 MUAAM-SoftmaxNo0.8560.1091.0640.1211.9820.19513.5713.7060.6087.23
Yes0.7820.1001.0160.1151.8880.18718.6410.2711.000-12.52
6.69UAAM-Softmax inter-intraNo0.9360.1021.0500.1221.9780.19513.4013.9740.5818.48
Yes0.8400.0860.9650.1101.8330.18921.2210.7810.835-1.16
6.19 MAM-SoftmaxNo1.0050.1071.2060.1332.2540.221Benchmark14.1620.611Benchmark
6.69 MUAM-Softmax inter-intraNo0.8880.0991.0760.1191.9730.18611.4612.4360.55310.84
Yes0.8080.0840.9910.1091.7940.17819.469.4111.000-15.03
6.19 MSphereFace2No0.9630.1081.1210.1251.9670.199Benchmark12.5820.573Benchmark
6.69 MUSphereFace2 inter-intraNo0.8560.1041.0350.1191.9180.1965.2112.2650.5503.27
Yes0.7390.1020.9650.1081.7710.17812.8110.5600.6243.59
ResNet346.63 MAAM-SoftmaxNo0.8670.0911.0490.1211.9600.192Benchmark11.0900.488Benchmark
7.92 MUAAM-SoftmaxNo0.8880.0850.9000.0991.7120.1759.6811.7320.513-5.46
Yes0.8670.0780.8680.0951.6410.17213.2910.0820.541-0.89
7.92 MUAAM-Softmax inter-intraNo0.9040.0700.9330.0981.6580.16513.0612.1160.505-6.37
Yes0.8130.0750.8470.0911.5320.16717.129.6310.5391.35
7.92 MUSphereFace2No1.4830.1481.4510.1562.1120.206-36.0011.4410.512-4.04
Yes1.3400.1561.3570.1501.9860.193-30.1910.9490.499-0.49
ReDimNet-B24.89 MAAM-SoftmaxNo0.7820.0640.9070.0971.6670.162Benchmark12.3850.552Benchmark
5.46 MUAAM-SoftmaxNo0.6490.0730.8010.0891.5320.1536.0913.4640.552-4.36
Yes0.6060.0650.7790.0911.4940.1579.129.4791.000-28.85
5.46 MUAAM-Softmax inter-intraNo0.6860.0700.8020.0901.5360.1516.0612.1320.5164.28
Yes0.6270.0640.7580.0881.4340.15310.848.6070.838-10.65
5.46 MUSphereFace2 inter-intraNo0.6220.0520.7760.0851.4400.14614.9212.0810.5154.58
Yes0.6220.0510.7740.0841.4330.14515.5611.8990.5066.13

3. A Unified Uncertainty-Aware Back-End for Speaker Verification: Scoring, Normalization, and Calibration

arXiv:2609.01221

This work carries embedding uncertainty through the speaker verification back-end, from similarity scoring to normalization and calibration. Each utterance is represented by a speaker embedding and its covariance estimate.

The unified pipeline includes three components:

  • Uncertainty-aware cosine scoring: adjusts similarity scoring using embedding uncertainty.
  • Uncertainty-aware AS-Norm (UAS-Norm): incorporates uncertainty into score normalization.
  • Uncertainty-aware quality measure function calibration (UQMF): uses uncertainty-related features for score calibration.

Experiments

The VoxCeleb recipe includes uncertainty extraction, scoring, normalization, and calibration. The corresponding scripts are listed in pipeline order below.

Results

The table reports EER (%) and minDCF on Vox1-O, Vox1-E, and Vox1-H. Lower values are better. Bold and underlined values denote the best and second-best results, respectively, within each architecture group. RI is the average relative improvement over the corresponding architecture-specific benchmark across the six VoxCeleb metrics. The suffixes -o and -u denote the original and uncertainty-aware variants, respectively; ①–③ identify the UAS-Norm components evaluated in the paper.

RowModel# Param.Cosine scoreAS-NormQMFsVox1-OVox1-EVox1-HRI (%)
EERminDCFEERminDCFEERminDCF
1ECAPA†6.19 Mscos-o1.0690.1221.2090.1362.3100.226Benchmark
2ECAPA+xi6.69 Mscos-o0.9360.1001.0540.1231.9780.19513.49
3scos-osAS-o0.8400.1110.9790.1151.8090.17318.34
4scos-osAS-osQMF-o0.7660.1060.9320.1081.6930.16722.96
5scos-u0.8400.0860.9650.1101.8330.18921.21
6scos-usAS-u (①)0.7610.1010.9130.1041.6770.16724.59
7scos-usAS-u (①+②)0.7610.1000.9120.1041.6750.16624.83
8scos-usAS-u (①+②+③)0.7500.0980.9060.1001.6630.16725.86
9scos-usAS-usQMF-u0.7500.0950.8920.0901.6290.16628.01
10ResNet†6.63 Mscos-o0.8670.0911.0490.1211.9600.192Benchmark
11ResNet+xi7.92 Mscos-o0.9040.0700.9330.0981.6580.16513.06
12scos-osAS-o0.8880.0790.9220.0961.6180.16312.68
13scos-osAS-osQMF-o0.7820.0650.8420.0901.4890.15121.52
14scos-u0.8130.0750.8470.0911.5320.16717.12
15scos-usAS-u0.7710.0520.8050.0861.4000.14526.53
16scos-usAS-usQMF-u0.7450.0530.7930.0861.3860.14527.15

† Results obtained directly from pretrained WeSpeaker models.

Citation

If you use this repository in your research, please cite the papers relevant to the methods you use. BibTeX entries are provided below; select an IEEE bibliography style in your manuscript if required.

@inproceedings{li2026xiplus,
  author    = {Junjie Li and Kong Aik Lee and Duc-Tuan Truong and Tianchi Liu and Man-Wai Mak},
  title     = {{XI+}: Uncertainty Supervision for Robust Speaker Embedding},
  booktitle = {2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
  year      = {2026},
  pages     = {18847--18851},
  doi       = {10.1109/ICASSP55912.2026.11463369}
}

@article{li2026u3xi,
  author  = {Junjie Li and Kong Aik Lee},
  title   = {{U3}-xi: Pushing the Boundaries of Speaker Recognition by Incorporating Uncertainty},
  journal = {arXiv preprint arXiv:2601.15719},
  year    = {2026},
  doi     = {10.48550/arXiv.2601.15719},
  url     = {https://arxiv.org/abs/2601.15719}
}

@article{li2026robust,
  author  = {Junjie Li and Yang Xiao and Kong Aik Lee},
  title   = {Towards Robust Uncertainty-Aware Speaker Modeling},
  journal = {arXiv preprint arXiv:2607.04937},
  year    = {2026},
  doi     = {10.48550/arXiv.2607.04937},
  url     = {https://arxiv.org/abs/2607.04937}
}

@article{li2026unified,
  author  = {Junjie Li and Kong Aik Lee},
  title   = {A Unified Uncertainty-Aware Back-End for Speaker Verification: Scoring, Normalization, and Calibration},
  journal = {arXiv preprint arXiv:2609.01221},
  year    = {2026},
  doi     = {10.48550/arXiv.2609.01221},
  url     = {https://arxiv.org/abs/2609.01221}
}

If you use the RecXi implementation, please also cite the original RecXi paper:

@inproceedings{liu2023disentangling,
  author    = {Tianchi Liu and Kong Aik Lee and Qiongqiong Wang and Haizhou Li},
  title     = {Disentangling Voice and Content with Self-Supervision for Speaker Recognition},
  booktitle = {Advances in Neural Information Processing Systems},
  volume    = {36},
  year      = {2023},
  url       = {https://proceedings.neurips.cc/paper_files/paper/2023/hash/9d276b0a087efdd2404f3295b26c24c1-Abstract-Conference.html}
}

Please also acknowledge the underlying toolkit using the WeSpeaker citations.

aam-softmax
disentangle
ecapa-tdnn
recxi
speaker-recognition
speaker-verification
u3-xi
uncertianty-aware
uncertianty-estimation
voxceleb
wespeaker

Languages

Python

86.4%

C++

7.9%

Shell

4.1%

CMake

1.0%