This repository hosts DiCoW v3.3, a Target-Speaker ASR (TS-ASR) model developed by BUT Speech@FIT. It is designed to transcribe the speech of a specific speaker within a multi-talker mixture by conditioning on speaker diarization outputs.
This model version incorporates the refinements and training strategies described in the paper SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper.
This version represents a significant stabilization and enhancement over the original DiCoW (v1):
The easiest way to use this model is via the DiCoW inference repository. We provide a Gradio app that handles diarization and STNO mask generation automatically:
python app.py
If you want to download and load the model manually for your own scripts:
from transformers import AutoModelForSpeechSeq2Seq
# Load the model (requires remote code for custom FDDT layers)
model = AutoModelForSpeechSeq2Seq.from_pretrained(
"BUT-FIT/DiCoW_v3_3",
trust_remote_code=True
)
# Note: The model expects specific STNO conditioning inputs.
# See inference.py in the GitHub repo for the full pipeline.
It's all yours with just two commands! This model is fully open-source and reproducible using our toolkit.
1. Data Preparation Clone the mt-asr-data-prep repository and run the setup script to generate the required manifests:
./prepare.sh --single-mic-only --root-dir /path/to/workdir
2. Training
Clone the training repository TS-ASR-Whisper and launch the experiment using the pre-configured dicow_v3 recipe:
sbatch --export SRC_ROOT=$PWD scripts/submit_slurm.sh +train=dicow_v3
If you use this model, please cite the following papers:
@article{polok2026sedicow,
title={SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper},
author={Alexander Polok and Dominik Klement and Samuele Cornell and Matthew Wiesner and Jan Černocký and Sanjeev Khudanpur and Lukáš Burget},
journal={arXiv preprint arXiv:2601.19194},
year={2026}
}
@article{POLOK2026101841,
title = {DiCoW: Diarization-conditioned Whisper for target speaker automatic speech recognition},
journal = {Computer Speech & Language},
volume = {95},
year = {2026},
doi = {10.1016/j.csl.2025.101841},
author = {Alexander Polok et al.}
}
@INPROCEEDINGS{10887683,
title={Target Speaker ASR with Whisper},
author={Polok, Alexander et al.},
booktitle={ICASSP 2025},
year={2025},
doi={10.1109/ICASSP49660.2025.10887683}
}
7 commits
1 commits
This repository hosts DiCoW v3.3, a Target-Speaker ASR (TS-ASR) model developed by BUT Speech@FIT. It is designed to transcribe the speech of a specific speaker within a multi-talker mixture by conditioning on speaker diarization outputs.
This model version incorporates the refinements and training strategies described in the paper SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper.
This version represents a significant stabilization and enhancement over the original DiCoW (v1):
The easiest way to use this model is via the DiCoW inference repository. We provide a Gradio app that handles diarization and STNO mask generation automatically:
python app.py
If you want to download and load the model manually for your own scripts:
from transformers import AutoModelForSpeechSeq2Seq
# Load the model (requires remote code for custom FDDT layers)
model = AutoModelForSpeechSeq2Seq.from_pretrained(
"BUT-FIT/DiCoW_v3_3",
trust_remote_code=True
)
# Note: The model expects specific STNO conditioning inputs.
# See inference.py in the GitHub repo for the full pipeline.
It's all yours with just two commands! This model is fully open-source and reproducible using our toolkit.
1. Data Preparation Clone the mt-asr-data-prep repository and run the setup script to generate the required manifests:
./prepare.sh --single-mic-only --root-dir /path/to/workdir
2. Training
Clone the training repository TS-ASR-Whisper and launch the experiment using the pre-configured dicow_v3 recipe:
sbatch --export SRC_ROOT=$PWD scripts/submit_slurm.sh +train=dicow_v3
If you use this model, please cite the following papers:
@article{polok2026sedicow,
title={SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper},
author={Alexander Polok and Dominik Klement and Samuele Cornell and Matthew Wiesner and Jan Černocký and Sanjeev Khudanpur and Lukáš Burget},
journal={arXiv preprint arXiv:2601.19194},
year={2026}
}
@article{POLOK2026101841,
title = {DiCoW: Diarization-conditioned Whisper for target speaker automatic speech recognition},
journal = {Computer Speech & Language},
volume = {95},
year = {2026},
doi = {10.1016/j.csl.2025.101841},
author = {Alexander Polok et al.}
}
@INPROCEEDINGS{10887683,
title={Target Speaker ASR with Whisper},
author={Polok, Alexander et al.},
booktitle={ICASSP 2025},
year={2025},
doi={10.1109/ICASSP49660.2025.10887683}
}
7 commits
1 commits