Bilingual Kazakh–English target-speaker ASR for overlapping speech. Datasets and checkpoints on Hugging Face (issai).
Python
1
2 commits
updated Sep 14, 2026
Bilingual Kazakh–English target-speaker speech recognition for overlapping speech.
Code for the paper:
R. Meiramov, T. Rakhimzhanova, A. Taibassarov, Z. Makhataeva, and H. A. Varol. Persona-ASR: Bilingual Target-Speaker Speech Recognition for Kazakh–English Overlapping Speech. Machine Learning and Knowledge Extraction, 8(8), 246, 2026. doi:10.3390/make8080246
Persona-ASR transcribes an enrolled speaker from a multi-talker mixture and emits <no_target> instead of a transcript when that speaker is absent. A target-presence gate first decides whether the enrolled speaker is in the mixture; an enrollment-conditioned recognizer then transcribes them, with a frozen ECAPA-TDNN speaker embedding modulating a WavLM-Base-Plus encoder through FiLM and language-specific CTC heads for Kazakh and English.
| Resource | Where |
|---|---|
| ASR backbone and presence gate checkpoints, Libri3Mix manifests | huggingface.co/issai/Persona-ASR |
| KazMix-3: Kazakh three-speaker manifests and clip map | huggingface.co/datasets/issai/KazMix-3 |
| PersonaMix: controlled bilingual benchmark | huggingface.co/datasets/issai/PersonaMix |
hf download issai/Persona-ASR --local-dir checkpoints
Three-speaker test sets (Table 3). Raw and gated WER are computed on target-present samples; balanced accuracy (BAcc) and F1 on target-present and target-absent samples.
| System | Test set | Raw WER (%) ↓ | Gated WER (%) ↓ | BAcc (%) ↑ | F1 (%) ↑ |
|---|---|---|---|---|---|
| Cascade (SepFormer → ECAPA-TDNN → OmniASR-7B) | English | 42.71 | – | – | – |
| Kazakh | 68.09 | – | – | – | |
| Persona-ASR | English (Libri3Mix-100h) | 29.41 | 35.45 | 81.33 | 86.92 |
| Kazakh (KazMix3-100h) | 43.49 | 50.77 | 86.28 | 88.78 | |
| Overall | 36.06 | 42.68 | 83.81 | 87.83 |
python3.10 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt --extra-index-url https://download.pytorch.org/whl/cu128
persona_asr/ model, training and evaluation
baselines/ separate-then-recognize cascade baseline
data_generation/ KazMix-3, Libri3Mix manifests, PersonaMix, manifest path remapping
scripts/ training launchers and the evaluation commands behind each table
The files in persona_asr/ keep the names of the experiment scripts that produced the paper results, so their imports are unchanged:
| File | Role | Paper |
|---|---|---|
maeda_twohead_ctc_baseline.py | Trains the ASR backbone: joint CTC and target-speaker VAD, following the TS-ASR-AD recipe of Maeda et al. (Interspeech 2025) | §4.1, §4.3 |
maeda_twohead_ctc.py | Backbone model and dataset used for gate training and evaluation | |
finetune_gate_attentive_stats_v7.py | Trains the target-presence gate; --paths A is the published gate, both and B are the Table 6 variants | §4.2, §5.4 |
maeda_tsasr_ad_librimix_multilang_presence.py | Shared data loading, VAD labels and WER | |
diagnose_vad_frame_predictions.py | Presence detection from the backbone VAD head | Table 2 |
eval_option7_gate_on_test.py | Test-set evaluation; --sweep-thresholds gives the threshold analysis | Tables 3, 5 |
diagnose_gate_paths.py, probe_pathB.py | Analysis of the two evidence paths in the gate | §5.4, Table 6 |
eval_personamix_target_detection.py | PersonaMix target-presence detection | Table 7 |
eval_personamix_asr_v2.py | PersonaMix target-speaker ASR | Table 8 |
benchmark_compute.py | Parameters, latency, real-time factor, peak memory | Table 9 |
baselines/eval_cascade_librimix.py | Cascade baseline (requires pip install omnilingual-asr) | Table 3 |
The manifests store absolute paths from the machine they were built on. Point them at your storage first:
python data_generation/remap_manifest_paths.py path/to/*.json \
--from /workspace/LibriMix/storage_dir --to /data/LibriMix/storage_dir --out-dir manifests_local
Libri3Mix (English). Generate the official Libri3Mix (16 kHz, max, clean) with LibriMix, then use the manifests in manifests/libri3mix/ on issai/Persona-ASR. data_generation/libri3mix/build_manifests.py shows how they were built; the released manifests were produced by that script followed by a re-selection of every enrollment utterance, which is why the exact files are released rather than regenerated.
KazMix-3 (Kazakh). The manifests are on issai/KazMix-3; the audio is rebuilt from KSD (OpenSLR 140):
python data_generation/kazmix3/extract_ksd_clips.py --manifests manifests_local/*_lufsfix_posneg.json \
--mapping kazmix3_clip_to_openslr140.json --openslr-root /data/openslr140 --storage-root /data/LibriMix/storage_dir
python data_generation/kazmix3/rerender_lufs_uniform.py --use-stored-lufs \
--in-manifest manifests_local/test_clean_kazakh3mix_lufsfix_posneg.json \
--out-audio-root /data/LibriMix/storage_dir/Kazakh3Mix_lufsfix/wav16k/max/test --out-manifest rendered/test.json
With the original clean clips, the re-render reproduces the released mixtures to within one 16-bit sample value. KazMix-3 was built from a 44.1 kHz copy of KSD, while OpenSLR 140 distributes at least part of the corpus at 16 kHz, so clips rebuilt from OpenSLR are very close to the originals but not bit-identical. generate_kazakh3mix.py and fix_enrollment_leakage.py are the original generation steps, kept for transparency.
PersonaMix. Download issai/PersonaMix. data_generation/personamix/generate_personamix.py regenerates its mixtures and manifests from the source recordings; the docstring describes the mixing recipe.
scripts/reproduce_tables.sh lists the evaluation command behind each table. For example, Tables 3 and 5:
python persona_asr/eval_option7_gate_on_test.py \
--asr-checkpoint checkpoints/asr_backbone.pt --gate-checkpoint checkpoints/presence_gate.pt \
--data-root /data/LibriMix/storage_dir --en-root manifests_local --kk-root manifests_local \
--en-test manifests_local/test_clean_official_libri3mix_clean_max_16k_posneg.json \
--kk-test manifests_local/test_clean_kazakh3mix_lufsfix_posneg.json \
--batch-size 16 --use-checkpoint-thresholds --sweep-thresholds 0.1 0.2 0.3 0.4 0.7 0.9 --output results/test.json
--batch-size 16. WavLM's feature encoder normalises over the padded batch, so batch composition slightly changes the outputs; with batch size 16 this evaluation reproduces the paper's per-sample outputs exactly.--use-checkpoint-thresholds applies the per-language thresholds selected on the validation set during gate training and stored in presence_gate.pt (τ_EN = 0.497, τ_KK = 0.609).--seed for repeatable numbers.STORAGE=/data/LibriMix/storage_dir EN=manifests_local KK=manifests_local scripts/train_backbone.sh # 2 GPUs
STORAGE=/data/LibriMix/storage_dir EN=manifests_local KK=manifests_local scripts/train_presence_gate.sh # 1 GPU
@article{meiramov2026personaasr,
author = {Meiramov, Rakhat and Rakhimzhanova, Tomiris and Taibassarov, Adil and Makhataeva, Zhanat and Varol, Huseyin Atakan},
title = {Persona-ASR: Bilingual Target-Speaker Speech Recognition for Kazakh--English Overlapping Speech},
journal = {Machine Learning and Knowledge Extraction},
year = {2026},
volume = {8},
number = {8},
pages = {246},
doi = {10.3390/make8080246}
}
This research is funded by the Committee of Science of the Ministry of Science and Higher Education of the Republic of Kazakhstan (Grant No. BR24993001). We thank the supervisor of the research grant, Madina Mansurova, for her coordination and support, and Galammadin Askar for contributing to the collection of the PersonaMix dataset.
The work was carried out at the Institute of Smart Systems and Artificial Intelligence (ISSAI), Nazarbayev University, and the Department of AI & Big Data, Al-Farabi Kazakh National University.
Python
98.4%
Shell
1.6%
Bilingual Kazakh–English target-speaker ASR for overlapping speech. Datasets and checkpoints on Hugging Face (issai).
Python
1
2 commits
updated Sep 14, 2026
Bilingual Kazakh–English target-speaker speech recognition for overlapping speech.
Code for the paper:
R. Meiramov, T. Rakhimzhanova, A. Taibassarov, Z. Makhataeva, and H. A. Varol. Persona-ASR: Bilingual Target-Speaker Speech Recognition for Kazakh–English Overlapping Speech. Machine Learning and Knowledge Extraction, 8(8), 246, 2026. doi:10.3390/make8080246
Persona-ASR transcribes an enrolled speaker from a multi-talker mixture and emits <no_target> instead of a transcript when that speaker is absent. A target-presence gate first decides whether the enrolled speaker is in the mixture; an enrollment-conditioned recognizer then transcribes them, with a frozen ECAPA-TDNN speaker embedding modulating a WavLM-Base-Plus encoder through FiLM and language-specific CTC heads for Kazakh and English.
| Resource | Where |
|---|---|
| ASR backbone and presence gate checkpoints, Libri3Mix manifests | huggingface.co/issai/Persona-ASR |
| KazMix-3: Kazakh three-speaker manifests and clip map | huggingface.co/datasets/issai/KazMix-3 |
| PersonaMix: controlled bilingual benchmark | huggingface.co/datasets/issai/PersonaMix |
hf download issai/Persona-ASR --local-dir checkpoints
Three-speaker test sets (Table 3). Raw and gated WER are computed on target-present samples; balanced accuracy (BAcc) and F1 on target-present and target-absent samples.
| System | Test set | Raw WER (%) ↓ | Gated WER (%) ↓ | BAcc (%) ↑ | F1 (%) ↑ |
|---|---|---|---|---|---|
| Cascade (SepFormer → ECAPA-TDNN → OmniASR-7B) | English | 42.71 | – | – | – |
| Kazakh | 68.09 | – | – | – | |
| Persona-ASR | English (Libri3Mix-100h) | 29.41 | 35.45 | 81.33 | 86.92 |
| Kazakh (KazMix3-100h) | 43.49 | 50.77 | 86.28 | 88.78 | |
| Overall | 36.06 | 42.68 | 83.81 | 87.83 |
python3.10 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt --extra-index-url https://download.pytorch.org/whl/cu128
persona_asr/ model, training and evaluation
baselines/ separate-then-recognize cascade baseline
data_generation/ KazMix-3, Libri3Mix manifests, PersonaMix, manifest path remapping
scripts/ training launchers and the evaluation commands behind each table
The files in persona_asr/ keep the names of the experiment scripts that produced the paper results, so their imports are unchanged:
| File | Role | Paper |
|---|---|---|
maeda_twohead_ctc_baseline.py | Trains the ASR backbone: joint CTC and target-speaker VAD, following the TS-ASR-AD recipe of Maeda et al. (Interspeech 2025) | §4.1, §4.3 |
maeda_twohead_ctc.py | Backbone model and dataset used for gate training and evaluation | |
finetune_gate_attentive_stats_v7.py | Trains the target-presence gate; --paths A is the published gate, both and B are the Table 6 variants | §4.2, §5.4 |
maeda_tsasr_ad_librimix_multilang_presence.py | Shared data loading, VAD labels and WER | |
diagnose_vad_frame_predictions.py | Presence detection from the backbone VAD head | Table 2 |
eval_option7_gate_on_test.py | Test-set evaluation; --sweep-thresholds gives the threshold analysis | Tables 3, 5 |
diagnose_gate_paths.py, probe_pathB.py | Analysis of the two evidence paths in the gate | §5.4, Table 6 |
eval_personamix_target_detection.py | PersonaMix target-presence detection | Table 7 |
eval_personamix_asr_v2.py | PersonaMix target-speaker ASR | Table 8 |
benchmark_compute.py | Parameters, latency, real-time factor, peak memory | Table 9 |
baselines/eval_cascade_librimix.py | Cascade baseline (requires pip install omnilingual-asr) | Table 3 |
The manifests store absolute paths from the machine they were built on. Point them at your storage first:
python data_generation/remap_manifest_paths.py path/to/*.json \
--from /workspace/LibriMix/storage_dir --to /data/LibriMix/storage_dir --out-dir manifests_local
Libri3Mix (English). Generate the official Libri3Mix (16 kHz, max, clean) with LibriMix, then use the manifests in manifests/libri3mix/ on issai/Persona-ASR. data_generation/libri3mix/build_manifests.py shows how they were built; the released manifests were produced by that script followed by a re-selection of every enrollment utterance, which is why the exact files are released rather than regenerated.
KazMix-3 (Kazakh). The manifests are on issai/KazMix-3; the audio is rebuilt from KSD (OpenSLR 140):
python data_generation/kazmix3/extract_ksd_clips.py --manifests manifests_local/*_lufsfix_posneg.json \
--mapping kazmix3_clip_to_openslr140.json --openslr-root /data/openslr140 --storage-root /data/LibriMix/storage_dir
python data_generation/kazmix3/rerender_lufs_uniform.py --use-stored-lufs \
--in-manifest manifests_local/test_clean_kazakh3mix_lufsfix_posneg.json \
--out-audio-root /data/LibriMix/storage_dir/Kazakh3Mix_lufsfix/wav16k/max/test --out-manifest rendered/test.json
With the original clean clips, the re-render reproduces the released mixtures to within one 16-bit sample value. KazMix-3 was built from a 44.1 kHz copy of KSD, while OpenSLR 140 distributes at least part of the corpus at 16 kHz, so clips rebuilt from OpenSLR are very close to the originals but not bit-identical. generate_kazakh3mix.py and fix_enrollment_leakage.py are the original generation steps, kept for transparency.
PersonaMix. Download issai/PersonaMix. data_generation/personamix/generate_personamix.py regenerates its mixtures and manifests from the source recordings; the docstring describes the mixing recipe.
scripts/reproduce_tables.sh lists the evaluation command behind each table. For example, Tables 3 and 5:
python persona_asr/eval_option7_gate_on_test.py \
--asr-checkpoint checkpoints/asr_backbone.pt --gate-checkpoint checkpoints/presence_gate.pt \
--data-root /data/LibriMix/storage_dir --en-root manifests_local --kk-root manifests_local \
--en-test manifests_local/test_clean_official_libri3mix_clean_max_16k_posneg.json \
--kk-test manifests_local/test_clean_kazakh3mix_lufsfix_posneg.json \
--batch-size 16 --use-checkpoint-thresholds --sweep-thresholds 0.1 0.2 0.3 0.4 0.7 0.9 --output results/test.json
--batch-size 16. WavLM's feature encoder normalises over the padded batch, so batch composition slightly changes the outputs; with batch size 16 this evaluation reproduces the paper's per-sample outputs exactly.--use-checkpoint-thresholds applies the per-language thresholds selected on the validation set during gate training and stored in presence_gate.pt (τ_EN = 0.497, τ_KK = 0.609).--seed for repeatable numbers.STORAGE=/data/LibriMix/storage_dir EN=manifests_local KK=manifests_local scripts/train_backbone.sh # 2 GPUs
STORAGE=/data/LibriMix/storage_dir EN=manifests_local KK=manifests_local scripts/train_presence_gate.sh # 1 GPU
@article{meiramov2026personaasr,
author = {Meiramov, Rakhat and Rakhimzhanova, Tomiris and Taibassarov, Adil and Makhataeva, Zhanat and Varol, Huseyin Atakan},
title = {Persona-ASR: Bilingual Target-Speaker Speech Recognition for Kazakh--English Overlapping Speech},
journal = {Machine Learning and Knowledge Extraction},
year = {2026},
volume = {8},
number = {8},
pages = {246},
doi = {10.3390/make8080246}
}
This research is funded by the Committee of Science of the Ministry of Science and Higher Education of the Republic of Kazakhstan (Grant No. BR24993001). We thank the supervisor of the research grant, Madina Mansurova, for her coordination and support, and Galammadin Askar for contributing to the collection of the PersonaMix dataset.
The work was carried out at the Institute of Smart Systems and Artificial Intelligence (ISSAI), Nazarbayev University, and the Department of AI & Big Data, Al-Farabi Kazakh National University.
Python
98.4%
Shell
1.6%