IS2AI/Persona_ASR

Bilingual Kazakh–English target-speaker ASR for overlapping speech. Datasets and checkpoints on Hugging Face (issai).

Python

1

2 commits

updated Sep 14, 2026

See the code

README

Persona-ASR

Bilingual Kazakh–English target-speaker speech recognition for overlapping speech.

Code for the paper:

R. Meiramov, T. Rakhimzhanova, A. Taibassarov, Z. Makhataeva, and H. A. Varol. Persona-ASR: Bilingual Target-Speaker Speech Recognition for Kazakh–English Overlapping Speech. Machine Learning and Knowledge Extraction, 8(8), 246, 2026. doi:10.3390/make8080246

Persona-ASR transcribes an enrolled speaker from a multi-talker mixture and emits <no_target> instead of a transcript when that speaker is absent. A target-presence gate first decides whether the enrolled speaker is in the mixture; an enrollment-conditioned recognizer then transcribes them, with a frozen ECAPA-TDNN speaker embedding modulating a WavLM-Base-Plus encoder through FiLM and language-specific CTC heads for Kazakh and English.

Checkpoints and data

ResourceWhere
ASR backbone and presence gate checkpoints, Libri3Mix manifestshuggingface.co/issai/Persona-ASR
KazMix-3: Kazakh three-speaker manifests and clip maphuggingface.co/datasets/issai/KazMix-3
PersonaMix: controlled bilingual benchmarkhuggingface.co/datasets/issai/PersonaMix
hf download issai/Persona-ASR --local-dir checkpoints

Results

Three-speaker test sets (Table 3). Raw and gated WER are computed on target-present samples; balanced accuracy (BAcc) and F1 on target-present and target-absent samples.

SystemTest setRaw WER (%) ↓Gated WER (%) ↓BAcc (%) ↑F1 (%) ↑
Cascade (SepFormer → ECAPA-TDNN → OmniASR-7B)English42.71–––
Kazakh68.09–––
Persona-ASREnglish (Libri3Mix-100h)29.4135.4581.3386.92
Kazakh (KazMix3-100h)43.4950.7786.2888.78
Overall36.0642.6883.8187.83

Setup

python3.10 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt --extra-index-url https://download.pytorch.org/whl/cu128

Repository layout

persona_asr/       model, training and evaluation
baselines/         separate-then-recognize cascade baseline
data_generation/   KazMix-3, Libri3Mix manifests, PersonaMix, manifest path remapping
scripts/           training launchers and the evaluation commands behind each table

The files in persona_asr/ keep the names of the experiment scripts that produced the paper results, so their imports are unchanged:

FileRolePaper
maeda_twohead_ctc_baseline.pyTrains the ASR backbone: joint CTC and target-speaker VAD, following the TS-ASR-AD recipe of Maeda et al. (Interspeech 2025)§4.1, §4.3
maeda_twohead_ctc.pyBackbone model and dataset used for gate training and evaluation
finetune_gate_attentive_stats_v7.pyTrains the target-presence gate; --paths A is the published gate, both and B are the Table 6 variants§4.2, §5.4
maeda_tsasr_ad_librimix_multilang_presence.pyShared data loading, VAD labels and WER
diagnose_vad_frame_predictions.pyPresence detection from the backbone VAD headTable 2
eval_option7_gate_on_test.pyTest-set evaluation; --sweep-thresholds gives the threshold analysisTables 3, 5
diagnose_gate_paths.py, probe_pathB.pyAnalysis of the two evidence paths in the gate§5.4, Table 6
eval_personamix_target_detection.pyPersonaMix target-presence detectionTable 7
eval_personamix_asr_v2.pyPersonaMix target-speaker ASRTable 8
benchmark_compute.pyParameters, latency, real-time factor, peak memoryTable 9
baselines/eval_cascade_librimix.pyCascade baseline (requires pip install omnilingual-asr)Table 3

Data

The manifests store absolute paths from the machine they were built on. Point them at your storage first:

python data_generation/remap_manifest_paths.py path/to/*.json \
  --from /workspace/LibriMix/storage_dir --to /data/LibriMix/storage_dir --out-dir manifests_local

Libri3Mix (English). Generate the official Libri3Mix (16 kHz, max, clean) with LibriMix, then use the manifests in manifests/libri3mix/ on issai/Persona-ASR. data_generation/libri3mix/build_manifests.py shows how they were built; the released manifests were produced by that script followed by a re-selection of every enrollment utterance, which is why the exact files are released rather than regenerated.

KazMix-3 (Kazakh). The manifests are on issai/KazMix-3; the audio is rebuilt from KSD (OpenSLR 140):

python data_generation/kazmix3/extract_ksd_clips.py --manifests manifests_local/*_lufsfix_posneg.json \
  --mapping kazmix3_clip_to_openslr140.json --openslr-root /data/openslr140 --storage-root /data/LibriMix/storage_dir
python data_generation/kazmix3/rerender_lufs_uniform.py --use-stored-lufs \
  --in-manifest manifests_local/test_clean_kazakh3mix_lufsfix_posneg.json \
  --out-audio-root /data/LibriMix/storage_dir/Kazakh3Mix_lufsfix/wav16k/max/test --out-manifest rendered/test.json

With the original clean clips, the re-render reproduces the released mixtures to within one 16-bit sample value. KazMix-3 was built from a 44.1 kHz copy of KSD, while OpenSLR 140 distributes at least part of the corpus at 16 kHz, so clips rebuilt from OpenSLR are very close to the originals but not bit-identical. generate_kazakh3mix.py and fix_enrollment_leakage.py are the original generation steps, kept for transparency.

PersonaMix. Download issai/PersonaMix. data_generation/personamix/generate_personamix.py regenerates its mixtures and manifests from the source recordings; the docstring describes the mixing recipe.

Reproducing the results

scripts/reproduce_tables.sh lists the evaluation command behind each table. For example, Tables 3 and 5:

python persona_asr/eval_option7_gate_on_test.py \
  --asr-checkpoint checkpoints/asr_backbone.pt --gate-checkpoint checkpoints/presence_gate.pt \
  --data-root /data/LibriMix/storage_dir --en-root manifests_local --kk-root manifests_local \
  --en-test manifests_local/test_clean_official_libri3mix_clean_max_16k_posneg.json \
  --kk-test manifests_local/test_clean_kazakh3mix_lufsfix_posneg.json \
  --batch-size 16 --use-checkpoint-thresholds --sweep-thresholds 0.1 0.2 0.3 0.4 0.7 0.9 --output results/test.json
  • Keep --batch-size 16. WavLM's feature encoder normalises over the padded batch, so batch composition slightly changes the outputs; with batch size 16 this evaluation reproduces the paper's per-sample outputs exactly.
  • --use-checkpoint-thresholds applies the per-language thresholds selected on the validation set during gate training and stored in presence_gate.pt (τ_EN = 0.497, τ_KK = 0.609).
  • Table 2 was computed on the validation split. The VAD evaluation samples enrollment utterances at random, so results vary by about one point between runs; pass --seed for repeatable numbers.

Training

STORAGE=/data/LibriMix/storage_dir EN=manifests_local KK=manifests_local scripts/train_backbone.sh        # 2 GPUs
STORAGE=/data/LibriMix/storage_dir EN=manifests_local KK=manifests_local scripts/train_presence_gate.sh   # 1 GPU

Citation

@article{meiramov2026personaasr,
  author  = {Meiramov, Rakhat and Rakhimzhanova, Tomiris and Taibassarov, Adil and Makhataeva, Zhanat and Varol, Huseyin Atakan},
  title   = {Persona-ASR: Bilingual Target-Speaker Speech Recognition for Kazakh--English Overlapping Speech},
  journal = {Machine Learning and Knowledge Extraction},
  year    = {2026},
  volume  = {8},
  number  = {8},
  pages   = {246},
  doi     = {10.3390/make8080246}
}

Funding and acknowledgments

This research is funded by the Committee of Science of the Ministry of Science and Higher Education of the Republic of Kazakhstan (Grant No. BR24993001). We thank the supervisor of the research grant, Madina Mansurova, for her coordination and support, and Galammadin Askar for contributing to the collection of the PersonaMix dataset.

The work was carried out at the Institute of Smart Systems and Artificial Intelligence (ISSAI), Nazarbayev University, and the Department of AI & Big Data, Al-Farabi Kazakh National University.

License

CC BY 4.0.

IS2AI/Persona_ASR

Bilingual Kazakh–English target-speaker ASR for overlapping speech. Datasets and checkpoints on Hugging Face (issai).

Python

1

2 commits

updated Sep 14, 2026

See the code

README

Persona-ASR

Bilingual Kazakh–English target-speaker speech recognition for overlapping speech.

Code for the paper:

R. Meiramov, T. Rakhimzhanova, A. Taibassarov, Z. Makhataeva, and H. A. Varol. Persona-ASR: Bilingual Target-Speaker Speech Recognition for Kazakh–English Overlapping Speech. Machine Learning and Knowledge Extraction, 8(8), 246, 2026. doi:10.3390/make8080246

Persona-ASR transcribes an enrolled speaker from a multi-talker mixture and emits <no_target> instead of a transcript when that speaker is absent. A target-presence gate first decides whether the enrolled speaker is in the mixture; an enrollment-conditioned recognizer then transcribes them, with a frozen ECAPA-TDNN speaker embedding modulating a WavLM-Base-Plus encoder through FiLM and language-specific CTC heads for Kazakh and English.

Checkpoints and data

ResourceWhere
ASR backbone and presence gate checkpoints, Libri3Mix manifestshuggingface.co/issai/Persona-ASR
KazMix-3: Kazakh three-speaker manifests and clip maphuggingface.co/datasets/issai/KazMix-3
PersonaMix: controlled bilingual benchmarkhuggingface.co/datasets/issai/PersonaMix
hf download issai/Persona-ASR --local-dir checkpoints

Results

Three-speaker test sets (Table 3). Raw and gated WER are computed on target-present samples; balanced accuracy (BAcc) and F1 on target-present and target-absent samples.

SystemTest setRaw WER (%) ↓Gated WER (%) ↓BAcc (%) ↑F1 (%) ↑
Cascade (SepFormer → ECAPA-TDNN → OmniASR-7B)English42.71–––
Kazakh68.09–––
Persona-ASREnglish (Libri3Mix-100h)29.4135.4581.3386.92
Kazakh (KazMix3-100h)43.4950.7786.2888.78
Overall36.0642.6883.8187.83

Setup

python3.10 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt --extra-index-url https://download.pytorch.org/whl/cu128

Repository layout

persona_asr/       model, training and evaluation
baselines/         separate-then-recognize cascade baseline
data_generation/   KazMix-3, Libri3Mix manifests, PersonaMix, manifest path remapping
scripts/           training launchers and the evaluation commands behind each table

The files in persona_asr/ keep the names of the experiment scripts that produced the paper results, so their imports are unchanged:

FileRolePaper
maeda_twohead_ctc_baseline.pyTrains the ASR backbone: joint CTC and target-speaker VAD, following the TS-ASR-AD recipe of Maeda et al. (Interspeech 2025)§4.1, §4.3
maeda_twohead_ctc.pyBackbone model and dataset used for gate training and evaluation
finetune_gate_attentive_stats_v7.pyTrains the target-presence gate; --paths A is the published gate, both and B are the Table 6 variants§4.2, §5.4
maeda_tsasr_ad_librimix_multilang_presence.pyShared data loading, VAD labels and WER
diagnose_vad_frame_predictions.pyPresence detection from the backbone VAD headTable 2
eval_option7_gate_on_test.pyTest-set evaluation; --sweep-thresholds gives the threshold analysisTables 3, 5
diagnose_gate_paths.py, probe_pathB.pyAnalysis of the two evidence paths in the gate§5.4, Table 6
eval_personamix_target_detection.pyPersonaMix target-presence detectionTable 7
eval_personamix_asr_v2.pyPersonaMix target-speaker ASRTable 8
benchmark_compute.pyParameters, latency, real-time factor, peak memoryTable 9
baselines/eval_cascade_librimix.pyCascade baseline (requires pip install omnilingual-asr)Table 3

Data

The manifests store absolute paths from the machine they were built on. Point them at your storage first:

python data_generation/remap_manifest_paths.py path/to/*.json \
  --from /workspace/LibriMix/storage_dir --to /data/LibriMix/storage_dir --out-dir manifests_local

Libri3Mix (English). Generate the official Libri3Mix (16 kHz, max, clean) with LibriMix, then use the manifests in manifests/libri3mix/ on issai/Persona-ASR. data_generation/libri3mix/build_manifests.py shows how they were built; the released manifests were produced by that script followed by a re-selection of every enrollment utterance, which is why the exact files are released rather than regenerated.

KazMix-3 (Kazakh). The manifests are on issai/KazMix-3; the audio is rebuilt from KSD (OpenSLR 140):

python data_generation/kazmix3/extract_ksd_clips.py --manifests manifests_local/*_lufsfix_posneg.json \
  --mapping kazmix3_clip_to_openslr140.json --openslr-root /data/openslr140 --storage-root /data/LibriMix/storage_dir
python data_generation/kazmix3/rerender_lufs_uniform.py --use-stored-lufs \
  --in-manifest manifests_local/test_clean_kazakh3mix_lufsfix_posneg.json \
  --out-audio-root /data/LibriMix/storage_dir/Kazakh3Mix_lufsfix/wav16k/max/test --out-manifest rendered/test.json

With the original clean clips, the re-render reproduces the released mixtures to within one 16-bit sample value. KazMix-3 was built from a 44.1 kHz copy of KSD, while OpenSLR 140 distributes at least part of the corpus at 16 kHz, so clips rebuilt from OpenSLR are very close to the originals but not bit-identical. generate_kazakh3mix.py and fix_enrollment_leakage.py are the original generation steps, kept for transparency.

PersonaMix. Download issai/PersonaMix. data_generation/personamix/generate_personamix.py regenerates its mixtures and manifests from the source recordings; the docstring describes the mixing recipe.

Reproducing the results

scripts/reproduce_tables.sh lists the evaluation command behind each table. For example, Tables 3 and 5:

python persona_asr/eval_option7_gate_on_test.py \
  --asr-checkpoint checkpoints/asr_backbone.pt --gate-checkpoint checkpoints/presence_gate.pt \
  --data-root /data/LibriMix/storage_dir --en-root manifests_local --kk-root manifests_local \
  --en-test manifests_local/test_clean_official_libri3mix_clean_max_16k_posneg.json \
  --kk-test manifests_local/test_clean_kazakh3mix_lufsfix_posneg.json \
  --batch-size 16 --use-checkpoint-thresholds --sweep-thresholds 0.1 0.2 0.3 0.4 0.7 0.9 --output results/test.json
  • Keep --batch-size 16. WavLM's feature encoder normalises over the padded batch, so batch composition slightly changes the outputs; with batch size 16 this evaluation reproduces the paper's per-sample outputs exactly.
  • --use-checkpoint-thresholds applies the per-language thresholds selected on the validation set during gate training and stored in presence_gate.pt (τ_EN = 0.497, τ_KK = 0.609).
  • Table 2 was computed on the validation split. The VAD evaluation samples enrollment utterances at random, so results vary by about one point between runs; pass --seed for repeatable numbers.

Training

STORAGE=/data/LibriMix/storage_dir EN=manifests_local KK=manifests_local scripts/train_backbone.sh        # 2 GPUs
STORAGE=/data/LibriMix/storage_dir EN=manifests_local KK=manifests_local scripts/train_presence_gate.sh   # 1 GPU

Citation

@article{meiramov2026personaasr,
  author  = {Meiramov, Rakhat and Rakhimzhanova, Tomiris and Taibassarov, Adil and Makhataeva, Zhanat and Varol, Huseyin Atakan},
  title   = {Persona-ASR: Bilingual Target-Speaker Speech Recognition for Kazakh--English Overlapping Speech},
  journal = {Machine Learning and Knowledge Extraction},
  year    = {2026},
  volume  = {8},
  number  = {8},
  pages   = {246},
  doi     = {10.3390/make8080246}
}

Funding and acknowledgments

This research is funded by the Committee of Science of the Ministry of Science and Higher Education of the Republic of Kazakhstan (Grant No. BR24993001). We thank the supervisor of the research grant, Madina Mansurova, for her coordination and support, and Galammadin Askar for contributing to the collection of the PersonaMix dataset.

The work was carried out at the Institute of Smart Systems and Artificial Intelligence (ISSAI), Nazarbayev University, and the Department of AI & Big Data, Al-Farabi Kazakh National University.

License

CC BY 4.0.

Languages

Python

98.4%

Shell

1.6%