Real-time multilingual ASR with code-switching orchestration. A lightweight, CPU-friendly coordinator routes between small monolingual Zipformer models (VAD + LID + rollback) instead of one large multilingual model.
git clone https://github.com/gladiaio/realtime_multilingual_asr_router.git
cd realtime_multilingual_asr_router
uv sync
# Download community Kroko ASR models (en, fr, es by default) plus VAD and LID
uv run scripts/download_models.py
Notes:
en, fr, es). Set HF_TOKEN for higher Hugging Face rate limits.GET /health, GET /languages (component status and discovered ASR languages).More languages and chunk size:
The list of available KrokoAI models is here. This solution is compatible with any online transducer model.
ASR languages: Kroko de, en, es, fr, it, iw (Hebrew), nl, pt, sv, tr (sv only supports chunk size 64), and ru (from AlphaCephei).
# All Kroko languages + Russian (~1.6 GB)
uv run scripts/download_models.py --languages de en es fr it iw nl pt sv tr ru
# Subset example
uv run scripts/download_models.py --languages en fr es de --chunk-size 64
Settings live in src/realtime_multilingual_asr_router/config.py and are loaded from:
LR_ prefix (e.g. LR_ASR_CHUNK_SIZE=64).env file in the project root (copy from .env.example)cp .env.example .env
ASR languages are discovered from models/streaming_transducers/<chunk_size>/<lang>/ (each folder needs encoder.onnx, decoder.onnx, joiner.onnx, and tokens.txt).
uv run uvicorn realtime_multilingual_asr_router.main:app
Open http://127.0.0.1:8000/ for the frontend. WebSocket: /ws/transcribe.
uv run pytest
uv run ruff check .
uv run ruff format .
Model-backed tests are marked slow and may be skipped without downloaded weights.
realtime_multilingual_asr_router/
├── src/realtime_multilingual_asr_router/ # Application package
├── frontend/ # Web UI
├── assets/ # Architecture diagram and benchmark charts
├── tests/
├── scripts/download_models.py
├── models/ # Downloaded weights (gitignored)
└── .env.example
| Model | Role | Source |
|---|---|---|
| Silero VAD | Voice activity | snakers4/silero-vad |
| Kroko ASR (community) | Streaming ASR (de, en, es, fr, it, iw, nl, pt, sv, tr) | Banafo/Kroko-ASR |
| Russian Zipformer | Streaming ASR (ru, pass --languages ru) | sherpa-onnx-streaming-zipformer-small-ru-vosk |
| VoxLingua107-ECAPA-TDNN | Language ID | speechbrain/lang-id-voxlingua107-ecapa |
scripts/download_models.py downloads only the languages you list. Kroko codes use .data extraction; ru pulls ONNX files from the Hub (no .data). Default: en, fr, es.
Stack: FastAPI, uvicorn, sherpa-onnx, silero-vad, speechbrain, torch, torchcodec, pytest, ruff.
| Component | Location |
|---|---|
| WebSocket server | src/realtime_multilingual_asr_router/api/websocket.py |
| Audio resampling (16 kHz) | src/realtime_multilingual_asr_router/audio/resampler.py |
| Sample-count master clock | src/realtime_multilingual_asr_router/coordinator/sample_clock.py |
| VAD (Silero v6) | src/realtime_multilingual_asr_router/vad/silero.py |
| ASR (Zipformer / sherpa-onnx) | src/realtime_multilingual_asr_router/asr/kroko.py |
| LID scheduler | src/realtime_multilingual_asr_router/lid/scheduler.py |
| Rollback engine | src/realtime_multilingual_asr_router/rollback/engine.py |
| Audio ring buffer | src/realtime_multilingual_asr_router/audio/buffer.py |
| REST API | src/realtime_multilingual_asr_router/api/rest.py |
| Frontend | frontend/ |
Wrong-language text may appear briefly in partials; finals have clean language boundaries.
Benchmarks on three datasets: monolingual baseline (FLEURS), inter-utterance code-switch (synthetic FLEURS blend), and intra-utterance code-switch (Bangor–Miami).
This work: 640 ms Zipformer chunks — Kroko for en/fr/es/de/it/pt/nl, Alphacephei ru.
Baselines: Deepgram Nova-3, ElevenLabs Scribe v2, Mistral Voxtral-Mini-4B-Realtime (480 ms), AssemblyAI u3-rt-pro (no nl/ru — excluded where those languages appear).
| Scenario | This work | Deepgram | ElevenLabs | Voxtral Mini |
|---|---|---|---|---|
| Inter-utterance (synthetic FLEURS) | ~13% | ~14% | ~14% | ~21% |
| Intra-utterance (Miami) | ~41% | ~29% | ~26% | ~76% |


| Dataset | Voxtral Mini | Deepgram Nova-3 | ElevenLabs Scribe v2 | AAI u3-rt-pro | This work |
|---|---|---|---|---|---|
| Italian (FLEURS) | 3.30 | 7.10 | 2.00 | 2.50 | 4.30 |
| Russian (FLEURS) | 5.20 | 10.20 | 7.40 | — | 14.30 |
| Portuguese (FLEURS) | 4.90 | 10.30 | 3.90 | 3.90 | 9.20 |
| English (FLEURS) | 12.40 | 10.20 | 4.30 | 3.40 | 13.20 |
| French (FLEURS) | 8.80 | 10.60 | 5.50 | 4.10 | 9.90 |
| Dutch (FLEURS) | 8.20 | 13.30 | 5.30 | — | 11.90 |
| German (FLEURS) | 6.00 | 10.00 | 3.90 | 3.40 | 11.00 |
| Spanish (FLEURS) | 3.10 | 5.90 | 2.70 | 2.40 | 4.70 |
| Code-switch (synthetic FLEURS) | 21.30 | 14.30 | 13.60 | — | 13.20 |
| Miami (intra-utterance) | 76.30 | 28.60 | 26.50 | 34.00 | 41.00 |
| Avg (FLEURS only) | 6.50 | 9.70 | 4.40 | 3.30 | 9.80 |
Bangor–Miami — Mozilla Data Collective. Deuchar et al. (2014), Advances in the Study of Bilingualism, Multilingual Matters.
FLEURS:
@article{fleurs2022arxiv,
title = {FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech},
author = {Conneau, Alexis and Ma, Min and Khanuja, Simran and Zhang, Yu and Axelrod, Vera and Dalmia, Siddharth and Riesa, Jason and Rivera, Clara and Bapna, Ankur},
journal = {arXiv preprint arXiv:2205.12446},
url = {https://arxiv.org/abs/2205.12446},
year = {2022}
}
MIT — see LICENSE.
Python
92.8%
JavaScript
6.3%
Real-time multilingual ASR with code-switching orchestration. A lightweight, CPU-friendly coordinator routes between small monolingual Zipformer models (VAD + LID + rollback) instead of one large multilingual model.
git clone https://github.com/gladiaio/realtime_multilingual_asr_router.git
cd realtime_multilingual_asr_router
uv sync
# Download community Kroko ASR models (en, fr, es by default) plus VAD and LID
uv run scripts/download_models.py
Notes:
en, fr, es). Set HF_TOKEN for higher Hugging Face rate limits.GET /health, GET /languages (component status and discovered ASR languages).More languages and chunk size:
The list of available KrokoAI models is here. This solution is compatible with any online transducer model.
ASR languages: Kroko de, en, es, fr, it, iw (Hebrew), nl, pt, sv, tr (sv only supports chunk size 64), and ru (from AlphaCephei).
# All Kroko languages + Russian (~1.6 GB)
uv run scripts/download_models.py --languages de en es fr it iw nl pt sv tr ru
# Subset example
uv run scripts/download_models.py --languages en fr es de --chunk-size 64
Settings live in src/realtime_multilingual_asr_router/config.py and are loaded from:
LR_ prefix (e.g. LR_ASR_CHUNK_SIZE=64).env file in the project root (copy from .env.example)cp .env.example .env
ASR languages are discovered from models/streaming_transducers/<chunk_size>/<lang>/ (each folder needs encoder.onnx, decoder.onnx, joiner.onnx, and tokens.txt).
uv run uvicorn realtime_multilingual_asr_router.main:app
Open http://127.0.0.1:8000/ for the frontend. WebSocket: /ws/transcribe.
uv run pytest
uv run ruff check .
uv run ruff format .
Model-backed tests are marked slow and may be skipped without downloaded weights.
realtime_multilingual_asr_router/
├── src/realtime_multilingual_asr_router/ # Application package
├── frontend/ # Web UI
├── assets/ # Architecture diagram and benchmark charts
├── tests/
├── scripts/download_models.py
├── models/ # Downloaded weights (gitignored)
└── .env.example
| Model | Role | Source |
|---|---|---|
| Silero VAD | Voice activity | snakers4/silero-vad |
| Kroko ASR (community) | Streaming ASR (de, en, es, fr, it, iw, nl, pt, sv, tr) | Banafo/Kroko-ASR |
| Russian Zipformer | Streaming ASR (ru, pass --languages ru) | sherpa-onnx-streaming-zipformer-small-ru-vosk |
| VoxLingua107-ECAPA-TDNN | Language ID | speechbrain/lang-id-voxlingua107-ecapa |
scripts/download_models.py downloads only the languages you list. Kroko codes use .data extraction; ru pulls ONNX files from the Hub (no .data). Default: en, fr, es.
Stack: FastAPI, uvicorn, sherpa-onnx, silero-vad, speechbrain, torch, torchcodec, pytest, ruff.
| Component | Location |
|---|---|
| WebSocket server | src/realtime_multilingual_asr_router/api/websocket.py |
| Audio resampling (16 kHz) | src/realtime_multilingual_asr_router/audio/resampler.py |
| Sample-count master clock | src/realtime_multilingual_asr_router/coordinator/sample_clock.py |
| VAD (Silero v6) | src/realtime_multilingual_asr_router/vad/silero.py |
| ASR (Zipformer / sherpa-onnx) | src/realtime_multilingual_asr_router/asr/kroko.py |
| LID scheduler | src/realtime_multilingual_asr_router/lid/scheduler.py |
| Rollback engine | src/realtime_multilingual_asr_router/rollback/engine.py |
| Audio ring buffer | src/realtime_multilingual_asr_router/audio/buffer.py |
| REST API | src/realtime_multilingual_asr_router/api/rest.py |
| Frontend | frontend/ |
Wrong-language text may appear briefly in partials; finals have clean language boundaries.
Benchmarks on three datasets: monolingual baseline (FLEURS), inter-utterance code-switch (synthetic FLEURS blend), and intra-utterance code-switch (Bangor–Miami).
This work: 640 ms Zipformer chunks — Kroko for en/fr/es/de/it/pt/nl, Alphacephei ru.
Baselines: Deepgram Nova-3, ElevenLabs Scribe v2, Mistral Voxtral-Mini-4B-Realtime (480 ms), AssemblyAI u3-rt-pro (no nl/ru — excluded where those languages appear).
| Scenario | This work | Deepgram | ElevenLabs | Voxtral Mini |
|---|---|---|---|---|
| Inter-utterance (synthetic FLEURS) | ~13% | ~14% | ~14% | ~21% |
| Intra-utterance (Miami) | ~41% | ~29% | ~26% | ~76% |


| Dataset | Voxtral Mini | Deepgram Nova-3 | ElevenLabs Scribe v2 | AAI u3-rt-pro | This work |
|---|---|---|---|---|---|
| Italian (FLEURS) | 3.30 | 7.10 | 2.00 | 2.50 | 4.30 |
| Russian (FLEURS) | 5.20 | 10.20 | 7.40 | — | 14.30 |
| Portuguese (FLEURS) | 4.90 | 10.30 | 3.90 | 3.90 | 9.20 |
| English (FLEURS) | 12.40 | 10.20 | 4.30 | 3.40 | 13.20 |
| French (FLEURS) | 8.80 | 10.60 | 5.50 | 4.10 | 9.90 |
| Dutch (FLEURS) | 8.20 | 13.30 | 5.30 | — | 11.90 |
| German (FLEURS) | 6.00 | 10.00 | 3.90 | 3.40 | 11.00 |
| Spanish (FLEURS) | 3.10 | 5.90 | 2.70 | 2.40 | 4.70 |
| Code-switch (synthetic FLEURS) | 21.30 | 14.30 | 13.60 | — | 13.20 |
| Miami (intra-utterance) | 76.30 | 28.60 | 26.50 | 34.00 | 41.00 |
| Avg (FLEURS only) | 6.50 | 9.70 | 4.40 | 3.30 | 9.80 |
Bangor–Miami — Mozilla Data Collective. Deuchar et al. (2014), Advances in the Study of Bilingualism, Multilingual Matters.
FLEURS:
@article{fleurs2022arxiv,
title = {FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech},
author = {Conneau, Alexis and Ma, Min and Khanuja, Simran and Zhang, Yu and Axelrod, Vera and Dalmia, Siddharth and Riesa, Jason and Rivera, Clara and Bapna, Ankur},
journal = {arXiv preprint arXiv:2205.12446},
url = {https://arxiv.org/abs/2205.12446},
year = {2022}
}
MIT — see LICENSE.
Python
92.8%
JavaScript
6.3%