gladiaio/realtime-multilingual-asr-router

Python

23

3 commits

updated May 29, 2026

See the code

README

Realtime Multilingual ASR Router

Real-time multilingual ASR with code-switching orchestration. A lightweight, CPU-friendly coordinator routes between small monolingual Zipformer models (VAD + LID + rollback) instead of one large multilingual model.

Installation

Prerequisites

  • Python 3.11+
  • uv

Setup

git clone https://github.com/gladiaio/realtime_multilingual_asr_router.git
cd realtime_multilingual_asr_router

uv sync

# Download community Kroko ASR models (en, fr, es by default) plus VAD and LID
uv run scripts/download_models.py

Notes:

  • Default Kroko download is ~450 MB (en, fr, es). Set HF_TOKEN for higher Hugging Face rate limits.
  • First server start preloads one ASR model per downloaded language (~2s per language).
  • REST: GET /health, GET /languages (component status and discovered ASR languages).

More languages and chunk size:

The list of available KrokoAI models is here. This solution is compatible with any online transducer model.

ASR languages: Kroko de, en, es, fr, it, iw (Hebrew), nl, pt, sv, tr (sv only supports chunk size 64), and ru (from AlphaCephei).

# All Kroko languages + Russian (~1.6 GB)
uv run scripts/download_models.py --languages de en es fr it iw nl pt sv tr ru

# Subset example
uv run scripts/download_models.py --languages en fr es de --chunk-size 64

Configuration (optional)

Settings live in src/realtime_multilingual_asr_router/config.py and are loaded from:

  1. Defaults in code
  2. Environment variables with the LR_ prefix (e.g. LR_ASR_CHUNK_SIZE=64)
  3. A .env file in the project root (copy from .env.example)
cp .env.example .env

ASR languages are discovered from models/streaming_transducers/<chunk_size>/<lang>/ (each folder needs encoder.onnx, decoder.onnx, joiner.onnx, and tokens.txt).

Run

uv run uvicorn realtime_multilingual_asr_router.main:app

Open http://127.0.0.1:8000/ for the frontend. WebSocket: /ws/transcribe.

Development

uv run pytest
uv run ruff check .
uv run ruff format .

Model-backed tests are marked slow and may be skipped without downloaded weights.

Project structure

realtime_multilingual_asr_router/
├── src/realtime_multilingual_asr_router/   # Application package
├── frontend/                   # Web UI
├── assets/                     # Architecture diagram and benchmark charts
├── tests/
├── scripts/download_models.py
├── models/                     # Downloaded weights (gitignored)
└── .env.example

Models

ModelRoleSource
Silero VADVoice activitysnakers4/silero-vad
Kroko ASR (community)Streaming ASR (de, en, es, fr, it, iw, nl, pt, sv, tr)Banafo/Kroko-ASR
Russian ZipformerStreaming ASR (ru, pass --languages ru)sherpa-onnx-streaming-zipformer-small-ru-vosk
VoxLingua107-ECAPA-TDNNLanguage IDspeechbrain/lang-id-voxlingua107-ecapa

scripts/download_models.py downloads only the languages you list. Kroko codes use .data extraction; ru pulls ONNX files from the Hub (no .data). Default: en, fr, es.

Stack: FastAPI, uvicorn, sherpa-onnx, silero-vad, speechbrain, torch, torchcodec, pytest, ruff.

Architecture

ComponentLocation
WebSocket serversrc/realtime_multilingual_asr_router/api/websocket.py
Audio resampling (16 kHz)src/realtime_multilingual_asr_router/audio/resampler.py
Sample-count master clocksrc/realtime_multilingual_asr_router/coordinator/sample_clock.py
VAD (Silero v6)src/realtime_multilingual_asr_router/vad/silero.py
ASR (Zipformer / sherpa-onnx)src/realtime_multilingual_asr_router/asr/kroko.py
LID schedulersrc/realtime_multilingual_asr_router/lid/scheduler.py
Rollback enginesrc/realtime_multilingual_asr_router/rollback/engine.py
Audio ring buffersrc/realtime_multilingual_asr_router/audio/buffer.py
REST APIsrc/realtime_multilingual_asr_router/api/rest.py
Frontendfrontend/

Asynchronous rollback pipeline

  1. Immediate transcription — Active monolingual ASR transcribes in real time.
  2. Asynchronous monitoring — After each VAD boundary, LID runs on expanding windows in the background.
  3. Rollback trigger — On a confident language switch, swap ASR model, roll the transcript back to the segment boundary, and re-infer buffered audio.

Wrong-language text may appear briefly in partials; finals have clean language boundaries.

Asynchronous rollback pipeline

Evaluation

Benchmarks on three datasets: monolingual baseline (FLEURS), inter-utterance code-switch (synthetic FLEURS blend), and intra-utterance code-switch (Bangor–Miami).

This work: 640 ms Zipformer chunks — Kroko for en/fr/es/de/it/pt/nl, Alphacephei ru.

Baselines: Deepgram Nova-3, ElevenLabs Scribe v2, Mistral Voxtral-Mini-4B-Realtime (480 ms), AssemblyAI u3-rt-pro (no nl/ru — excluded where those languages appear).

Results

ScenarioThis workDeepgramElevenLabsVoxtral Mini
Inter-utterance (synthetic FLEURS)~13%~14%~14%~21%
Intra-utterance (Miami)~41%~29%~26%~76%

Monolingual FLEURS WER by language

Code-switching WER — synthetic FLEURS vs Miami corpus

Detailed WER (%)

DatasetVoxtral MiniDeepgram Nova-3ElevenLabs Scribe v2AAI u3-rt-proThis work
Italian (FLEURS)3.307.102.002.504.30
Russian (FLEURS)5.2010.207.40—14.30
Portuguese (FLEURS)4.9010.303.903.909.20
English (FLEURS)12.4010.204.303.4013.20
French (FLEURS)8.8010.605.504.109.90
Dutch (FLEURS)8.2013.305.30—11.90
German (FLEURS)6.0010.003.903.4011.00
Spanish (FLEURS)3.105.902.702.404.70
Code-switch (synthetic FLEURS)21.3014.3013.60—13.20
Miami (intra-utterance)76.3028.6026.5034.0041.00
Avg (FLEURS only)6.509.704.403.309.80

Limitations

  • Intra-utterance gap — Mid-sentence mixing (Spanglish, etc.) outpaces VAD boundaries; rollback rarely fires.
  • Per-language ceilings — Accuracy depends on open-source Zipformer quality per language.

References

Bangor–Miami — Mozilla Data Collective. Deuchar et al. (2014), Advances in the Study of Bilingualism, Multilingual Matters.

FLEURS:

@article{fleurs2022arxiv,
  title = {FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech},
  author = {Conneau, Alexis and Ma, Min and Khanuja, Simran and Zhang, Yu and Axelrod, Vera and Dalmia, Siddharth and Riesa, Jason and Rivera, Clara and Bapna, Ankur},
  journal = {arXiv preprint arXiv:2205.12446},
  url = {https://arxiv.org/abs/2205.12446},
  year = {2022}
}

License

MIT — see LICENSE.

gladiaio/realtime-multilingual-asr-router

Python

23

3 commits

updated May 29, 2026

See the code

README

Realtime Multilingual ASR Router

Real-time multilingual ASR with code-switching orchestration. A lightweight, CPU-friendly coordinator routes between small monolingual Zipformer models (VAD + LID + rollback) instead of one large multilingual model.

Installation

Prerequisites

  • Python 3.11+
  • uv

Setup

git clone https://github.com/gladiaio/realtime_multilingual_asr_router.git
cd realtime_multilingual_asr_router

uv sync

# Download community Kroko ASR models (en, fr, es by default) plus VAD and LID
uv run scripts/download_models.py

Notes:

  • Default Kroko download is ~450 MB (en, fr, es). Set HF_TOKEN for higher Hugging Face rate limits.
  • First server start preloads one ASR model per downloaded language (~2s per language).
  • REST: GET /health, GET /languages (component status and discovered ASR languages).

More languages and chunk size:

The list of available KrokoAI models is here. This solution is compatible with any online transducer model.

ASR languages: Kroko de, en, es, fr, it, iw (Hebrew), nl, pt, sv, tr (sv only supports chunk size 64), and ru (from AlphaCephei).

# All Kroko languages + Russian (~1.6 GB)
uv run scripts/download_models.py --languages de en es fr it iw nl pt sv tr ru

# Subset example
uv run scripts/download_models.py --languages en fr es de --chunk-size 64

Configuration (optional)

Settings live in src/realtime_multilingual_asr_router/config.py and are loaded from:

  1. Defaults in code
  2. Environment variables with the LR_ prefix (e.g. LR_ASR_CHUNK_SIZE=64)
  3. A .env file in the project root (copy from .env.example)
cp .env.example .env

ASR languages are discovered from models/streaming_transducers/<chunk_size>/<lang>/ (each folder needs encoder.onnx, decoder.onnx, joiner.onnx, and tokens.txt).

Run

uv run uvicorn realtime_multilingual_asr_router.main:app

Open http://127.0.0.1:8000/ for the frontend. WebSocket: /ws/transcribe.

Development

uv run pytest
uv run ruff check .
uv run ruff format .

Model-backed tests are marked slow and may be skipped without downloaded weights.

Project structure

realtime_multilingual_asr_router/
├── src/realtime_multilingual_asr_router/   # Application package
├── frontend/                   # Web UI
├── assets/                     # Architecture diagram and benchmark charts
├── tests/
├── scripts/download_models.py
├── models/                     # Downloaded weights (gitignored)
└── .env.example

Models

ModelRoleSource
Silero VADVoice activitysnakers4/silero-vad
Kroko ASR (community)Streaming ASR (de, en, es, fr, it, iw, nl, pt, sv, tr)Banafo/Kroko-ASR
Russian ZipformerStreaming ASR (ru, pass --languages ru)sherpa-onnx-streaming-zipformer-small-ru-vosk
VoxLingua107-ECAPA-TDNNLanguage IDspeechbrain/lang-id-voxlingua107-ecapa

scripts/download_models.py downloads only the languages you list. Kroko codes use .data extraction; ru pulls ONNX files from the Hub (no .data). Default: en, fr, es.

Stack: FastAPI, uvicorn, sherpa-onnx, silero-vad, speechbrain, torch, torchcodec, pytest, ruff.

Architecture

ComponentLocation
WebSocket serversrc/realtime_multilingual_asr_router/api/websocket.py
Audio resampling (16 kHz)src/realtime_multilingual_asr_router/audio/resampler.py
Sample-count master clocksrc/realtime_multilingual_asr_router/coordinator/sample_clock.py
VAD (Silero v6)src/realtime_multilingual_asr_router/vad/silero.py
ASR (Zipformer / sherpa-onnx)src/realtime_multilingual_asr_router/asr/kroko.py
LID schedulersrc/realtime_multilingual_asr_router/lid/scheduler.py
Rollback enginesrc/realtime_multilingual_asr_router/rollback/engine.py
Audio ring buffersrc/realtime_multilingual_asr_router/audio/buffer.py
REST APIsrc/realtime_multilingual_asr_router/api/rest.py
Frontendfrontend/

Asynchronous rollback pipeline

  1. Immediate transcription — Active monolingual ASR transcribes in real time.
  2. Asynchronous monitoring — After each VAD boundary, LID runs on expanding windows in the background.
  3. Rollback trigger — On a confident language switch, swap ASR model, roll the transcript back to the segment boundary, and re-infer buffered audio.

Wrong-language text may appear briefly in partials; finals have clean language boundaries.

Asynchronous rollback pipeline

Evaluation

Benchmarks on three datasets: monolingual baseline (FLEURS), inter-utterance code-switch (synthetic FLEURS blend), and intra-utterance code-switch (Bangor–Miami).

This work: 640 ms Zipformer chunks — Kroko for en/fr/es/de/it/pt/nl, Alphacephei ru.

Baselines: Deepgram Nova-3, ElevenLabs Scribe v2, Mistral Voxtral-Mini-4B-Realtime (480 ms), AssemblyAI u3-rt-pro (no nl/ru — excluded where those languages appear).

Results

ScenarioThis workDeepgramElevenLabsVoxtral Mini
Inter-utterance (synthetic FLEURS)~13%~14%~14%~21%
Intra-utterance (Miami)~41%~29%~26%~76%

Monolingual FLEURS WER by language

Code-switching WER — synthetic FLEURS vs Miami corpus

Detailed WER (%)

DatasetVoxtral MiniDeepgram Nova-3ElevenLabs Scribe v2AAI u3-rt-proThis work
Italian (FLEURS)3.307.102.002.504.30
Russian (FLEURS)5.2010.207.40—14.30
Portuguese (FLEURS)4.9010.303.903.909.20
English (FLEURS)12.4010.204.303.4013.20
French (FLEURS)8.8010.605.504.109.90
Dutch (FLEURS)8.2013.305.30—11.90
German (FLEURS)6.0010.003.903.4011.00
Spanish (FLEURS)3.105.902.702.404.70
Code-switch (synthetic FLEURS)21.3014.3013.60—13.20
Miami (intra-utterance)76.3028.6026.5034.0041.00
Avg (FLEURS only)6.509.704.403.309.80

Limitations

  • Intra-utterance gap — Mid-sentence mixing (Spanglish, etc.) outpaces VAD boundaries; rollback rarely fires.
  • Per-language ceilings — Accuracy depends on open-source Zipformer quality per language.

References

Bangor–Miami — Mozilla Data Collective. Deuchar et al. (2014), Advances in the Study of Bilingualism, Multilingual Matters.

FLEURS:

@article{fleurs2022arxiv,
  title = {FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech},
  author = {Conneau, Alexis and Ma, Min and Khanuja, Simran and Zhang, Yu and Axelrod, Vera and Dalmia, Siddharth and Riesa, Jason and Rivera, Clara and Bapna, Ankur},
  journal = {arXiv preprint arXiv:2205.12446},
  url = {https://arxiv.org/abs/2205.12446},
  year = {2022}
}

License

MIT — see LICENSE.

Languages

Python

92.8%

JavaScript

6.3%