GilgameshWind/X-ASR-zh-en

Model

15

stars

51

commits

1

repos using this model

3

linked in READMEs

Jul 29, 2026

updated

automatic-speech-recognition
chinese
english
icefall
k2
offline-asr
onnx
sherpa-onnx
streaming-asr
transducer
x-asr-zipformer-transducer
zipformer

README

πŸŽ™οΈ X-ASR-zh-en

Chinese-English offline-streaming unified ASR model artifacts for low-latency deployment.

Shanghai Jiao Tong University Shanghai Innovation Institute Fudan University Huazhong University of Science and Technology

Participating Institutions

🌐 GitHub Project | πŸ€— Hugging Face Hub | 🧩 ModelScope | πŸͺ Hugging Face Space | 🎧 Online Demo | πŸš€ Deployment Guide

πŸ“„ X-ASR-zh-en Technical Report: Coming Soon

Model released Languages Streaming Deployment License

πŸ” Model Card Scope | πŸ“¦ Repository Contents | πŸ“Š Evaluation | ⬇️ Download | πŸš€ Deployment


πŸ” Model Card Scope

🧩 X-ASR Series

X-ASR is a series of automatic speech recognition models built with the icefall framework. The series focuses on streaming ASR and low-latency deployment, while also supporting offline recognition. The broader project roadmap, source organization, issue tracking, and bilingual documentation are maintained on the GitHub project page.

πŸ€– X-ASR-zh-en

X-ASR-zh-en is trained on approximately 1 million hours of open-source and collected speech data. It is designed as an offline-streaming unified transducer ASR model with the Zipformer architecture, supporting both offline decoding and true streaming decoding. The model provides multiple streaming chunk sizes: 160 ms, 480 ms, 960 ms, and 1920 ms, supports punctuation and casing, and can be deployed with sherpa-onnx.

Zipformer architecture

✨ Artifact Page Notes

This repository is the model artifact page for X-ASR-zh-en.

What this artifact page providesWhat the GitHub project provides
Downloadable model artifactsProject-level overview
ONNX encoder / decoder / joiner filesBilingual README and release notes
sherpa-onnx deployment entry pointSource layout and issue tracking
Model-card metadata, tags, license, and metricsDevelopment history and contribution workflow

πŸ“¦ Repository Contents

PathPurpose
deployment/Deployment-ready sherpa-onnx runtime files and examples
deployment/models/Exported streaming ONNX model variants
deployment/infer_and_client/WebSocket server, inference wrapper, and test client
figure/Architecture figure and demo preview media
demo/Demo video asset
applications/vibe-xasr/Vibe XASR desktop application package, manifest, and download notes
streaming_exp/Averaged/pretrained checkpoint artifact for research reference

Directory Layout

.
|-- README.md
|-- config.json
|-- demo/
|   `-- demo.mov
|-- applications/
|   `-- vibe-xasr/
|       |-- README.md
|       |-- download_manifest.json
|       `-- VibeXASR-1.1.2-macos-universal.dmg
|-- deployment/
|   |-- README.md
|   |-- infer_and_client/
|   |   |-- README.md
|   |   |-- sherpa_streaming_client.py
|   |   |-- sherpa_streaming_infer.py
|   |   `-- sherpa_streaming_server.py
|   `-- models/
|       |-- chunk-160ms-model/
|       |   |-- encoder-160ms.onnx
|       |   |-- decoder-160ms.onnx
|       |   |-- joiner-160ms.onnx
|       |   `-- tokens.txt
|       |-- chunk-480ms-model/
|       |-- chunk-960ms-model/
|       `-- chunk-1920ms-model/
|-- figure/
|   |-- zipformer.png
|   |-- demo-preview.png
|   `-- institutions/
`-- streaming_exp/
    `-- pretrained.pt

🧩 Model Variants

Each streaming variant contains a matched encoder, decoder, joiner, and tokens.txt. Do not mix files across model folders.

DirectoryEncoderDecoderJoinerIntended chunk
deployment/models/chunk-160ms-modelencoder-160ms.onnxdecoder-160ms.onnxjoiner-160ms.onnx160 ms
deployment/models/chunk-480ms-modelencoder-480ms.onnxdecoder-480ms.onnxjoiner-480ms.onnx480 ms
deployment/models/chunk-960ms-modelencoder-960ms.onnxdecoder-960ms.onnxjoiner-960ms.onnx960 ms
deployment/models/chunk-1920ms-modelencoder-1920ms.onnxdecoder-1920ms.onnxjoiner-1920ms.onnx1920 ms

⭐ Highlights

CategoryDescription
Frameworkicefall / k2
ArchitectureZipformer transducer
Runtimesherpa-onnx
LanguagesChinese and English
Training scaleApproximately 1 million hours of open-source and collected speech data
Recognition modesOffline decoding and true streaming decoding
Streaming chunks160 ms, 480 ms, 960 ms, 1920 ms
Text outputSupports punctuation and casing

πŸ“Š Evaluation

The following results are for the current X-ASR-zh-en release. Values are WER/CER percentages; lower is better. All results are reported with greedy search.

ModeChunk sizeLibriSpeechGigaSpeechWenetSpeech
cleanothernetmeeting
Streaming160 ms3.9110.1710.979.4512.04
Streaming480 ms3.147.579.777.389.31
Streaming960 ms3.127.229.626.968.84
Streaming1920 ms2.846.479.466.428.03
Offline-2.695.769.235.967.20

Note: Bold numbers indicate the best result among the listed modes for each benchmark column.

Public Benchmark Model Comparison

The following table compares representative ASR models on the same public benchmark columns. Ranks are computed by AVG across the five listed columns; lower is better. Parameter sizes are shown when provided by the source sheet.

RankModelParamsLibriSpeechGigaSpeechWenetSpeechAVG
cleanothernetmeeting
1Qwen3-ASR1.7B1.653.458.565.295.464.882
2Qwen3-ASR0.6B2.184.548.945.976.885.702
3X-ASR-zh-en (offline)0.16B2.565.569.175.837.066.036
4SenseVoice-small234M3.167.2111.245.736.476.762
5VibeVoice-ASR9B2.185.659.4914.4517.199.792

GigaSpeechBench Vertical Domain Evaluation

The following results report GigaSpeechBench vertical-domain performance for the current X-ASR-zh-en release. Values are WER/CER percentages; lower is better. Domain abbreviations follow the GigaSpeechBench vertical-domain labels.

CH

ModeChunk sizeARGAITARTBIOECMENGENTFINHUMLAWMEDMIL
Streaming160 ms9.886.764.397.324.133.588.453.2310.426.584.252.55
Streaming480 ms8.676.173.606.223.783.047.042.789.435.843.762.11
Streaming960 ms8.005.693.446.103.692.886.712.729.075.583.692.11
Streaming1920 ms7.245.583.275.823.482.746.552.578.594.973.531.94
Offline-6.564.542.775.042.992.326.021.947.644.202.901.68

EN

ModeChunk sizeARGAITARTBIOECMENGENTFINHUMLAWMEDMIL
Streaming160 ms5.298.578.557.314.335.0116.255.587.3613.396.036.20
Streaming480 ms4.628.407.736.124.194.6514.505.216.7911.515.596.02
Streaming960 ms4.588.357.456.004.134.4413.995.126.5810.865.526.04
Streaming1920 ms4.338.326.905.894.004.3713.614.986.3910.525.455.78
Offline-4.098.286.735.484.124.3012.304.946.1710.415.355.61

🎧 Demo

A sherpa-onnx based online demo is available here:

Demo video:

X-ASR demo video preview

Open demo video

⬇️ Download

GitHub

Use GitHub when you want the full project repository, bilingual documentation, training references, deployment examples, and issue-tracking context.

git lfs install
git clone https://github.com/Gilgamesh-J/X-ASR.git
cd X-ASR
git lfs pull

Hugging Face

Use Hugging Face when you want the model artifact page and standard HF Hub download tooling.

hf download GilgameshWind/X-ASR-zh-en \
  --local-dir ./X-ASR-zh-en

You can also clone the Hugging Face repository with Git LFS:

git lfs install
git clone https://huggingface.co/GilgameshWind/X-ASR-zh-en
cd X-ASR-zh-en
git lfs pull

ModelScope

Use ModelScope when you prefer the ModelScope mirror or Git LFS clone from ModelScope.

git lfs install
git clone https://www.modelscope.ai/Gilgamesh-J/X-ASR-zh-en.git
cd X-ASR-zh-en
git lfs pull

πŸš€ Deployment

The recommended runtime is sherpa-onnx. The shortest path is to use the deployment package in this repository.

cd deployment
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt

Start a CPU streaming server with the 160 ms model:

python infer_and_client/sherpa_streaming_server.py \
  --host 0.0.0.0 \
  --port 8766 \
  --tokens models/chunk-160ms-model/tokens.txt \
  --encoder models/chunk-160ms-model/encoder-160ms.onnx \
  --decoder models/chunk-160ms-model/decoder-160ms.onnx \
  --joiner models/chunk-160ms-model/joiner-160ms.onnx \
  --provider cpu \
  --sample-rate 16000 \
  --feature-dim 80 \
  --num-threads 1 \
  --decoding-method greedy_search \
  --model-type zipformer2 \
  --enable-endpoint-detection 0 \
  --text-format none

Optional interactive tail-probe mode:

python infer_and_client/sherpa_streaming_server.py \
  --host 0.0.0.0 \
  --port 8766 \
  --tokens models/chunk-160ms-model/tokens.txt \
  --encoder models/chunk-160ms-model/encoder-160ms.onnx \
  --decoder models/chunk-160ms-model/decoder-160ms.onnx \
  --joiner models/chunk-160ms-model/joiner-160ms.onnx \
  --provider cpu \
  --sample-rate 16000 \
  --feature-dim 80 \
  --num-threads 1 \
  --decoding-method greedy_search \
  --model-type zipformer2 \
  --enable-endpoint-detection 0 \
  --text-format none \
  --enable-energy-tail-probe 1 \
  --low-energy-rms 0.003 \
  --speech-rms 0.010 \
  --min-speech-ms 200 \
  --min-silence-ms 500 \
  --tail-probe-ms 500 \
  --tail-probe-cooldown-ms 1000

The default mode keeps --enable-energy-tail-probe 0 and decodes only from client audio chunks. Tail-probe mode is useful for interactive voice-input demos where trailing partial results should refresh after the user pauses. Tune the RMS and silence thresholds according to microphone gain, background noise, and frontend chunking behavior.

Test it with a WAV file:

python infer_and_client/sherpa_streaming_client.py \
  --server-uri ws://127.0.0.1:8766 \
  --wav /path/to/test.wav \
  --chunk-ms 100 \
  --simulate-realtime 1

For complete runtime options, see deployment/README.md. For the script-level server/client guide and full parameter reference, see deployment/infer_and_client/README.md.

⚠️ Intended Use and Limitations

  • This release is intended for Chinese-English ASR research, evaluation, demos, and deployment experiments.
  • The current release focuses on streaming and offline-streaming unified recognition.
  • Production latency depends on hardware, concurrency, audio chunking, endpointing, and server configuration.
  • The technical report with training details, evaluation protocol, ablations, and additional analysis is coming soon.

πŸ“„ Citation

The X-ASR-zh-en technical report is coming soon. Please cite the report once it is released. For now, refer to this model card and the GitHub project page.

πŸ“œ License

This model is released under the Apache-2.0 License.

πŸ™ Acknowledgements

This model is trained with icefall and deployed with sherpa-onnx.

Contributors

GilgameshWind

39 commits

GI
Gilgamesh

11 commits

chenxie95

1 commits

GilgameshWind/X-ASR-zh-en

Model

15

stars

51

commits

1

repos using this model

3

linked in READMEs

Jul 29, 2026

updated

automatic-speech-recognition
chinese
english
icefall
k2
offline-asr
onnx
sherpa-onnx
streaming-asr
transducer
x-asr-zipformer-transducer
zipformer

README

πŸŽ™οΈ X-ASR-zh-en

Chinese-English offline-streaming unified ASR model artifacts for low-latency deployment.

Shanghai Jiao Tong University Shanghai Innovation Institute Fudan University Huazhong University of Science and Technology

Participating Institutions

🌐 GitHub Project | πŸ€— Hugging Face Hub | 🧩 ModelScope | πŸͺ Hugging Face Space | 🎧 Online Demo | πŸš€ Deployment Guide

πŸ“„ X-ASR-zh-en Technical Report: Coming Soon

Model released Languages Streaming Deployment License

πŸ” Model Card Scope | πŸ“¦ Repository Contents | πŸ“Š Evaluation | ⬇️ Download | πŸš€ Deployment


πŸ” Model Card Scope

🧩 X-ASR Series

X-ASR is a series of automatic speech recognition models built with the icefall framework. The series focuses on streaming ASR and low-latency deployment, while also supporting offline recognition. The broader project roadmap, source organization, issue tracking, and bilingual documentation are maintained on the GitHub project page.

πŸ€– X-ASR-zh-en

X-ASR-zh-en is trained on approximately 1 million hours of open-source and collected speech data. It is designed as an offline-streaming unified transducer ASR model with the Zipformer architecture, supporting both offline decoding and true streaming decoding. The model provides multiple streaming chunk sizes: 160 ms, 480 ms, 960 ms, and 1920 ms, supports punctuation and casing, and can be deployed with sherpa-onnx.

Zipformer architecture

✨ Artifact Page Notes

This repository is the model artifact page for X-ASR-zh-en.

What this artifact page providesWhat the GitHub project provides
Downloadable model artifactsProject-level overview
ONNX encoder / decoder / joiner filesBilingual README and release notes
sherpa-onnx deployment entry pointSource layout and issue tracking
Model-card metadata, tags, license, and metricsDevelopment history and contribution workflow

πŸ“¦ Repository Contents

PathPurpose
deployment/Deployment-ready sherpa-onnx runtime files and examples
deployment/models/Exported streaming ONNX model variants
deployment/infer_and_client/WebSocket server, inference wrapper, and test client
figure/Architecture figure and demo preview media
demo/Demo video asset
applications/vibe-xasr/Vibe XASR desktop application package, manifest, and download notes
streaming_exp/Averaged/pretrained checkpoint artifact for research reference

Directory Layout

.
|-- README.md
|-- config.json
|-- demo/
|   `-- demo.mov
|-- applications/
|   `-- vibe-xasr/
|       |-- README.md
|       |-- download_manifest.json
|       `-- VibeXASR-1.1.2-macos-universal.dmg
|-- deployment/
|   |-- README.md
|   |-- infer_and_client/
|   |   |-- README.md
|   |   |-- sherpa_streaming_client.py
|   |   |-- sherpa_streaming_infer.py
|   |   `-- sherpa_streaming_server.py
|   `-- models/
|       |-- chunk-160ms-model/
|       |   |-- encoder-160ms.onnx
|       |   |-- decoder-160ms.onnx
|       |   |-- joiner-160ms.onnx
|       |   `-- tokens.txt
|       |-- chunk-480ms-model/
|       |-- chunk-960ms-model/
|       `-- chunk-1920ms-model/
|-- figure/
|   |-- zipformer.png
|   |-- demo-preview.png
|   `-- institutions/
`-- streaming_exp/
    `-- pretrained.pt

🧩 Model Variants

Each streaming variant contains a matched encoder, decoder, joiner, and tokens.txt. Do not mix files across model folders.

DirectoryEncoderDecoderJoinerIntended chunk
deployment/models/chunk-160ms-modelencoder-160ms.onnxdecoder-160ms.onnxjoiner-160ms.onnx160 ms
deployment/models/chunk-480ms-modelencoder-480ms.onnxdecoder-480ms.onnxjoiner-480ms.onnx480 ms
deployment/models/chunk-960ms-modelencoder-960ms.onnxdecoder-960ms.onnxjoiner-960ms.onnx960 ms
deployment/models/chunk-1920ms-modelencoder-1920ms.onnxdecoder-1920ms.onnxjoiner-1920ms.onnx1920 ms

⭐ Highlights

CategoryDescription
Frameworkicefall / k2
ArchitectureZipformer transducer
Runtimesherpa-onnx
LanguagesChinese and English
Training scaleApproximately 1 million hours of open-source and collected speech data
Recognition modesOffline decoding and true streaming decoding
Streaming chunks160 ms, 480 ms, 960 ms, 1920 ms
Text outputSupports punctuation and casing

πŸ“Š Evaluation

The following results are for the current X-ASR-zh-en release. Values are WER/CER percentages; lower is better. All results are reported with greedy search.

ModeChunk sizeLibriSpeechGigaSpeechWenetSpeech
cleanothernetmeeting
Streaming160 ms3.9110.1710.979.4512.04
Streaming480 ms3.147.579.777.389.31
Streaming960 ms3.127.229.626.968.84
Streaming1920 ms2.846.479.466.428.03
Offline-2.695.769.235.967.20

Note: Bold numbers indicate the best result among the listed modes for each benchmark column.

Public Benchmark Model Comparison

The following table compares representative ASR models on the same public benchmark columns. Ranks are computed by AVG across the five listed columns; lower is better. Parameter sizes are shown when provided by the source sheet.

RankModelParamsLibriSpeechGigaSpeechWenetSpeechAVG
cleanothernetmeeting
1Qwen3-ASR1.7B1.653.458.565.295.464.882
2Qwen3-ASR0.6B2.184.548.945.976.885.702
3X-ASR-zh-en (offline)0.16B2.565.569.175.837.066.036
4SenseVoice-small234M3.167.2111.245.736.476.762
5VibeVoice-ASR9B2.185.659.4914.4517.199.792

GigaSpeechBench Vertical Domain Evaluation

The following results report GigaSpeechBench vertical-domain performance for the current X-ASR-zh-en release. Values are WER/CER percentages; lower is better. Domain abbreviations follow the GigaSpeechBench vertical-domain labels.

CH

ModeChunk sizeARGAITARTBIOECMENGENTFINHUMLAWMEDMIL
Streaming160 ms9.886.764.397.324.133.588.453.2310.426.584.252.55
Streaming480 ms8.676.173.606.223.783.047.042.789.435.843.762.11
Streaming960 ms8.005.693.446.103.692.886.712.729.075.583.692.11
Streaming1920 ms7.245.583.275.823.482.746.552.578.594.973.531.94
Offline-6.564.542.775.042.992.326.021.947.644.202.901.68

EN

ModeChunk sizeARGAITARTBIOECMENGENTFINHUMLAWMEDMIL
Streaming160 ms5.298.578.557.314.335.0116.255.587.3613.396.036.20
Streaming480 ms4.628.407.736.124.194.6514.505.216.7911.515.596.02
Streaming960 ms4.588.357.456.004.134.4413.995.126.5810.865.526.04
Streaming1920 ms4.338.326.905.894.004.3713.614.986.3910.525.455.78
Offline-4.098.286.735.484.124.3012.304.946.1710.415.355.61

🎧 Demo

A sherpa-onnx based online demo is available here:

Demo video:

X-ASR demo video preview

Open demo video

⬇️ Download

GitHub

Use GitHub when you want the full project repository, bilingual documentation, training references, deployment examples, and issue-tracking context.

git lfs install
git clone https://github.com/Gilgamesh-J/X-ASR.git
cd X-ASR
git lfs pull

Hugging Face

Use Hugging Face when you want the model artifact page and standard HF Hub download tooling.

hf download GilgameshWind/X-ASR-zh-en \
  --local-dir ./X-ASR-zh-en

You can also clone the Hugging Face repository with Git LFS:

git lfs install
git clone https://huggingface.co/GilgameshWind/X-ASR-zh-en
cd X-ASR-zh-en
git lfs pull

ModelScope

Use ModelScope when you prefer the ModelScope mirror or Git LFS clone from ModelScope.

git lfs install
git clone https://www.modelscope.ai/Gilgamesh-J/X-ASR-zh-en.git
cd X-ASR-zh-en
git lfs pull

πŸš€ Deployment

The recommended runtime is sherpa-onnx. The shortest path is to use the deployment package in this repository.

cd deployment
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt

Start a CPU streaming server with the 160 ms model:

python infer_and_client/sherpa_streaming_server.py \
  --host 0.0.0.0 \
  --port 8766 \
  --tokens models/chunk-160ms-model/tokens.txt \
  --encoder models/chunk-160ms-model/encoder-160ms.onnx \
  --decoder models/chunk-160ms-model/decoder-160ms.onnx \
  --joiner models/chunk-160ms-model/joiner-160ms.onnx \
  --provider cpu \
  --sample-rate 16000 \
  --feature-dim 80 \
  --num-threads 1 \
  --decoding-method greedy_search \
  --model-type zipformer2 \
  --enable-endpoint-detection 0 \
  --text-format none

Optional interactive tail-probe mode:

python infer_and_client/sherpa_streaming_server.py \
  --host 0.0.0.0 \
  --port 8766 \
  --tokens models/chunk-160ms-model/tokens.txt \
  --encoder models/chunk-160ms-model/encoder-160ms.onnx \
  --decoder models/chunk-160ms-model/decoder-160ms.onnx \
  --joiner models/chunk-160ms-model/joiner-160ms.onnx \
  --provider cpu \
  --sample-rate 16000 \
  --feature-dim 80 \
  --num-threads 1 \
  --decoding-method greedy_search \
  --model-type zipformer2 \
  --enable-endpoint-detection 0 \
  --text-format none \
  --enable-energy-tail-probe 1 \
  --low-energy-rms 0.003 \
  --speech-rms 0.010 \
  --min-speech-ms 200 \
  --min-silence-ms 500 \
  --tail-probe-ms 500 \
  --tail-probe-cooldown-ms 1000

The default mode keeps --enable-energy-tail-probe 0 and decodes only from client audio chunks. Tail-probe mode is useful for interactive voice-input demos where trailing partial results should refresh after the user pauses. Tune the RMS and silence thresholds according to microphone gain, background noise, and frontend chunking behavior.

Test it with a WAV file:

python infer_and_client/sherpa_streaming_client.py \
  --server-uri ws://127.0.0.1:8766 \
  --wav /path/to/test.wav \
  --chunk-ms 100 \
  --simulate-realtime 1

For complete runtime options, see deployment/README.md. For the script-level server/client guide and full parameter reference, see deployment/infer_and_client/README.md.

⚠️ Intended Use and Limitations

  • This release is intended for Chinese-English ASR research, evaluation, demos, and deployment experiments.
  • The current release focuses on streaming and offline-streaming unified recognition.
  • Production latency depends on hardware, concurrency, audio chunking, endpointing, and server configuration.
  • The technical report with training details, evaluation protocol, ablations, and additional analysis is coming soon.

πŸ“„ Citation

The X-ASR-zh-en technical report is coming soon. Please cite the report once it is released. For now, refer to this model card and the GitHub project page.

πŸ“œ License

This model is released under the Apache-2.0 License.

πŸ™ Acknowledgements

This model is trained with icefall and deployed with sherpa-onnx.

Contributors

GilgameshWind

39 commits

GI
Gilgamesh

11 commits

chenxie95

1 commits