IndexTeam/Index-Echo-S2TT-2B

Model

Index-Echo-S2TT-2B

5

10 commits

1 linked in READMEs

updated Sep 30, 2026

See the code

README

Index-Echo-S2TT-2B

Online demo Β· GitHub Β· Technical report Β· Hugging Face collection Β· ModelScope collection

Index-Echo-S2TT-2B translates speech into text using a Qwen3-Omni AuT audio encoder, an audio connector, and an Index-Translate-2B decoder. The released package accepts audio or video files and writes bilingual subtitles with sentence timestamps. Its packaged inference interface supports Chinese β†’ English, Japanese, or Spanish.

This is the speech-to-text member of the Index-Translate family. The 9B sibling provides the same packaged interface at a different model size.

Architecture

Index-Echo speech-to-text and speech-to-speech architecture

The report's upper path shows the S2TT model: source audio β†’ Qwen3-Omni AuT encoder β†’ audio connector β†’ Index-Translate decoder β†’ translated target text. This card covers that speech-to-text path; the lower path adds speech generation for S2ST.

Model and training

The technical report describes end-to-end training on speech translation: acoustic input, target-language instructions, and translated text are represented in a single sequence. The model learns to interpret source audio directly while using the multilingual text decoder for translation. It produces the source transcript and target translation together with per-sentence timestamps.

The self-contained export includes the trained audio encoder and connector as safetensors, plus the Qwen3.5-family decoder in Hugging Face format under llm/. See MODEL_INFO.json for the export's checkpoint details. The name's 2B refers to the decoder size; the package also contains the audio components.

Inference

Install the Hugging Face CLI, download the package, and install its inference requirements. System ffmpeg must be available on PATH.

pip install -U huggingface_hub
hf download IndexTeam/Index-Echo-S2TT-2B --local-dir ./Index-Echo-S2TT-2B
pip install -r Index-Echo-S2TT-2B/requirements.txt
python3 Index-Echo-S2TT-2B/infer.py input.mp4 --target-lang en --out out_dir

The official inference guide uses torch==2.11.0 and transformers==5.6.0, with safetensors, librosa, soundfile, and silero-vad. Keep the packaged requirements and a CUDA-compatible PyTorch build when setting up the environment. See the S2TT guide for the GitHub wrapper and additional examples.

Outputs are out_dir/input.srt (Chinese transcript and target translation in each cue) and out_dir/input.windows.jsonl (raw per-window output and a final summary).

Default inference settings

Setting / CLI optionReleased defaultMeaning
inputsRequired, one or more filesAudio/video paths, processed sequentially
--outRequiredDirectory for SRT and per-window JSONL outputs
--temperature0.0Greedy decoding (do_sample=False); positive values enable sampling
--max-new-tokens2000Output-token budget per audio window
--max-win60.0 secondsVAD audio-window cap
--ctx-k5Maximum number of previous output windows used as context
--target-langenTarget choices: en, ja, es
--glossaryEmpty stringInline terminology; for example 原名:Translated name
--glossary-fileEmpty stringUTF-8 terminology file; takes priority over --glossary
--devicecuda:0Device used to load the model
Model dtypetorch.bfloat16Set by the package loader, not a CLI option
Text stoppingTokenizer EOS or `<im_end

These defaults are identical for the 2B and 9B subtitle scripts. The package does not explicitly override top_p, top_k, or repetition penalties; values for those settings follow the decoder generation configuration. The prompt prefills an empty <think> block, as in the released script. The greedy speech-transcription/translation settings belong to the speech package; they are distinct from the text-model client settings.

python3 Index-Echo-S2TT-2B/infer.py input.mp4 \
  --target-lang ja --out out_ja \
  --temperature 0 --max-new-tokens 2000 --max-win 60 --ctx-k 5 \
  --glossary "原名1:訳名1,原名2:訳名2" --device cuda:0

ffmpeg converts input to 16 kHz mono audio. Silero VAD divides it into windows; inference runs sequentially and feeds previous transcript/translation windows into the context slot. --glossary supplies terminology, and --glossary-file reads the same format from a file with priority over the inline value. Context and glossary instructions help with consistency but do not guarantee exact compliance.

Evaluation

On the report's in-house video translation set, Index-Echo-2B achieves 0.825 MT judge, 0.088 median ASR error, and 0.395 s start-time MAE. Its ASR and timestamp errors are the lowest in this comparison.

The set contains 140 total windows, each 50–60 seconds long: 20 windows per direction for Chinese β†’ English/Japanese/Korean/Spanish/Portuguese/Arabic and English β†’ Chinese. This broader evaluation setup is separate from the released infer.py interface, which exposes Chinese β†’ English/Japanese/Spanish.

ModelMT judge ↑ASR error (median) ↓Start-time MAE (s) ↓
Qwen3.8-Omni-Flash0.8870.1121.004
Index-Echo-9B0.8570.1020.488
Gemini-3.1-Pro (thinking)0.8490.1691.824
Qwen3.8-LiveTranslate0.8330.186β€”
Index-Echo-2B0.8250.0880.395
Gemini-2.5-Flash (thinking)0.8170.2012.402
Gemini-2.5-Flash0.7570.4791.746
Qwen3-Omni0.7000.1391.145
FireRed Audio0.4480.1280.765
SeamlessM4T-v20.0601.000β€”

The MT judge averages Gemini-3.1-Pro's per-line quality ratings on a 0/0.5/1 scale. ASR error compares the source transcription with the reference transcript; the table reports its median. Start-time MAE is the mean absolute difference between predicted and reference sentence start times, in seconds. Bold metric values mark the best result in each column.

Qwen3.8-LiveTranslate was evaluated in streaming mode. Dashes indicate unavailable timestamp outputs. SeamlessM4T-v2 received unchunked inputs of approximately 57 seconds, beyond its approximately 30-second effective range, and does not support the line-level timestamp instructions used here. Its score should be interpreted under those input conditions. These are in-house results, not a universal ranking across speech tasks or deployment settings. See the report and evaluation tables for the protocol.

Package files and limitations

File or directoryPurpose
infer.py, requirements.txtPackaged subtitle inference and dependencies
llm/Index-Translate-derived decoder and tokenizer
audio_tower.safetensors, audio_config.jsonTrained AuT audio encoder
connector.safetensorsAudio-to-decoder connector
MODEL_INFO.jsonSource checkpoint and export information

The packaged 2B script reports approximately 10 GB VRAM with bf16; allow additional memory headroom for your input and runtime. Processing is sequential within each input file; separate files can be assigned to separate GPUs.

Timestamp accuracy and transcript/translation quality can vary with audio conditions and content. Greedy decoding may repeat on out-of-distribution inputs; the script accepts a positive --temperature to enable sampling, which can change both translation and timing. Review subtitles when precise alignment or terminology is required. The package is a file-based subtitle interface; its windowing does not establish a real-time latency guarantee.

For speech generation and long-video dubbing, use the separate S2ST packages or the video dubbing pipeline.

FamilyTaskReleased sizes
Index-TranslateText translation and translation instructions across 150 languages2B, 9B, 35B-A3B (preview)
Index-Echo S2TTSpeech-to-text translation and subtitles2B, 9B
Index-Echo S2STSpeech-to-speech translation with source-voice conditioning2B, 9B
Index-HomuraTranslation with a target syllable count2B, 9B
Index-NativeLongNative long-document translation; released as Index-Nailong2B, 9B

The text foundation's 150-language coverage does not describe the released speech interfaces. Use the task-specific directions documented above.

Citation

@techreport{indextranslate2026,
  author={Tianjiao Li and Mengran Yu and Chenyu Shi and Lusheng Zhang and
          Qisi Chen and Yanshan Zhou and Ji Qi and Jingying Liu and
          Yuang Feng and Ziang Cui and Tianxing Yan},
  title={Index-Translate: A Multilingual Translation Model Family --- Text, Speech, Controlled Dubbing, and Long-Document Translation},
  institution={Index LLM Team},
  year={2026},
  month={September}
}

License and feedback

Apache-2.0. Questions and feedback are welcome through GitHub Issues.

audio-translation
index
safetensors
speech-translation

IndexTeam/Index-Echo-S2TT-2B

Model

Index-Echo-S2TT-2B

5

10 commits

1 linked in READMEs

updated Sep 30, 2026

See the code

README

Index-Echo-S2TT-2B

Online demo Β· GitHub Β· Technical report Β· Hugging Face collection Β· ModelScope collection

Index-Echo-S2TT-2B translates speech into text using a Qwen3-Omni AuT audio encoder, an audio connector, and an Index-Translate-2B decoder. The released package accepts audio or video files and writes bilingual subtitles with sentence timestamps. Its packaged inference interface supports Chinese β†’ English, Japanese, or Spanish.

This is the speech-to-text member of the Index-Translate family. The 9B sibling provides the same packaged interface at a different model size.

Architecture

Index-Echo speech-to-text and speech-to-speech architecture

The report's upper path shows the S2TT model: source audio β†’ Qwen3-Omni AuT encoder β†’ audio connector β†’ Index-Translate decoder β†’ translated target text. This card covers that speech-to-text path; the lower path adds speech generation for S2ST.

Model and training

The technical report describes end-to-end training on speech translation: acoustic input, target-language instructions, and translated text are represented in a single sequence. The model learns to interpret source audio directly while using the multilingual text decoder for translation. It produces the source transcript and target translation together with per-sentence timestamps.

The self-contained export includes the trained audio encoder and connector as safetensors, plus the Qwen3.5-family decoder in Hugging Face format under llm/. See MODEL_INFO.json for the export's checkpoint details. The name's 2B refers to the decoder size; the package also contains the audio components.

Inference

Install the Hugging Face CLI, download the package, and install its inference requirements. System ffmpeg must be available on PATH.

pip install -U huggingface_hub
hf download IndexTeam/Index-Echo-S2TT-2B --local-dir ./Index-Echo-S2TT-2B
pip install -r Index-Echo-S2TT-2B/requirements.txt
python3 Index-Echo-S2TT-2B/infer.py input.mp4 --target-lang en --out out_dir

The official inference guide uses torch==2.11.0 and transformers==5.6.0, with safetensors, librosa, soundfile, and silero-vad. Keep the packaged requirements and a CUDA-compatible PyTorch build when setting up the environment. See the S2TT guide for the GitHub wrapper and additional examples.

Outputs are out_dir/input.srt (Chinese transcript and target translation in each cue) and out_dir/input.windows.jsonl (raw per-window output and a final summary).

Default inference settings

Setting / CLI optionReleased defaultMeaning
inputsRequired, one or more filesAudio/video paths, processed sequentially
--outRequiredDirectory for SRT and per-window JSONL outputs
--temperature0.0Greedy decoding (do_sample=False); positive values enable sampling
--max-new-tokens2000Output-token budget per audio window
--max-win60.0 secondsVAD audio-window cap
--ctx-k5Maximum number of previous output windows used as context
--target-langenTarget choices: en, ja, es
--glossaryEmpty stringInline terminology; for example 原名:Translated name
--glossary-fileEmpty stringUTF-8 terminology file; takes priority over --glossary
--devicecuda:0Device used to load the model
Model dtypetorch.bfloat16Set by the package loader, not a CLI option
Text stoppingTokenizer EOS or `<im_end

These defaults are identical for the 2B and 9B subtitle scripts. The package does not explicitly override top_p, top_k, or repetition penalties; values for those settings follow the decoder generation configuration. The prompt prefills an empty <think> block, as in the released script. The greedy speech-transcription/translation settings belong to the speech package; they are distinct from the text-model client settings.

python3 Index-Echo-S2TT-2B/infer.py input.mp4 \
  --target-lang ja --out out_ja \
  --temperature 0 --max-new-tokens 2000 --max-win 60 --ctx-k 5 \
  --glossary "原名1:訳名1,原名2:訳名2" --device cuda:0

ffmpeg converts input to 16 kHz mono audio. Silero VAD divides it into windows; inference runs sequentially and feeds previous transcript/translation windows into the context slot. --glossary supplies terminology, and --glossary-file reads the same format from a file with priority over the inline value. Context and glossary instructions help with consistency but do not guarantee exact compliance.

Evaluation

On the report's in-house video translation set, Index-Echo-2B achieves 0.825 MT judge, 0.088 median ASR error, and 0.395 s start-time MAE. Its ASR and timestamp errors are the lowest in this comparison.

The set contains 140 total windows, each 50–60 seconds long: 20 windows per direction for Chinese β†’ English/Japanese/Korean/Spanish/Portuguese/Arabic and English β†’ Chinese. This broader evaluation setup is separate from the released infer.py interface, which exposes Chinese β†’ English/Japanese/Spanish.

ModelMT judge ↑ASR error (median) ↓Start-time MAE (s) ↓
Qwen3.8-Omni-Flash0.8870.1121.004
Index-Echo-9B0.8570.1020.488
Gemini-3.1-Pro (thinking)0.8490.1691.824
Qwen3.8-LiveTranslate0.8330.186β€”
Index-Echo-2B0.8250.0880.395
Gemini-2.5-Flash (thinking)0.8170.2012.402
Gemini-2.5-Flash0.7570.4791.746
Qwen3-Omni0.7000.1391.145
FireRed Audio0.4480.1280.765
SeamlessM4T-v20.0601.000β€”

The MT judge averages Gemini-3.1-Pro's per-line quality ratings on a 0/0.5/1 scale. ASR error compares the source transcription with the reference transcript; the table reports its median. Start-time MAE is the mean absolute difference between predicted and reference sentence start times, in seconds. Bold metric values mark the best result in each column.

Qwen3.8-LiveTranslate was evaluated in streaming mode. Dashes indicate unavailable timestamp outputs. SeamlessM4T-v2 received unchunked inputs of approximately 57 seconds, beyond its approximately 30-second effective range, and does not support the line-level timestamp instructions used here. Its score should be interpreted under those input conditions. These are in-house results, not a universal ranking across speech tasks or deployment settings. See the report and evaluation tables for the protocol.

Package files and limitations

File or directoryPurpose
infer.py, requirements.txtPackaged subtitle inference and dependencies
llm/Index-Translate-derived decoder and tokenizer
audio_tower.safetensors, audio_config.jsonTrained AuT audio encoder
connector.safetensorsAudio-to-decoder connector
MODEL_INFO.jsonSource checkpoint and export information

The packaged 2B script reports approximately 10 GB VRAM with bf16; allow additional memory headroom for your input and runtime. Processing is sequential within each input file; separate files can be assigned to separate GPUs.

Timestamp accuracy and transcript/translation quality can vary with audio conditions and content. Greedy decoding may repeat on out-of-distribution inputs; the script accepts a positive --temperature to enable sampling, which can change both translation and timing. Review subtitles when precise alignment or terminology is required. The package is a file-based subtitle interface; its windowing does not establish a real-time latency guarantee.

For speech generation and long-video dubbing, use the separate S2ST packages or the video dubbing pipeline.

FamilyTaskReleased sizes
Index-TranslateText translation and translation instructions across 150 languages2B, 9B, 35B-A3B (preview)
Index-Echo S2TTSpeech-to-text translation and subtitles2B, 9B
Index-Echo S2STSpeech-to-speech translation with source-voice conditioning2B, 9B
Index-HomuraTranslation with a target syllable count2B, 9B
Index-NativeLongNative long-document translation; released as Index-Nailong2B, 9B

The text foundation's 150-language coverage does not describe the released speech interfaces. Use the task-specific directions documented above.

Citation

@techreport{indextranslate2026,
  author={Tianjiao Li and Mengran Yu and Chenyu Shi and Lusheng Zhang and
          Qisi Chen and Yanshan Zhou and Ji Qi and Jingying Liu and
          Yuang Feng and Ziang Cui and Tianxing Yan},
  title={Index-Translate: A Multilingual Translation Model Family --- Text, Speech, Controlled Dubbing, and Long-Document Translation},
  institution={Index LLM Team},
  year={2026},
  month={September}
}

License and feedback

Apache-2.0. Questions and feedback are welcome through GitHub Issues.

audio-translation
index
safetensors
speech-translation