IndexTeam/Index-Echo-S2ST-2B

Model

Index-Echo-S2ST-2B

5

12 commits

1 linked in READMEs

updated Sep 30, 2026

See the code

README

Index-Echo-S2ST-2B

Online demo · GitHub · Technical report · Hugging Face collection · ModelScope collection

Index-Echo-S2ST-2B translates a Chinese or English speech clip into another language and generates speech conditioned on the source speaker's voice. It combines a frozen Index-Echo-S2TT-2B backbone, a learned Hidden2CV mapper, and CosyVoice3 speech generation in one self-contained package and Python process.

The release supports six directions: Chinese → English/Spanish/Japanese and English → Chinese/Spanish/Japanese. The caller specifies the target language; the package infers the source language from its source transcript. The 9B sibling exposes the same interface. 2B denotes the speech translator's decoder size, rather than the total size of all packaged components.

Architecture

Index-Echo speech-to-text and speech-to-speech architecture

The report's lower path extends S2TT to S2ST: the frozen speech translator supplies its final hidden states to Hidden2CV, which connects to CosyVoice3 speech generation. Source audio supplies the reference prompt and CampPlus speaker embedding for voice conditioning. The diagram also shows mapper alignment and DiffRO training; the S2TT backbone remains frozen throughout S2ST training.

Model and training

The technical report initializes S2ST from the trained Index-Echo S2TT model, whose Qwen3-Omni AuT audio encoder and connector feed an Index-Translate decoder. The speech-generation path uses a roughly 30M-parameter Hidden2CV mapper to connect the translator's final hidden states to the CosyVoice3 semantic layer.

Training first distills the mapper with both the S2TT backbone and speech generator frozen, then applies differentiable reward optimization (DiffRO). A Token2Text content-consistency reward is backpropagated through Gumbel-Softmax speech-token samples. The S2TT backbone remains frozen throughout S2ST training. At inference, source audio supplies the reference prompt and CampPlus speaker embedding for voice conditioning. Voice similarity varies with the source, target language, and synthesis result.

Inference

Download the complete model package before importing its local modeling_dubbing.py:

pip install -U huggingface_hub
hf download IndexTeam/Index-Echo-S2ST-2B --local-dir ./Index-Echo-S2ST-2B
cd Index-Echo-S2ST-2B

Pinned environment

The release documents the following tested dependency pins. Use CUDA-compatible PyTorch wheels; the +cu129 builds require a wheel source that provides that CUDA build. System ffmpeg must be on PATH.

pip install torch==2.11.0+cu129 torchaudio==2.11.0+cu129 transformers==5.6.0 \
    tokenizers==0.22.2 librosa==1.0.0 onnxruntime==1.30.0 wetext==0.0.4 \
    kaldifst==1.8.0 conformer==0.3.2 hydra-core==1.3.7 HyperPyYAML==1.2.3 \
    lightning==2.6.6 x-transformers==2.11.24 einx==0.4.3 frozendict==2.4.7 \
    pyworld==0.3.5 soundfile==0.13.1 modelscope==1.37.1 safetensors==0.7.0 \
    numpy==2.4.6 openai-whisper==20250625

openai-whisper is used only for the package's ASR-based evaluation/self-check. The Japanese frontend (pyopenjtalk==0.4.1, including its dictionary, and pykakasi==2.3.0) is vendored under code/ja_ext/; the pipeline adds it to the import path automatically. All components run in one Python environment and process.

The package automatically applies two CosyVoice3 compatibility patches for Transformers 5.6 in code/tts/pipeline.py: keeping the CosyVoice3 LLM in fp32, and providing the full visible attention mask during autoregressive decoding. Keep the bundled code and dependency pins together.

Python API

Run this from the downloaded package directory, with input clips available at the indicated paths:

from modeling_dubbing import DubbingBridgeModel

model = DubbingBridgeModel.from_pretrained('.', device='cuda')

# Chinese source: three supported targets.
wav, sr = model.dub('input_zh.wav', lang='en', out_wav='dub_en.wav')
model.dub('input_zh.wav', lang='es', out_wav='dub_es.wav')
model.dub('input_zh.wav', lang='ja', out_wav='dub_ja.wav')

# English source: three supported targets.
model.dub('input_en.wav', lang='zh', out_wav='dub_zh.wav')
model.dub('input_en.wav', lang='es', out_wav='dub_es2.wav')
model.dub('input_en.wav', lang='ja', out_wav='dub_ja2.wav')

wav, sr, info = model.dub('input_en.wav', lang='zh', return_info=True)
print(sr)               # 24000 Hz
print(info['zh'])       # Source transcript; the key name is retained for both sources.
print(info['tgt_raw'])  # Model translation, before the speech frontend's normalization.
print(info['tgt_cv'])   # Speech-frontend text; Japanese uses a katakana representation.
print(info['src_lang']) # 'zh' or 'en'

lang always means the target language: en, es, ja, or zh. The source-language heuristic treats a source transcript containing CJK characters as Chinese and otherwise as English; it is not a general language identifier. Source and target must differ. The waveform is a float32 tensor of shape (1, T) at 24 kHz. out_wav saves it to a file. With return_info=True, the third return value includes the transcript, translation, source language, and synthesis diagnostics.

At the lower pipeline layer, extract(audio_path, lang) returns the translated text and aligned hidden states separately from synthesis. See the S2ST inference guide for the dub.py wrapper, and the video pipeline for segmentation and video timeline alignment.

Default inference settings

Stage or settingReleased default
Transcription and translationGreedy, temperature=0.0 / do_sample=False
Transcription/translation budgetmax_new_tokens=1024
Speech modeh2cv (the only supported mode in this package)
Speech-token budgetmax_tokens=1500; actual cap min(20 * m, 1500), where m is aligned target-text token count
Minimum speech length2 * m tokens before stop tokens are allowed
Speech samplingsampling=25; release notes describe ras_sampling, top-k 25 / top-p 0.8
Speech-generation seed42
Synthesis speed1.0
Output sample rate24000 Hz
CosyVoice3 executionstream=False, load_trt=False, load_vllm=False, fp16=False; LLM kept in fp32
Chunked orchestrationDisabled (chunk=False)

Both sizes use these settings. Greedy transcription/translation does not make speech synthesis greedy. The public DubbingBridgeModel.dub wrapper does not expose a seed argument; seed changes are available through the lower-level pipeline's dub/synth APIs. Sampling may vary between calls even with the default seed, as documented in the release notes.

Configurable API parameters

The public wrapper and lower pipeline expose different controls. The defaults below are shared by both sizes.

InterfaceParameterDefaultBehavior
from_pretrainedmodel_dirRequiredLocal directory containing the full package
from_pretraineddevicecudaCUDA device for the loaded components
from_pretrainedmodeNone → config default h2cvSets the default speech mode
Public dubaudio_pathRequiredChinese or English source clip
Public dublangenTarget language: en, es, ja, zh; must differ from detected source
Public dubout_wavNoneSave path; None creates a temporary WAV
Public dubmodeNone → loaded defaultOnly h2cv is accepted
Public dubchunkFalseTrue is unsupported in this release
Public dubreturn_infoFalseAdds transcript, translation, and synthesis diagnostics to the return value
Pipeline extractlangenTarget language for transcription/translation and alignment
Pipeline extractmax_new_tokens1024Transcription/translation output budget
Pipeline synthout_wavNoneOptional output save path
Pipeline synthmax_tokens1500Upper bound on generated speech tokens
Pipeline synthseed42Seed set before speech-token generation and acoustic synthesis
Pipeline synthprompt_wavNoneUses the source audio unless a reference clip is supplied
Pipeline synthprompt_textNoneUses the source transcript unless reference text is supplied

For explicit text/speech budgets or seed control, use the lower pipeline after loading model as above:

pipe = model._pipe
ext = pipe.extract('input_zh.wav', lang='en', max_new_tokens=1024)
wav, info = pipe.synth(ext, out_wav='dub_en.wav', max_tokens=1500, seed=42)

The lower pipeline's dub(..., **kw) forwards synthesis options such as seed and max_tokens; changing the translation budget requires calling extract separately. Sampling argument 25 and synthesis speed 1.0 are fixed in the released pipeline code, not keyword arguments of the public wrapper. Translation stops at tokenizer EOS or <|im_end|>; speech generation uses CosyVoice3's defined stop tokens and the length limits above. The script does not explicitly override other text-generation knobs such as top_p, top_k, or repetition penalties.

Evaluation

The report compares end-to-end dubbing with a matched Index-Echo-S2TT + CosyVoice3 pipeline on in-house video dubbing data. Each pair shares the same frozen S2TT model and text translation, so the comparison measures the speech-generation path under a shared translation. For 2B, end-to-end dubbing lowers mean content error in four of six directions: zh→es, zh→ja, en→zh, and en→es. Across both sizes it improves eight of twelve size–direction pairs; Chinese → English favors the pipeline.

DirectionTranslator sizeContent metricText judge ↑Pipeline error ↓
mean / median
E2E error ↓
mean / median
Pipeline speaker ↑E2E speaker ↑
zh→en2BWER0.8300.0450 / 0.0000.0709 / 0.0290.7220.743
zh→en9BWER0.8400.0477 / 0.0000.0615 / 0.0000.7180.739
zh→es2BWER0.7280.0811 / 0.0700.0704 / 0.0510.7470.743
zh→es9BWER0.7930.0916 / 0.0710.0773 / 0.0570.7460.746
zh→ja2BKata CER0.7600.0575 / 0.0340.0493 / 0.0270.7770.780
zh→ja9BKata CER0.8150.0552 / 0.0290.0436 / 0.0210.7780.782
en→zh2BCER0.8250.0614 / 0.0000.0372 / 0.0000.6300.623
en→zh9BCER0.8000.0463 / 0.0000.0485 / 0.0000.6320.623
en→es2BWER0.8050.0867 / 0.0700.0702 / 0.0000.7320.735
en→es9BWER0.8200.0819 / 0.0000.0660 / 0.0000.7360.729
en→ja2BKata CER0.7900.0342 / 0.0000.0345 / 0.0000.7070.698
en→ja9BKata CER0.8030.0347 / 0.0000.0337 / 0.0000.7050.710

The table includes both Index-Echo sizes, with this card's 2B translator rows emphasized. Every pipeline uses the same translator size and text translation as its E2E counterpart. Bold content-error values mark the lower mean within each matched row. Content error compares synthesized speech with the intended translated text using WER for English and Spanish, normalized CER for Chinese, and katakana CER for Japanese. These are different metrics and test sets; do not compare their values directly across target languages. Text translation is judged by Gemini-3.1-Pro and is shared within each pair. Speaker similarity is cosine similarity against the source speaker, rather than a guarantee of perfect voice preservation.

Chinese-source tests contain 100 examples for English and 200 each for Spanish and Japanese. English-source tests start with 200 examples per target, with 146–198 retained after filtering in the reported evaluation. Speaker similarity differs by 0.021 for Chinese → English and by at most 0.01 elsewhere in the complete matched table. These results use the report's stated protocol and are separate from the website's deployed-system comparison or earlier Chinese-source evaluations. See the technical report for details.

Earlier Chinese-source S2ST evaluation (2B)

The report also records an external comparison for the 2B model on an earlier Chinese-source test set. Its data and protocol differ from the matched study above; the scores should be read within this comparison.

TargetIndex-Echo-2B S2ST MT judge ↑SeamlessM4T-v2 MT judge ↑
English0.9050.370
Spanish0.8230.343
Japanese0.8380.310

SeamlessM4T-v2 is an external speech-translation baseline; this comparison does not establish equal total parameter counts. The report does not provide 9B results for this earlier test. The current 2B/9B comparison is the matched six-direction table above.

Package details

File or directoryPurpose
config.json, modeling_dubbing.pyComponent configuration and DubbingBridgeModel interface
stlm_llm/Qwen3.5-family translator decoder and tokenizer
stlm_ckpt/Trained audio encoder/connector checkpoint
stlm_omni/Qwen3-Omni configuration and AuT source components
cosyvoice3/CosyVoice3 LLM, flow, HiFT, and ONNX assets
bridge/mapper.safetensors, mapper_config.json, and checkpoint archive
wetext_en_tn/, wetext_repo/Offline text-normalization assets
code/tts/, code/stlm/, code/bridge/Dubbing, speech translation, and mapper implementation
code/ja_ext/, code/cosyvoice_repo/Vendored Japanese frontend and CosyVoice inference code
_legacy/Original pre-conversion checkpoint archives, when present
samples/, verify_export.pyInput clips, reference dubs, and export/self-check assets

The loader prefers safetensors for the translator, mapper, and CosyVoice3 components, with original-format fallback. The export records version v4.0.0-fulldir-2b (2026-09-27), translator EXP-53-2B iter12613, and mapper rft_fd2b53_sw_lam20/mapper_best. Keep bundled assets together for offline loading. Use materialized copies (cp -rL or an appropriate rsync configuration) when copying training-side packages whose weight files are hard links.

Text normalization and alignment

DirectionTranslator text used for alignmentSpeech-frontend textHidden-state alignment
zh→enNormalized EnglishSame normalized textCharacter-overlap mean pooling
zh→esNormalized SpanishSame normalized textCharacter-overlap mean pooling
zh→jaOriginal mixed-script JapaneseKatakana frontend representationWord mapping to original text, then pooling
en→zhOriginal translationWetext-normalized ChineseCharacter-overlap mean pooling
en→esOriginal translationNormalized SpanishCharacter-overlap mean pooling
en→jaOriginal mixed-script JapaneseKatakana frontend representationWord mapping to original text, then pooling

The source transcript is the first nonempty line and the translation is the last nonempty line of translator output. The mapper's training language list is ['en', 'es', 'ja', 'zh', 'ja', 'es'], with first-occurrence indices en=0, es=1, ja=2, and zh=3; the duplicate trailing slots are unused. The loader asserts this mapping. Do not reorder it when modifying or converting the package.

Limitations

  • Use one CUDA GPU with at least 12 GB VRAM. CPU inference is unverified; memory use depends on the runtime and input.
  • The release is tuned for utterance-level dubbing. Keep individual utterances at or below approximately 30 seconds; use the full video pipeline for longer recordings. chunk=True raises NotImplementedError in this release.
  • The release notes report occasional failures on roughly 1–2% of tested short English-source inputs, including a CosyVoice convolution-kernel error or a one-token collapse producing approximately 0.04 seconds of audio. Check the generated audio and retry through the lower-level pipeline with a different seed, or discard the failed result.
  • Translation errors propagate into speech. Inspect tgt_raw and tgt_cv when fidelity is important.
  • Source-voice conditioning does not guarantee identical timbre, pronunciation, or delivery. The source-language heuristic can misclassify atypical or mixed-language transcripts.
  • Reported reward-optimization gains for English-source synthesis are small and direction-dependent; quality still depends strongly on the underlying translator and speech generator.
FamilyTaskReleased sizes
Index-TranslateText translation and translation instructions across 150 languages2B, 9B, 35B-A3B (preview)
Index-Echo S2TTSpeech-to-text translation and subtitles2B, 9B
Index-Echo S2STSpeech-to-speech translation with source-voice conditioning2B, 9B
Index-HomuraTranslation with a target syllable count2B, 9B
Index-NativeLongNative long-document translation; released as Index-Nailong2B, 9B

The text foundation's 150-language coverage does not describe the released speech interfaces. Use the task-specific directions documented above.

Citation

@techreport{indextranslate2026,
  author={Tianjiao Li and Mengran Yu and Chenyu Shi and Lusheng Zhang and
          Qisi Chen and Yanshan Zhou and Ji Qi and Jingying Liu and
          Yuang Feng and Ziang Cui and Tianxing Yan},
  title={Index-Translate: A Multilingual Translation Model Family --- Text, Speech, Controlled Dubbing, and Long-Document Translation},
  institution={Index LLM Team},
  year={2026},
  month={September}
}

License and feedback

Apache-2.0. Questions and feedback are welcome through GitHub Issues.

dubbing
dubbing_bridge
index
onnx
safetensors
speech-translation

IndexTeam/Index-Echo-S2ST-2B

Model

Index-Echo-S2ST-2B

5

12 commits

1 linked in READMEs

updated Sep 30, 2026

See the code

README

Index-Echo-S2ST-2B

Online demo · GitHub · Technical report · Hugging Face collection · ModelScope collection

Index-Echo-S2ST-2B translates a Chinese or English speech clip into another language and generates speech conditioned on the source speaker's voice. It combines a frozen Index-Echo-S2TT-2B backbone, a learned Hidden2CV mapper, and CosyVoice3 speech generation in one self-contained package and Python process.

The release supports six directions: Chinese → English/Spanish/Japanese and English → Chinese/Spanish/Japanese. The caller specifies the target language; the package infers the source language from its source transcript. The 9B sibling exposes the same interface. 2B denotes the speech translator's decoder size, rather than the total size of all packaged components.

Architecture

Index-Echo speech-to-text and speech-to-speech architecture

The report's lower path extends S2TT to S2ST: the frozen speech translator supplies its final hidden states to Hidden2CV, which connects to CosyVoice3 speech generation. Source audio supplies the reference prompt and CampPlus speaker embedding for voice conditioning. The diagram also shows mapper alignment and DiffRO training; the S2TT backbone remains frozen throughout S2ST training.

Model and training

The technical report initializes S2ST from the trained Index-Echo S2TT model, whose Qwen3-Omni AuT audio encoder and connector feed an Index-Translate decoder. The speech-generation path uses a roughly 30M-parameter Hidden2CV mapper to connect the translator's final hidden states to the CosyVoice3 semantic layer.

Training first distills the mapper with both the S2TT backbone and speech generator frozen, then applies differentiable reward optimization (DiffRO). A Token2Text content-consistency reward is backpropagated through Gumbel-Softmax speech-token samples. The S2TT backbone remains frozen throughout S2ST training. At inference, source audio supplies the reference prompt and CampPlus speaker embedding for voice conditioning. Voice similarity varies with the source, target language, and synthesis result.

Inference

Download the complete model package before importing its local modeling_dubbing.py:

pip install -U huggingface_hub
hf download IndexTeam/Index-Echo-S2ST-2B --local-dir ./Index-Echo-S2ST-2B
cd Index-Echo-S2ST-2B

Pinned environment

The release documents the following tested dependency pins. Use CUDA-compatible PyTorch wheels; the +cu129 builds require a wheel source that provides that CUDA build. System ffmpeg must be on PATH.

pip install torch==2.11.0+cu129 torchaudio==2.11.0+cu129 transformers==5.6.0 \
    tokenizers==0.22.2 librosa==1.0.0 onnxruntime==1.30.0 wetext==0.0.4 \
    kaldifst==1.8.0 conformer==0.3.2 hydra-core==1.3.7 HyperPyYAML==1.2.3 \
    lightning==2.6.6 x-transformers==2.11.24 einx==0.4.3 frozendict==2.4.7 \
    pyworld==0.3.5 soundfile==0.13.1 modelscope==1.37.1 safetensors==0.7.0 \
    numpy==2.4.6 openai-whisper==20250625

openai-whisper is used only for the package's ASR-based evaluation/self-check. The Japanese frontend (pyopenjtalk==0.4.1, including its dictionary, and pykakasi==2.3.0) is vendored under code/ja_ext/; the pipeline adds it to the import path automatically. All components run in one Python environment and process.

The package automatically applies two CosyVoice3 compatibility patches for Transformers 5.6 in code/tts/pipeline.py: keeping the CosyVoice3 LLM in fp32, and providing the full visible attention mask during autoregressive decoding. Keep the bundled code and dependency pins together.

Python API

Run this from the downloaded package directory, with input clips available at the indicated paths:

from modeling_dubbing import DubbingBridgeModel

model = DubbingBridgeModel.from_pretrained('.', device='cuda')

# Chinese source: three supported targets.
wav, sr = model.dub('input_zh.wav', lang='en', out_wav='dub_en.wav')
model.dub('input_zh.wav', lang='es', out_wav='dub_es.wav')
model.dub('input_zh.wav', lang='ja', out_wav='dub_ja.wav')

# English source: three supported targets.
model.dub('input_en.wav', lang='zh', out_wav='dub_zh.wav')
model.dub('input_en.wav', lang='es', out_wav='dub_es2.wav')
model.dub('input_en.wav', lang='ja', out_wav='dub_ja2.wav')

wav, sr, info = model.dub('input_en.wav', lang='zh', return_info=True)
print(sr)               # 24000 Hz
print(info['zh'])       # Source transcript; the key name is retained for both sources.
print(info['tgt_raw'])  # Model translation, before the speech frontend's normalization.
print(info['tgt_cv'])   # Speech-frontend text; Japanese uses a katakana representation.
print(info['src_lang']) # 'zh' or 'en'

lang always means the target language: en, es, ja, or zh. The source-language heuristic treats a source transcript containing CJK characters as Chinese and otherwise as English; it is not a general language identifier. Source and target must differ. The waveform is a float32 tensor of shape (1, T) at 24 kHz. out_wav saves it to a file. With return_info=True, the third return value includes the transcript, translation, source language, and synthesis diagnostics.

At the lower pipeline layer, extract(audio_path, lang) returns the translated text and aligned hidden states separately from synthesis. See the S2ST inference guide for the dub.py wrapper, and the video pipeline for segmentation and video timeline alignment.

Default inference settings

Stage or settingReleased default
Transcription and translationGreedy, temperature=0.0 / do_sample=False
Transcription/translation budgetmax_new_tokens=1024
Speech modeh2cv (the only supported mode in this package)
Speech-token budgetmax_tokens=1500; actual cap min(20 * m, 1500), where m is aligned target-text token count
Minimum speech length2 * m tokens before stop tokens are allowed
Speech samplingsampling=25; release notes describe ras_sampling, top-k 25 / top-p 0.8
Speech-generation seed42
Synthesis speed1.0
Output sample rate24000 Hz
CosyVoice3 executionstream=False, load_trt=False, load_vllm=False, fp16=False; LLM kept in fp32
Chunked orchestrationDisabled (chunk=False)

Both sizes use these settings. Greedy transcription/translation does not make speech synthesis greedy. The public DubbingBridgeModel.dub wrapper does not expose a seed argument; seed changes are available through the lower-level pipeline's dub/synth APIs. Sampling may vary between calls even with the default seed, as documented in the release notes.

Configurable API parameters

The public wrapper and lower pipeline expose different controls. The defaults below are shared by both sizes.

InterfaceParameterDefaultBehavior
from_pretrainedmodel_dirRequiredLocal directory containing the full package
from_pretraineddevicecudaCUDA device for the loaded components
from_pretrainedmodeNone → config default h2cvSets the default speech mode
Public dubaudio_pathRequiredChinese or English source clip
Public dublangenTarget language: en, es, ja, zh; must differ from detected source
Public dubout_wavNoneSave path; None creates a temporary WAV
Public dubmodeNone → loaded defaultOnly h2cv is accepted
Public dubchunkFalseTrue is unsupported in this release
Public dubreturn_infoFalseAdds transcript, translation, and synthesis diagnostics to the return value
Pipeline extractlangenTarget language for transcription/translation and alignment
Pipeline extractmax_new_tokens1024Transcription/translation output budget
Pipeline synthout_wavNoneOptional output save path
Pipeline synthmax_tokens1500Upper bound on generated speech tokens
Pipeline synthseed42Seed set before speech-token generation and acoustic synthesis
Pipeline synthprompt_wavNoneUses the source audio unless a reference clip is supplied
Pipeline synthprompt_textNoneUses the source transcript unless reference text is supplied

For explicit text/speech budgets or seed control, use the lower pipeline after loading model as above:

pipe = model._pipe
ext = pipe.extract('input_zh.wav', lang='en', max_new_tokens=1024)
wav, info = pipe.synth(ext, out_wav='dub_en.wav', max_tokens=1500, seed=42)

The lower pipeline's dub(..., **kw) forwards synthesis options such as seed and max_tokens; changing the translation budget requires calling extract separately. Sampling argument 25 and synthesis speed 1.0 are fixed in the released pipeline code, not keyword arguments of the public wrapper. Translation stops at tokenizer EOS or <|im_end|>; speech generation uses CosyVoice3's defined stop tokens and the length limits above. The script does not explicitly override other text-generation knobs such as top_p, top_k, or repetition penalties.

Evaluation

The report compares end-to-end dubbing with a matched Index-Echo-S2TT + CosyVoice3 pipeline on in-house video dubbing data. Each pair shares the same frozen S2TT model and text translation, so the comparison measures the speech-generation path under a shared translation. For 2B, end-to-end dubbing lowers mean content error in four of six directions: zh→es, zh→ja, en→zh, and en→es. Across both sizes it improves eight of twelve size–direction pairs; Chinese → English favors the pipeline.

DirectionTranslator sizeContent metricText judge ↑Pipeline error ↓
mean / median
E2E error ↓
mean / median
Pipeline speaker ↑E2E speaker ↑
zh→en2BWER0.8300.0450 / 0.0000.0709 / 0.0290.7220.743
zh→en9BWER0.8400.0477 / 0.0000.0615 / 0.0000.7180.739
zh→es2BWER0.7280.0811 / 0.0700.0704 / 0.0510.7470.743
zh→es9BWER0.7930.0916 / 0.0710.0773 / 0.0570.7460.746
zh→ja2BKata CER0.7600.0575 / 0.0340.0493 / 0.0270.7770.780
zh→ja9BKata CER0.8150.0552 / 0.0290.0436 / 0.0210.7780.782
en→zh2BCER0.8250.0614 / 0.0000.0372 / 0.0000.6300.623
en→zh9BCER0.8000.0463 / 0.0000.0485 / 0.0000.6320.623
en→es2BWER0.8050.0867 / 0.0700.0702 / 0.0000.7320.735
en→es9BWER0.8200.0819 / 0.0000.0660 / 0.0000.7360.729
en→ja2BKata CER0.7900.0342 / 0.0000.0345 / 0.0000.7070.698
en→ja9BKata CER0.8030.0347 / 0.0000.0337 / 0.0000.7050.710

The table includes both Index-Echo sizes, with this card's 2B translator rows emphasized. Every pipeline uses the same translator size and text translation as its E2E counterpart. Bold content-error values mark the lower mean within each matched row. Content error compares synthesized speech with the intended translated text using WER for English and Spanish, normalized CER for Chinese, and katakana CER for Japanese. These are different metrics and test sets; do not compare their values directly across target languages. Text translation is judged by Gemini-3.1-Pro and is shared within each pair. Speaker similarity is cosine similarity against the source speaker, rather than a guarantee of perfect voice preservation.

Chinese-source tests contain 100 examples for English and 200 each for Spanish and Japanese. English-source tests start with 200 examples per target, with 146–198 retained after filtering in the reported evaluation. Speaker similarity differs by 0.021 for Chinese → English and by at most 0.01 elsewhere in the complete matched table. These results use the report's stated protocol and are separate from the website's deployed-system comparison or earlier Chinese-source evaluations. See the technical report for details.

Earlier Chinese-source S2ST evaluation (2B)

The report also records an external comparison for the 2B model on an earlier Chinese-source test set. Its data and protocol differ from the matched study above; the scores should be read within this comparison.

TargetIndex-Echo-2B S2ST MT judge ↑SeamlessM4T-v2 MT judge ↑
English0.9050.370
Spanish0.8230.343
Japanese0.8380.310

SeamlessM4T-v2 is an external speech-translation baseline; this comparison does not establish equal total parameter counts. The report does not provide 9B results for this earlier test. The current 2B/9B comparison is the matched six-direction table above.

Package details

File or directoryPurpose
config.json, modeling_dubbing.pyComponent configuration and DubbingBridgeModel interface
stlm_llm/Qwen3.5-family translator decoder and tokenizer
stlm_ckpt/Trained audio encoder/connector checkpoint
stlm_omni/Qwen3-Omni configuration and AuT source components
cosyvoice3/CosyVoice3 LLM, flow, HiFT, and ONNX assets
bridge/mapper.safetensors, mapper_config.json, and checkpoint archive
wetext_en_tn/, wetext_repo/Offline text-normalization assets
code/tts/, code/stlm/, code/bridge/Dubbing, speech translation, and mapper implementation
code/ja_ext/, code/cosyvoice_repo/Vendored Japanese frontend and CosyVoice inference code
_legacy/Original pre-conversion checkpoint archives, when present
samples/, verify_export.pyInput clips, reference dubs, and export/self-check assets

The loader prefers safetensors for the translator, mapper, and CosyVoice3 components, with original-format fallback. The export records version v4.0.0-fulldir-2b (2026-09-27), translator EXP-53-2B iter12613, and mapper rft_fd2b53_sw_lam20/mapper_best. Keep bundled assets together for offline loading. Use materialized copies (cp -rL or an appropriate rsync configuration) when copying training-side packages whose weight files are hard links.

Text normalization and alignment

DirectionTranslator text used for alignmentSpeech-frontend textHidden-state alignment
zh→enNormalized EnglishSame normalized textCharacter-overlap mean pooling
zh→esNormalized SpanishSame normalized textCharacter-overlap mean pooling
zh→jaOriginal mixed-script JapaneseKatakana frontend representationWord mapping to original text, then pooling
en→zhOriginal translationWetext-normalized ChineseCharacter-overlap mean pooling
en→esOriginal translationNormalized SpanishCharacter-overlap mean pooling
en→jaOriginal mixed-script JapaneseKatakana frontend representationWord mapping to original text, then pooling

The source transcript is the first nonempty line and the translation is the last nonempty line of translator output. The mapper's training language list is ['en', 'es', 'ja', 'zh', 'ja', 'es'], with first-occurrence indices en=0, es=1, ja=2, and zh=3; the duplicate trailing slots are unused. The loader asserts this mapping. Do not reorder it when modifying or converting the package.

Limitations

  • Use one CUDA GPU with at least 12 GB VRAM. CPU inference is unverified; memory use depends on the runtime and input.
  • The release is tuned for utterance-level dubbing. Keep individual utterances at or below approximately 30 seconds; use the full video pipeline for longer recordings. chunk=True raises NotImplementedError in this release.
  • The release notes report occasional failures on roughly 1–2% of tested short English-source inputs, including a CosyVoice convolution-kernel error or a one-token collapse producing approximately 0.04 seconds of audio. Check the generated audio and retry through the lower-level pipeline with a different seed, or discard the failed result.
  • Translation errors propagate into speech. Inspect tgt_raw and tgt_cv when fidelity is important.
  • Source-voice conditioning does not guarantee identical timbre, pronunciation, or delivery. The source-language heuristic can misclassify atypical or mixed-language transcripts.
  • Reported reward-optimization gains for English-source synthesis are small and direction-dependent; quality still depends strongly on the underlying translator and speech generator.
FamilyTaskReleased sizes
Index-TranslateText translation and translation instructions across 150 languages2B, 9B, 35B-A3B (preview)
Index-Echo S2TTSpeech-to-text translation and subtitles2B, 9B
Index-Echo S2STSpeech-to-speech translation with source-voice conditioning2B, 9B
Index-HomuraTranslation with a target syllable count2B, 9B
Index-NativeLongNative long-document translation; released as Index-Nailong2B, 9B

The text foundation's 150-language coverage does not describe the released speech interfaces. Use the task-specific directions documented above.

Citation

@techreport{indextranslate2026,
  author={Tianjiao Li and Mengran Yu and Chenyu Shi and Lusheng Zhang and
          Qisi Chen and Yanshan Zhou and Ji Qi and Jingying Liu and
          Yuang Feng and Ziang Cui and Tianxing Yan},
  title={Index-Translate: A Multilingual Translation Model Family --- Text, Speech, Controlled Dubbing, and Long-Document Translation},
  institution={Index LLM Team},
  year={2026},
  month={September}
}

License and feedback

Apache-2.0. Questions and feedback are welcome through GitHub Issues.

dubbing
dubbing_bridge
index
onnx
safetensors
speech-translation