IndexTeam/Index-Homura-2B

Model

Index-Homura-2B

5

10 commits

1 linked in READMEs

updated Sep 30, 2026

See the code

README

Index-Homura-2B

Online demo · GitHub · Technical report · Hugging Face collection · ModelScope collection

Index-Homura-2B is the 2B syllable-controlled translation specialist in the Index-Translate family. It takes source text together with an explicit target syllable count, then adjusts the translation's wording toward that count while retaining meaning and natural expression. This is useful when preparing translated lines for dubbing or timed subtitles.

On SandGlass, the released RL checkpoint achieves 63.08% within 10% of the target count, 54.39% within one syllable, and a separate translation-quality score of 0.7615. Length adherence and translation quality measure different aspects of the task; the count is an approximate control objective.

Training and task

The family builds on Qwen3.5 with multilingual mid-training and translation post-training. Index-Homura adapts the 2B text foundation to an explicitly specified target count. Following the HOMURA reinforcement-learning approach described in the technical report, GRPO jointly optimizes translation quality and a length reward measuring deviation from the requested syllable count.

Use this specialist for explicit count control. The general Index-Translate models also follow subtitle-related instructions, but their source-relative subtitle constraint is a separate task. SandGlass evaluates English, Japanese, Arabic, and Spanish targets; these results do not establish equal control precision across every language in the general text model's 150-language inventory.

Inference

Use a CUDA GPU and a vLLM build with Qwen3.5 support; the official guide records testing with vLLM 0.29. The bf16 memory guide is approximately 8 GB for 2B and 24 GB for 9B, plus the KV cache. The released client uses an OpenAI-compatible chat-completions endpoint.

git clone https://github.com/bilibili/Index-Translate.git
cd Index-Translate
pip install -U vllm
pip install -r inference/llm/requirements.txt
bash inference/llm/serve_vllm.sh homura-2b --host 127.0.0.1 --port 8000

From the repository directory in another terminal:

python inference/llm/syllable_translate.py \
  "说到底聊天群的规则一句话就能总结" \
  --syllables 14 --target en --model IndexTeam/Index-Homura-2B

The official client sends this single user prompt:

请将以下文本翻译为英语,译文严格控制在 14 个音节。直接输出翻译结果,不要进行任何解释。

说到底聊天群的规则一句话就能总结

Released client defaults

SettingDefault
Temperaturetemperature=0.3; exposed as --temperature
ThinkingDisabled through chat_template_kwargs={"enable_thinking": False}
Output budgetmax_tokens=max(512, 3 * len(text)) after stripping the input text
Top-pNot explicitly set by this client; use the inference server's default
Top-k / min-pNot explicitly set by this client; use the inference server's defaults
Presence / frequency / repetition penaltiesNot explicitly set by this client; use the inference server's defaults
SeedNot explicitly set by this client; use the inference server's default
Stop strings / stop token IDs / EOS handlingNot explicitly set by this client; use the model and server defaults
StreamingNot requested; the client prints the returned translation after completion
Chat messagesOne user message with the target language and explicit syllable count

Complete client arguments

These are the actual defaults in syllable_translate.py, shared by both checkpoint sizes. The example above selects this card's checkpoint explicitly.

ArgumentDefault and behavior
textOptional positional text; reads standard input if omitted, strips outer whitespace, rejects empty input
--syllables, -nRequired integer target syllable count; no default
--target, -ten; mapped language codes become Chinese language names in the prompt, other values are used as supplied
--model, -mINDEX_MODEL environment variable, otherwise IndexTeam/Index-Homura-9B; pass IndexTeam/Index-Homura-2B to select this checkpoint
--base-urlOPENAI_BASE_URL environment variable, otherwise http://127.0.0.1:8000/v1
--api-keyOPENAI_API_KEY environment variable, otherwise EMPTY
--max-tokens0; a positive value overrides the adaptive output budget
--temperature0.3
--help, -hPrints the command-line help

The serving preset uses --max-model-len 32768 and --served-model-name IndexTeam/Index-Homura-2B; the example binds the server to 127.0.0.1:8000. Extra preset arguments pass through to vllm serve, including deployment options such as --tensor-parallel-size. The client bypasses environment proxies for localhost endpoints.

len(text) is the Python character count used to choose the output-token budget. The syllable target must be supplied separately with --syllables; it is distinct from max_tokens. Override the output limit with --max-tokens, or sampling with --temperature, if needed. The same defaults apply to both Homura sizes. The temperature-0 setting for Index-Translate text models does not change Homura's released temperature of 0.3.

See the inference guide and syllable client for endpoint options and full usage.

SandGlass evaluation

SandGlass contains 300 subtitle sentences, with 60 each from animation, film and television, travel, gaming, and knowledge. Each sentence is translated into four target languages at three requested lengths, yielding 3,600 cases per model. The central count comes from subtitle duration and a language-specific speaking rate; short and long budgets scale it by 0.75 and 1.25.

The report evaluates local models with greedy decoding and a 512-token output limit. Those benchmark settings differ from the released client's default temperature of 0.3 and adaptive output budget above. API baselines use their recorded settings; GPT-5.6-Sol uses a gateway configuration that may inject additional system prompts.

Translation quality is a reference-free Gemini-2.5-Flash judgment on a 0/0.5/1 scale, averaged across cases. Relative deviation is the absolute syllable-count error divided by the target count; lower is better. The two adherence columns report the proportions within one syllable or 10% of the target. Regression slope compares output counts with target counts across the three-budget groups; a value close to 1 indicates a stronger response to changes in the requested length.

The comparison includes both Index-Homura sizes and their SFT-stage ablations, same-size Qwen3.5 2B/9B baselines, nearby-size Hy-MT2 1.8B/7B and Hunyuan-MT-7B models, and larger/API systems. Every row uses the same SandGlass task and metrics.

ModelTranslation quality ↑Mean relative deviation ↓Within ±1 syllable ↑Within 10% ↑Slope ≈ 1
Index-Homura-9B (RL)0.78630.069374.42%81.92%0.968
Index-Homura-9B (SFT)0.85810.175241.39%45.75%0.680
Index-Homura-2B (RL; this checkpoint)0.76150.096254.39%63.08%0.947
Index-Homura-2B (SFT)0.79650.159437.89%43.06%0.672
GPT-5.6-Sol (low)0.87150.322342.53%47.64%0.891
DeepSeek-V4-Flash (no thinking)0.87420.336816.58%16.36%0.153
Hy-MT2-7B0.85600.317418.14%18.86%0.195
Hy-MT2-30B-A3B0.75540.274320.97%22.39%0.340
Hy-MT2-1.8B0.59040.708014.58%16.72%0.222
Hunyuan-MT-7B0.71881.057912.31%15.11%0.031
Qwen3.5-9B0.71830.324715.61%16.39%0.430
Qwen3.5-35B-A3B0.78630.346013.31%13.25%0.249
Qwen3.5-2B0.49170.515112.50%12.25%0.036

Source: the report's expanded SandGlass evaluation and target-language breakdown; full evaluation tables.

Limits and practical use

Syllable control is approximate. Verify the count in the generated line, especially before recording or synthesizing speech. Exact count does not guarantee exact duration, since pronunciation, pauses, and delivery affect timing.

The report shows a quality–control trade-off: the external GPT and DeepSeek baselines score higher on translation quality while Homura achieves stronger count adherence. In a separate 2B reward-weight experiment, increasing the syllable reward raises the within-10% rate from 63.08% to 74.19% while reducing quality from 0.7615 to 0.6900. That ablation is separate from the released 2B checkpoint's row above.

SandGlass release is listed as planned in the GitHub repository; the report provides the protocol and results. Individual demo examples that hit their targets illustrate possible outputs rather than guarantees for new requests.

Model family

ModelReleased checkpointsTask
Index-Translate2B · 9B · 35B-A3B (preview)Text translation and translation instructions across 150 languages
Index-Echo S2TT2B · 9BSpeech-to-text translation
Index-Echo S2ST2B · 9BSpeech-to-speech translation with voice conditioning
Index-Homura2B · 9BTranslation with a target syllable count
Index-NativeLong2B · 9BFull-document translation; released templates support zh↔en and zh↔ja

The 150-language coverage refers to the Index-Translate text models. Speech and long-document packages have their own language interfaces. NativeLong retains the IndexTeam/Index-Nailong-* repository IDs.

Citation

@techreport{indextranslate2026,
  author={Tianjiao Li and Mengran Yu and Chenyu Shi and Lusheng Zhang and
          Qisi Chen and Yanshan Zhou and Ji Qi and Jingying Liu and
          Yuang Feng and Ziang Cui and Tianxing Yan},
  title={Index-Translate: A Multilingual Translation Model Family --- Text, Speech, Controlled Dubbing, and Long-Document Translation},
  institution={Index LLM Team},
  year={2026},
  month={September}
}

License and feedback

Apache-2.0. Please use GitHub Issues for questions and feedback.

dubbing
index
qwen3_5
safetensors
translation

IndexTeam/Index-Homura-2B

Model

Index-Homura-2B

5

10 commits

1 linked in READMEs

updated Sep 30, 2026

See the code

README

Index-Homura-2B

Online demo · GitHub · Technical report · Hugging Face collection · ModelScope collection

Index-Homura-2B is the 2B syllable-controlled translation specialist in the Index-Translate family. It takes source text together with an explicit target syllable count, then adjusts the translation's wording toward that count while retaining meaning and natural expression. This is useful when preparing translated lines for dubbing or timed subtitles.

On SandGlass, the released RL checkpoint achieves 63.08% within 10% of the target count, 54.39% within one syllable, and a separate translation-quality score of 0.7615. Length adherence and translation quality measure different aspects of the task; the count is an approximate control objective.

Training and task

The family builds on Qwen3.5 with multilingual mid-training and translation post-training. Index-Homura adapts the 2B text foundation to an explicitly specified target count. Following the HOMURA reinforcement-learning approach described in the technical report, GRPO jointly optimizes translation quality and a length reward measuring deviation from the requested syllable count.

Use this specialist for explicit count control. The general Index-Translate models also follow subtitle-related instructions, but their source-relative subtitle constraint is a separate task. SandGlass evaluates English, Japanese, Arabic, and Spanish targets; these results do not establish equal control precision across every language in the general text model's 150-language inventory.

Inference

Use a CUDA GPU and a vLLM build with Qwen3.5 support; the official guide records testing with vLLM 0.29. The bf16 memory guide is approximately 8 GB for 2B and 24 GB for 9B, plus the KV cache. The released client uses an OpenAI-compatible chat-completions endpoint.

git clone https://github.com/bilibili/Index-Translate.git
cd Index-Translate
pip install -U vllm
pip install -r inference/llm/requirements.txt
bash inference/llm/serve_vllm.sh homura-2b --host 127.0.0.1 --port 8000

From the repository directory in another terminal:

python inference/llm/syllable_translate.py \
  "说到底聊天群的规则一句话就能总结" \
  --syllables 14 --target en --model IndexTeam/Index-Homura-2B

The official client sends this single user prompt:

请将以下文本翻译为英语,译文严格控制在 14 个音节。直接输出翻译结果,不要进行任何解释。

说到底聊天群的规则一句话就能总结

Released client defaults

SettingDefault
Temperaturetemperature=0.3; exposed as --temperature
ThinkingDisabled through chat_template_kwargs={"enable_thinking": False}
Output budgetmax_tokens=max(512, 3 * len(text)) after stripping the input text
Top-pNot explicitly set by this client; use the inference server's default
Top-k / min-pNot explicitly set by this client; use the inference server's defaults
Presence / frequency / repetition penaltiesNot explicitly set by this client; use the inference server's defaults
SeedNot explicitly set by this client; use the inference server's default
Stop strings / stop token IDs / EOS handlingNot explicitly set by this client; use the model and server defaults
StreamingNot requested; the client prints the returned translation after completion
Chat messagesOne user message with the target language and explicit syllable count

Complete client arguments

These are the actual defaults in syllable_translate.py, shared by both checkpoint sizes. The example above selects this card's checkpoint explicitly.

ArgumentDefault and behavior
textOptional positional text; reads standard input if omitted, strips outer whitespace, rejects empty input
--syllables, -nRequired integer target syllable count; no default
--target, -ten; mapped language codes become Chinese language names in the prompt, other values are used as supplied
--model, -mINDEX_MODEL environment variable, otherwise IndexTeam/Index-Homura-9B; pass IndexTeam/Index-Homura-2B to select this checkpoint
--base-urlOPENAI_BASE_URL environment variable, otherwise http://127.0.0.1:8000/v1
--api-keyOPENAI_API_KEY environment variable, otherwise EMPTY
--max-tokens0; a positive value overrides the adaptive output budget
--temperature0.3
--help, -hPrints the command-line help

The serving preset uses --max-model-len 32768 and --served-model-name IndexTeam/Index-Homura-2B; the example binds the server to 127.0.0.1:8000. Extra preset arguments pass through to vllm serve, including deployment options such as --tensor-parallel-size. The client bypasses environment proxies for localhost endpoints.

len(text) is the Python character count used to choose the output-token budget. The syllable target must be supplied separately with --syllables; it is distinct from max_tokens. Override the output limit with --max-tokens, or sampling with --temperature, if needed. The same defaults apply to both Homura sizes. The temperature-0 setting for Index-Translate text models does not change Homura's released temperature of 0.3.

See the inference guide and syllable client for endpoint options and full usage.

SandGlass evaluation

SandGlass contains 300 subtitle sentences, with 60 each from animation, film and television, travel, gaming, and knowledge. Each sentence is translated into four target languages at three requested lengths, yielding 3,600 cases per model. The central count comes from subtitle duration and a language-specific speaking rate; short and long budgets scale it by 0.75 and 1.25.

The report evaluates local models with greedy decoding and a 512-token output limit. Those benchmark settings differ from the released client's default temperature of 0.3 and adaptive output budget above. API baselines use their recorded settings; GPT-5.6-Sol uses a gateway configuration that may inject additional system prompts.

Translation quality is a reference-free Gemini-2.5-Flash judgment on a 0/0.5/1 scale, averaged across cases. Relative deviation is the absolute syllable-count error divided by the target count; lower is better. The two adherence columns report the proportions within one syllable or 10% of the target. Regression slope compares output counts with target counts across the three-budget groups; a value close to 1 indicates a stronger response to changes in the requested length.

The comparison includes both Index-Homura sizes and their SFT-stage ablations, same-size Qwen3.5 2B/9B baselines, nearby-size Hy-MT2 1.8B/7B and Hunyuan-MT-7B models, and larger/API systems. Every row uses the same SandGlass task and metrics.

ModelTranslation quality ↑Mean relative deviation ↓Within ±1 syllable ↑Within 10% ↑Slope ≈ 1
Index-Homura-9B (RL)0.78630.069374.42%81.92%0.968
Index-Homura-9B (SFT)0.85810.175241.39%45.75%0.680
Index-Homura-2B (RL; this checkpoint)0.76150.096254.39%63.08%0.947
Index-Homura-2B (SFT)0.79650.159437.89%43.06%0.672
GPT-5.6-Sol (low)0.87150.322342.53%47.64%0.891
DeepSeek-V4-Flash (no thinking)0.87420.336816.58%16.36%0.153
Hy-MT2-7B0.85600.317418.14%18.86%0.195
Hy-MT2-30B-A3B0.75540.274320.97%22.39%0.340
Hy-MT2-1.8B0.59040.708014.58%16.72%0.222
Hunyuan-MT-7B0.71881.057912.31%15.11%0.031
Qwen3.5-9B0.71830.324715.61%16.39%0.430
Qwen3.5-35B-A3B0.78630.346013.31%13.25%0.249
Qwen3.5-2B0.49170.515112.50%12.25%0.036

Source: the report's expanded SandGlass evaluation and target-language breakdown; full evaluation tables.

Limits and practical use

Syllable control is approximate. Verify the count in the generated line, especially before recording or synthesizing speech. Exact count does not guarantee exact duration, since pronunciation, pauses, and delivery affect timing.

The report shows a quality–control trade-off: the external GPT and DeepSeek baselines score higher on translation quality while Homura achieves stronger count adherence. In a separate 2B reward-weight experiment, increasing the syllable reward raises the within-10% rate from 63.08% to 74.19% while reducing quality from 0.7615 to 0.6900. That ablation is separate from the released 2B checkpoint's row above.

SandGlass release is listed as planned in the GitHub repository; the report provides the protocol and results. Individual demo examples that hit their targets illustrate possible outputs rather than guarantees for new requests.

Model family

ModelReleased checkpointsTask
Index-Translate2B · 9B · 35B-A3B (preview)Text translation and translation instructions across 150 languages
Index-Echo S2TT2B · 9BSpeech-to-text translation
Index-Echo S2ST2B · 9BSpeech-to-speech translation with voice conditioning
Index-Homura2B · 9BTranslation with a target syllable count
Index-NativeLong2B · 9BFull-document translation; released templates support zh↔en and zh↔ja

The 150-language coverage refers to the Index-Translate text models. Speech and long-document packages have their own language interfaces. NativeLong retains the IndexTeam/Index-Nailong-* repository IDs.

Citation

@techreport{indextranslate2026,
  author={Tianjiao Li and Mengran Yu and Chenyu Shi and Lusheng Zhang and
          Qisi Chen and Yanshan Zhou and Ji Qi and Jingying Liu and
          Yuang Feng and Ziang Cui and Tianxing Yan},
  title={Index-Translate: A Multilingual Translation Model Family --- Text, Speech, Controlled Dubbing, and Long-Document Translation},
  institution={Index LLM Team},
  year={2026},
  month={September}
}

License and feedback

Apache-2.0. Please use GitHub Issues for questions and feedback.

dubbing
index
qwen3_5
safetensors
translation