IndexTeam/Index-Nailong-2B

Model

Index-NativeLong-2B

5

10 commits

1 linked in READMEs

updated Sep 30, 2026

See the code

README

Index-NativeLong-2B

Online demo · GitHub · Technical report · Hugging Face collection · ModelScope collection

Index-NativeLong-2B, published as IndexTeam/Index-Nailong-2B, is the 2B long-document translation specialist in the Index-Translate family. It receives the full source document and produces its translation in a single generation, keeping distant source passages and the translation history together. This supports content coverage, consistent names and terminology, and continuity across chapters or scenes.

The 2B checkpoint scores 0.6701 / 0.7300 / 0.8651 on GuoFeng / BWB / General Books. It exceeds the external models in the report at both 32K and 64K on BWB. The shipped serving configuration supports a 262,144-token context window, shared by the prompt, source, and generated translation. Use the published Nailong model ID in commands.

Training and supported tasks

The family builds on Qwen3.5 with multilingual mid-training and translation post-training. This 2B model follows 128K-sequence mid-training decay and general translation SFT before full-parameter document SFT. Document supervision consists of aligned book passages and complete target translations, with bidirectional training over approximately 4K–64K Chinese-side tokens. Training sequence length reaches 128K in the long-document adaptation; this is separate from the shipped serving context limit.

The technical report studies long-sequence decay and RoPE/NoPE positional-encoding choices in separate 2B ablations. After document SFT, their overall scores are similar, with strengths varying by length. These experiments describe training alternatives rather than a runtime switch in the released checkpoint.

The released document client supplies fixed prompt templates for Chinese→English, English→Chinese, Chinese→Japanese, and Japanese→Chinese (zh-en, en-zh, zh-ja, ja-zh). Its instructions require full coverage, consistent entity names and terminology, and preservation of chapter, paragraph, and content order. These four directions are the packaged interface; the general text model's 150-language coverage does not establish NativeLong performance in 150 languages.

Inference

Use a CUDA GPU and a vLLM build with Qwen3.5 support; the official guide records testing with vLLM 0.29. The bf16 model-memory guide is approximately 8 GB for 2B and 24 GB for 9B, plus KV cache for the chosen context length. Full-context serving can require substantially more memory or tensor parallelism.

git clone https://github.com/bilibili/Index-Translate.git
cd Index-Translate
pip install -U vllm
pip install -r inference/llm/requirements.txt
bash inference/llm/serve_vllm.sh nailong-2b --host 127.0.0.1 --port 8000

From the repository directory in another terminal, translate a UTF-8 document using the official direction template:

python inference/llm/doc_translate.py novel.zh.txt \
  --direction zh-en --model IndexTeam/Index-Nailong-2B \
  --output novel.en.txt

python inference/llm/doc_translate.py paper.en.txt \
  --direction en-zh --model IndexTeam/Index-Nailong-2B \
  --output paper.zh.txt

Replace the input filenames with your documents. For Japanese, use --direction zh-ja or --direction ja-zh. The client streams the translation into the output file and loads the prompt from inference/llm/prompts/nailong_<direction>.txt.

Released client defaults

SettingDefault
Temperaturetemperature=0 (greedy)
Top-p / top-k / min-ptop_p=1, top_k=-1, min_p=0
Seedseed=42
Presence / repetition penaltypresence_penalty=0, repetition_penalty=1
Frequency penaltyNot explicitly set by this client; use the inference server's default
Stop token IDsstop_token_ids=[248044, 248046]
Stop stringsNot explicitly set by this client; use the inference server's default
Ignore EOSignore_eos=False
ThinkingDisabled through chat_template_kwargs={"enable_thinking": False}
Output budgetmax_tokens omitted by default; generation uses the remaining context window
Streamingstream=True; writes each received content delta to the output
Chat messagesOne user message rendered from the fixed direction template and complete source
Serving context limit for this checkpoint262144 tokens

Complete client arguments

These are the actual defaults in doc_translate.py, shared by both checkpoint sizes. The example above selects this card's checkpoint explicitly. Sampling, stop-token, thinking, and streaming settings are fixed in the client; its command-line interface exposes the following controls:

ArgumentDefault and behavior
inputRequired UTF-8 source filename; - reads standard input; normalizes CRLF line endings, strips outer whitespace, rejects empty input
--output, -oStandard output if omitted; otherwise writes the streamed translation to the specified UTF-8 file
--direction, -dUnset; if supplied, selects one of zh-en, en-zh, zh-ja, ja-zh and takes priority over source/target flags
--source, -szh; used to derive the direction when --direction is omitted
--target, -ten; used with --source, so the default derived direction is zh-en
--model, -mINDEX_MODEL environment variable, otherwise IndexTeam/Index-Nailong-9B; pass IndexTeam/Index-Nailong-2B to select this checkpoint
--base-urlOPENAI_BASE_URL environment variable, otherwise http://127.0.0.1:8000/v1
--api-keyOPENAI_API_KEY environment variable, otherwise EMPTY
--max-tokens0; omits API max_tokens; a positive value sets an explicit output-token cap
--help, -hPrints the command-line help

The serving preset sets --served-model-name IndexTeam/Index-Nailong-2B and --max-model-len 262144; the example binds it to 127.0.0.1:8000. Extra preset arguments pass through to vllm serve, including deployment options such as --tensor-parallel-size. The client bypasses environment proxies for localhost endpoints.

The context window includes the complete rendered prompt, source text, chat-template tokens, and generated translation. A source that fits on its own may leave too little room for a complete translation. Estimate both source and expected output sizes with the model tokenizer, and select an input size that leaves adequate output capacity. --max-tokens adds an explicit output cap when desired; it does not extend the total window.

The 2B and 9B shipped limits are 262,144 and 229,376 tokens, respectively. The report's native benchmark uses a 256K-token context setting; the public 9B serving preset follows its shipped configuration rather than overriding it to 262,144.

See the inference guide, document client, and fixed templates for full endpoint and prompt details.

NativeLongBench evaluation

NativeLongBench evaluates continuous passages from held-out books in three corpora. GuoFeng and BWB contain Chinese web fiction with English translations. General Books contains English originals from literature, history, social science, science and technology, paired with model-generated Chinese translations. Japanese originals with model-generated Chinese translations provide additional training supervision.

The report divides books by work and removes identified cross-corpus copies of test works from training:

CorpusBooksTrain / validation / test worksEvaluated direction
GuoFeng140130 / 4 / 6zh→en
BWB337327 / 4 / 6zh→en
General Books306292 / 3 / 11en→zh

The main comparison evaluates native translation: every model receives the complete test passage and is asked to translate it in one generation. The GitHub evaluation summary reports 274 document windows across five length groups: 4K, 8K, 16K, 32K, and 64K Chinese-side tokens. For GuoFeng/BWB the length is measured on the Chinese source; for General Books it is measured on the Chinese reference translation. K means 1,024 tokens. These are document-length groups, distinct from the total context window; the General Books 64K group corresponds to approximately 68K–82K English input tokens.

SEGALE aligns contiguous source and output blocks before COMET scores them against references. Unmatched blocks receive zero. Scores are averaged within each document, then within each length group; the corpus mean equally weights the five groups. The table uses document-SEGALE/COMET (max_size=8), on a 0–1 scale where higher is better.

The comparison includes both released Index-NativeLong sizes, the 7B and 30B-A3B translation baselines available in the report, and frontier/API systems under the same native-document protocol.

ModelGuoFeng zh→en ↑BWB zh→en ↑General Books en→zh ↑
Index-NativeLong-9B0.78910.76830.8848
Index-NativeLong-2B (this checkpoint)0.67010.73000.8651
Qwen3.8 Flash0.70530.68230.8206
GLM-5.3-Flash*0.70540.66910.8795
North-Small-Translate-1.00.61600.62820.8350
Hy-MT2-7B0.20820.19830.2360
Hy-MT2-30B-A3B0.29040.29980.3306

*Some GLM requests were refused; its General Books scores average the remaining samples. North's long-input evaluation uses an experimental extension beyond its supported 16K input/output range. Gemini 3.8 Flash is excluded from this main comparison because its output limit is below the 128K output budget reserved for the longest group. See the report for these comparison conditions.

This checkpoint's results by length are:

Corpus / direction4K ↑8K ↑16K ↑32K ↑64K ↑Mean ↑
GuoFeng zh→en0.77040.77960.65080.60360.54600.6701
BWB zh→en0.76880.76690.72010.71000.68430.7300
General Books en→zh0.89080.87360.87350.84350.84390.8651

Source: the report's main native long-document comparison; GitHub evaluation summary.

Limits and practical use

Native generation helps preserve document-level context, but does not guarantee complete coverage, accurate facts, or consistent terminology on every input. The report observes repeated passages in some 2B outputs. Review long translations for omissions, repetition, untranslated spans, and EOS or context-budget truncation before use.

The fixed evaluation sets have also been used during model development. GuoFeng test works are drawn from its public training corpus because the official test set is unavailable, while being held out from this study's training. General Books uses model-generated references. Interpret the scores with these protocol limits; they do not measure every aspect of a publishable book translation.

The report's case studies compare native 9B with glossary-assisted chunked workflows on conceptual distinctions and recurring names. They illustrate document-level benefits on selected passages, without establishing universal superiority for every workflow. The report's Japanese–Chinese table presents external baselines rather than NativeLong checkpoint scores, so no Japanese quality result is claimed here.

The team plans to release shareable portions of NativeLongBench. Benchmark release remains listed as planned in the GitHub repository; this card links to the current protocol and results.

Model family

ModelReleased checkpointsTask
Index-Translate2B · 9B · 35B-A3B (preview)Text translation and translation instructions across 150 languages
Index-Echo S2TT2B · 9BSpeech-to-text translation
Index-Echo S2ST2B · 9BSpeech-to-speech translation with voice conditioning
Index-Homura2B · 9BTranslation with a target syllable count
Index-NativeLong2B · 9BFull-document translation; released templates support zh↔en and zh↔ja

The 150-language coverage refers to the Index-Translate text models. Speech and long-document packages have their own language interfaces. NativeLong retains the IndexTeam/Index-Nailong-* repository IDs.

Citation

@techreport{indextranslate2026,
  author={Tianjiao Li and Mengran Yu and Chenyu Shi and Lusheng Zhang and
          Qisi Chen and Yanshan Zhou and Ji Qi and Jingying Liu and
          Yuang Feng and Ziang Cui and Tianxing Yan},
  title={Index-Translate: A Multilingual Translation Model Family --- Text, Speech, Controlled Dubbing, and Long-Document Translation},
  institution={Index LLM Team},
  year={2026},
  month={September}
}

License and feedback

Apache-2.0. Please use GitHub Issues for questions and feedback.

index
long-context
qwen3_5
safetensors
translation

IndexTeam/Index-Nailong-2B

Model

Index-NativeLong-2B

5

10 commits

1 linked in READMEs

updated Sep 30, 2026

See the code

README

Index-NativeLong-2B

Online demo · GitHub · Technical report · Hugging Face collection · ModelScope collection

Index-NativeLong-2B, published as IndexTeam/Index-Nailong-2B, is the 2B long-document translation specialist in the Index-Translate family. It receives the full source document and produces its translation in a single generation, keeping distant source passages and the translation history together. This supports content coverage, consistent names and terminology, and continuity across chapters or scenes.

The 2B checkpoint scores 0.6701 / 0.7300 / 0.8651 on GuoFeng / BWB / General Books. It exceeds the external models in the report at both 32K and 64K on BWB. The shipped serving configuration supports a 262,144-token context window, shared by the prompt, source, and generated translation. Use the published Nailong model ID in commands.

Training and supported tasks

The family builds on Qwen3.5 with multilingual mid-training and translation post-training. This 2B model follows 128K-sequence mid-training decay and general translation SFT before full-parameter document SFT. Document supervision consists of aligned book passages and complete target translations, with bidirectional training over approximately 4K–64K Chinese-side tokens. Training sequence length reaches 128K in the long-document adaptation; this is separate from the shipped serving context limit.

The technical report studies long-sequence decay and RoPE/NoPE positional-encoding choices in separate 2B ablations. After document SFT, their overall scores are similar, with strengths varying by length. These experiments describe training alternatives rather than a runtime switch in the released checkpoint.

The released document client supplies fixed prompt templates for Chinese→English, English→Chinese, Chinese→Japanese, and Japanese→Chinese (zh-en, en-zh, zh-ja, ja-zh). Its instructions require full coverage, consistent entity names and terminology, and preservation of chapter, paragraph, and content order. These four directions are the packaged interface; the general text model's 150-language coverage does not establish NativeLong performance in 150 languages.

Inference

Use a CUDA GPU and a vLLM build with Qwen3.5 support; the official guide records testing with vLLM 0.29. The bf16 model-memory guide is approximately 8 GB for 2B and 24 GB for 9B, plus KV cache for the chosen context length. Full-context serving can require substantially more memory or tensor parallelism.

git clone https://github.com/bilibili/Index-Translate.git
cd Index-Translate
pip install -U vllm
pip install -r inference/llm/requirements.txt
bash inference/llm/serve_vllm.sh nailong-2b --host 127.0.0.1 --port 8000

From the repository directory in another terminal, translate a UTF-8 document using the official direction template:

python inference/llm/doc_translate.py novel.zh.txt \
  --direction zh-en --model IndexTeam/Index-Nailong-2B \
  --output novel.en.txt

python inference/llm/doc_translate.py paper.en.txt \
  --direction en-zh --model IndexTeam/Index-Nailong-2B \
  --output paper.zh.txt

Replace the input filenames with your documents. For Japanese, use --direction zh-ja or --direction ja-zh. The client streams the translation into the output file and loads the prompt from inference/llm/prompts/nailong_<direction>.txt.

Released client defaults

SettingDefault
Temperaturetemperature=0 (greedy)
Top-p / top-k / min-ptop_p=1, top_k=-1, min_p=0
Seedseed=42
Presence / repetition penaltypresence_penalty=0, repetition_penalty=1
Frequency penaltyNot explicitly set by this client; use the inference server's default
Stop token IDsstop_token_ids=[248044, 248046]
Stop stringsNot explicitly set by this client; use the inference server's default
Ignore EOSignore_eos=False
ThinkingDisabled through chat_template_kwargs={"enable_thinking": False}
Output budgetmax_tokens omitted by default; generation uses the remaining context window
Streamingstream=True; writes each received content delta to the output
Chat messagesOne user message rendered from the fixed direction template and complete source
Serving context limit for this checkpoint262144 tokens

Complete client arguments

These are the actual defaults in doc_translate.py, shared by both checkpoint sizes. The example above selects this card's checkpoint explicitly. Sampling, stop-token, thinking, and streaming settings are fixed in the client; its command-line interface exposes the following controls:

ArgumentDefault and behavior
inputRequired UTF-8 source filename; - reads standard input; normalizes CRLF line endings, strips outer whitespace, rejects empty input
--output, -oStandard output if omitted; otherwise writes the streamed translation to the specified UTF-8 file
--direction, -dUnset; if supplied, selects one of zh-en, en-zh, zh-ja, ja-zh and takes priority over source/target flags
--source, -szh; used to derive the direction when --direction is omitted
--target, -ten; used with --source, so the default derived direction is zh-en
--model, -mINDEX_MODEL environment variable, otherwise IndexTeam/Index-Nailong-9B; pass IndexTeam/Index-Nailong-2B to select this checkpoint
--base-urlOPENAI_BASE_URL environment variable, otherwise http://127.0.0.1:8000/v1
--api-keyOPENAI_API_KEY environment variable, otherwise EMPTY
--max-tokens0; omits API max_tokens; a positive value sets an explicit output-token cap
--help, -hPrints the command-line help

The serving preset sets --served-model-name IndexTeam/Index-Nailong-2B and --max-model-len 262144; the example binds it to 127.0.0.1:8000. Extra preset arguments pass through to vllm serve, including deployment options such as --tensor-parallel-size. The client bypasses environment proxies for localhost endpoints.

The context window includes the complete rendered prompt, source text, chat-template tokens, and generated translation. A source that fits on its own may leave too little room for a complete translation. Estimate both source and expected output sizes with the model tokenizer, and select an input size that leaves adequate output capacity. --max-tokens adds an explicit output cap when desired; it does not extend the total window.

The 2B and 9B shipped limits are 262,144 and 229,376 tokens, respectively. The report's native benchmark uses a 256K-token context setting; the public 9B serving preset follows its shipped configuration rather than overriding it to 262,144.

See the inference guide, document client, and fixed templates for full endpoint and prompt details.

NativeLongBench evaluation

NativeLongBench evaluates continuous passages from held-out books in three corpora. GuoFeng and BWB contain Chinese web fiction with English translations. General Books contains English originals from literature, history, social science, science and technology, paired with model-generated Chinese translations. Japanese originals with model-generated Chinese translations provide additional training supervision.

The report divides books by work and removes identified cross-corpus copies of test works from training:

CorpusBooksTrain / validation / test worksEvaluated direction
GuoFeng140130 / 4 / 6zh→en
BWB337327 / 4 / 6zh→en
General Books306292 / 3 / 11en→zh

The main comparison evaluates native translation: every model receives the complete test passage and is asked to translate it in one generation. The GitHub evaluation summary reports 274 document windows across five length groups: 4K, 8K, 16K, 32K, and 64K Chinese-side tokens. For GuoFeng/BWB the length is measured on the Chinese source; for General Books it is measured on the Chinese reference translation. K means 1,024 tokens. These are document-length groups, distinct from the total context window; the General Books 64K group corresponds to approximately 68K–82K English input tokens.

SEGALE aligns contiguous source and output blocks before COMET scores them against references. Unmatched blocks receive zero. Scores are averaged within each document, then within each length group; the corpus mean equally weights the five groups. The table uses document-SEGALE/COMET (max_size=8), on a 0–1 scale where higher is better.

The comparison includes both released Index-NativeLong sizes, the 7B and 30B-A3B translation baselines available in the report, and frontier/API systems under the same native-document protocol.

ModelGuoFeng zh→en ↑BWB zh→en ↑General Books en→zh ↑
Index-NativeLong-9B0.78910.76830.8848
Index-NativeLong-2B (this checkpoint)0.67010.73000.8651
Qwen3.8 Flash0.70530.68230.8206
GLM-5.3-Flash*0.70540.66910.8795
North-Small-Translate-1.00.61600.62820.8350
Hy-MT2-7B0.20820.19830.2360
Hy-MT2-30B-A3B0.29040.29980.3306

*Some GLM requests were refused; its General Books scores average the remaining samples. North's long-input evaluation uses an experimental extension beyond its supported 16K input/output range. Gemini 3.8 Flash is excluded from this main comparison because its output limit is below the 128K output budget reserved for the longest group. See the report for these comparison conditions.

This checkpoint's results by length are:

Corpus / direction4K ↑8K ↑16K ↑32K ↑64K ↑Mean ↑
GuoFeng zh→en0.77040.77960.65080.60360.54600.6701
BWB zh→en0.76880.76690.72010.71000.68430.7300
General Books en→zh0.89080.87360.87350.84350.84390.8651

Source: the report's main native long-document comparison; GitHub evaluation summary.

Limits and practical use

Native generation helps preserve document-level context, but does not guarantee complete coverage, accurate facts, or consistent terminology on every input. The report observes repeated passages in some 2B outputs. Review long translations for omissions, repetition, untranslated spans, and EOS or context-budget truncation before use.

The fixed evaluation sets have also been used during model development. GuoFeng test works are drawn from its public training corpus because the official test set is unavailable, while being held out from this study's training. General Books uses model-generated references. Interpret the scores with these protocol limits; they do not measure every aspect of a publishable book translation.

The report's case studies compare native 9B with glossary-assisted chunked workflows on conceptual distinctions and recurring names. They illustrate document-level benefits on selected passages, without establishing universal superiority for every workflow. The report's Japanese–Chinese table presents external baselines rather than NativeLong checkpoint scores, so no Japanese quality result is claimed here.

The team plans to release shareable portions of NativeLongBench. Benchmark release remains listed as planned in the GitHub repository; this card links to the current protocol and results.

Model family

ModelReleased checkpointsTask
Index-Translate2B · 9B · 35B-A3B (preview)Text translation and translation instructions across 150 languages
Index-Echo S2TT2B · 9BSpeech-to-text translation
Index-Echo S2ST2B · 9BSpeech-to-speech translation with voice conditioning
Index-Homura2B · 9BTranslation with a target syllable count
Index-NativeLong2B · 9BFull-document translation; released templates support zh↔en and zh↔ja

The 150-language coverage refers to the Index-Translate text models. Speech and long-document packages have their own language interfaces. NativeLong retains the IndexTeam/Index-Nailong-* repository IDs.

Citation

@techreport{indextranslate2026,
  author={Tianjiao Li and Mengran Yu and Chenyu Shi and Lusheng Zhang and
          Qisi Chen and Yanshan Zhou and Ji Qi and Jingying Liu and
          Yuang Feng and Ziang Cui and Tianxing Yan},
  title={Index-Translate: A Multilingual Translation Model Family --- Text, Speech, Controlled Dubbing, and Long-Document Translation},
  institution={Index LLM Team},
  year={2026},
  month={September}
}

License and feedback

Apache-2.0. Please use GitHub Issues for questions and feedback.

index
long-context
qwen3_5
safetensors
translation