nvidia/Nemotron-SFT-Multilingual-v1

Dataset

Dataset Description:

18

8 commits

2 linked in READMEs

updated Mar 11, 2026

See the code

README

Dataset Description:

Nemotron-Multilingual-v1 is a multilingual reasoning dataset made by translating a subsample of SFT data from Nemotron-Math-v2, Nemotron-Competitive-Programming-v1, and Nemotron-Science-v1 into to 6 languages (German, French, Japanese, German, Italian, Japanese, Chinese).
The original datasets were translated with Qwen2.5-14B-Instruct, then filtered with heuristics to remove translation failures and hallucinations. The STEM subsets are further post-edited with an LLM Qwen3-4B-Thinking-2507 to fix format mismatching problems.
We provide the prompt and the final answer in the target language, while the reasoning trace is kept in English. This is due to architectural decisions for Nemotron 3 series.

For details of dataset in each domain, please refer to the original data cards mentioned before.

This dataset is ready for commercial use.

Dataset Owner(s):

NVIDIA Corporation

Dataset Creation Date:

Created on: Jan 28, 2026 Last Modified on: Jan 28, 2026

License/Terms of Use:

This dataset is governed by the Creative Commons Attribution 4.0 International License (CC BY 4.0), except for the StackOverflow and MathGenSelect data, which is governed by the Creative Commons Attribution-ShareAlike 4.0 International License (CC BY-SA 4.0)

Intended Usage:

This dataset is intended for post-training large language models with multilingual capabilities, with a special focus on improving model's capability of handling STEM, math and coding applications under multilingual setup.

Because all the examples are translated from English-sourced datasets, it is not intended to instill local or regional knowledge specific to where a specific language is spoken.

Dataset Characterization

Data Collection Method

  • [Hybrid]

Labeling Method

  • [Synthetic]

Dataset Format

Modality: Text
Format: JSONL
Structure: Text + Metadata

Dataset Quantification

The samples below represent the number of Q&A prompts, along with the reasoning trace in English.

SubsetSamples
code_de133322
code_es131578
code_fr136045
code_it143122
code_ja126393
code_zh154653
math_de128846
math_es102866
math_fr115916
math_it130388
math_ja101820
math_zh88594
stem_de261205
stem_es262353
stem_fr264221
stem_it269240
stem_ja256624
stem_zh258069
Total3065255

Measurement of Total Data Storage: ~90GB

Reference(s):

You can find our recipe for translation here.

Ethical Considerations:

NVIDIA believes Trustworthy AI is a shared responsibility and we have NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications.  When downloaded or used in accordance with our terms of service, developers should work with their internal developer teams to ensure this dataset meets requirements for the relevant industry and use case and addresses unforeseen product misuse.  Please report quality, risk, security vulnerabilities or NVIDIA AI Concerns here

Contributors

soarescmsa

8 commits

nvidia/Nemotron-SFT-Multilingual-v1

Dataset

Dataset Description:

18

8 commits

2 linked in READMEs

updated Mar 11, 2026

See the code

README

Dataset Description:

Nemotron-Multilingual-v1 is a multilingual reasoning dataset made by translating a subsample of SFT data from Nemotron-Math-v2, Nemotron-Competitive-Programming-v1, and Nemotron-Science-v1 into to 6 languages (German, French, Japanese, German, Italian, Japanese, Chinese).
The original datasets were translated with Qwen2.5-14B-Instruct, then filtered with heuristics to remove translation failures and hallucinations. The STEM subsets are further post-edited with an LLM Qwen3-4B-Thinking-2507 to fix format mismatching problems.
We provide the prompt and the final answer in the target language, while the reasoning trace is kept in English. This is due to architectural decisions for Nemotron 3 series.

For details of dataset in each domain, please refer to the original data cards mentioned before.

This dataset is ready for commercial use.

Dataset Owner(s):

NVIDIA Corporation

Dataset Creation Date:

Created on: Jan 28, 2026 Last Modified on: Jan 28, 2026

License/Terms of Use:

This dataset is governed by the Creative Commons Attribution 4.0 International License (CC BY 4.0), except for the StackOverflow and MathGenSelect data, which is governed by the Creative Commons Attribution-ShareAlike 4.0 International License (CC BY-SA 4.0)

Intended Usage:

This dataset is intended for post-training large language models with multilingual capabilities, with a special focus on improving model's capability of handling STEM, math and coding applications under multilingual setup.

Because all the examples are translated from English-sourced datasets, it is not intended to instill local or regional knowledge specific to where a specific language is spoken.

Dataset Characterization

Data Collection Method

  • [Hybrid]

Labeling Method

  • [Synthetic]

Dataset Format

Modality: Text
Format: JSONL
Structure: Text + Metadata

Dataset Quantification

The samples below represent the number of Q&A prompts, along with the reasoning trace in English.

SubsetSamples
code_de133322
code_es131578
code_fr136045
code_it143122
code_ja126393
code_zh154653
math_de128846
math_es102866
math_fr115916
math_it130388
math_ja101820
math_zh88594
stem_de261205
stem_es262353
stem_fr264221
stem_it269240
stem_ja256624
stem_zh258069
Total3065255

Measurement of Total Data Storage: ~90GB

Reference(s):

You can find our recipe for translation here.

Ethical Considerations:

NVIDIA believes Trustworthy AI is a shared responsibility and we have NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications.  When downloaded or used in accordance with our terms of service, developers should work with their internal developer teams to ensure this dataset meets requirements for the relevant industry and use case and addresses unforeseen product misuse.  Please report quality, risk, security vulnerabilities or NVIDIA AI Concerns here

Contributors

soarescmsa

8 commits