Nemotron-Multilingual-v1 is a multilingual reasoning dataset made by translating a subsample of SFT data from Nemotron-Math-v2, Nemotron-Competitive-Programming-v1, and Nemotron-Science-v1 into to 6 languages (German, French, Japanese, German, Italian, Japanese, Chinese).
The original datasets were translated with Qwen2.5-14B-Instruct, then filtered with heuristics to remove translation failures and hallucinations. The STEM subsets are further post-edited with an LLM Qwen3-4B-Thinking-2507 to fix format mismatching problems.
We provide the prompt and the final answer in the target language, while the reasoning trace is kept in English. This is due to architectural decisions for Nemotron 3 series.
For details of dataset in each domain, please refer to the original data cards mentioned before.
This dataset is ready for commercial use.
NVIDIA Corporation
Created on: Jan 28, 2026 Last Modified on: Jan 28, 2026
This dataset is governed by the Creative Commons Attribution 4.0 International License (CC BY 4.0), except for the StackOverflow and MathGenSelect data, which is governed by the Creative Commons Attribution-ShareAlike 4.0 International License (CC BY-SA 4.0)
This dataset is intended for post-training large language models with multilingual capabilities, with a special focus on improving model's capability of handling STEM, math and coding applications under multilingual setup.
Because all the examples are translated from English-sourced datasets, it is not intended to instill local or regional knowledge specific to where a specific language is spoken.
Data Collection Method
Labeling Method
Modality: Text
Format: JSONL
Structure: Text + Metadata
The samples below represent the number of Q&A prompts, along with the reasoning trace in English.
| Subset | Samples |
|---|---|
| code_de | 133322 |
| code_es | 131578 |
| code_fr | 136045 |
| code_it | 143122 |
| code_ja | 126393 |
| code_zh | 154653 |
| math_de | 128846 |
| math_es | 102866 |
| math_fr | 115916 |
| math_it | 130388 |
| math_ja | 101820 |
| math_zh | 88594 |
| stem_de | 261205 |
| stem_es | 262353 |
| stem_fr | 264221 |
| stem_it | 269240 |
| stem_ja | 256624 |
| stem_zh | 258069 |
| Total | 3065255 |
Measurement of Total Data Storage: ~90GB
You can find our recipe for translation here.
NVIDIA believes Trustworthy AI is a shared responsibility and we have NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal developer teams to ensure this dataset meets requirements for the relevant industry and use case and addresses unforeseen product misuse. Please report quality, risk, security vulnerabilities or NVIDIA AI Concerns here
8 commits
Nemotron-Multilingual-v1 is a multilingual reasoning dataset made by translating a subsample of SFT data from Nemotron-Math-v2, Nemotron-Competitive-Programming-v1, and Nemotron-Science-v1 into to 6 languages (German, French, Japanese, German, Italian, Japanese, Chinese).
The original datasets were translated with Qwen2.5-14B-Instruct, then filtered with heuristics to remove translation failures and hallucinations. The STEM subsets are further post-edited with an LLM Qwen3-4B-Thinking-2507 to fix format mismatching problems.
We provide the prompt and the final answer in the target language, while the reasoning trace is kept in English. This is due to architectural decisions for Nemotron 3 series.
For details of dataset in each domain, please refer to the original data cards mentioned before.
This dataset is ready for commercial use.
NVIDIA Corporation
Created on: Jan 28, 2026 Last Modified on: Jan 28, 2026
This dataset is governed by the Creative Commons Attribution 4.0 International License (CC BY 4.0), except for the StackOverflow and MathGenSelect data, which is governed by the Creative Commons Attribution-ShareAlike 4.0 International License (CC BY-SA 4.0)
This dataset is intended for post-training large language models with multilingual capabilities, with a special focus on improving model's capability of handling STEM, math and coding applications under multilingual setup.
Because all the examples are translated from English-sourced datasets, it is not intended to instill local or regional knowledge specific to where a specific language is spoken.
Data Collection Method
Labeling Method
Modality: Text
Format: JSONL
Structure: Text + Metadata
The samples below represent the number of Q&A prompts, along with the reasoning trace in English.
| Subset | Samples |
|---|---|
| code_de | 133322 |
| code_es | 131578 |
| code_fr | 136045 |
| code_it | 143122 |
| code_ja | 126393 |
| code_zh | 154653 |
| math_de | 128846 |
| math_es | 102866 |
| math_fr | 115916 |
| math_it | 130388 |
| math_ja | 101820 |
| math_zh | 88594 |
| stem_de | 261205 |
| stem_es | 262353 |
| stem_fr | 264221 |
| stem_it | 269240 |
| stem_ja | 256624 |
| stem_zh | 258069 |
| Total | 3065255 |
Measurement of Total Data Storage: ~90GB
You can find our recipe for translation here.
NVIDIA believes Trustworthy AI is a shared responsibility and we have NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal developer teams to ensure this dataset meets requirements for the relevant industry and use case and addresses unforeseen product misuse. Please report quality, risk, security vulnerabilities or NVIDIA AI Concerns here
8 commits