openbmb/UltraData-RL-2609

Dataset

106

stars

11

commits

1

linked in READMEs

Sep 7, 2026

updated

code
llm
long-context
math
minicpm
post-training
reinforcement-learning
rlvr
stem
verifiable-rewards
Browse cluster: Math, Code, and Reasoning in LLMs

README

UltraData-RL-2609

📦 UltraData Collection | 🌐 UltraData | 🤗 MiniCPM5 Series

English | 中文

📚 Introduction

UltraData-RL-2609 is the L3 refined data for reinforcement learning within UltraData's L0-L4 tiered data management framework. Built for the RL stage of MiniCPM5-2B post-training, it complements UltraData-SFT-2605 with verifiable-reward tasks. It is also the training corpus used by JustRL II (Scaling Small LLMs to 128K Reasoning with a Critic), which extends the minimalist JustRL recipe with a critic and 128K-scale reasoning RL.

The release contains more than 85,000 training samples spanning mathematical reasoning, scientific / knowledge reasoning, long-context understanding, and code generation. Each sample is a verifiable task, normalized into JSONL for RL training, and constructed around three goals — verifiable outcomes, trustworthy rewards, and calibrated difficulty — so that the policy receives a stable, traceable training signal.

📢 What's New

  • [2026.09.07] The UltraData-RL-2609 dataset is released! Verifiable-reward RL data for the post-training of MiniCPM5-2B, about 86K samples across Math, Knowledge (STEM), Long-Context, and Code. 🚀🚀🚀
  • [2026.09.07] MiniCPM5-2B is released!, the second model in the MiniCPM5 series after MiniCPM5-1B. It is a dense 2B Transformer that scales up the same training recipe, built for on-device, local deployment, and resource-constrained scenarios. It reaches 2B-class open-source SOTA, remains competitive with 4B-class models, and shows particular advantages in coding, mathematics, long-context understanding, tool use, and agentic tasks. UltraData-RL-2609 serves as the core RL dataset for MiniCPM5-2B. 🚀🚀🚀
  • [2026.02.08] The UltraData platform is now live, introducing the L0-L4 tiered data management framework. 🔍🔍🔍

🎯 Dataset Statistics and Capability Coverage

The release contains 85,995 samples across four RL directions. Knowledge corresponds to the STEM / science-reasoning slice in the construction write-up.

DirectionSamplesShareTaskHow outcomes are verified
Math32,41237.7%Competition- and textbook-style problems with a single extractable answerAnswer match against ground_truth
Code23,66527.5%Program synthesis from a natural-language specificationExecute the submission against test cases in ground_truth
Long-Context18,04621.0%Long-document multi-hop QA (context is included in query)Answer match against ground_truth; context is guaranteed to support the answer
Knowledge11,87213.8%Short-answer science / knowledge reasoningAnswer match against ground_truth
Total85,995100%

🌟 Dataset Characteristics

  • Complete domain coverage. Math, Knowledge (STEM), Long-Context, and Code target mathematical reasoning, scientific reasoning, long-context understanding, and code generation, giving a complementary distribution of RL signals.
  • Verifiable outcomes. Every domain keeps an explicit reference target and a defined way to judge the result: Math and Knowledge use extractable, checkable answers; Long-Context guarantees that the context supports the answer; Code judges outputs by program execution against test cases. Multiple-choice, true/false, proof, multi-part, and image-dependent items were removed from the construction pipeline.
  • Reliable reward signal. Reference answers, gold labels, and execution verdicts passed consistency checks (LLM judge, multi-model consensus, test-case cross-validation). Items whose label could not be confirmed were dropped, not guessed.
  • Difficulty matched to RL training. Difficulty is controlled around the RL initialization model: items that are already fully mastered (pass rate 1) are removed; learnable items are kept; hard-but-valid items (pass rate 0 with a confirmed label) are retained and scheduled by online dynamic sampling. Difficulty filtering never modifies a label.

🧪 Usage Example: JustRL II

JustRL II uses UltraData-RL-2609 as its verifiable-reward training set for scaling small LLMs to 128K reasoning with a critic. Building on the minimalist JustRL recipe, it further adopts length-adaptive advantage estimation and a critic for more stable long-horizon RL.

In the JustRL II experimental setting with this dataset, AIME 2025 rises from 61 to 81 within about 300 RL steps. The final MiniCPM5-2B checkpoint that consumes UltraData-RL-2609 in post-training further reaches 86 on AIME 2025. See the JustRL II technical blog for the full ablations, and additional benchmarks.

🏗️ Data Construction Pipeline

Stages 1–2 prepare the data; stages 3–5 enforce verifiable → trustworthy → calibrated; stage 6 packages the release. All four domains share the same six stages with domain-specific verifiers and filters.

  1. Data collection & integration. Math: union of DAPO, DeepScaler, and DeepMath. Knowledge / STEM: OpenScienceReasoning-2. Long-Context: HotpotQA, Qasper, and MuSiQue extended with longer contexts, plus in-house synthetic data. Code: OpenCodeReasoning, OpenCodeReasoning-2, and HardTests with synthesized test cases.
  2. Task standardization & rewriting. Unify task format, answer field, and metadata for the trainer and verifiers. Math / Knowledge are rewritten into short-answer form; Long-Context is extended while preserving the question–answer correspondence; Code queries, I/O conventions, and submission format are unified.
  3. Verifiability filtering. Keep only items with a unique, automatically checkable answer and a unified extraction format. Long-Context items must show an explicit context–question–answer link; Code items must run in a sandbox against their test cases.
  4. Reward reliability check. An LLM judge audits question–answer consistency. Math / Knowledge answers are re-solved by several independent models and relabeled by consensus (no consensus → dropped). Long-Context answers are checked by reference comparison plus multi-model quality review. Code test cases, reference solution, and expected outputs are cross-validated; invalid or non-executable tests are removed.
  5. Difficulty calibration. Repeated rollouts on the RL initialization checkpoint estimate an empirical pass rate per item. Pass rate 1 (already mastered, no gradient) is removed; the learnable band is kept; pass rate 0 with a confirmed-valid label is kept and scheduled with online dynamic sampling. This stage changes difficulty and sampling weight only — never the reference label.
  6. Formatting & quality review. Export to unified JSONL; check parsability, required fields, unique IDs, extraction rules, and verifier configs; remove exact / near duplicates and items overlapping public evaluation benchmarks; re-check unique answers (Math / Knowledge), context linkage (Long-Context), and test-case executability (Code); record source, filter version, and verifier version.

What each domain does at stages 2–4:

DomainStage 2 · standardizationStage 3 · verifiabilityStage 4 · reward check
Math / KnowledgeRewritten as short-answerAnswer extraction and matchMulti-model consensus
Long-ContextContext extended, Q–A link keptExplicit context–answer link requiredAnswer compare + model review
CodeUnified query and I/O contractSandboxed execution against testsTest / solution / output cross-check

📦 Data Format

Each JSONL line is one RL training sample. The release uses five fields: uuid, query, ground_truth, source, and domain.

{
  "uuid": "{Domain}_00001",
  "query": "…the full problem statement shown to the policy…",
  "ground_truth": "…reference answer…",
  "source": "{Original Source} | UltraData-RL-2609",
  "domain": "{Domain}"
}
FieldTypeDescription
uuidstringGlobally unique sample ID (domain prefix + index).
querystringThe task the model must answer or complete; for long-context tasks the full context and the question are both in this field.
ground_truthstring | objectReference answer or verifiable target: a string for Math / Knowledge / Long-Context; a test-case object for Code (see below).
sourcestringUltraData-RL-2609
domainstringOne of Math, Knowledge, Long_Context, Code.

Code problems are standard-I/O tasks (call_type = std, fn_name = null). Their ground_truth holds two equal-length arrays, inputs and outputs: the i-th entries are the complete stdin and expected stdout of the i-th test case. Verification is execution-based: pipe inputs[i] to the candidate program's stdin and compare stdout with outputs[i] after trailing-whitespace normalization; a submission is correct only if every test case passes. Example with two tests:

{"inputs": ["3\n1 2 3\n", "1\n5\n"], "outputs": ["6\n", "5\n"]}

🚀 Quick Start

from datasets import load_dataset

ds = load_dataset("openbmb/UltraData-RL-2609", "Math", split="train")
ds = load_dataset("openbmb/UltraData-RL-2609", "Knowledge", split="train")
ds = load_dataset("openbmb/UltraData-RL-2609", "Long-Context", split="train")
ds = load_dataset("openbmb/UltraData-RL-2609", "Code", split="train")

print(ds[0]["query"][:300])
print(ds[0]["ground_truth"])

Available configs: Math, Knowledge, Long-Context, Code.

💡 Intended Uses

  • RL / RLVR post-training for MiniCPM-style on-device models.
  • Domain slices: mathematical reasoning (Math), science / knowledge reasoning (Knowledge), long-context multi-hop QA (Long-Context), program synthesis with executable tests (Code).
  • Mix-ratio studies of RL data versus SFT resources such as UltraData-SFT-2605.

⚠️ Notes and Limitations

  • Verifier required for Code. The release includes test cases, not a sandbox. Users must run submissions themselves to compute rewards.
  • Static labels. Difficulty filtering and online sampling weights used at construction time are not stored as fields; only the kept items are released.
  • Decontamination scope. Screening covered evaluation sets known at construction time; run a new check before introducing a new benchmark.

📂 Data Sources

📜 License and Data Sources

This project is released under the Apache 2.0 license. Upstream datasets are licensed under MIT, CC BY 4.0, and CC BY-SA 4.0, which continue to apply to content derived from them. Apache 2.0 does not override those terms.

No unauthorized unchanged redistribution: Without prior written permission from the original authors (or this organization), any institution, organization, or third-party platform is strictly prohibited from directly reposting, mirroring, re-hosting, or commercially repackaging and republishing any artifacts of this project in any form.

📖 Citation

If you find UltraData-RL-2609 useful in your research, please consider citing:

@misc{ultradata_rl_2609,
  title        = {UltraData-RL-2609},
  author       = {MiniCPM Team},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/datasets/openbmb/UltraData-RL-2609}}
}

Contributors

BigDong

10 commits

chuyue

1 commits

openbmb/UltraData-RL-2609

Dataset

106

stars

11

commits

1

linked in READMEs

Sep 7, 2026

updated

code
llm
long-context
math
minicpm
post-training
reinforcement-learning
rlvr
stem
verifiable-rewards
Browse cluster: Math, Code, and Reasoning in LLMs

README

UltraData-RL-2609

📦 UltraData Collection | 🌐 UltraData | 🤗 MiniCPM5 Series

English | 中文

📚 Introduction

UltraData-RL-2609 is the L3 refined data for reinforcement learning within UltraData's L0-L4 tiered data management framework. Built for the RL stage of MiniCPM5-2B post-training, it complements UltraData-SFT-2605 with verifiable-reward tasks. It is also the training corpus used by JustRL II (Scaling Small LLMs to 128K Reasoning with a Critic), which extends the minimalist JustRL recipe with a critic and 128K-scale reasoning RL.

The release contains more than 85,000 training samples spanning mathematical reasoning, scientific / knowledge reasoning, long-context understanding, and code generation. Each sample is a verifiable task, normalized into JSONL for RL training, and constructed around three goals — verifiable outcomes, trustworthy rewards, and calibrated difficulty — so that the policy receives a stable, traceable training signal.

📢 What's New

  • [2026.09.07] The UltraData-RL-2609 dataset is released! Verifiable-reward RL data for the post-training of MiniCPM5-2B, about 86K samples across Math, Knowledge (STEM), Long-Context, and Code. 🚀🚀🚀
  • [2026.09.07] MiniCPM5-2B is released!, the second model in the MiniCPM5 series after MiniCPM5-1B. It is a dense 2B Transformer that scales up the same training recipe, built for on-device, local deployment, and resource-constrained scenarios. It reaches 2B-class open-source SOTA, remains competitive with 4B-class models, and shows particular advantages in coding, mathematics, long-context understanding, tool use, and agentic tasks. UltraData-RL-2609 serves as the core RL dataset for MiniCPM5-2B. 🚀🚀🚀
  • [2026.02.08] The UltraData platform is now live, introducing the L0-L4 tiered data management framework. 🔍🔍🔍

🎯 Dataset Statistics and Capability Coverage

The release contains 85,995 samples across four RL directions. Knowledge corresponds to the STEM / science-reasoning slice in the construction write-up.

DirectionSamplesShareTaskHow outcomes are verified
Math32,41237.7%Competition- and textbook-style problems with a single extractable answerAnswer match against ground_truth
Code23,66527.5%Program synthesis from a natural-language specificationExecute the submission against test cases in ground_truth
Long-Context18,04621.0%Long-document multi-hop QA (context is included in query)Answer match against ground_truth; context is guaranteed to support the answer
Knowledge11,87213.8%Short-answer science / knowledge reasoningAnswer match against ground_truth
Total85,995100%

🌟 Dataset Characteristics

  • Complete domain coverage. Math, Knowledge (STEM), Long-Context, and Code target mathematical reasoning, scientific reasoning, long-context understanding, and code generation, giving a complementary distribution of RL signals.
  • Verifiable outcomes. Every domain keeps an explicit reference target and a defined way to judge the result: Math and Knowledge use extractable, checkable answers; Long-Context guarantees that the context supports the answer; Code judges outputs by program execution against test cases. Multiple-choice, true/false, proof, multi-part, and image-dependent items were removed from the construction pipeline.
  • Reliable reward signal. Reference answers, gold labels, and execution verdicts passed consistency checks (LLM judge, multi-model consensus, test-case cross-validation). Items whose label could not be confirmed were dropped, not guessed.
  • Difficulty matched to RL training. Difficulty is controlled around the RL initialization model: items that are already fully mastered (pass rate 1) are removed; learnable items are kept; hard-but-valid items (pass rate 0 with a confirmed label) are retained and scheduled by online dynamic sampling. Difficulty filtering never modifies a label.

🧪 Usage Example: JustRL II

JustRL II uses UltraData-RL-2609 as its verifiable-reward training set for scaling small LLMs to 128K reasoning with a critic. Building on the minimalist JustRL recipe, it further adopts length-adaptive advantage estimation and a critic for more stable long-horizon RL.

In the JustRL II experimental setting with this dataset, AIME 2025 rises from 61 to 81 within about 300 RL steps. The final MiniCPM5-2B checkpoint that consumes UltraData-RL-2609 in post-training further reaches 86 on AIME 2025. See the JustRL II technical blog for the full ablations, and additional benchmarks.

🏗️ Data Construction Pipeline

Stages 1–2 prepare the data; stages 3–5 enforce verifiable → trustworthy → calibrated; stage 6 packages the release. All four domains share the same six stages with domain-specific verifiers and filters.

  1. Data collection & integration. Math: union of DAPO, DeepScaler, and DeepMath. Knowledge / STEM: OpenScienceReasoning-2. Long-Context: HotpotQA, Qasper, and MuSiQue extended with longer contexts, plus in-house synthetic data. Code: OpenCodeReasoning, OpenCodeReasoning-2, and HardTests with synthesized test cases.
  2. Task standardization & rewriting. Unify task format, answer field, and metadata for the trainer and verifiers. Math / Knowledge are rewritten into short-answer form; Long-Context is extended while preserving the question–answer correspondence; Code queries, I/O conventions, and submission format are unified.
  3. Verifiability filtering. Keep only items with a unique, automatically checkable answer and a unified extraction format. Long-Context items must show an explicit context–question–answer link; Code items must run in a sandbox against their test cases.
  4. Reward reliability check. An LLM judge audits question–answer consistency. Math / Knowledge answers are re-solved by several independent models and relabeled by consensus (no consensus → dropped). Long-Context answers are checked by reference comparison plus multi-model quality review. Code test cases, reference solution, and expected outputs are cross-validated; invalid or non-executable tests are removed.
  5. Difficulty calibration. Repeated rollouts on the RL initialization checkpoint estimate an empirical pass rate per item. Pass rate 1 (already mastered, no gradient) is removed; the learnable band is kept; pass rate 0 with a confirmed-valid label is kept and scheduled with online dynamic sampling. This stage changes difficulty and sampling weight only — never the reference label.
  6. Formatting & quality review. Export to unified JSONL; check parsability, required fields, unique IDs, extraction rules, and verifier configs; remove exact / near duplicates and items overlapping public evaluation benchmarks; re-check unique answers (Math / Knowledge), context linkage (Long-Context), and test-case executability (Code); record source, filter version, and verifier version.

What each domain does at stages 2–4:

DomainStage 2 · standardizationStage 3 · verifiabilityStage 4 · reward check
Math / KnowledgeRewritten as short-answerAnswer extraction and matchMulti-model consensus
Long-ContextContext extended, Q–A link keptExplicit context–answer link requiredAnswer compare + model review
CodeUnified query and I/O contractSandboxed execution against testsTest / solution / output cross-check

📦 Data Format

Each JSONL line is one RL training sample. The release uses five fields: uuid, query, ground_truth, source, and domain.

{
  "uuid": "{Domain}_00001",
  "query": "…the full problem statement shown to the policy…",
  "ground_truth": "…reference answer…",
  "source": "{Original Source} | UltraData-RL-2609",
  "domain": "{Domain}"
}
FieldTypeDescription
uuidstringGlobally unique sample ID (domain prefix + index).
querystringThe task the model must answer or complete; for long-context tasks the full context and the question are both in this field.
ground_truthstring | objectReference answer or verifiable target: a string for Math / Knowledge / Long-Context; a test-case object for Code (see below).
sourcestringUltraData-RL-2609
domainstringOne of Math, Knowledge, Long_Context, Code.

Code problems are standard-I/O tasks (call_type = std, fn_name = null). Their ground_truth holds two equal-length arrays, inputs and outputs: the i-th entries are the complete stdin and expected stdout of the i-th test case. Verification is execution-based: pipe inputs[i] to the candidate program's stdin and compare stdout with outputs[i] after trailing-whitespace normalization; a submission is correct only if every test case passes. Example with two tests:

{"inputs": ["3\n1 2 3\n", "1\n5\n"], "outputs": ["6\n", "5\n"]}

🚀 Quick Start

from datasets import load_dataset

ds = load_dataset("openbmb/UltraData-RL-2609", "Math", split="train")
ds = load_dataset("openbmb/UltraData-RL-2609", "Knowledge", split="train")
ds = load_dataset("openbmb/UltraData-RL-2609", "Long-Context", split="train")
ds = load_dataset("openbmb/UltraData-RL-2609", "Code", split="train")

print(ds[0]["query"][:300])
print(ds[0]["ground_truth"])

Available configs: Math, Knowledge, Long-Context, Code.

💡 Intended Uses

  • RL / RLVR post-training for MiniCPM-style on-device models.
  • Domain slices: mathematical reasoning (Math), science / knowledge reasoning (Knowledge), long-context multi-hop QA (Long-Context), program synthesis with executable tests (Code).
  • Mix-ratio studies of RL data versus SFT resources such as UltraData-SFT-2605.

⚠️ Notes and Limitations

  • Verifier required for Code. The release includes test cases, not a sandbox. Users must run submissions themselves to compute rewards.
  • Static labels. Difficulty filtering and online sampling weights used at construction time are not stored as fields; only the kept items are released.
  • Decontamination scope. Screening covered evaluation sets known at construction time; run a new check before introducing a new benchmark.

📂 Data Sources

📜 License and Data Sources

This project is released under the Apache 2.0 license. Upstream datasets are licensed under MIT, CC BY 4.0, and CC BY-SA 4.0, which continue to apply to content derived from them. Apache 2.0 does not override those terms.

No unauthorized unchanged redistribution: Without prior written permission from the original authors (or this organization), any institution, organization, or third-party platform is strictly prohibited from directly reposting, mirroring, re-hosting, or commercially repackaging and republishing any artifacts of this project in any form.

📖 Citation

If you find UltraData-RL-2609 useful in your research, please consider citing:

@misc{ultradata_rl_2609,
  title        = {UltraData-RL-2609},
  author       = {MiniCPM Team},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/datasets/openbmb/UltraData-RL-2609}}
}

Contributors

BigDong

10 commits

chuyue

1 commits