PrimeIntellect/SYNTHETIC-2-RL

Dataset

SYNTHETIC-2

5

3 commits

1 linked in READMEs

updated Jul 10, 2025

See the code

README

SYNTHETIC-2

SYNTHETIC-2 is an open reasoning dataset spanning a variety of math, coding and general reasoning tasks along with reasoning traces generated in a collaborative manner. The dataset contains both high quality reasoning traces from Deepseek-R1-0528 ideally suited for SFT, as well as multiple reasoning traces from smaller models which can be used for difficulty estimation.

To read more about our data collection approach, check out our blog post.

image/png

We release the following final dataset splits on Huggingface:

  • SYNTHETIC-2: The full SYNTHETIC-2 dataset consisting of all prompts and completions along with rewards
  • SYNTHETIC-2-SFT-verified: The SFT split of SYNTHETIC-2 with responses from Deepseek-R1-0528 verified as correct (rewards of 1 for binary rewards and over 0.7 for non-binary rewards)
  • SYNTHETIC-2-SFT-unverified: The SFT split of SYNTHETIC-2 with all responses, including those not verified as correct
  • SYNTHETIC-2-RL: The RL subset of SYNTHETIC-2 with difficulty annotations from Qwen3-32B, Qwen3-4B and DeepSeek-R1-0528-Qwen3-8B

Contributors

justus27

3 commits

PrimeIntellect/SYNTHETIC-2-RL

Dataset

SYNTHETIC-2

5

3 commits

1 linked in READMEs

updated Jul 10, 2025

See the code

README

SYNTHETIC-2

SYNTHETIC-2 is an open reasoning dataset spanning a variety of math, coding and general reasoning tasks along with reasoning traces generated in a collaborative manner. The dataset contains both high quality reasoning traces from Deepseek-R1-0528 ideally suited for SFT, as well as multiple reasoning traces from smaller models which can be used for difficulty estimation.

To read more about our data collection approach, check out our blog post.

image/png

We release the following final dataset splits on Huggingface:

  • SYNTHETIC-2: The full SYNTHETIC-2 dataset consisting of all prompts and completions along with rewards
  • SYNTHETIC-2-SFT-verified: The SFT split of SYNTHETIC-2 with responses from Deepseek-R1-0528 verified as correct (rewards of 1 for binary rewards and over 0.7 for non-binary rewards)
  • SYNTHETIC-2-SFT-unverified: The SFT split of SYNTHETIC-2 with all responses, including those not verified as correct
  • SYNTHETIC-2-RL: The RL subset of SYNTHETIC-2 with difficulty annotations from Qwen3-32B, Qwen3-4B and DeepSeek-R1-0528-Qwen3-8B

Contributors

justus27

3 commits