PrimeIntellect/INTELLECT-MATH

Model

INTELLECT-MATH: Frontier Mathematical Reasoning through Better Initializations for Reinforcement Learning

8

6 commits

2 linked in READMEs

updated Jan 22, 2025

See the code

README

INTELLECT-MATH: Frontier Mathematical Reasoning through Better Initializations for Reinforcement Learning

INTELLECT-MATH is a 7B parameter model optimized for mathematical reasoning. It was trained in two stages, an SFT stage, in which the model was fine-tuned on verified QwQ outputs, and an RL stage, in which the model was trained using the PRIME-RL recipe.

We demonstrate that the quality of our SFT data can impact the performance and training speed of the RL stage: Due to its better synthetic SFT dataset that encourages the model to imitate the reasoning behavior of a strong teacher model, INTELLECT-MATH outperforms Eurus-2-PRIME, the previous state-of-the-art trained with PRIME-RL, and matches its performance with 10x faster training.

Intellect-Math (Step 255)Intellect-Math (Step 47)Eurus-2-Prime (Step 592)Intellect-Math-SFTEurus-2-SFTQwen-2.5-Math
MATH-50082.081.679.272.865.179.8
OLYMPIADBENCH49.546.742.139.129.840.7
AIME 202426.726.726.716.63.313.3
AMC60.257.857.845.830.150.6
MINERVA MATH39.737.838.633.832.734.6
AVG51.650.148.941.632.243.8
qwen2
safetensors

Contributors

justus27

6 commits

PrimeIntellect/INTELLECT-MATH

Model

INTELLECT-MATH: Frontier Mathematical Reasoning through Better Initializations for Reinforcement Learning

8

6 commits

2 linked in READMEs

updated Jan 22, 2025

See the code

README

INTELLECT-MATH: Frontier Mathematical Reasoning through Better Initializations for Reinforcement Learning

INTELLECT-MATH is a 7B parameter model optimized for mathematical reasoning. It was trained in two stages, an SFT stage, in which the model was fine-tuned on verified QwQ outputs, and an RL stage, in which the model was trained using the PRIME-RL recipe.

We demonstrate that the quality of our SFT data can impact the performance and training speed of the RL stage: Due to its better synthetic SFT dataset that encourages the model to imitate the reasoning behavior of a strong teacher model, INTELLECT-MATH outperforms Eurus-2-PRIME, the previous state-of-the-art trained with PRIME-RL, and matches its performance with 10x faster training.

Intellect-Math (Step 255)Intellect-Math (Step 47)Eurus-2-Prime (Step 592)Intellect-Math-SFTEurus-2-SFTQwen-2.5-Math
MATH-50082.081.679.272.865.179.8
OLYMPIADBENCH49.546.742.139.129.840.7
AIME 202426.726.726.716.63.313.3
AMC60.257.857.845.830.150.6
MINERVA MATH39.737.838.633.832.734.6
AVG51.650.148.941.632.243.8
qwen2
safetensors

Contributors

justus27

6 commits