NovaSky-AI/Sky-T1-32B-Flash

Model

Model Details

67

8 commits

1 linked in READMEs

updated Feb 2, 2025

See the code

README

Model Details

Model Description

This is a 32B reasoning model preference optimized on top of Sky-T1-32B-Preview to significantly reduce generation lengths while maintaining accuracy. The performance is on par with o1-preview model in both math and coding, while reducing generation lengths by up to 57% relative to Sky-T1-32B-Preview. Please see our blog post for more details.

  • Developed by: NovaSky Team from Sky Computing Lab at UC Berkeley.

Training Details

Training Data

10K preference pairs in math and coding domains, generated by Sky-T1-32B-Preview.

Training Procedure

We perform Simple Policy Optimization (SimPO) with a batch size of 96, learning rate of 5e-7, gamma of 0.3, and beta of 2.0.

Speeds

We use Llama-Factory for training. On 8xH100, the SimPO training takes ~2.5 hours with DeepSpeed Zero-3 Offload.

Evaluation

Sky-T1-32B-PreviewSky-T1-32B-FlashQwen2.5-32B-InstructQwQ-32B- BaseDeepSeek-R1-Distill-Qwen-32B
Math500Acc88.688.676.289.290.8
Avg Len21241417 (-33%)52220892010
AIME24Acc43.343.316.75066.7
Avg Len68814365 (-37%)97073799173
LCB EasyAcc87.48984.690.791.2
Avg Len34152265 (-34%)41432552775
LCB MediumAcc56.856.340.856.376.7
Avg Len82634389 (-47%)53567426324
LCB HardAcc17.917.99.817.138.2
Avg Len145646199 (-57%)6181045010448
MMLUAcc82.481.780.185.282.1
Avg Len1087799 (-17%)3121041774
GPQA DiamondAcc56.856.645.552.562.6
Avg Len35032148 (-39%)60033025108

Acknowledgement

We would like to thanks the compute resources from Lambda Lab and AnyScale.

License

Apache-2.0

Citation

Please considering citing our blog post if you found it useful for your research. Thank you!

@misc{reduce_overthinking_2025,
  author       = {NovaSky Team},
  title        = {Think Less, Achieve More: Cut Reasoning Costs by 50% Without Sacrificing Accuracy},
  howpublished = {https://novasky-ai.github.io/posts/reduce-overthinking},
  note         = {Accessed: 2025-01-23},
  year         = {2025}
}
conversational
endpoints_compatible
qwen2
safetensors
text-generation
text-generation-inference
transformers

Contributors

NovaSkyAI

5 commits

tgriggs

3 commits

NovaSky-AI/Sky-T1-32B-Flash

Model

Model Details

67

8 commits

1 linked in READMEs

updated Feb 2, 2025

See the code

README

Model Details

Model Description

This is a 32B reasoning model preference optimized on top of Sky-T1-32B-Preview to significantly reduce generation lengths while maintaining accuracy. The performance is on par with o1-preview model in both math and coding, while reducing generation lengths by up to 57% relative to Sky-T1-32B-Preview. Please see our blog post for more details.

  • Developed by: NovaSky Team from Sky Computing Lab at UC Berkeley.

Training Details

Training Data

10K preference pairs in math and coding domains, generated by Sky-T1-32B-Preview.

Training Procedure

We perform Simple Policy Optimization (SimPO) with a batch size of 96, learning rate of 5e-7, gamma of 0.3, and beta of 2.0.

Speeds

We use Llama-Factory for training. On 8xH100, the SimPO training takes ~2.5 hours with DeepSpeed Zero-3 Offload.

Evaluation

Sky-T1-32B-PreviewSky-T1-32B-FlashQwen2.5-32B-InstructQwQ-32B- BaseDeepSeek-R1-Distill-Qwen-32B
Math500Acc88.688.676.289.290.8
Avg Len21241417 (-33%)52220892010
AIME24Acc43.343.316.75066.7
Avg Len68814365 (-37%)97073799173
LCB EasyAcc87.48984.690.791.2
Avg Len34152265 (-34%)41432552775
LCB MediumAcc56.856.340.856.376.7
Avg Len82634389 (-47%)53567426324
LCB HardAcc17.917.99.817.138.2
Avg Len145646199 (-57%)6181045010448
MMLUAcc82.481.780.185.282.1
Avg Len1087799 (-17%)3121041774
GPQA DiamondAcc56.856.645.552.562.6
Avg Len35032148 (-39%)60033025108

Acknowledgement

We would like to thanks the compute resources from Lambda Lab and AnyScale.

License

Apache-2.0

Citation

Please considering citing our blog post if you found it useful for your research. Thank you!

@misc{reduce_overthinking_2025,
  author       = {NovaSky Team},
  title        = {Think Less, Achieve More: Cut Reasoning Costs by 50% Without Sacrificing Accuracy},
  howpublished = {https://novasky-ai.github.io/posts/reduce-overthinking},
  note         = {Accessed: 2025-01-23},
  year         = {2025}
}
conversational
endpoints_compatible
qwen2
safetensors
text-generation
text-generation-inference
transformers

Contributors

NovaSkyAI

5 commits

tgriggs

3 commits