simplescaling/s1.1-32B

Model

99

stars

15

commits

2

linked in READMEs

Feb 27, 2025

updated

conversational
endpoints_compatible
qwen2
safetensors
tensorboard
text-generation
text-generation-inference
transformers

README

Model Summary

s1.1 is our sucessor of s1 with better reasoning performance by leveraging reasoning traces from r1 instead of Gemini.

This model is a successor of s1-32B with slightly better performance. Thanks to Bespoke Labs (Ryan Marten) for helping generate r1 traces for s1K with Curator.

Use

The model usage is documented here.

Evaluation

Metrics1-32Bs1.1-32Bo1-previewo1DeepSeek-R1DeepSeek-R1-Distill-Qwen-32B
# examples1K1K??>800K800K
AIME202456.756.740.074.479.872.6
AIME2025 I26.760.037.5?65.046.1
MATH50093.095.481.494.897.394.3
GPQA-Diamond59.663.675.277.371.562.1

Note that s1-32B and s1.1-32B use budget forcing in this table; specifically ignoring end-of-thinking and appending "Wait" up to four times.

Contributors

Muennighoff

13 commits

madiator

1 commits

simplescaling/s1.1-32B

Model

99

stars

15

commits

2

linked in READMEs

Feb 27, 2025

updated

conversational
endpoints_compatible
qwen2
safetensors
tensorboard
text-generation
text-generation-inference
transformers

README

Model Summary

s1.1 is our sucessor of s1 with better reasoning performance by leveraging reasoning traces from r1 instead of Gemini.

This model is a successor of s1-32B with slightly better performance. Thanks to Bespoke Labs (Ryan Marten) for helping generate r1 traces for s1K with Curator.

Use

The model usage is documented here.

Evaluation

Metrics1-32Bs1.1-32Bo1-previewo1DeepSeek-R1DeepSeek-R1-Distill-Qwen-32B
# examples1K1K??>800K800K
AIME202456.756.740.074.479.872.6
AIME2025 I26.760.037.5?65.046.1
MATH50093.095.481.494.897.394.3
GPQA-Diamond59.663.675.277.371.562.1

Note that s1-32B and s1.1-32B use budget forcing in this table; specifically ignoring end-of-thinking and appending "Wait" up to four times.

Contributors

Muennighoff

13 commits

madiator

1 commits