TypeSafeAI/Step-5-Preview-BF16).
Try it via our API, or deploy locally with vLLM / SGLang.
Step-5-Preview is StepFun's flagship foundation model, designed from the ground up for real-world agentic tasks. It targets professional domains such as AI coding, software engineering, professional knowledge work, and financial analysis.
StepFun's core philosophy for Step 5 is the "Pareto Frontier" β achieving the optimal balance between intelligence and cost. While previous scaling efforts focused on trading more compute for stronger intelligence, the next phase requires improving the efficiency of converting compute into intelligence.
Step-5-Preview represents a generational leap, with StepFun skipping the entire Step 4.x line entirely, going directly from Step-3.7-Flash to Step 5. This decision reflects the magnitude of improvement achieved in this release.
low, medium, high / xhigh.TypeSafeAI/Step-5-Preview-BF16.
Step-5-Preview uses a 92-layer Transformer with a narrow-deep configuration. This design is specifically intended to create longer information propagation paths for implicit multi-hop reasoning during long prefill operations.
To handle the 1M-token context window efficiently, Step-5-Preview introduces Sparse GQA with block-wise token merging. This mechanism uses sparse indexing to select only historical information relevant to the current task, reducing the number of tokens that actually enter attention computation. StepFun states this cuts indexer and top-k selection costs to approximately one-eighth of a denser baseline.
The model incorporates a unified multimodal encoder that processes text, images, and video frames into a shared latent space. Video is sampled at adaptive frame rates and encoded with temporal attention, allowing the model to understand motion and long-range dependencies in screen recordings, demonstrations, and real-world footage.
| Category | Specification |
|---|---|
| Model Name | Step-5-Preview |
| Developer | StepFun |
| Architecture | Sparse Mixture-of-Experts (MoE) |
| Total Parameters | 600B |
| Active Parameters | 27B per token (~4.5% sparsity) |
| Layers | 92 (narrow-deep Transformer) |
| Context Window | 1,000,000 tokens |
| Attention | Sparse GQA with block-wise token merging |
| Input Modalities | Text, Image, Video |
| Output Modalities | Text |
| Video Formats | MP4, QuickTime, Matroska (β€128 MB, β€5 min recommended) |
| Reasoning Effort | low / medium / high (xhigh) |
| Tool Calling | Parallel, strict JSON schema |
| Intelligence Index | 44 (Artificial Analysis v4.3.2) |
| Open Weights | BF16 checkpoint available now |
| API Availability | Immediate (OpenAI-compatible) |
| License | StepFun Community License |
Step-5-Preview was trained on a massive, carefully curated corpus spanning:
The data mixture was optimized for long-horizon reasoning and tool use, with a strong emphasis on real-world professional tasks. All data was filtered for quality, safety, and license compliance. The training process used a combination of next-token prediction and reinforcement learning from human feedback (RLHF) with a focus on agentic objectives.
Overall Score: 44 (Intelligence Index v4.3.2, recalibrated September 7, 2026)
This places Step-5-Preview among the top three open-weight models globally, on par with models like Kimi K3 Max (approximately 5Γ larger at 2.8T parameters) and Qwen3.8 Max. The index covers 10 evaluations including AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exam.
| Benchmark | Step-5-Preview (High) | Kimi K3 (Max) | GLM-5.3 (Max) | Claude Opus 5 (Max) | GPT-6 Astra (Max) |
|---|---|---|---|---|---|
| DeepSWE v1.1 | 67.7 | 67.5 | 66.9 | 74.0 | 74.1 |
| StepCodeBench | 49.0 | 43.9 | 40.2 | 63.9 | 61.0 |
| ProgramBench | 80.5 | 77.8 | 72.0 | 82.3 | 85.4 |
| Terminal-Bench v4 | 33.3 | 12.6 | 41.9 | 52.3 | 57.9 |
| Agents' Last Exam (ALE-CLI) | 29.5 | 27.6 | 28.6 | 28.6 | 33.3 |
| GDPval-AA v2 | 1571 | 1548 | 1634 | 1735 | 1580 |
| FrontierFinance | 66.4 | 62.6 | 64.1 | 69.7 | 55.0 |
| DRACO | 83.3 | 78.5 | 82.3 | 87.6 | 76.8 |
temperature=1.0 and top_p=0.95.
In a landmark demonstration of sustained agentic execution, Step-5-Preview was tasked with autonomously optimizing an H100 GPU kernel for up to 24 consecutive hours. The model:
For comparison, Claude Opus 5 achieved 493 TFLOPS in the same experiment. This demonstrates Step-5-Preview's ability to sustain productive work over extended periods without human intervention.
In another 24-hour experiment, Step-5-Preview autonomously improved the accuracy of Qwen3-30B-A3B on AIME24 from 53.3% to 60% through automated post-training experiments. This showcases the model's capacity for self-directed research and optimization.
The model is specifically optimized for agent workflows that require:
StepFun demonstrated the model's capabilities across several complex, real-world projects:
pip install transformers>=4.56.0
pip install torch>=2.4.0
pip install accelerate
For video/image support:
pip install av pillow
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "TypeSafeAI/Step-5-Preview-BF16"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
device_map="auto",
torch_dtype="bfloat16",
)
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain the significance of the Pareto Frontier in AI scaling."},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(
inputs,
max_new_tokens=1024,
temperature=0.7,
top_p=0.95,
reasoning_effort="high", # low / medium / high / xhigh
)
response = tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True)
print(response)
from transformers import AutoProcessor
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
messages = [
{
"role": "user",
"content": [
{"type": "image", "url": "https://example.com/image.jpg"},
{"type": "video", "url": "https://example.com/video.mp4"},
{"type": "text", "text": "Describe the scene and summarize the video."},
],
}
]
inputs = processor.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt")
# ... generate as above
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}
]
messages = [{"role": "user", "content": "What's the weather in Tokyo?"}]
inputs = tokenizer.apply_chat_template(
messages,
tools=tools,
add_generation_prompt=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(inputs, max_new_tokens=256, reasoning_effort="medium")
print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))
vllm serve TypeSafeAI/Step-5-Preview-BF16 \
--trust-remote-code \
--tensor-parallel-size 8 \
--max-model-len 1000000 \
--enable-reasoning \
--reasoning-parser stepfun
python -m sglang.launch_server \
--model-path TypeSafeAI/Step-5-Preview-BF16 \
--trust-remote-code \
--tp 8 \
--context-length 1000000 \
--reasoning-parser stepfun
from openai import OpenAI
client = OpenAI(
api_key="YOUR_STEP_API_KEY",
base_url="https://api.stepfun.com/v1",
)
response = client.chat.completions.create(
model="step-5-preview",
messages=[{"role": "user", "content": "Write a Python function to merge two sorted lists."}],
reasoning_effort="high",
max_tokens=2048,
)
print(response.choices[0].message.content)
stepfun for vLLM/SGLang
Step-5-Preview was evaluated on a comprehensive suite of public and internal benchmarks.
All evaluations used the model's high reasoning effort setting unless otherwise noted.
| Benchmark | Score | Notes |
|---|---|---|
| DeepSWE v1.1 | 67.7 | SWE-agent harness, temp=1.0, top_p=0.95 |
| StepCodeBench | 49.0 | avg@4 |
| ProgramBench | 80.5 | β |
| Terminal-Bench v4 | 33.3 | β |
| Agents' Last Exam (ALE-CLI) | 29.5 | β |
| GDPval-AA v2 | 1571 | Artificial Analysis, Sep 19, 2026 |
| FrontierFinance | 66.4 | β |
| DRACO | 83.3 | β |
| SciCode | Higher than Kimi K3 | β |
| Output Speed | 99.8 tokens/sec | GLM-5.3: 72.1 tokens/sec |
| Time to First Token | 2.96s | GLM-5.3: 2.99s; Claude Opus 5: 56.84s (max effort) |
StepFun is committed to the responsible development and deployment of AI. We have taken the following measures:
We urge all users to consider the ethical implications of their applications and to implement appropriate safeguards.
| Precision | Minimum GPU Memory | Recommended GPU Configuration |
|---|---|---|
| BF16 | 1.2 TB | 8Γ H100 80GB (tensor parallel) |
| FP8 | 600 GB | 4Γ H100 80GB (tensor parallel) |
| INT4 | 300 GB | 4Γ A100 80GB (tensor parallel) |
For inference with 1M context, additional memory is required for KV cache. We recommend using paged attention and offloading techniques available in vLLM and SGLang.
| Metric | Value |
|---|---|
| Output Speed | 99.8 tokens/sec |
| Time to First Token (TTFT) | 2.96 seconds |
| Context Window | 1,000,000 tokens |
| Max Output Tokens | 32,768 (default), configurable up to 131,072 |
| Reasoning Effort Modes | low, medium, high, xhigh |
| Tool Calling Latency | < 500 ms for simple calls |
Measured on 8Γ H100 80GB with vLLM, batch size 1, BF16.
If you use Step-5-Preview in your research, please cite:
@misc{stepfun2026step5preview,
title = {Step-5-Preview: A 600B Sparse MoE Foundation Model for Real-World Agentic Work},
author = {StepFun Team},
year = {2026},
howpublished = {\url{https://huggingface.co/TypeSafeAI/Step-5-Preview-BF16}},
note = {Released September 20, 2026}
}
Step-5-Preview is released under the StepFun Community License. See the LICENSE file for full terms.
TypeSafeAI/Step-5-Preview-BF16).
Try it via our API, or deploy locally with vLLM / SGLang.
Step-5-Preview is StepFun's flagship foundation model, designed from the ground up for real-world agentic tasks. It targets professional domains such as AI coding, software engineering, professional knowledge work, and financial analysis.
StepFun's core philosophy for Step 5 is the "Pareto Frontier" β achieving the optimal balance between intelligence and cost. While previous scaling efforts focused on trading more compute for stronger intelligence, the next phase requires improving the efficiency of converting compute into intelligence.
Step-5-Preview represents a generational leap, with StepFun skipping the entire Step 4.x line entirely, going directly from Step-3.7-Flash to Step 5. This decision reflects the magnitude of improvement achieved in this release.
low, medium, high / xhigh.TypeSafeAI/Step-5-Preview-BF16.
Step-5-Preview uses a 92-layer Transformer with a narrow-deep configuration. This design is specifically intended to create longer information propagation paths for implicit multi-hop reasoning during long prefill operations.
To handle the 1M-token context window efficiently, Step-5-Preview introduces Sparse GQA with block-wise token merging. This mechanism uses sparse indexing to select only historical information relevant to the current task, reducing the number of tokens that actually enter attention computation. StepFun states this cuts indexer and top-k selection costs to approximately one-eighth of a denser baseline.
The model incorporates a unified multimodal encoder that processes text, images, and video frames into a shared latent space. Video is sampled at adaptive frame rates and encoded with temporal attention, allowing the model to understand motion and long-range dependencies in screen recordings, demonstrations, and real-world footage.
| Category | Specification |
|---|---|
| Model Name | Step-5-Preview |
| Developer | StepFun |
| Architecture | Sparse Mixture-of-Experts (MoE) |
| Total Parameters | 600B |
| Active Parameters | 27B per token (~4.5% sparsity) |
| Layers | 92 (narrow-deep Transformer) |
| Context Window | 1,000,000 tokens |
| Attention | Sparse GQA with block-wise token merging |
| Input Modalities | Text, Image, Video |
| Output Modalities | Text |
| Video Formats | MP4, QuickTime, Matroska (β€128 MB, β€5 min recommended) |
| Reasoning Effort | low / medium / high (xhigh) |
| Tool Calling | Parallel, strict JSON schema |
| Intelligence Index | 44 (Artificial Analysis v4.3.2) |
| Open Weights | BF16 checkpoint available now |
| API Availability | Immediate (OpenAI-compatible) |
| License | StepFun Community License |
Step-5-Preview was trained on a massive, carefully curated corpus spanning:
The data mixture was optimized for long-horizon reasoning and tool use, with a strong emphasis on real-world professional tasks. All data was filtered for quality, safety, and license compliance. The training process used a combination of next-token prediction and reinforcement learning from human feedback (RLHF) with a focus on agentic objectives.
Overall Score: 44 (Intelligence Index v4.3.2, recalibrated September 7, 2026)
This places Step-5-Preview among the top three open-weight models globally, on par with models like Kimi K3 Max (approximately 5Γ larger at 2.8T parameters) and Qwen3.8 Max. The index covers 10 evaluations including AA-Briefcase, GDPval-AA v2, Terminal-Bench 4.0, SciCode, and Humanity's Last Exam.
| Benchmark | Step-5-Preview (High) | Kimi K3 (Max) | GLM-5.3 (Max) | Claude Opus 5 (Max) | GPT-6 Astra (Max) |
|---|---|---|---|---|---|
| DeepSWE v1.1 | 67.7 | 67.5 | 66.9 | 74.0 | 74.1 |
| StepCodeBench | 49.0 | 43.9 | 40.2 | 63.9 | 61.0 |
| ProgramBench | 80.5 | 77.8 | 72.0 | 82.3 | 85.4 |
| Terminal-Bench v4 | 33.3 | 12.6 | 41.9 | 52.3 | 57.9 |
| Agents' Last Exam (ALE-CLI) | 29.5 | 27.6 | 28.6 | 28.6 | 33.3 |
| GDPval-AA v2 | 1571 | 1548 | 1634 | 1735 | 1580 |
| FrontierFinance | 66.4 | 62.6 | 64.1 | 69.7 | 55.0 |
| DRACO | 83.3 | 78.5 | 82.3 | 87.6 | 76.8 |
temperature=1.0 and top_p=0.95.
In a landmark demonstration of sustained agentic execution, Step-5-Preview was tasked with autonomously optimizing an H100 GPU kernel for up to 24 consecutive hours. The model:
For comparison, Claude Opus 5 achieved 493 TFLOPS in the same experiment. This demonstrates Step-5-Preview's ability to sustain productive work over extended periods without human intervention.
In another 24-hour experiment, Step-5-Preview autonomously improved the accuracy of Qwen3-30B-A3B on AIME24 from 53.3% to 60% through automated post-training experiments. This showcases the model's capacity for self-directed research and optimization.
The model is specifically optimized for agent workflows that require:
StepFun demonstrated the model's capabilities across several complex, real-world projects:
pip install transformers>=4.56.0
pip install torch>=2.4.0
pip install accelerate
For video/image support:
pip install av pillow
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "TypeSafeAI/Step-5-Preview-BF16"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
device_map="auto",
torch_dtype="bfloat16",
)
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain the significance of the Pareto Frontier in AI scaling."},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(
inputs,
max_new_tokens=1024,
temperature=0.7,
top_p=0.95,
reasoning_effort="high", # low / medium / high / xhigh
)
response = tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True)
print(response)
from transformers import AutoProcessor
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
messages = [
{
"role": "user",
"content": [
{"type": "image", "url": "https://example.com/image.jpg"},
{"type": "video", "url": "https://example.com/video.mp4"},
{"type": "text", "text": "Describe the scene and summarize the video."},
],
}
]
inputs = processor.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt")
# ... generate as above
tools = [
{
"type": "function",
"function": {
"name": "get_weather",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}
]
messages = [{"role": "user", "content": "What's the weather in Tokyo?"}]
inputs = tokenizer.apply_chat_template(
messages,
tools=tools,
add_generation_prompt=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(inputs, max_new_tokens=256, reasoning_effort="medium")
print(tokenizer.decode(outputs[0][inputs.shape[1]:], skip_special_tokens=True))
vllm serve TypeSafeAI/Step-5-Preview-BF16 \
--trust-remote-code \
--tensor-parallel-size 8 \
--max-model-len 1000000 \
--enable-reasoning \
--reasoning-parser stepfun
python -m sglang.launch_server \
--model-path TypeSafeAI/Step-5-Preview-BF16 \
--trust-remote-code \
--tp 8 \
--context-length 1000000 \
--reasoning-parser stepfun
from openai import OpenAI
client = OpenAI(
api_key="YOUR_STEP_API_KEY",
base_url="https://api.stepfun.com/v1",
)
response = client.chat.completions.create(
model="step-5-preview",
messages=[{"role": "user", "content": "Write a Python function to merge two sorted lists."}],
reasoning_effort="high",
max_tokens=2048,
)
print(response.choices[0].message.content)
stepfun for vLLM/SGLang
Step-5-Preview was evaluated on a comprehensive suite of public and internal benchmarks.
All evaluations used the model's high reasoning effort setting unless otherwise noted.
| Benchmark | Score | Notes |
|---|---|---|
| DeepSWE v1.1 | 67.7 | SWE-agent harness, temp=1.0, top_p=0.95 |
| StepCodeBench | 49.0 | avg@4 |
| ProgramBench | 80.5 | β |
| Terminal-Bench v4 | 33.3 | β |
| Agents' Last Exam (ALE-CLI) | 29.5 | β |
| GDPval-AA v2 | 1571 | Artificial Analysis, Sep 19, 2026 |
| FrontierFinance | 66.4 | β |
| DRACO | 83.3 | β |
| SciCode | Higher than Kimi K3 | β |
| Output Speed | 99.8 tokens/sec | GLM-5.3: 72.1 tokens/sec |
| Time to First Token | 2.96s | GLM-5.3: 2.99s; Claude Opus 5: 56.84s (max effort) |
StepFun is committed to the responsible development and deployment of AI. We have taken the following measures:
We urge all users to consider the ethical implications of their applications and to implement appropriate safeguards.
| Precision | Minimum GPU Memory | Recommended GPU Configuration |
|---|---|---|
| BF16 | 1.2 TB | 8Γ H100 80GB (tensor parallel) |
| FP8 | 600 GB | 4Γ H100 80GB (tensor parallel) |
| INT4 | 300 GB | 4Γ A100 80GB (tensor parallel) |
For inference with 1M context, additional memory is required for KV cache. We recommend using paged attention and offloading techniques available in vLLM and SGLang.
| Metric | Value |
|---|---|
| Output Speed | 99.8 tokens/sec |
| Time to First Token (TTFT) | 2.96 seconds |
| Context Window | 1,000,000 tokens |
| Max Output Tokens | 32,768 (default), configurable up to 131,072 |
| Reasoning Effort Modes | low, medium, high, xhigh |
| Tool Calling Latency | < 500 ms for simple calls |
Measured on 8Γ H100 80GB with vLLM, batch size 1, BF16.
If you use Step-5-Preview in your research, please cite:
@misc{stepfun2026step5preview,
title = {Step-5-Preview: A 600B Sparse MoE Foundation Model for Real-World Agentic Work},
author = {StepFun Team},
year = {2026},
howpublished = {\url{https://huggingface.co/TypeSafeAI/Step-5-Preview-BF16}},
note = {Released September 20, 2026}
}
Step-5-Preview is released under the StepFun Community License. See the LICENSE file for full terms.