stepfun-ai/Step-Audio-2-mini-Base

Model

26

stars

10

commits

1

repos using this model

1

linked in READMEs

Sep 1, 2025

updated

custom_code
onnx
safetensors
step_audio_2

README

GitHubHomepageTwitter FollowDiscord
License

Introduction

Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation.

  • Advanced Speech and Audio Understanding: Promising performance in ASR and audio understanding by comprehending and reasoning semantic information, para-linguistic and non-vocal information.

  • Intelligent Speech Conversation: Achieving natural and intelligent interactions that are contextually appropriate for various conversational scenarios and paralinguistic information.

  • Tool Calling and Multimodal RAG: By leveraging tool calling and RAG to access real-world knowledge (both textual and acoustic), Step-Audio 2 can generate responses with fewer hallucinations for diverse scenarios, while also having the ability to switch timbres based on retrieved speech.

  • State-of-the-Art Performance: Achieving state-of-the-art performance on various audio understanding and conversational benchmarks compared to other open-source and commercial solutions. (See Evaluation and Technical Report).

Model Download

Huggingface

Models🤗 Hugging Face
Step-Audio 2 ministepfun-ai/Step-Audio-2-mini
Step-Audio 2 mini Basestepfun-ai/Step-Audio-2-mini-Base

Model Usage

🔧 Dependencies and Installation

conda create -n stepaudio2 python=3.10
conda activate stepaudio2
pip install transformers==4.49.0 torchaudio librosa onnxruntime s3tokenizer diffusers hyperpyyaml

git clone https://github.com/stepfun-ai/Step-Audio2.git
cd Step-Audio2
git lfs install
git clone https://huggingface.co/stepfun-ai/Step-Audio-2-mini-Base

🚀 Inference Scripts

python examples-base.py

Online demonstration

StepFun realtime console

StepFun AI Assistant

  • Step-Audio 2 is also available in our StepFun AI Assistant mobile App with both web and audio search tools enabled.
  • Please scan the following QR code to download it from your app store then tap the phone icon in the top-right corner.
QR code

WeChat group

You can scan the following QR code to join our WeChat group for communication and discussion.

QR code

Evaluation

Architecture

Automatic speech recognition

CER for Chinese, Cantonese and Japanese and WER for Arabian and English. N/A indicates that the language is not supported.

CategoryTest setDoubao LLM ASRGPT-4o TranscribeKimi-AudioQwen-OmniStep-Audio 2Step-Audio 2 mini
EnglishCommon Voice9.209.307.838.335.956.76
FLEURS English7.222.714.475.053.033.05
LibriSpeech clean2.921.751.492.931.171.33
LibriSpeech other5.324.232.915.072.422.86
Average6.174.504.185.353.143.50
ChineseAISHELL0.983.520.641.170.630.78
AISHELL-23.104.262.672.402.102.16
FLEURS Chinese2.922.622.917.012.682.53
KeSpeech phase16.4826.805.116.453.633.97
WenetSpeech meeting4.9031.405.216.614.754.87
WenetSpeech net4.4615.715.935.244.674.82
Average3.8114.053.754.813.083.19
Multilingual FLEURS ArabianN/A11.72N/A25.1314.2216.46
Common Voice yue9.2011.1038.907.897.908.32
FLEURS JapaneseN/A3.27N/A10.493.184.67
In-houseAnhui accent8.8350.5522.1718.7310.6111.65
Guangdong accent4.997.833.764.033.814.44
Guangxi accent3.377.094.293.354.113.51
Shanxi accent20.2655.0334.7125.9512.4415.60
Sichuan dialect3.0132.855.265.614.354.57
Shanghai dialect47.4989.5882.9058.7417.7719.30
Average14.6640.4925.5219.408.859.85

Paralinguistic information understanding

StepEval-Audio-Paralinguistic

ModelAvg.GenderAgeTimbreScenarioEventEmotionPitchRhythmSpeedStyleVocal
GPT-4o Audio43.451842342214824060586444
Kimi-Audio49.649450103048665640445454
Qwen-Omni44.184050162842763254505048
Step-Audio-AQAA36.91706618141440384854440
Step-Audio 283.0910096827860868286888868
Step-Audio 2 mini80.0010094807860828268748676

Audio understanding and reasoning

MMAU

ModelAvg.SoundSpeechMusic
Audio Flamingo 373.176.966.173.9
Gemini 2.5 Pro71.675.171.568.3
GPT-4o Audio58.158.064.651.8
Kimi-Audio69.679.065.564.4
Omni-R177.081.776.073.4
Qwen2.5-Omni71.578.170.665.9
Step-Audio-AQAA49.750.551.447.3
Step-Audio 278.083.576.973.7
Step-Audio 2 mini73.276.671.571.6

Speech translation

ModelCoVoST 2 (S2TT)
Avg.English-to-ChineseChinese-to-English
GPT-4o Audio29.6140.2019.01
Qwen2.5-Omni35.4041.4029.40
Step-Audio-AQAA28.5737.7119.43
Step-Audio 239.2649.0129.51
Step-Audio 2 mini39.2949.1229.47
ModelCVSS (S2ST)
Avg.English-to-ChineseChinese-to-English
GPT-4o Audio23.6820.0727.29
Qwen-Omni15.358.0422.66
Step-Audio-AQAA27.3630.7423.98
Step-Audio 230.8734.8326.92
Step-Audio 2 mini29.0832.8125.35

Tool calling

StepEval-Audio-Toolcall. Date and time tools have no parameter.

ModelObjectiveMetricAudio searchDate & TimeWeatherWeb search
Qwen3-32BTriggerPrecision / Recall67.5 / 98.598.4 / 100.090.1 / 100.086.8 / 98.5
TypeAccuracy100.0100.098.598.5
ParameterAccuracy100.0N/A100.0100.0
Step-Audio 2TriggerPrecision / Recall86.8 / 99.596.9 / 98.492.2 / 100.088.4 / 95.5
TypeAccuracy100.0100.090.598.4
ParameterAccuracy100.0N/A100.0100.0

Speech-to-speech conversation

URO-Bench. U. R. O. stands for understanding, reasoning, and oral conversation, respectively.

ModelLanguageBasicPro
Avg.U.R.O.Avg.U.R.O.
GPT-4o AudioChinese78.5989.4065.4885.2467.1070.6057.2270.20
Kimi-Audio73.5979.3464.6679.7566.0760.4459.2976.21
Qwen-Omni68.9859.6669.7477.2759.1159.0159.8258.74
Step-Audio-AQAA74.7187.6159.6381.9365.6174.7647.2968.97
Step-Audio 283.3291.0575.4586.0868.2574.7863.1865.10
Step-Audio 2 mini77.8189.1964.5384.1269.5776.8458.9069.42
GPT-4o AudioEnglish84.5490.1875.9090.4167.5160.6564.3678.46
Kimi-Audio60.0483.3642.3160.3649.7950.3240.5956.04
Qwen-Omni70.5866.2969.6276.1650.9944.5163.8849.41
Step-Audio-AQAA71.1190.1556.1272.0652.0144.2554.5459.81
Step-Audio 283.9092.7276.5184.9266.0764.8667.7566.33
Step-Audio 2 mini74.3690.0760.1277.6561.2558.7961.9463.80

License

The model and code in the repository is licensed under Apache 2.0 License.

Citation

@misc{wu2025stepaudio2technicalreport,
      title={Step-Audio 2 Technical Report},
      author={Boyong Wu and Chao Yan and Chen Hu and Cheng Yi and Chengli Feng and Fei Tian and Feiyu Shen and Gang Yu and Haoyang Zhang and Jingbei Li and Mingrui Chen and Peng Liu and Wang You and Xiangyu Tony Zhang and Xingyuan Li and Xuerui Yang and Yayue Deng and Yechang Huang and Yuxin Li and Yuxin Zhang and Zhao You and Brian Li and Changyi Wan and Hanpeng Hu and Jiangjie Zhen and Siyu Chen and Song Yuan and Xuelin Zhang and Yimin Jiang and Yu Zhou and Yuxiang Yang and Bingxin Li and Buyun Ma and Changhe Song and Dongqing Pang and Guoqiang Hu and Haiyang Sun and Kang An and Na Wang and Shuli Gao and Wei Ji and Wen Li and Wen Sun and Xuan Wen and Yong Ren and Yuankai Ma and Yufan Lu and Bin Wang and Bo Li and Changxin Miao and Che Liu and Chen Xu and Dapeng Shi and Dingyuan Hu and Donghang Wu and Enle Liu and Guanzhe Huang and Gulin Yan and Han Zhang and Hao Nie and Haonan Jia and Hongyu Zhou and Jianjian Sun and Jiaoren Wu and Jie Wu and Jie Yang and Jin Yang and Junzhe Lin and Kaixiang Li and Lei Yang and Liying Shi and Li Zhou and Longlong Gu and Ming Li and Mingliang Li and Mingxiao Li and Nan Wu and Qi Han and Qinyuan Tan and Shaoliang Pang and Shengjie Fan and Siqi Liu and Tiancheng Cao and Wanying Lu and Wenqing He and Wuxun Xie and Xu Zhao and Xueqi Li and Yanbo Yu and Yang Yang and Yi Liu and Yifan Lu and Yilei Wang and Yuanhao Ding and Yuanwei Liang and Yuanwei Lu and Yuchu Luo and Yuhe Yin and Yumeng Zhan and Yuxiang Zhang and Zidong Yang and Zixin Zhang and Binxing Jiao and Daxin Jiang and Heung-Yeung Shum and Jiansheng Chen and Jing Li and Xiangyu Zhang and Yibo Zhu},
      year={2025},
      eprint={2507.16632},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2507.16632},
}

Contributors

petronny

8 commits

LI
lijingbei

2 commits

stepfun-ai/Step-Audio-2-mini-Base

Model

26

stars

10

commits

1

repos using this model

1

linked in READMEs

Sep 1, 2025

updated

custom_code
onnx
safetensors
step_audio_2

README

GitHubHomepageTwitter FollowDiscord
License

Introduction

Step-Audio 2 is an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation.

  • Advanced Speech and Audio Understanding: Promising performance in ASR and audio understanding by comprehending and reasoning semantic information, para-linguistic and non-vocal information.

  • Intelligent Speech Conversation: Achieving natural and intelligent interactions that are contextually appropriate for various conversational scenarios and paralinguistic information.

  • Tool Calling and Multimodal RAG: By leveraging tool calling and RAG to access real-world knowledge (both textual and acoustic), Step-Audio 2 can generate responses with fewer hallucinations for diverse scenarios, while also having the ability to switch timbres based on retrieved speech.

  • State-of-the-Art Performance: Achieving state-of-the-art performance on various audio understanding and conversational benchmarks compared to other open-source and commercial solutions. (See Evaluation and Technical Report).

Model Download

Huggingface

Models🤗 Hugging Face
Step-Audio 2 ministepfun-ai/Step-Audio-2-mini
Step-Audio 2 mini Basestepfun-ai/Step-Audio-2-mini-Base

Model Usage

🔧 Dependencies and Installation

conda create -n stepaudio2 python=3.10
conda activate stepaudio2
pip install transformers==4.49.0 torchaudio librosa onnxruntime s3tokenizer diffusers hyperpyyaml

git clone https://github.com/stepfun-ai/Step-Audio2.git
cd Step-Audio2
git lfs install
git clone https://huggingface.co/stepfun-ai/Step-Audio-2-mini-Base

🚀 Inference Scripts

python examples-base.py

Online demonstration

StepFun realtime console

StepFun AI Assistant

  • Step-Audio 2 is also available in our StepFun AI Assistant mobile App with both web and audio search tools enabled.
  • Please scan the following QR code to download it from your app store then tap the phone icon in the top-right corner.
QR code

WeChat group

You can scan the following QR code to join our WeChat group for communication and discussion.

QR code

Evaluation

Architecture

Automatic speech recognition

CER for Chinese, Cantonese and Japanese and WER for Arabian and English. N/A indicates that the language is not supported.

CategoryTest setDoubao LLM ASRGPT-4o TranscribeKimi-AudioQwen-OmniStep-Audio 2Step-Audio 2 mini
EnglishCommon Voice9.209.307.838.335.956.76
FLEURS English7.222.714.475.053.033.05
LibriSpeech clean2.921.751.492.931.171.33
LibriSpeech other5.324.232.915.072.422.86
Average6.174.504.185.353.143.50
ChineseAISHELL0.983.520.641.170.630.78
AISHELL-23.104.262.672.402.102.16
FLEURS Chinese2.922.622.917.012.682.53
KeSpeech phase16.4826.805.116.453.633.97
WenetSpeech meeting4.9031.405.216.614.754.87
WenetSpeech net4.4615.715.935.244.674.82
Average3.8114.053.754.813.083.19
Multilingual FLEURS ArabianN/A11.72N/A25.1314.2216.46
Common Voice yue9.2011.1038.907.897.908.32
FLEURS JapaneseN/A3.27N/A10.493.184.67
In-houseAnhui accent8.8350.5522.1718.7310.6111.65
Guangdong accent4.997.833.764.033.814.44
Guangxi accent3.377.094.293.354.113.51
Shanxi accent20.2655.0334.7125.9512.4415.60
Sichuan dialect3.0132.855.265.614.354.57
Shanghai dialect47.4989.5882.9058.7417.7719.30
Average14.6640.4925.5219.408.859.85

Paralinguistic information understanding

StepEval-Audio-Paralinguistic

ModelAvg.GenderAgeTimbreScenarioEventEmotionPitchRhythmSpeedStyleVocal
GPT-4o Audio43.451842342214824060586444
Kimi-Audio49.649450103048665640445454
Qwen-Omni44.184050162842763254505048
Step-Audio-AQAA36.91706618141440384854440
Step-Audio 283.0910096827860868286888868
Step-Audio 2 mini80.0010094807860828268748676

Audio understanding and reasoning

MMAU

ModelAvg.SoundSpeechMusic
Audio Flamingo 373.176.966.173.9
Gemini 2.5 Pro71.675.171.568.3
GPT-4o Audio58.158.064.651.8
Kimi-Audio69.679.065.564.4
Omni-R177.081.776.073.4
Qwen2.5-Omni71.578.170.665.9
Step-Audio-AQAA49.750.551.447.3
Step-Audio 278.083.576.973.7
Step-Audio 2 mini73.276.671.571.6

Speech translation

ModelCoVoST 2 (S2TT)
Avg.English-to-ChineseChinese-to-English
GPT-4o Audio29.6140.2019.01
Qwen2.5-Omni35.4041.4029.40
Step-Audio-AQAA28.5737.7119.43
Step-Audio 239.2649.0129.51
Step-Audio 2 mini39.2949.1229.47
ModelCVSS (S2ST)
Avg.English-to-ChineseChinese-to-English
GPT-4o Audio23.6820.0727.29
Qwen-Omni15.358.0422.66
Step-Audio-AQAA27.3630.7423.98
Step-Audio 230.8734.8326.92
Step-Audio 2 mini29.0832.8125.35

Tool calling

StepEval-Audio-Toolcall. Date and time tools have no parameter.

ModelObjectiveMetricAudio searchDate & TimeWeatherWeb search
Qwen3-32BTriggerPrecision / Recall67.5 / 98.598.4 / 100.090.1 / 100.086.8 / 98.5
TypeAccuracy100.0100.098.598.5
ParameterAccuracy100.0N/A100.0100.0
Step-Audio 2TriggerPrecision / Recall86.8 / 99.596.9 / 98.492.2 / 100.088.4 / 95.5
TypeAccuracy100.0100.090.598.4
ParameterAccuracy100.0N/A100.0100.0

Speech-to-speech conversation

URO-Bench. U. R. O. stands for understanding, reasoning, and oral conversation, respectively.

ModelLanguageBasicPro
Avg.U.R.O.Avg.U.R.O.
GPT-4o AudioChinese78.5989.4065.4885.2467.1070.6057.2270.20
Kimi-Audio73.5979.3464.6679.7566.0760.4459.2976.21
Qwen-Omni68.9859.6669.7477.2759.1159.0159.8258.74
Step-Audio-AQAA74.7187.6159.6381.9365.6174.7647.2968.97
Step-Audio 283.3291.0575.4586.0868.2574.7863.1865.10
Step-Audio 2 mini77.8189.1964.5384.1269.5776.8458.9069.42
GPT-4o AudioEnglish84.5490.1875.9090.4167.5160.6564.3678.46
Kimi-Audio60.0483.3642.3160.3649.7950.3240.5956.04
Qwen-Omni70.5866.2969.6276.1650.9944.5163.8849.41
Step-Audio-AQAA71.1190.1556.1272.0652.0144.2554.5459.81
Step-Audio 283.9092.7276.5184.9266.0764.8667.7566.33
Step-Audio 2 mini74.3690.0760.1277.6561.2558.7961.9463.80

License

The model and code in the repository is licensed under Apache 2.0 License.

Citation

@misc{wu2025stepaudio2technicalreport,
      title={Step-Audio 2 Technical Report},
      author={Boyong Wu and Chao Yan and Chen Hu and Cheng Yi and Chengli Feng and Fei Tian and Feiyu Shen and Gang Yu and Haoyang Zhang and Jingbei Li and Mingrui Chen and Peng Liu and Wang You and Xiangyu Tony Zhang and Xingyuan Li and Xuerui Yang and Yayue Deng and Yechang Huang and Yuxin Li and Yuxin Zhang and Zhao You and Brian Li and Changyi Wan and Hanpeng Hu and Jiangjie Zhen and Siyu Chen and Song Yuan and Xuelin Zhang and Yimin Jiang and Yu Zhou and Yuxiang Yang and Bingxin Li and Buyun Ma and Changhe Song and Dongqing Pang and Guoqiang Hu and Haiyang Sun and Kang An and Na Wang and Shuli Gao and Wei Ji and Wen Li and Wen Sun and Xuan Wen and Yong Ren and Yuankai Ma and Yufan Lu and Bin Wang and Bo Li and Changxin Miao and Che Liu and Chen Xu and Dapeng Shi and Dingyuan Hu and Donghang Wu and Enle Liu and Guanzhe Huang and Gulin Yan and Han Zhang and Hao Nie and Haonan Jia and Hongyu Zhou and Jianjian Sun and Jiaoren Wu and Jie Wu and Jie Yang and Jin Yang and Junzhe Lin and Kaixiang Li and Lei Yang and Liying Shi and Li Zhou and Longlong Gu and Ming Li and Mingliang Li and Mingxiao Li and Nan Wu and Qi Han and Qinyuan Tan and Shaoliang Pang and Shengjie Fan and Siqi Liu and Tiancheng Cao and Wanying Lu and Wenqing He and Wuxun Xie and Xu Zhao and Xueqi Li and Yanbo Yu and Yang Yang and Yi Liu and Yifan Lu and Yilei Wang and Yuanhao Ding and Yuanwei Liang and Yuanwei Lu and Yuchu Luo and Yuhe Yin and Yumeng Zhan and Yuxiang Zhang and Zidong Yang and Zixin Zhang and Binxing Jiao and Daxin Jiang and Heung-Yeung Shum and Jiansheng Chen and Jing Li and Xiangyu Zhang and Yibo Zhu},
      year={2025},
      eprint={2507.16632},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2507.16632},
}

Contributors

petronny

8 commits

LI
lijingbei

2 commits