stepfun-ai/Step-Audio-Chat

Model

1. Step-Audio-Chat

461

12 commits

2 linked in READMEs

updated Feb 17, 2025

See the code

README

1. Step-Audio-Chat

This repository contains the Multimodal Large Language Model (LLM) component of Step-Audio. It is a 130 billion parameter multimodal LLM that is responsible for understanding and generating human speech. The model is specifically designed to seamlessly integrate functions such as speech recognition, semantic understanding, dialogue management, voice cloning, and speech generation.

2. Evaluation

2.1 LLM judge metrics(GPT-4o) on StepEval-Audio-360

Comparison of fundamental capabilities of voice chat on the StepEval-Audio-360.
ModelFactuality (% ↑)Relevance (% ↑)Chat Score ↑
GLM4-Voice54.766.43.49
Qwen2-Audio22.626.32.27
Moshi*1.001.49
Step-Audio-Chat66.475.24.11

*Note: Moshi are marked with "*" and should be considered for reference only.

2.2 Public Test Set

ModelLlama QuestionWeb QuestionsTriviaQA*ComplexBenchHSK-6
GLM4-Voice64.732.239.166.074.0
Moshi62.326.622.8--
Freeze-Omni72.044.753.9--
LUCY59.729.327.0--
MinMo78.955.048.3--
Qwen2-Audio52.027.037.354.0-
Step-Audio-Chat81.075.158.074.086.0

Note: Results marked with "*" on TriviaQA dataset are considered for reference only.

TriviaQA dataset marked with "*" indicates results are for reference only.

2.3 Audio instruction following

CategoryInstruction FollowingAudio Quality
GLM-4-VoiceStep-AudioGLM-4-VoiceStep-Audio
Languages1.93.82.93.3
Role-playing3.84.23.23.6
Singing / RAP2.12.42.44
Voice Control3.64.43.34.1

3. More information

For more information, please refer to our repository: Step-Audio.

audio-text-to-text
custom_code
safetensors
step1

Contributors

buyun

9 commits

jckhang

1 commits

reach-vb

1 commits

ryanmiao

1 commits

stepfun-ai/Step-Audio-Chat

Model

1. Step-Audio-Chat

461

12 commits

2 linked in READMEs

updated Feb 17, 2025

See the code

README

1. Step-Audio-Chat

This repository contains the Multimodal Large Language Model (LLM) component of Step-Audio. It is a 130 billion parameter multimodal LLM that is responsible for understanding and generating human speech. The model is specifically designed to seamlessly integrate functions such as speech recognition, semantic understanding, dialogue management, voice cloning, and speech generation.

2. Evaluation

2.1 LLM judge metrics(GPT-4o) on StepEval-Audio-360

Comparison of fundamental capabilities of voice chat on the StepEval-Audio-360.
ModelFactuality (% ↑)Relevance (% ↑)Chat Score ↑
GLM4-Voice54.766.43.49
Qwen2-Audio22.626.32.27
Moshi*1.001.49
Step-Audio-Chat66.475.24.11

*Note: Moshi are marked with "*" and should be considered for reference only.

2.2 Public Test Set

ModelLlama QuestionWeb QuestionsTriviaQA*ComplexBenchHSK-6
GLM4-Voice64.732.239.166.074.0
Moshi62.326.622.8--
Freeze-Omni72.044.753.9--
LUCY59.729.327.0--
MinMo78.955.048.3--
Qwen2-Audio52.027.037.354.0-
Step-Audio-Chat81.075.158.074.086.0

Note: Results marked with "*" on TriviaQA dataset are considered for reference only.

TriviaQA dataset marked with "*" indicates results are for reference only.

2.3 Audio instruction following

CategoryInstruction FollowingAudio Quality
GLM-4-VoiceStep-AudioGLM-4-VoiceStep-Audio
Languages1.93.82.93.3
Role-playing3.84.23.23.6
Singing / RAP2.12.42.44
Voice Control3.64.43.34.1

3. More information

For more information, please refer to our repository: Step-Audio.

audio-text-to-text
custom_code
safetensors
step1

Contributors

buyun

9 commits

jckhang

1 commits

reach-vb

1 commits

ryanmiao

1 commits