Paper: Step-Audio 2 Technical Report
Code: https://github.com/stepfun-ai/Step-Audio2
Project Page: https://www.stepfun.com/docs/en/step-audio2
StepEval-Audio-Paralinguistic is a speech-to-speech benchmark designed to evaluate AI models' understanding of paralinguistic information in speech across 11 distinct dimensions. The dataset contains 550 carefully curated and annotated speech samples for assessing capabilities beyond semantic understanding.
Basic Attributes
Speech Characteristics
Environmental Sounds
| Category | Task Description | Label Distribution | Total Samples |
|---|---|---|---|
| Gender | Identify speaker's gender | Male: 25, Female: 25 | 50 |
| Age | Classify speaker's age | 20y:6, 25y:6, 30y:5, 35y:5, 40y:5, 45y:4, 50y:4 + Child:7, Elderly:8 | 50 |
| Speed | Categorize speaking speed | Slow:10, Medium-slow:10, Medium:10, Medium-fast:10, Fast:10 | 50 |
| Emotion | Recognize emotional states | Anger, Joy, Sadness, Surprise, Sarcasm, etc. (50 manually annotated) | 50 |
| Scenarios | Detect background scenes | Indoor:14, Outdoor:12, Restaurant:6, Kitchen:6, Park:6, Subway:6 | 50 |
| Vocal | Identify non-speech vocal effects | Cough:14, Sniff:8, Sneeze:7, Throat-clearing:6, Laugh:5, Sigh:5, Other:5 | 50 |
| Style | Distinguish speaking styles | Dialogue:4, Discussion:4, Narration:8, Commentary:8, Colloquial:8, Speech:8, Other:10 | 50 |
| Rhythm | Characterize rhythm patterns | Steady:10, Fluent:10, Paused:10, Hurried:10, Fluctuating:10 | 50 |
| Pitch | Classify dominant pitch ranges | Mid:12, Mid-high:14, High:12, Mid-low:12 | 50 |
| Event | Recognize non-vocal audio events | Music:8, Other events:42 (from AudioSet) | 50 |
Dataset Notes:
The benchmark evaluation follows a standardized three-phase process:
Audio-in/audio-out models are queried through their APIs using the original audio files as input. Each 24kHz audio sample (≤30s duration) generates a corresponding response audio, saved with matching filenames for traceability.
All model response audios are transcribed using a ASR system. Transcripts undergo automatic text normalization and are stored.
The evaluation script (LLM_judge.py) compares ASR transcripts against ground truth annotations using an LLM judge. Scoring considers semantic similarity rather than exact matches, with partial credit for partially correct responses. The final metrics include per-category accuracy scores.
| Model | Avg | Gender | Age | Timbre | Scenario | Event | Emotion | Pitch | Rhythm | Speed | Style | Vocal |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-4o Audio | 43.45 | 18 | 42 | 34 | 22 | 14 | 82 | 40 | 60 | 58 | 64 | 44 |
| Kimi-Audio | 49.64 | 94 | 50 | 10 | 30 | 48 | 66 | 56 | 40 | 44 | 54 | 54 |
| Qwen-Omni | 44.18 | 40 | 50 | 16 | 28 | 42 | 76 | 32 | 54 | 50 | 50 | 48 |
| Step-Audio-AQAA | 36.91 | 70 | 66 | 18 | 14 | 14 | 40 | 38 | 48 | 54 | 44 | 0 |
| Step-Audio 2 | 76.55 | 98 | 92 | 78 | 64 | 46 | 72 | 78 | 70 | 78 | 84 | 82 |
Paper: Step-Audio 2 Technical Report
Code: https://github.com/stepfun-ai/Step-Audio2
Project Page: https://www.stepfun.com/docs/en/step-audio2
StepEval-Audio-Paralinguistic is a speech-to-speech benchmark designed to evaluate AI models' understanding of paralinguistic information in speech across 11 distinct dimensions. The dataset contains 550 carefully curated and annotated speech samples for assessing capabilities beyond semantic understanding.
Basic Attributes
Speech Characteristics
Environmental Sounds
| Category | Task Description | Label Distribution | Total Samples |
|---|---|---|---|
| Gender | Identify speaker's gender | Male: 25, Female: 25 | 50 |
| Age | Classify speaker's age | 20y:6, 25y:6, 30y:5, 35y:5, 40y:5, 45y:4, 50y:4 + Child:7, Elderly:8 | 50 |
| Speed | Categorize speaking speed | Slow:10, Medium-slow:10, Medium:10, Medium-fast:10, Fast:10 | 50 |
| Emotion | Recognize emotional states | Anger, Joy, Sadness, Surprise, Sarcasm, etc. (50 manually annotated) | 50 |
| Scenarios | Detect background scenes | Indoor:14, Outdoor:12, Restaurant:6, Kitchen:6, Park:6, Subway:6 | 50 |
| Vocal | Identify non-speech vocal effects | Cough:14, Sniff:8, Sneeze:7, Throat-clearing:6, Laugh:5, Sigh:5, Other:5 | 50 |
| Style | Distinguish speaking styles | Dialogue:4, Discussion:4, Narration:8, Commentary:8, Colloquial:8, Speech:8, Other:10 | 50 |
| Rhythm | Characterize rhythm patterns | Steady:10, Fluent:10, Paused:10, Hurried:10, Fluctuating:10 | 50 |
| Pitch | Classify dominant pitch ranges | Mid:12, Mid-high:14, High:12, Mid-low:12 | 50 |
| Event | Recognize non-vocal audio events | Music:8, Other events:42 (from AudioSet) | 50 |
Dataset Notes:
The benchmark evaluation follows a standardized three-phase process:
Audio-in/audio-out models are queried through their APIs using the original audio files as input. Each 24kHz audio sample (≤30s duration) generates a corresponding response audio, saved with matching filenames for traceability.
All model response audios are transcribed using a ASR system. Transcripts undergo automatic text normalization and are stored.
The evaluation script (LLM_judge.py) compares ASR transcripts against ground truth annotations using an LLM judge. Scoring considers semantic similarity rather than exact matches, with partial credit for partially correct responses. The final metrics include per-category accuracy scores.
| Model | Avg | Gender | Age | Timbre | Scenario | Event | Emotion | Pitch | Rhythm | Speed | Style | Vocal |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-4o Audio | 43.45 | 18 | 42 | 34 | 22 | 14 | 82 | 40 | 60 | 58 | 64 | 44 |
| Kimi-Audio | 49.64 | 94 | 50 | 10 | 30 | 48 | 66 | 56 | 40 | 44 | 54 | 54 |
| Qwen-Omni | 44.18 | 40 | 50 | 16 | 28 | 42 | 76 | 32 | 54 | 50 | 50 | 48 |
| Step-Audio-AQAA | 36.91 | 70 | 66 | 18 | 14 | 14 | 40 | 38 | 48 | 54 | 44 | 0 |
| Step-Audio 2 | 76.55 | 98 | 92 | 78 | 64 | 46 | 72 | 78 | 70 | 78 | 84 | 82 |