NARUTO-2024/WavBench

WavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue Models

Python

38

15 commits

updated Feb 13, 2026

See the code

README

WavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue Models

📖 Paper | 🏠 Website | 🤗 Dataset (HuggingFace)

Overview of WavBench


Figure 1: Examples of Colloquial Expression in WavBench, covering diverse cognitive domains across Basic and Pro subsets.


Figure 2: Examples of Acoustic Interaction in WavBench, demonstrating Explicit Understanding, Explicit Generation, and Implicit Dialogue.

News

  • 2026.02.11 Released the WavBench paper, code, and dataset.
  • 2026.02.11 Released the leaderboard evaluating 5 state-of-the-art E2E spoken dialogue models.

Table of Contents

Leaderboard

Below is the overall evaluation of WavBench across five panels: Colloquial Expression (Pro & Basic) and Acoustic Interaction (Explicit Understanding, Explicit Generation, and Implicit).

Metrics / TasksQwen3-OmniKimi-AudioMimo-AudioStep-Audio-2GPT-4o Audio
Panel A: Colloquial (Pro)
Code39.7530.2928.9631.2053.60
Creativity48.3931.7842.8635.0063.00
Instruction43.0129.8636.4429.4057.80
Logic33.2126.0327.5726.2042.60
Math38.5527.3025.6822.4050.20
QA50.9342.5441.2840.8072.80
Safety60.0056.1956.1952.4067.60
Avg (Pro)39.5330.7932.0230.4058.23
Panel B: Colloquial (Basic)
Code53.1040.6942.0737.2058.00
Creativity57.4441.5745.2947.2071.20
Instruction57.2944.4133.5636.6066.80
Logic52.3550.7449.9148.8067.00
Math51.0541.2738.7330.2062.40
QA57.5449.0749.1248.6075.60
Safety59.6758.8362.8360.2081.00
Avg (Basic)55.8049.2349.5748.5068.80
Panel C: Explicit Understanding
Accent37.5011.0027.0020.6715.67
Age64.3353.6753.0067.6720.33
Emotion92.8677.3377.3375.4385.90
Gender21.0044.5020.0068.0061.50
Language83.5091.0053.5096.5097.00
Pitch32.4423.1124.0034.2223.56
Speed46.6754.6748.8944.0048.00
Volume33.7838.2231.1150.6741.78
Audio Event61.7367.9019.7539.5159.26
Music22.2266.6755.5677.7833.33
Avg (Understand)49.6052.8041.0257.3648.70
Panel D: Explicit Generation
Accent37.503.5223.4422.0774.22
Age64.6546.8851.9531.6478.12
Emotion90.0450.2957.1366.5095.51
Gender72.2745.3167.5859.7798.83
Language89.8474.8051.5691.4187.89
Pitch76.5647.2780.2755.6685.74
Speed43.7547.2751.5669.1466.60
Volume56.2564.0659.9657.0382.42
Audio27.0310.819.4632.4345.95
Music62.5020.8316.6770.8377.08
Avg (Generation)62.0341.1046.9355.6579.23
Panel E: Implicit
Single-Turn (Text)1.851.842.231.122.43
Single-Turn (Audio)3.173.212.473.502.96
Multi-Turn (Text)4.884.574.614.384.48
Multi-Turn (Audio)1.251.081.041.211.23
Avg (Implicit)2.782.672.592.552.78

Setup

conda create -n wavbench python=3.10
conda activate wavbench
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu124
pip install -r requirements.txt

git clone https://github.com/NARUTO-2024/WavBench.git
cd WavBench

Dataset

The data used in this project is available at WavBench Dataset hosted on Hugging Face.

You can load the dataset directly using the Hugging Face datasets library:

from datasets import load_dataset

# Load the dataset directly from Hugging Face
ds = load_dataset("WavBench/WavBench")

Alternatively, you can download the dataset to your local directory and use it directly.

1. Colloquial Expression

This category is divided into Basic and Pro subsets. Each subset contains tasks across 7 diverse cognitive domains:

DomainDescription
CodeEvaluate the model's ability to explain code logic conversationally.
CreativeEvaluate creative writing without rigid formatting constraints.
InstructionEvaluate adherence to spoken instructions.
LogicEvaluate logical reasoning in a spoken context.
MathEvaluate the verbalization of mathematical reasoning.
QAEvaluate general knowledge answering capabilities.
SafetyEvaluate safety mechanisms in spoken interaction.

2. Acoustic Interaction

This category evaluates the model's paralinguistic capabilities across three dimensions: Explicit Understanding, Explicit Generation, and Implicit.

CategorySub-tasks / Attributes
Explicit Understanding10 Attributes: Accent, Age, Emotion, Gender, Language, Pitch, Speed, Volume, Audio, Music.
Explicit Generation10 Attributes: Accent, Age, Emotion, Gender, Language, Pitch, Speed, Volume, Audio, Music.
ImplicitSingle-turn Audio, Single-turn Text, Multi-turn Audio, Multi-turn Text.

Evaluation

Step 1: Run Inference

main.py is the unified entry point for all dataset types.

# Colloquial Inference (Basic) - With audio output
python main.py --model step_audio2 --data basic_code --audio_output

# Colloquial Inference (Pro) - With audio output
python main.py --model step_audio2 --data pro_math --audio_output

# Acoustic Single-turn Inference (With audio output)
python main.py --model step_audio2 --data acoustic_explicit_generation_emotion --audio_output

# Acoustic Multi-round Inference (With audio output)
python main.py --model step_audio2 --data acoustic_multi_round_generation --audio_output

# [Optional] Run with custom data directory
python main.py --model step_audio2 --data basic_code --data_dir /path/to/your/wavbench

Supported Arguments:

  • --model: Model name (e.g., step_audio2).
  • --data: Dataset name (e.g., basic_code, pro_math, acoustic_explicit_generation_emotion).
  • --data_dir: Optional. Base directory for WavBench data (Default: ./wavbench). Use this argument if you have downloaded the dataset to a specific location other than the default.
  • --audio_output: Important Flag. If set, the model generates audio files in addition to text.
    • Required for all Acoustic tasks (as evaluation relies on audio).
    • Optional for Colloquial tasks (useful if you want to check the TTS quality manually).

Step 2: Automatic Evaluation

evaluate.py uses LLMs (Gemini) to judge the responses based on the specific criteria of each subset.

# Option 1: Set API key via environment variable
export GOOGLE_API_KEY="your-api-key"

# Evaluate ALL Colloquial datasets
python evaluate.py --eval_type colloquial --dataset all

# Evaluate a SPECIFIC Colloquial dataset
python evaluate.py --eval_type colloquial --dataset basic_code

# Evaluate ALL Acoustic datasets
python evaluate.py --eval_type acoustic --dataset all

# Evaluate a SPECIFIC Acoustic dataset
python evaluate.py --eval_type acoustic --dataset explicit_generation_emotion

Supported Arguments:

  • --eval_type: Choose between colloquial or acoustic.
  • --dataset: Specific dataset name (e.g., basic_code) or use all to run the entire suite.
👇 Available Dataset Options (for --data / --dataset)
CategoryAvailable Values
Basicbasic_code, basic_creative, basic_instruction, basic_logic, basic_math, basic_qa, basic_satety
Propro_code, pro_creative, pro_instruction, pro_logic, pro_math, pro_qa, pro_satety
Explicit Generationacoustic_explicit_generation_accent, acoustic_explicit_generation_age, acoustic_explicit_generation_audio, acoustic_explicit_generation_emotion, acoustic_explicit_generation_gender, acoustic_explicit_generation_lang, acoustic_explicit_generation_music, acoustic_explicit_generation_pitch, acoustic_explicit_generation_speed, acoustic_explicit_generation_volume
Explicit Understandingacoustic_explicit_understanding_accent, acoustic_explicit_understanding_age, acoustic_explicit_understanding_audio, acoustic_explicit_understanding_emotion, acoustic_explicit_understanding_gender, acoustic_explicit_understanding_lang, acoustic_explicit_understanding_music, acoustic_explicit_understanding_pitch, acoustic_explicit_understanding_speed, acoustic_explicit_understanding_volume
Implicitacoustic_implicit_age_generation, acoustic_implicit_emotion_generation, acoustic_implicit_pitch_generation, acoustic_implicit_speed_generation, acoustic_implicit_understanding
Multi-roundacoustic_multi_round_generation, acoustic_multi_round_understanding

Step 3: Get Statistics

statistics.py aggregates the evaluation results into a final report.

# Basic usage: Output to TXT file
python statistics.py --eval_dir ./eval_results --output ./statistics.txt

# Advanced usage: Output to TXT and CSV format simultaneously
python statistics.py --eval_dir ./eval_results --output ./statistics.txt --csv

Citation

If you use WavBench in your research, please cite the following paper:

@misc{li2026wavbenchbenchmarkingreasoningcolloquialism,
      title={WavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue Models}, 
      author={Yangzhuo Li and Shengpeng Ji and Yifu Chen and Tianle Liang and Haorong Ying and Yule Wang and Junbo Li and Jun Fang and Zhou Zhao},
      year={2026},
      eprint={2602.12135},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2602.12135}, 
}

Contributors

NARUTO-2024

15 commits

NARUTO-2024/WavBench

WavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue Models

Python

38

15 commits

updated Feb 13, 2026

See the code

README

WavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue Models

📖 Paper | 🏠 Website | 🤗 Dataset (HuggingFace)

Overview of WavBench


Figure 1: Examples of Colloquial Expression in WavBench, covering diverse cognitive domains across Basic and Pro subsets.


Figure 2: Examples of Acoustic Interaction in WavBench, demonstrating Explicit Understanding, Explicit Generation, and Implicit Dialogue.

News

  • 2026.02.11 Released the WavBench paper, code, and dataset.
  • 2026.02.11 Released the leaderboard evaluating 5 state-of-the-art E2E spoken dialogue models.

Table of Contents

Leaderboard

Below is the overall evaluation of WavBench across five panels: Colloquial Expression (Pro & Basic) and Acoustic Interaction (Explicit Understanding, Explicit Generation, and Implicit).

Metrics / TasksQwen3-OmniKimi-AudioMimo-AudioStep-Audio-2GPT-4o Audio
Panel A: Colloquial (Pro)
Code39.7530.2928.9631.2053.60
Creativity48.3931.7842.8635.0063.00
Instruction43.0129.8636.4429.4057.80
Logic33.2126.0327.5726.2042.60
Math38.5527.3025.6822.4050.20
QA50.9342.5441.2840.8072.80
Safety60.0056.1956.1952.4067.60
Avg (Pro)39.5330.7932.0230.4058.23
Panel B: Colloquial (Basic)
Code53.1040.6942.0737.2058.00
Creativity57.4441.5745.2947.2071.20
Instruction57.2944.4133.5636.6066.80
Logic52.3550.7449.9148.8067.00
Math51.0541.2738.7330.2062.40
QA57.5449.0749.1248.6075.60
Safety59.6758.8362.8360.2081.00
Avg (Basic)55.8049.2349.5748.5068.80
Panel C: Explicit Understanding
Accent37.5011.0027.0020.6715.67
Age64.3353.6753.0067.6720.33
Emotion92.8677.3377.3375.4385.90
Gender21.0044.5020.0068.0061.50
Language83.5091.0053.5096.5097.00
Pitch32.4423.1124.0034.2223.56
Speed46.6754.6748.8944.0048.00
Volume33.7838.2231.1150.6741.78
Audio Event61.7367.9019.7539.5159.26
Music22.2266.6755.5677.7833.33
Avg (Understand)49.6052.8041.0257.3648.70
Panel D: Explicit Generation
Accent37.503.5223.4422.0774.22
Age64.6546.8851.9531.6478.12
Emotion90.0450.2957.1366.5095.51
Gender72.2745.3167.5859.7798.83
Language89.8474.8051.5691.4187.89
Pitch76.5647.2780.2755.6685.74
Speed43.7547.2751.5669.1466.60
Volume56.2564.0659.9657.0382.42
Audio27.0310.819.4632.4345.95
Music62.5020.8316.6770.8377.08
Avg (Generation)62.0341.1046.9355.6579.23
Panel E: Implicit
Single-Turn (Text)1.851.842.231.122.43
Single-Turn (Audio)3.173.212.473.502.96
Multi-Turn (Text)4.884.574.614.384.48
Multi-Turn (Audio)1.251.081.041.211.23
Avg (Implicit)2.782.672.592.552.78

Setup

conda create -n wavbench python=3.10
conda activate wavbench
pip install torch==2.6.0 torchvision==0.21.0 torchaudio==2.6.0 --index-url https://download.pytorch.org/whl/cu124
pip install -r requirements.txt

git clone https://github.com/NARUTO-2024/WavBench.git
cd WavBench

Dataset

The data used in this project is available at WavBench Dataset hosted on Hugging Face.

You can load the dataset directly using the Hugging Face datasets library:

from datasets import load_dataset

# Load the dataset directly from Hugging Face
ds = load_dataset("WavBench/WavBench")

Alternatively, you can download the dataset to your local directory and use it directly.

1. Colloquial Expression

This category is divided into Basic and Pro subsets. Each subset contains tasks across 7 diverse cognitive domains:

DomainDescription
CodeEvaluate the model's ability to explain code logic conversationally.
CreativeEvaluate creative writing without rigid formatting constraints.
InstructionEvaluate adherence to spoken instructions.
LogicEvaluate logical reasoning in a spoken context.
MathEvaluate the verbalization of mathematical reasoning.
QAEvaluate general knowledge answering capabilities.
SafetyEvaluate safety mechanisms in spoken interaction.

2. Acoustic Interaction

This category evaluates the model's paralinguistic capabilities across three dimensions: Explicit Understanding, Explicit Generation, and Implicit.

CategorySub-tasks / Attributes
Explicit Understanding10 Attributes: Accent, Age, Emotion, Gender, Language, Pitch, Speed, Volume, Audio, Music.
Explicit Generation10 Attributes: Accent, Age, Emotion, Gender, Language, Pitch, Speed, Volume, Audio, Music.
ImplicitSingle-turn Audio, Single-turn Text, Multi-turn Audio, Multi-turn Text.

Evaluation

Step 1: Run Inference

main.py is the unified entry point for all dataset types.

# Colloquial Inference (Basic) - With audio output
python main.py --model step_audio2 --data basic_code --audio_output

# Colloquial Inference (Pro) - With audio output
python main.py --model step_audio2 --data pro_math --audio_output

# Acoustic Single-turn Inference (With audio output)
python main.py --model step_audio2 --data acoustic_explicit_generation_emotion --audio_output

# Acoustic Multi-round Inference (With audio output)
python main.py --model step_audio2 --data acoustic_multi_round_generation --audio_output

# [Optional] Run with custom data directory
python main.py --model step_audio2 --data basic_code --data_dir /path/to/your/wavbench

Supported Arguments:

  • --model: Model name (e.g., step_audio2).
  • --data: Dataset name (e.g., basic_code, pro_math, acoustic_explicit_generation_emotion).
  • --data_dir: Optional. Base directory for WavBench data (Default: ./wavbench). Use this argument if you have downloaded the dataset to a specific location other than the default.
  • --audio_output: Important Flag. If set, the model generates audio files in addition to text.
    • Required for all Acoustic tasks (as evaluation relies on audio).
    • Optional for Colloquial tasks (useful if you want to check the TTS quality manually).

Step 2: Automatic Evaluation

evaluate.py uses LLMs (Gemini) to judge the responses based on the specific criteria of each subset.

# Option 1: Set API key via environment variable
export GOOGLE_API_KEY="your-api-key"

# Evaluate ALL Colloquial datasets
python evaluate.py --eval_type colloquial --dataset all

# Evaluate a SPECIFIC Colloquial dataset
python evaluate.py --eval_type colloquial --dataset basic_code

# Evaluate ALL Acoustic datasets
python evaluate.py --eval_type acoustic --dataset all

# Evaluate a SPECIFIC Acoustic dataset
python evaluate.py --eval_type acoustic --dataset explicit_generation_emotion

Supported Arguments:

  • --eval_type: Choose between colloquial or acoustic.
  • --dataset: Specific dataset name (e.g., basic_code) or use all to run the entire suite.
👇 Available Dataset Options (for --data / --dataset)
CategoryAvailable Values
Basicbasic_code, basic_creative, basic_instruction, basic_logic, basic_math, basic_qa, basic_satety
Propro_code, pro_creative, pro_instruction, pro_logic, pro_math, pro_qa, pro_satety
Explicit Generationacoustic_explicit_generation_accent, acoustic_explicit_generation_age, acoustic_explicit_generation_audio, acoustic_explicit_generation_emotion, acoustic_explicit_generation_gender, acoustic_explicit_generation_lang, acoustic_explicit_generation_music, acoustic_explicit_generation_pitch, acoustic_explicit_generation_speed, acoustic_explicit_generation_volume
Explicit Understandingacoustic_explicit_understanding_accent, acoustic_explicit_understanding_age, acoustic_explicit_understanding_audio, acoustic_explicit_understanding_emotion, acoustic_explicit_understanding_gender, acoustic_explicit_understanding_lang, acoustic_explicit_understanding_music, acoustic_explicit_understanding_pitch, acoustic_explicit_understanding_speed, acoustic_explicit_understanding_volume
Implicitacoustic_implicit_age_generation, acoustic_implicit_emotion_generation, acoustic_implicit_pitch_generation, acoustic_implicit_speed_generation, acoustic_implicit_understanding
Multi-roundacoustic_multi_round_generation, acoustic_multi_round_understanding

Step 3: Get Statistics

statistics.py aggregates the evaluation results into a final report.

# Basic usage: Output to TXT file
python statistics.py --eval_dir ./eval_results --output ./statistics.txt

# Advanced usage: Output to TXT and CSV format simultaneously
python statistics.py --eval_dir ./eval_results --output ./statistics.txt --csv

Citation

If you use WavBench in your research, please cite the following paper:

@misc{li2026wavbenchbenchmarkingreasoningcolloquialism,
      title={WavBench: Benchmarking Reasoning, Colloquialism, and Paralinguistics for End-to-End Spoken Dialogue Models}, 
      author={Yangzhuo Li and Shengpeng Ji and Yifu Chen and Tianle Liang and Haorong Ying and Yule Wang and Junbo Li and Jun Fang and Zhou Zhao},
      year={2026},
      eprint={2602.12135},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2602.12135}, 
}

Contributors

NARUTO-2024

15 commits

Languages

Python

99.9%