Ruiqi-Yan/URO-Bench

Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Python

58

16 commits

updated Sep 2, 2025

See the code

README

URO-Bench

Official code for evaluating spoken dialogue models with
URO-Bench: A Comprehensive Benchmark for End-to-End Spoken Dialogue Models

URO-Bench Logo

version version arXiv Hugging Face mit

News

  • [Update Aug. 03, 2025] We have reconstructed the 6 datasets in the Chinese Basic Track and added four new ones — Wildchat-zh, HSK5-zh, APE-zh, and SQuAD-zh — so that the Chinese and English datasets are now fully paired. Safety-en, Safety-zh, and SRT-zh were expanded to a total of 25 samples. All 40 datasets are available on HuggingFace, and we also provide a curated miniset of 1,000 samples for quick evaluation before full-scale assessment. The corresponding test results have also been updated.
  • [Update Feb. 25, 2025] 🔥🔥🔥 code and data of URO-Bench have been released!

Overview

This repo contains the code of URO-Bench: A Comprehensive Benchmark for End-to-End Spoken Dialogue Models.

Recent advances in large language models (LLMs) have driven significant progress in end-to-end spoken dialogue models (SDMs). In contrast to text-based LLMs, the evaluation framework for SDMs should encompass both cognitive dimensions (e.g., logical reasoning, knowledge) and speech-related aspects (e.g., paralinguistic cues, audio quality). However, there is still a lack of comprehensive evaluations for SDMs in speech-to-speech (S2S) scenarios.

To address this gap, we propose URO-Bench, an extensive benchmark for SDMs. Notably, URO-Bench is the first S2S benchmark that covers evaluations about multilingualism, multi-round dialogues, and paralinguistics. Our benchmark is divided into two difficulty levels: basic track and pro track, each comprising 20 test sets, evaluating the model's abilities in Understanding, Reasoning, and Oral conversation. We hope that URO-Bench can facilitate the development of spoken dialogue models by providing a multifaceted evaluation of existing models and helping to track progress in this area.

Representative Examples of URO-Bench
URO-Bench Benchmark Construction Pipeline

Contents

  1. Datasets
  2. Evaluation
  3. Leaderboard
  4. Acknowledge
  5. Citation

Datasets

Currently, we support 40 datasets, including 20 different tasks, available at HuggingFace | URO-Bench.
The test sets are divided into 2 tracks, basic and pro.

SubsetTrackLang# SamplesTask Type
Repeatbasicen252Repeat the user's words verbatim
Repeat-zhbasiczh127Repeat the user's words verbatim
Summarybasicen118Summarize a given story or statement
LCSTS-zhbasiczh119Summarize a given story or statement
GaokaoEvalbasicen303English listening exam
HSK5-zhbasiczh100Chinese listening exam
StoralEvalbasicen201Deduce morals from a given story
SQuAD-zhbasiczh153Answer extraction, contextual reasoning
TruthfulEvalbasicen470Fact QA
OpenbookQA-zhbasiczh189Single-choice QA
Gsm8kEvalbasicen582Math application problem
APE-zhbasiczh190Math application problem
MLCbasicen177Math, Logic, Commen sense
MLC-zhbasiczh145Math, Logic, Commen sense
AlpacaEvalbasicen199Open-Ended QA
AlpacaEval-zhbasiczh147Open-Ended QA
CommonEvalbasicen200Open-Ended QA
Claude-zhbasiczh222Open-Ended QA
WildchatEvalbasicen349Real-world conversation
Wildchat-zhbasiczh299Real-world conversation
CodeSwitching-enproen70Code switching QA
CodeSwitching-zhprozh70Code switching QA
GenEmotion-enproen54Speech emotion generation
GenEmotion-zhprozh43Speech emotion generation
GenStyle-enproen44Speech style generation
GenStyle-zhprozh39Speech style generation
MLCpro-enproen91Math, Logic, Commen sense
MLCpro-zhprozh64Math, Logic, Commen sense
Safety-enproen25Pravicy-related
Safety-zhprozh25Pravicy-related
SRT-enproen43Singing, Reciting, Tongue twister
SRT-zhprozh25Singing, Reciting, Tongue twister
UnderEmotion-enproen137Speech emotion understanding
UnderEmotion-zhprozh79Speech emotion understanding
Multilingualpromulti1108Multilingual QA
ClothoEval-enproen265Audio understanding
MuChoEval-enproen311Music understanding
MtBenchEval-enproen190Multi-round conversation
SpeakerAware-enproen55Speaker recognition
SpeakerAware-zhprozh49Speaker recognition

Evaluation

With just four simple steps, you can get all the test results in one go.
We provide some examples in folder examples and scripts.
We've tried our best to make it easy to use. If you encounter any issues, feel free to contact us through the 'Issues' section.

Step 0: setup

# get environment ready
git clone https://github.com/Ruiqi-Yan/URO-Bench
cd URO-Bench
conda create -n uro python=3.11
conda activate uro
pip install -r requirements.txt

# get data ready
cd ..
export HF_ENDPOINT=https://hf-mirror.com    # if you have trouble with the network
huggingface-cli download --repo-type dataset --resume-download Honggao/URO-Bench URO-Bench-data.zip --local-dir ./ --local-dir-use-symlinks False
unzip URO-Bench-data.zip

# download whisper-large-v3 (optional)
# please ignore this if your network is OK
modelscope download --model AI-ModelScope/whisper-large-v3 --local_dir ./whisper-large-v3

Step 1: modify the inference code

You can modify the code based on examples/example-test/inference_for_eval.py (single-round) and examples/example-test/inference_multi.py (multi-round). Just wrap the inference code of your SDM inside the load_sdm and respond functions. Please ensure the output file matches the required format.

Step 2: modify the scripts

Fill in scripts/config.sh according to the guidelines.
Complete the inference part of scripts/example.sh according to your inference code. Please modify line 20 and line 88.

Step 3: run the automatic evaluation pipeline

Run example.sh and get the results.
You need to pass the path of config.sh as a parameter to the bash script.

# bash scripts/example.sh /data/ruiqi.yan/URO-Bench/scripts/config.sh
bash scripts/example.sh scripts/config.sh

Leaderboard

We tested GPT-4o-Audio-Preview on the miniset. The scores of Whisper-large-v3 + LLMs are provided as reference.

Basic track

EN

Rank                       Model                       LLM ScaleOverall↑Avg.UTMOS↑Avg.ASR-WER↓Repeat↑Summary↑GaokaoEval↑StoralEval↑TruthfulEval↑Gsm8kEval↑MLC↑AlpacaEval↑CommonEval↑WildchatEval↑
-Whisper + GPT-4o-89.33--95.2496.1686.4786.9778.2490.7275.7198.2989.7795.74
-GPT-4o-Audio-Preview (on miniset)-87.48--97.1694.1372.0084.2782.6780.0080.0095.2094.1395.20
-Whisper + GLM-4-9B-Chat-HF-84.24--97.1893.4581.8577.6868.8178.6480.0492.5382.2789.99
-Whisper + Qwen2-7B-Instruct-78.13--96.8797.450.6682.3567.8988.2673.2695.9185.9392.72
-Whisper + Llama-3.1-8B-Instruct-71.78--58.4192.320.3374.1067.4287.2971.7594.4780.7390.96
1GLM-4-Voice9B69.094.15        12.71%        90.9591.0764.4773.8059.2830.9357.8280.7763.0778.76
-Whisper + Qwen2-0.5B-Instruct-49.71--60.1278.590.3349.8239.7335.1752.9258.9357.5063.97
2Freeze-Omni7B48.284.3716.32%70.8978.8726.2957.7446.952.8142.5652.2348.7055.80
3LLaMA-Omni8B48.144.0210.42%45.6280.6816.0650.6545.133.8944.4464.3658.4072.19
4SLAM-Omni0.5B31.594.454.54%12.2666.211.3236.9534.65021.8548.9841.0352.61
5Mini-Omni20.5B21.314.4310.24%8.1040.060.6628.4926.9206.9734.8130.7036.43
6Mini-Omni0.5B18.064.426.05%5.0732.20023.2525.0602.8230.9929.8031.42

ZH

Rank                        Model                        LLM ScaleOverall↑Avg.UTMOS↑Avg.ASR-CER↓Repeat-zh↑LCSTS-zh↑HSK5-zh↑SQuAD-zh↑OpenbookQA-zh↑APE-zh↑MLC-zh↑AlpacaEval-zh↑Claude-zh↑Whildchat-zh↑
-Whisper + GPT-4o-79.27--69.3585.3771.0049.2380.9584.7364.8296.0099.4591.82
-Whisper + GLM-4-9b-Chat-HF-76.49--76.7285.3972.0051.8571.1066.1469.6590.2394.5387.24
-GPT-4o-Audio-Preview (on miniset)-73.78--93.5081.6088.0042.6776.0025.3381.3386.4082.9380.00
-Whisper + Qwen2-7B-Instruct-71.60--26.2085.3877.0039.4370.3776.1459.7792.6598.5890.43
1GLM-4-Voice9B66.903.10        4.54%              92.64            77.08            96.00                 28.75                       56.96                  15.78            78.85                83.35                82.12              84.48        
-Whisper + LLaMA-3.1-8B-Instruct-65.50--15.9781.8570.0039.4365.2567.1951.4986.8091.6585.39
2Freeze-Omni7B37.373.606.36%4.9771.827.669.5816.4011.7547.3567.9864.8971.28
-Whisper + Qwen2-0.5B-Instruct-35.92--22.0560.2830.0021.3525.3915.6515.9631.7270.8465.95
3SLAM-Omni0.5B23.683.645.15%22.6034.674.007.185.821.0529.6543.8145.3442.72

Pro track

EN

Rank                        Model                        LLM ScaleOverall↑Avg.UTMOS↑Avg.ASR-WER↓UnderEmotion-en↑CodeSwitching-en↑Safety-en↑ClothoEval-en↑MuChoEval-en↑MLCpro-en↑MtBenchEval-en↑SpeakerAware-en↑SRT-en↑GenEmotion-en↑GenStyle-en↑Multilingual↑
-GPT-4o-Audio-Preview (on miniset)-66.92--48.5371.4785.2876.0056.0046.6773.8750.6762.4033.46100.0098.67
1GLM-4-Voice9B53.414.14         9.67%                     52.41                        58.00                   65.56                 17.36                    32.37                 65.20                   68.35                        50.30                 45.12               48.13                 94.55       43.53
2LLaMA-Omni8B32.973.997.13%36.3525.5243.8922.5215.9747.62--25.128.6283.0321.10
3Freeze-Omni7B30.424.3025.94%48.2737.9058.061.510.325.49--46.9818.9266.3620.42
4SLAM-Omni0.5B26.904.463.61%45.8421.1448.3310.942.6810.2632.8831.0326.518.4264.2420.54
5Mini-Omni20.5B21.154.427.62%42.5322.0056.940.380.320--20.473.7344.3920.70
6Mini-Omni0.5B18.054.425.63%29.0520.3858.89000--9.771.2940.3020.83

ZH

Rank                        Model                        LLM ScaleOverall↑Avg.UTMOS↑Avg.ASR-CER↓UnderEmotion-zh↑CodeSwitching-zh↑Safety-zh↑MLCpro-zh↑SpeakerAware-zh↑SRT-zh↑GenEmotion-zh↑GenStyle-zh↑
1GLM-4-Voice9B63.803.27        4.05%                    74.51                        72.00                   57.67              47.40                   52.52                 67.62               44.79                 93.85       
-GPT-4o-Audio-Preview (on miniset)-62.46--67.2061.0776.6760.0054.1353.3332.0995.20
2Freeze-Omni7B44.953.677.46%66.0854.6744.0022.40-41.907.8377.78
3SLAM-Omni0.5B33.943.744.55%27.5943.7135.0010.9438.5037.145.6772.99

Acknowledge

Citation

If you use URO-Bench in your research, please cite the following paper:

@article{yan2025uro,
  title={URO-Bench: A Comprehensive Benchmark for End-to-End Spoken Dialogue Models},
  author={Yan, Ruiqi and Li, Xiquan and Chen, Wenxi and Niu, Zhikang and Yang, Chen and Ma, Ziyang and Yu, Kai and Chen, Xie},
  journal={arXiv preprint arXiv:2502.17810},
  year={2025}
}

Ruiqi-Yan/URO-Bench

Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Python

58

16 commits

updated Sep 2, 2025

See the code

README

URO-Bench

Official code for evaluating spoken dialogue models with
URO-Bench: A Comprehensive Benchmark for End-to-End Spoken Dialogue Models

URO-Bench Logo

version version arXiv Hugging Face mit

News

  • [Update Aug. 03, 2025] We have reconstructed the 6 datasets in the Chinese Basic Track and added four new ones — Wildchat-zh, HSK5-zh, APE-zh, and SQuAD-zh — so that the Chinese and English datasets are now fully paired. Safety-en, Safety-zh, and SRT-zh were expanded to a total of 25 samples. All 40 datasets are available on HuggingFace, and we also provide a curated miniset of 1,000 samples for quick evaluation before full-scale assessment. The corresponding test results have also been updated.
  • [Update Feb. 25, 2025] 🔥🔥🔥 code and data of URO-Bench have been released!

Overview

This repo contains the code of URO-Bench: A Comprehensive Benchmark for End-to-End Spoken Dialogue Models.

Recent advances in large language models (LLMs) have driven significant progress in end-to-end spoken dialogue models (SDMs). In contrast to text-based LLMs, the evaluation framework for SDMs should encompass both cognitive dimensions (e.g., logical reasoning, knowledge) and speech-related aspects (e.g., paralinguistic cues, audio quality). However, there is still a lack of comprehensive evaluations for SDMs in speech-to-speech (S2S) scenarios.

To address this gap, we propose URO-Bench, an extensive benchmark for SDMs. Notably, URO-Bench is the first S2S benchmark that covers evaluations about multilingualism, multi-round dialogues, and paralinguistics. Our benchmark is divided into two difficulty levels: basic track and pro track, each comprising 20 test sets, evaluating the model's abilities in Understanding, Reasoning, and Oral conversation. We hope that URO-Bench can facilitate the development of spoken dialogue models by providing a multifaceted evaluation of existing models and helping to track progress in this area.

Representative Examples of URO-Bench
URO-Bench Benchmark Construction Pipeline

Contents

  1. Datasets
  2. Evaluation
  3. Leaderboard
  4. Acknowledge
  5. Citation

Datasets

Currently, we support 40 datasets, including 20 different tasks, available at HuggingFace | URO-Bench.
The test sets are divided into 2 tracks, basic and pro.

SubsetTrackLang# SamplesTask Type
Repeatbasicen252Repeat the user's words verbatim
Repeat-zhbasiczh127Repeat the user's words verbatim
Summarybasicen118Summarize a given story or statement
LCSTS-zhbasiczh119Summarize a given story or statement
GaokaoEvalbasicen303English listening exam
HSK5-zhbasiczh100Chinese listening exam
StoralEvalbasicen201Deduce morals from a given story
SQuAD-zhbasiczh153Answer extraction, contextual reasoning
TruthfulEvalbasicen470Fact QA
OpenbookQA-zhbasiczh189Single-choice QA
Gsm8kEvalbasicen582Math application problem
APE-zhbasiczh190Math application problem
MLCbasicen177Math, Logic, Commen sense
MLC-zhbasiczh145Math, Logic, Commen sense
AlpacaEvalbasicen199Open-Ended QA
AlpacaEval-zhbasiczh147Open-Ended QA
CommonEvalbasicen200Open-Ended QA
Claude-zhbasiczh222Open-Ended QA
WildchatEvalbasicen349Real-world conversation
Wildchat-zhbasiczh299Real-world conversation
CodeSwitching-enproen70Code switching QA
CodeSwitching-zhprozh70Code switching QA
GenEmotion-enproen54Speech emotion generation
GenEmotion-zhprozh43Speech emotion generation
GenStyle-enproen44Speech style generation
GenStyle-zhprozh39Speech style generation
MLCpro-enproen91Math, Logic, Commen sense
MLCpro-zhprozh64Math, Logic, Commen sense
Safety-enproen25Pravicy-related
Safety-zhprozh25Pravicy-related
SRT-enproen43Singing, Reciting, Tongue twister
SRT-zhprozh25Singing, Reciting, Tongue twister
UnderEmotion-enproen137Speech emotion understanding
UnderEmotion-zhprozh79Speech emotion understanding
Multilingualpromulti1108Multilingual QA
ClothoEval-enproen265Audio understanding
MuChoEval-enproen311Music understanding
MtBenchEval-enproen190Multi-round conversation
SpeakerAware-enproen55Speaker recognition
SpeakerAware-zhprozh49Speaker recognition

Evaluation

With just four simple steps, you can get all the test results in one go.
We provide some examples in folder examples and scripts.
We've tried our best to make it easy to use. If you encounter any issues, feel free to contact us through the 'Issues' section.

Step 0: setup

# get environment ready
git clone https://github.com/Ruiqi-Yan/URO-Bench
cd URO-Bench
conda create -n uro python=3.11
conda activate uro
pip install -r requirements.txt

# get data ready
cd ..
export HF_ENDPOINT=https://hf-mirror.com    # if you have trouble with the network
huggingface-cli download --repo-type dataset --resume-download Honggao/URO-Bench URO-Bench-data.zip --local-dir ./ --local-dir-use-symlinks False
unzip URO-Bench-data.zip

# download whisper-large-v3 (optional)
# please ignore this if your network is OK
modelscope download --model AI-ModelScope/whisper-large-v3 --local_dir ./whisper-large-v3

Step 1: modify the inference code

You can modify the code based on examples/example-test/inference_for_eval.py (single-round) and examples/example-test/inference_multi.py (multi-round). Just wrap the inference code of your SDM inside the load_sdm and respond functions. Please ensure the output file matches the required format.

Step 2: modify the scripts

Fill in scripts/config.sh according to the guidelines.
Complete the inference part of scripts/example.sh according to your inference code. Please modify line 20 and line 88.

Step 3: run the automatic evaluation pipeline

Run example.sh and get the results.
You need to pass the path of config.sh as a parameter to the bash script.

# bash scripts/example.sh /data/ruiqi.yan/URO-Bench/scripts/config.sh
bash scripts/example.sh scripts/config.sh

Leaderboard

We tested GPT-4o-Audio-Preview on the miniset. The scores of Whisper-large-v3 + LLMs are provided as reference.

Basic track

EN

Rank                       Model                       LLM ScaleOverall↑Avg.UTMOS↑Avg.ASR-WER↓Repeat↑Summary↑GaokaoEval↑StoralEval↑TruthfulEval↑Gsm8kEval↑MLC↑AlpacaEval↑CommonEval↑WildchatEval↑
-Whisper + GPT-4o-89.33--95.2496.1686.4786.9778.2490.7275.7198.2989.7795.74
-GPT-4o-Audio-Preview (on miniset)-87.48--97.1694.1372.0084.2782.6780.0080.0095.2094.1395.20
-Whisper + GLM-4-9B-Chat-HF-84.24--97.1893.4581.8577.6868.8178.6480.0492.5382.2789.99
-Whisper + Qwen2-7B-Instruct-78.13--96.8797.450.6682.3567.8988.2673.2695.9185.9392.72
-Whisper + Llama-3.1-8B-Instruct-71.78--58.4192.320.3374.1067.4287.2971.7594.4780.7390.96
1GLM-4-Voice9B69.094.15        12.71%        90.9591.0764.4773.8059.2830.9357.8280.7763.0778.76
-Whisper + Qwen2-0.5B-Instruct-49.71--60.1278.590.3349.8239.7335.1752.9258.9357.5063.97
2Freeze-Omni7B48.284.3716.32%70.8978.8726.2957.7446.952.8142.5652.2348.7055.80
3LLaMA-Omni8B48.144.0210.42%45.6280.6816.0650.6545.133.8944.4464.3658.4072.19
4SLAM-Omni0.5B31.594.454.54%12.2666.211.3236.9534.65021.8548.9841.0352.61
5Mini-Omni20.5B21.314.4310.24%8.1040.060.6628.4926.9206.9734.8130.7036.43
6Mini-Omni0.5B18.064.426.05%5.0732.20023.2525.0602.8230.9929.8031.42

ZH

Rank                        Model                        LLM ScaleOverall↑Avg.UTMOS↑Avg.ASR-CER↓Repeat-zh↑LCSTS-zh↑HSK5-zh↑SQuAD-zh↑OpenbookQA-zh↑APE-zh↑MLC-zh↑AlpacaEval-zh↑Claude-zh↑Whildchat-zh↑
-Whisper + GPT-4o-79.27--69.3585.3771.0049.2380.9584.7364.8296.0099.4591.82
-Whisper + GLM-4-9b-Chat-HF-76.49--76.7285.3972.0051.8571.1066.1469.6590.2394.5387.24
-GPT-4o-Audio-Preview (on miniset)-73.78--93.5081.6088.0042.6776.0025.3381.3386.4082.9380.00
-Whisper + Qwen2-7B-Instruct-71.60--26.2085.3877.0039.4370.3776.1459.7792.6598.5890.43
1GLM-4-Voice9B66.903.10        4.54%              92.64            77.08            96.00                 28.75                       56.96                  15.78            78.85                83.35                82.12              84.48        
-Whisper + LLaMA-3.1-8B-Instruct-65.50--15.9781.8570.0039.4365.2567.1951.4986.8091.6585.39
2Freeze-Omni7B37.373.606.36%4.9771.827.669.5816.4011.7547.3567.9864.8971.28
-Whisper + Qwen2-0.5B-Instruct-35.92--22.0560.2830.0021.3525.3915.6515.9631.7270.8465.95
3SLAM-Omni0.5B23.683.645.15%22.6034.674.007.185.821.0529.6543.8145.3442.72

Pro track

EN

Rank                        Model                        LLM ScaleOverall↑Avg.UTMOS↑Avg.ASR-WER↓UnderEmotion-en↑CodeSwitching-en↑Safety-en↑ClothoEval-en↑MuChoEval-en↑MLCpro-en↑MtBenchEval-en↑SpeakerAware-en↑SRT-en↑GenEmotion-en↑GenStyle-en↑Multilingual↑
-GPT-4o-Audio-Preview (on miniset)-66.92--48.5371.4785.2876.0056.0046.6773.8750.6762.4033.46100.0098.67
1GLM-4-Voice9B53.414.14         9.67%                     52.41                        58.00                   65.56                 17.36                    32.37                 65.20                   68.35                        50.30                 45.12               48.13                 94.55       43.53
2LLaMA-Omni8B32.973.997.13%36.3525.5243.8922.5215.9747.62--25.128.6283.0321.10
3Freeze-Omni7B30.424.3025.94%48.2737.9058.061.510.325.49--46.9818.9266.3620.42
4SLAM-Omni0.5B26.904.463.61%45.8421.1448.3310.942.6810.2632.8831.0326.518.4264.2420.54
5Mini-Omni20.5B21.154.427.62%42.5322.0056.940.380.320--20.473.7344.3920.70
6Mini-Omni0.5B18.054.425.63%29.0520.3858.89000--9.771.2940.3020.83

ZH

Rank                        Model                        LLM ScaleOverall↑Avg.UTMOS↑Avg.ASR-CER↓UnderEmotion-zh↑CodeSwitching-zh↑Safety-zh↑MLCpro-zh↑SpeakerAware-zh↑SRT-zh↑GenEmotion-zh↑GenStyle-zh↑
1GLM-4-Voice9B63.803.27        4.05%                    74.51                        72.00                   57.67              47.40                   52.52                 67.62               44.79                 93.85       
-GPT-4o-Audio-Preview (on miniset)-62.46--67.2061.0776.6760.0054.1353.3332.0995.20
2Freeze-Omni7B44.953.677.46%66.0854.6744.0022.40-41.907.8377.78
3SLAM-Omni0.5B33.943.744.55%27.5943.7135.0010.9438.5037.145.6772.99

Acknowledge

Citation

If you use URO-Bench in your research, please cite the following paper:

@article{yan2025uro,
  title={URO-Bench: A Comprehensive Benchmark for End-to-End Spoken Dialogue Models},
  author={Yan, Ruiqi and Li, Xiquan and Chen, Wenxi and Niu, Zhikang and Yang, Chen and Ma, Ziyang and Yu, Kai and Chen, Xie},
  journal={arXiv preprint arXiv:2502.17810},
  year={2025}
}

Languages

Python

68.7%

Shell

31.3%