qqjz/VoxEval

Dataset

3

stars

9

commits

2

linked in READMEs

Jul 25, 2025

updated

README

VoxEval

GitHub arXiv

Github repository for paper: VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models

Also check out our survey paper at Recent Advances in Speech Language Models: A Survey!

VoxEval is a novel speech question-answering benchmark specifically designed to assess SLMs' knowledge understanding through purely speech-based interactions.

Below are the three highlights of our VoxEval benchmark:

  • End-to-end speech-based evaluation: Both input and output are audio-based.
  • Diverse audio conditions: VoxEval includes audio files featuring a variety of speakers, speaking styles, and audio qualities.
  • Complex spoken evaluations: It supports advanced spoken assessments, including spoken math tasks.

Download Data

You can access our VoxEval Dataset Repository on πŸ€— Hugging Face to directly download the dataset.

Below is the layout of the dataset folder. all_fewshot_examples folder contains the few shot examples for the evaluation. math_CoT_fewshot folder contains the few shot examples for evaluating the math subjects via Chain-of-Thought prompting. test folder contains the actual test data in VoxEval.

  # β”œβ”€β”€ Root Folder
  # β”‚   β”œβ”€β”€ all_fewshot_examples
  # β”‚       β”œβ”€β”€ alloy (different speaker voice)
  # β”‚           β”œβ”€β”€ abstract_algebra_4o (different subjects)
  # β”‚               β”œβ”€β”€ XXX.mp3
  # β”‚               β”œβ”€β”€ ...
  # β”‚           β”œβ”€β”€ ...
  # β”‚       β”œβ”€β”€ echo
  # β”‚       β”œβ”€β”€ ...
  # β”‚   β”œβ”€β”€ test
  # β”‚       β”œβ”€β”€ alloy (different settings)
  # β”‚           β”œβ”€β”€ abstract_algebra_4o (different subjects)
  # β”‚               β”œβ”€β”€ XXX.mp3
  # β”‚               β”œβ”€β”€ ...
  # β”‚           β”œβ”€β”€ ...
  # β”‚       β”œβ”€β”€ echo
  # β”‚       β”œβ”€β”€ ...

Evaluation results of existing end-to-end Spoken Language Models

SLMsSpeechGPTTWISTSPIRIT-LMMoshiGLM-4-Voice
Speakers
Alloy0.00010.04800.20840.12160.3763
Echo0.00010.05580.20960.12210.3764
Fable0.00000.01160.20840.11530.3642
Nova0.00010.03320.20700.12980.3677
Onyx0.00020.02750.19660.11920.3764
Shimmer0.00000.05160.20760.12050.3815
Speaking Styles
Linguistic0.00010.04880.20440.11870.3643
Speed0.00010.05030.19110.10130.3469
Pitch0.00000.05440.17880.06090.3345
Audio Qualities
Noise0.00000.03680.19500.10180.3695
Other Env Acoustics0.00010.04340.20190.10510.3728
Underlying Text LMsLlama-7BLlama-7BLlama-2-7BHelium-7BGLM-4-9B
Text MMLU0.35100.35100.45300.54300.7470

What's Next

  • The VoxEval evaluation code will be released soon.

License

The dataset is licensed under the Creative Commons Attribution 4.0.

Citation

@article{cui2025voxeval,
  title={VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models},
  author={Cui, Wenqian and Jiao, Xiaoqi and Meng, Ziqiao and King, Irwin},
  journal={arXiv preprint arXiv:2501.04962},
  year={2025}
}

Contributors

UN
unknown

5 commits

qqjz

4 commits

qqjz/VoxEval

Dataset

3

stars

9

commits

2

linked in READMEs

Jul 25, 2025

updated

README

VoxEval

GitHub arXiv

Github repository for paper: VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models

Also check out our survey paper at Recent Advances in Speech Language Models: A Survey!

VoxEval is a novel speech question-answering benchmark specifically designed to assess SLMs' knowledge understanding through purely speech-based interactions.

Below are the three highlights of our VoxEval benchmark:

  • End-to-end speech-based evaluation: Both input and output are audio-based.
  • Diverse audio conditions: VoxEval includes audio files featuring a variety of speakers, speaking styles, and audio qualities.
  • Complex spoken evaluations: It supports advanced spoken assessments, including spoken math tasks.

Download Data

You can access our VoxEval Dataset Repository on πŸ€— Hugging Face to directly download the dataset.

Below is the layout of the dataset folder. all_fewshot_examples folder contains the few shot examples for the evaluation. math_CoT_fewshot folder contains the few shot examples for evaluating the math subjects via Chain-of-Thought prompting. test folder contains the actual test data in VoxEval.

  # β”œβ”€β”€ Root Folder
  # β”‚   β”œβ”€β”€ all_fewshot_examples
  # β”‚       β”œβ”€β”€ alloy (different speaker voice)
  # β”‚           β”œβ”€β”€ abstract_algebra_4o (different subjects)
  # β”‚               β”œβ”€β”€ XXX.mp3
  # β”‚               β”œβ”€β”€ ...
  # β”‚           β”œβ”€β”€ ...
  # β”‚       β”œβ”€β”€ echo
  # β”‚       β”œβ”€β”€ ...
  # β”‚   β”œβ”€β”€ test
  # β”‚       β”œβ”€β”€ alloy (different settings)
  # β”‚           β”œβ”€β”€ abstract_algebra_4o (different subjects)
  # β”‚               β”œβ”€β”€ XXX.mp3
  # β”‚               β”œβ”€β”€ ...
  # β”‚           β”œβ”€β”€ ...
  # β”‚       β”œβ”€β”€ echo
  # β”‚       β”œβ”€β”€ ...

Evaluation results of existing end-to-end Spoken Language Models

SLMsSpeechGPTTWISTSPIRIT-LMMoshiGLM-4-Voice
Speakers
Alloy0.00010.04800.20840.12160.3763
Echo0.00010.05580.20960.12210.3764
Fable0.00000.01160.20840.11530.3642
Nova0.00010.03320.20700.12980.3677
Onyx0.00020.02750.19660.11920.3764
Shimmer0.00000.05160.20760.12050.3815
Speaking Styles
Linguistic0.00010.04880.20440.11870.3643
Speed0.00010.05030.19110.10130.3469
Pitch0.00000.05440.17880.06090.3345
Audio Qualities
Noise0.00000.03680.19500.10180.3695
Other Env Acoustics0.00010.04340.20190.10510.3728
Underlying Text LMsLlama-7BLlama-7BLlama-2-7BHelium-7BGLM-4-9B
Text MMLU0.35100.35100.45300.54300.7470

What's Next

  • The VoxEval evaluation code will be released soon.

License

The dataset is licensed under the Creative Commons Attribution 4.0.

Citation

@article{cui2025voxeval,
  title={VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models},
  author={Cui, Wenqian and Jiao, Xiaoqi and Meng, Ziqiao and King, Irwin},
  journal={arXiv preprint arXiv:2501.04962},
  year={2025}
}

Contributors

UN
unknown

5 commits

qqjz

4 commits