๐ฅ VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
12
39 commits
4 linked in READMEs
updated Oct 21, 2025
[๐ Homepage] [๐ฎ Visualization] [๐ป Github] [๐ Paper] [๐ Leaderboard ] [๐ Detailed Leaderboard ] [๐ Roleplay Leaderboard ]

from datasets import load_dataset
for split in ['listening_general', 'listening_music', 'listening_sound', 'listening_speech',
'speaking_assistant', 'speaking_emotion', 'speaking_instruction_following',
'speaking_multi_round', 'speaking_reasoning', 'speaking_robustness',
'speaking_roleplay', 'speaking_safety', 'viewing_multi_discipline']:
data = load_dataset("MathLLMs/VoiceAssistant-Eval", split)
print(data)
# load user_audio_0 directly with torchaudio
import torchaudio
waveform, sample_rate = torchaudio.load(data["test"][0]["user_audio_0"])
print(waveform.shape, sample_rate)
# load user_audio_0 directly with soundfile
import soundfile as sf
import io
audio_bytes = data["test"][0]["user_audio_0"]
waveform, sample_rate = sf.read(io.BytesIO(audio_bytes))
print(waveform.shape, sample_rate)
# save user_audio_0 to disk
data = load_dataset("MathLLMs/VoiceAssistant-Eval", 'listening_general')
def save_to_file(data, output_file):
with open(output_file, "wb") as f:
f.write(data)
user_audio_0 = data["test"][0]["user_audio_0"]
save_to_file(user_audio_0, "user_audio_0.wav")
The growing capabilities of large language models and multimodal systems have spurred interest in voice-first AI assistants, yet existing benchmarks are inadequate for evaluating the full range of these systems' capabilities. We summarize four key weaknesses of current benchmarks, highlighting the urgent need for a new evaluation framework:
W1: Lack of voice personalization evaluation.
Current benchmarks rarely test how well models mimic specific voices, which is key for personalized assistants (e.g., in healthcare). Without this, models may fail in real-world personalized applications.
W2: Limited focus on hands-free interaction.
Benchmarks often use text-based instructions, ignoring true voice-first, hands-free use. This limits reliability in critical contexts like driving or accessibility for visually impaired users.
W3: Neglect of real-world audio contexts.
Datasets seldom cover varied, realistic audio environments. Models aren't tested on understanding beyond speech (e.g., music, nature sounds), reducing their everyday usefulness.
W4: Insufficient multi-modal (vision + audio) assessment.
Benchmarks rarely test joint speech and visual input, missing key scenarios like smart tutors. This gap means benchmarks don't reflect real-world multimodal needs.
We introduce
VoiceAssistant-Eval, a comprehensive benchmark designed to assess AI assistants across listening, speaking, and viewing. VoiceAssistant-Eval comprises 10,497 curated examples spanning 13 task categories. These tasks include natural sounds, music, and spoken dialogue for listening; multi-turn dialogue, role-play imitation, and various scenarios for speaking; and highly heterogeneous images for viewing.
To demonstrate its utility, we evaluate 21 open-source models and GPT-4o-Audio, measuring the quality of the response content and speech, as well as their consistency. The results reveal three key findings: (1) proprietary models do not universally outperform open-source models; (2) most models excel at speaking tasks but lag in audio understanding; and (3) well-designed smaller models can rival much larger ones. Notably, the mid-sized Step-Audio-2-mini (7B) achieves more than double the listening accuracy of LLaMA-Omni2-32B-Bilingual. However, challenges remain: multimodal (audio+visual) input and role-play voice imitation tasks are difficult for current models, and significant gaps persist in robustness and safety alignment. VoiceAssistant-Eval identifies these gaps and establishes a rigorous framework for evaluating and guiding the development of next-generation multimodal voice assistants.
Figure 1: (a) Scores of six prominent omni-models across 13 tasks. (b) Examples from four newly designed tasks for voice assistants: I. Example from the role-play task with reference audio. II. A truly voice-based multi-turn conversation, instead of providing multi-round context in text. III. Multi-modal (vision + audio) integration understanding. IV. An audio question with music context.
Please refer to our project homepage and the paper for more details.
![]() | ![]() |
|---|---|
| Overview of principal statistics for VoiceAssistant-Eval. | Proportional distribution of tasks and the corresponding weaknesses addressed in VoiceAssistant-Eval. |
Explore the comprehensive evaluation results of AI assistants across multiple dimensions:
See [๐ป Github] for details.
| Dimension | Method | Models Used | Output Range |
|---|---|---|---|
| Emotion | Emotion Classification | emotion2vec | Probability distribution |
| Speaker Similarity | Voice Verification | WeSpeaker | 0-1 similarity score |
| Content Quality | LLM Judgment | gpt-oss-20b | 0-100% |
| Speech Quality | MOS Prediction | UTMOS22 | 0-100 (MOSร20) |
| Consistency | Modified WER | Whisper-Large-v3 | 0-100% (100-WER) |
This comprehensive evaluation framework enables thorough assessment of multimodal AI assistants across listening, speaking, and viewing capabilities, providing both granular insights and unified performance metrics.
If you find this benchmark useful in your research, please consider citing this BibTex:
@misc{wang2025voiceassistantevalbenchmarkingaiassistants,
title={VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing},
author={Ke Wang and Houxing Ren and Zimu Lu and Mingjie Zhan and Hongsheng Li},
year={2025},
eprint={2509.22651},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2509.22651},
}
๐ฅ VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
12
39 commits
4 linked in READMEs
updated Oct 21, 2025
[๐ Homepage] [๐ฎ Visualization] [๐ป Github] [๐ Paper] [๐ Leaderboard ] [๐ Detailed Leaderboard ] [๐ Roleplay Leaderboard ]

from datasets import load_dataset
for split in ['listening_general', 'listening_music', 'listening_sound', 'listening_speech',
'speaking_assistant', 'speaking_emotion', 'speaking_instruction_following',
'speaking_multi_round', 'speaking_reasoning', 'speaking_robustness',
'speaking_roleplay', 'speaking_safety', 'viewing_multi_discipline']:
data = load_dataset("MathLLMs/VoiceAssistant-Eval", split)
print(data)
# load user_audio_0 directly with torchaudio
import torchaudio
waveform, sample_rate = torchaudio.load(data["test"][0]["user_audio_0"])
print(waveform.shape, sample_rate)
# load user_audio_0 directly with soundfile
import soundfile as sf
import io
audio_bytes = data["test"][0]["user_audio_0"]
waveform, sample_rate = sf.read(io.BytesIO(audio_bytes))
print(waveform.shape, sample_rate)
# save user_audio_0 to disk
data = load_dataset("MathLLMs/VoiceAssistant-Eval", 'listening_general')
def save_to_file(data, output_file):
with open(output_file, "wb") as f:
f.write(data)
user_audio_0 = data["test"][0]["user_audio_0"]
save_to_file(user_audio_0, "user_audio_0.wav")
The growing capabilities of large language models and multimodal systems have spurred interest in voice-first AI assistants, yet existing benchmarks are inadequate for evaluating the full range of these systems' capabilities. We summarize four key weaknesses of current benchmarks, highlighting the urgent need for a new evaluation framework:
W1: Lack of voice personalization evaluation.
Current benchmarks rarely test how well models mimic specific voices, which is key for personalized assistants (e.g., in healthcare). Without this, models may fail in real-world personalized applications.
W2: Limited focus on hands-free interaction.
Benchmarks often use text-based instructions, ignoring true voice-first, hands-free use. This limits reliability in critical contexts like driving or accessibility for visually impaired users.
W3: Neglect of real-world audio contexts.
Datasets seldom cover varied, realistic audio environments. Models aren't tested on understanding beyond speech (e.g., music, nature sounds), reducing their everyday usefulness.
W4: Insufficient multi-modal (vision + audio) assessment.
Benchmarks rarely test joint speech and visual input, missing key scenarios like smart tutors. This gap means benchmarks don't reflect real-world multimodal needs.
We introduce
VoiceAssistant-Eval, a comprehensive benchmark designed to assess AI assistants across listening, speaking, and viewing. VoiceAssistant-Eval comprises 10,497 curated examples spanning 13 task categories. These tasks include natural sounds, music, and spoken dialogue for listening; multi-turn dialogue, role-play imitation, and various scenarios for speaking; and highly heterogeneous images for viewing.
To demonstrate its utility, we evaluate 21 open-source models and GPT-4o-Audio, measuring the quality of the response content and speech, as well as their consistency. The results reveal three key findings: (1) proprietary models do not universally outperform open-source models; (2) most models excel at speaking tasks but lag in audio understanding; and (3) well-designed smaller models can rival much larger ones. Notably, the mid-sized Step-Audio-2-mini (7B) achieves more than double the listening accuracy of LLaMA-Omni2-32B-Bilingual. However, challenges remain: multimodal (audio+visual) input and role-play voice imitation tasks are difficult for current models, and significant gaps persist in robustness and safety alignment. VoiceAssistant-Eval identifies these gaps and establishes a rigorous framework for evaluating and guiding the development of next-generation multimodal voice assistants.
Figure 1: (a) Scores of six prominent omni-models across 13 tasks. (b) Examples from four newly designed tasks for voice assistants: I. Example from the role-play task with reference audio. II. A truly voice-based multi-turn conversation, instead of providing multi-round context in text. III. Multi-modal (vision + audio) integration understanding. IV. An audio question with music context.
Please refer to our project homepage and the paper for more details.
![]() | ![]() |
|---|---|
| Overview of principal statistics for VoiceAssistant-Eval. | Proportional distribution of tasks and the corresponding weaknesses addressed in VoiceAssistant-Eval. |
Explore the comprehensive evaluation results of AI assistants across multiple dimensions:
See [๐ป Github] for details.
| Dimension | Method | Models Used | Output Range |
|---|---|---|---|
| Emotion | Emotion Classification | emotion2vec | Probability distribution |
| Speaker Similarity | Voice Verification | WeSpeaker | 0-1 similarity score |
| Content Quality | LLM Judgment | gpt-oss-20b | 0-100% |
| Speech Quality | MOS Prediction | UTMOS22 | 0-100 (MOSร20) |
| Consistency | Modified WER | Whisper-Large-v3 | 0-100% (100-WER) |
This comprehensive evaluation framework enables thorough assessment of multimodal AI assistants across listening, speaking, and viewing capabilities, providing both granular insights and unified performance metrics.
If you find this benchmark useful in your research, please consider citing this BibTex:
@misc{wang2025voiceassistantevalbenchmarkingaiassistants,
title={VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing},
author={Ke Wang and Houxing Ren and Zimu Lu and Mingjie Zhan and Hongsheng Li},
year={2025},
eprint={2509.22651},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2509.22651},
}