ckyang1124/LALM-Evaluation-Survey

Collection of works for evaluating (and analyzing) large audio-language models (LALMs)

41

16 commits

updated Aug 11, 2025

See the code

README

Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey


Overview

Abstract

With advancements in large audio-language models (LALMs), which enhance large language models (LLMs) with auditory capabilities, these models are expected to demonstrate universal proficiency across various auditory tasks. While numerous benchmarks have emerged to assess LALMs' performance, they remain fragmented and lack a structured taxonomy. To bridge this gap, we conduct a comprehensive survey and propose a systematic taxonomy for LALM evaluations, categorizing them into four dimensions based on their objectives: (1) General Auditory Awareness and Processing, (2) Knowledge and Reasoning, (3) Dialogue-oriented Ability, and (4) Fairness, Safety, and Trustworthiness. We provide detailed overviews within each category and highlight challenges in this field, offering insights into promising future directions. To the best of our knowledge, this is the first survey specifically focused on the evaluations of LALMs, providing clear guidelines for the community.

We will release the collection of the surveyed papers and actively maintain it to support ongoing advancements in the field.

News

  • [2025/07/17] Our paper collection is now available on on Hugging Face! We will continue to actively maintain and update it. Stay tuned!
  • [2025/05/23] Our paper is now available on arXiv

Taxonomy and Paper List

🔊 General Auditory Awareness and Processing

Auditory Awareness
Auditory Processing
YearAuthorsVenuePaper
2023Huang et al.ICASSP 2024Dynamic-SUPERB: Towards A Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark for Speech
2024Huang et al.ICLR 2025Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
2024Yang et al.ACL 2024 (Main)AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
2024Wang et al.NAACL 2025 (Main)AudioBench: A Universal Benchmark for Audio Large Language Models
2024Weck et al.ISMIR 2024MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
2025Cao et al.PreprintFinAudio: A Benchmark for Audio Large Language Models in Financial Applications
2024Wu et al.SLT 2024Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify and Understand Speaker in Spoken Dialogue
2024Bu et al.PreprintRoadmap towards Superhuman Speech Understanding using Large Language Models
2024Chen et al.EMNLP 2024 (Findings)Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models
2025Zang et al.PreprintAre you really listening? Boosting Perceptual Awareness in Music-QA Benchmarks
2024Zhao et al.PreprintOpenMU: Your Swiss Army Knife for Music Understanding
2025Wang et al.PreprintAdvancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models
2024Gong et al.PreprintAV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
2025Xue et al.PreprintAudio-FLAN: A Preliminary Release
2025Wang et al.PreprintQualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions
2025Pandey et al.PreprintSIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning
2023Gong et al.ICLR 2024Listen, Think, and Understand
2022Lipping et al.EUSIPCO 2022Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering
2025Huang et al.ICASSP 2025SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style Captioning
2024Wei et al.PreprintASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction
2024Li et al.SLT 2024WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding
2025Robinson et al.PreprintNatureLM-audio: an Audio-Language Foundation Model for Bioacoustics
2025Ma et al.ISMIR 2025CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following
2025Beyene et al.PreprintmSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks
2025Wang et al.PreprintMMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
2025Hou et al.PreprintSOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
2025Ahia et al.PreprintBLAB: Brutally Long Audio Bench
2025Wan et al.ACL 2025 (Main)SpeechIQ: Speech Intelligence Quotient Across Cognitive Levels in Voice Understanding Large Language Models
2025Jiang et al.PreprintAdvancing the Foundation Model for Music Understanding

🧠 Knowledge and Reasoning

Linguistic Knowledge
World Knowledge Assessment
Reasoning
YearAuthorsVenuePaper
2024Ghosh et al.ICLR 2024CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
2025Sakshi et al.ICLR 2025MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
2025Cui et al.PreprintVoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
2025Yang et al.Interspeech 2025SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
2025Yan et al.PreprintURO-Bench: A Comprehensive Benchmark for End-to-End Spoken Dialogue Models
2025Deshmukh et al.AAAI 2025Audio Entailment: Assessing Deductive Reasoning for Audio Understanding
2024Gao et al.PreprintBenchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models
2024Zhao et al.PreprintOpenMU: Your Swiss Army Knife for Music Understanding
2024Gong et al.PreprintAV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
2024Ghosh et al.EMNLP 2024 (Main)GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
2023Gong et al.ICLR 2024Listen, Think, and Understand
2022Lipping et al.EUSIPCO 2022Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering
2024Li et al.SLT 2024WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding
2025Huang et al.ICASSP 2025SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style Captioning
2025Wang et al.ICASSP 2025What Are They Doing? Joint Audio-Speech Co-Reasoning
2025Deshmukh et al.ICLR 2025ADIFF: Explaining audio difference using natural language
2025Wang et al.PreprintMMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
2025Hou et al.PreprintSOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
2025Yosha et al.PreprintStressTest: Can YOUR Speech LM Handle the Stress?
2025Wei et al.PreprintTowards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems
2025Fang et al.PreprintS2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models
2025Bhattacharya et al.Interspeech 2025Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
2025Ma et al.PreprintMMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
2025Yang et al.DCASE 2025 Audio QA ChallengeMulti-Domain Audio Question Answering Toward Acoustic Content Reasoning in The DCASE 2025 Challenge
2025Ahia et al.PreprintBLAB: Brutally Long Audio Bench
2025Yang et al.PreprintSpeechR: A Benchmark for Speech Reasoning in Large Audio-Language Models

🗣️ Dialogue-oriented Ability

Conversation Ability
Instruction Following

🛡️ Fairness, Safety, and Trustworthiness

Fairness and Bias
Safety
Hallucination

How to Contribute

If you know of any interesting papers that aren’t listed yet, we welcome your contributions! Please open an issue using the format below:

We’ll review your suggestion and update the list as soon as possible. Thank you for helping us keep this resource up to date!

Citations

If you find this survey helpful for your research, please consider to cite our paper.

@article{yang2025towards,
  title={Towards holistic evaluation of large audio-language models: A comprehensive survey},
  author={Yang, Chih-Kai and Ho, Neo S and Lee, Hung-yi},
  journal={arXiv preprint arXiv:2505.15957},
  year={2025}
}

Contributors

wind-apprentice

11 commits

ckyang1124

5 commits

ckyang1124/LALM-Evaluation-Survey

Collection of works for evaluating (and analyzing) large audio-language models (LALMs)

41

16 commits

updated Aug 11, 2025

See the code

README

Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey


Overview

Abstract

With advancements in large audio-language models (LALMs), which enhance large language models (LLMs) with auditory capabilities, these models are expected to demonstrate universal proficiency across various auditory tasks. While numerous benchmarks have emerged to assess LALMs' performance, they remain fragmented and lack a structured taxonomy. To bridge this gap, we conduct a comprehensive survey and propose a systematic taxonomy for LALM evaluations, categorizing them into four dimensions based on their objectives: (1) General Auditory Awareness and Processing, (2) Knowledge and Reasoning, (3) Dialogue-oriented Ability, and (4) Fairness, Safety, and Trustworthiness. We provide detailed overviews within each category and highlight challenges in this field, offering insights into promising future directions. To the best of our knowledge, this is the first survey specifically focused on the evaluations of LALMs, providing clear guidelines for the community.

We will release the collection of the surveyed papers and actively maintain it to support ongoing advancements in the field.

News

  • [2025/07/17] Our paper collection is now available on on Hugging Face! We will continue to actively maintain and update it. Stay tuned!
  • [2025/05/23] Our paper is now available on arXiv

Taxonomy and Paper List

🔊 General Auditory Awareness and Processing

Auditory Awareness
Auditory Processing
YearAuthorsVenuePaper
2023Huang et al.ICASSP 2024Dynamic-SUPERB: Towards A Dynamic, Collaborative, and Comprehensive Instruction-Tuning Benchmark for Speech
2024Huang et al.ICLR 2025Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks
2024Yang et al.ACL 2024 (Main)AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension
2024Wang et al.NAACL 2025 (Main)AudioBench: A Universal Benchmark for Audio Large Language Models
2024Weck et al.ISMIR 2024MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
2025Cao et al.PreprintFinAudio: A Benchmark for Audio Large Language Models in Financial Applications
2024Wu et al.SLT 2024Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify and Understand Speaker in Spoken Dialogue
2024Bu et al.PreprintRoadmap towards Superhuman Speech Understanding using Large Language Models
2024Chen et al.EMNLP 2024 (Findings)Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models
2025Zang et al.PreprintAre you really listening? Boosting Perceptual Awareness in Music-QA Benchmarks
2024Zhao et al.PreprintOpenMU: Your Swiss Army Knife for Music Understanding
2025Wang et al.PreprintAdvancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models
2024Gong et al.PreprintAV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
2025Xue et al.PreprintAudio-FLAN: A Preliminary Release
2025Wang et al.PreprintQualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions
2025Pandey et al.PreprintSIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning
2023Gong et al.ICLR 2024Listen, Think, and Understand
2022Lipping et al.EUSIPCO 2022Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering
2025Huang et al.ICASSP 2025SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style Captioning
2024Wei et al.PreprintASR-EC Benchmark: Evaluating Large Language Models on Chinese ASR Error Correction
2024Li et al.SLT 2024WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding
2025Robinson et al.PreprintNatureLM-audio: an Audio-Language Foundation Model for Bioacoustics
2025Ma et al.ISMIR 2025CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following
2025Beyene et al.PreprintmSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks
2025Wang et al.PreprintMMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
2025Hou et al.PreprintSOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
2025Ahia et al.PreprintBLAB: Brutally Long Audio Bench
2025Wan et al.ACL 2025 (Main)SpeechIQ: Speech Intelligence Quotient Across Cognitive Levels in Voice Understanding Large Language Models
2025Jiang et al.PreprintAdvancing the Foundation Model for Music Understanding

🧠 Knowledge and Reasoning

Linguistic Knowledge
World Knowledge Assessment
Reasoning
YearAuthorsVenuePaper
2024Ghosh et al.ICLR 2024CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
2025Sakshi et al.ICLR 2025MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
2025Cui et al.PreprintVoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
2025Yang et al.Interspeech 2025SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
2025Yan et al.PreprintURO-Bench: A Comprehensive Benchmark for End-to-End Spoken Dialogue Models
2025Deshmukh et al.AAAI 2025Audio Entailment: Assessing Deductive Reasoning for Audio Understanding
2024Gao et al.PreprintBenchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models
2024Zhao et al.PreprintOpenMU: Your Swiss Army Knife for Music Understanding
2024Gong et al.PreprintAV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
2024Ghosh et al.EMNLP 2024 (Main)GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
2023Gong et al.ICLR 2024Listen, Think, and Understand
2022Lipping et al.EUSIPCO 2022Clotho-AQA: A Crowdsourced Dataset for Audio Question Answering
2024Li et al.SLT 2024WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding
2025Huang et al.ICASSP 2025SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style Captioning
2025Wang et al.ICASSP 2025What Are They Doing? Joint Audio-Speech Co-Reasoning
2025Deshmukh et al.ICLR 2025ADIFF: Explaining audio difference using natural language
2025Wang et al.PreprintMMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
2025Hou et al.PreprintSOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
2025Yosha et al.PreprintStressTest: Can YOUR Speech LM Handle the Stress?
2025Wei et al.PreprintTowards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems
2025Fang et al.PreprintS2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models
2025Bhattacharya et al.Interspeech 2025Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
2025Ma et al.PreprintMMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
2025Yang et al.DCASE 2025 Audio QA ChallengeMulti-Domain Audio Question Answering Toward Acoustic Content Reasoning in The DCASE 2025 Challenge
2025Ahia et al.PreprintBLAB: Brutally Long Audio Bench
2025Yang et al.PreprintSpeechR: A Benchmark for Speech Reasoning in Large Audio-Language Models

🗣️ Dialogue-oriented Ability

Conversation Ability
Instruction Following

🛡️ Fairness, Safety, and Trustworthiness

Fairness and Bias
Safety
Hallucination

How to Contribute

If you know of any interesting papers that aren’t listed yet, we welcome your contributions! Please open an issue using the format below:

We’ll review your suggestion and update the list as soon as possible. Thank you for helping us keep this resource up to date!

Citations

If you find this survey helpful for your research, please consider to cite our paper.

@article{yang2025towards,
  title={Towards holistic evaluation of large audio-language models: A comprehensive survey},
  author={Yang, Chih-Kai and Ho, Neo S and Lee, Hung-yi},
  journal={arXiv preprint arXiv:2505.15957},
  year={2025}
}

Contributors

wind-apprentice

11 commits

ckyang1124

5 commits