The AudSem dataset (audsem) is a novel, high-quality, and diverse audio-language dataset designed to enhance the reasoning capabilities of Audio-Language Models (ALMs), particularly by enabling structured thinking over the fine-grained semantics of sound. It provides a carefully curated collection of audio samples paired with rich, synthetically generated captions.
The dataset includes four task-specific configurations:
aac (Audio Captioning): Audio captioning tasks with detailed descriptionscreative_qa (Creative Question Answering): Creative writing and story generation based on audiomc_qa (Multiple Choice QA): Multiple-choice questions about audio contentqa (Open-ended QA): Open-ended questions requiring free-form answersA simplified version without semantic elements is available at: https://huggingface.co/datasets/gijs/audsem-simple
All configurations are derived from the same rigorously filtered audio-visual data and are designed to minimize overlap with existing benchmarks like AudioSet, AudioCaps, and WavCaps, addressing a critical challenge of data contamination in zero-shot evaluations.
Traditional audio-language models often struggle with complex reasoning over sound events, primarily due to:
AudSem directly addresses these issues by:
Each configuration in the AudSem dataset has the following structure:
Common fields across all configurations:
key: Unique identifier for the examplefile_name: Audio file name/identifierthinking: The model's reasoning process before answeringsemantic_elements: Detailed breakdown of sound components (e.g., sound-generating agents, acoustic properties)answer: The final response to the promptConfiguration-specific fields:
aac: Contains question (the audio captioning prompt)creative_qa: Contains question (creative writing prompt)mc_qa: Contains question (multiple choice question) and choices (answer options as a string)qa: Contains question (open-ended question)When loaded with the Hugging Face datasets library, you must specify a configuration:
{
'__key__': '4ac66a37-19f1-4ada-a7c9-3655f4bdb61e',
'__url__': './cache/hub/datasets--gijs--audsem/snapshots/3a8ca917ebc41a45834ab1ffb970e159f7a856ce/creative_qa/train/0000.tar',
'flac': {
'path': '4ac66a37-19f1-4ada-a7c9-3655f4bdb61e.flac',
'array': array([0.03851318, 0.03240967, 0.0223999 , ..., 0.04953003, 0.04876709, 0.046875 ]),
'sampling_rate': 48000
},
'json': {
'__key__': '4ac66a37-19f1-4ada-a7c9-3655f4bdb61e',
'answer': 'The grand hall was alive with the rich, resonant notes of a symphony, each melody wea... connection it forged between all present.',
'file_name': '4ac66a37-19f1-4ada-a7c9-3655f4bdb61e',
'question': 'Imagine you are a writer tasked with capturing the essence of this moment in a short ...pheric sounds and the emotions they evoke.',
'semantic_elements': 'continuous music, rhythmic applause, periodic bursts, harmonious, steady, celebratory, grand, well-received performance',
'thinking': "To answer this question, I will focus on the continuous music and the rhythmic applau...hat immerses the reader in the experience."
}
}
The dataset includes four types of tasks, generated for audsem-semantic configuration:
The dataset examples have the following fields depending on configuration:
All configurations include:
key: Unique identifier for the examplefile_name: Audio file name/identifierthinking: The model's reasoning process (detailed thought process about the audio)semantic_elements: Structured breakdown of sound componentsanswer: The final response to the promptConfiguration-specific fields:
aac: Includes question (audio captioning prompt)creative_qa: Includes question (creative writing prompt)mc_qa: Includes question (multiple choice question) and choices (answer options as string)qa: Includes question (open-ended question)audsem-semantic configuration)The audsem-semantic configuration explicitly defines the following semantic descriptors within the <semantic_elements> tag, guiding the model's reasoning process:
| Semantic Descriptor | Description |
|---|---|
| Sound-Generating Agents (Who) | Animated beings that generate sounds (e.g., people, birds, animals). |
| Physical Sound Sources (What) | Physical objects and substances that generate sounds (e.g., bells, cars, shoes). |
| Sound Generation Mechanisms (How) | Actions and mechanisms of sound generation, including abstract nouns and verbs (e.g., chirping, walking, ringing). |
| Temporal Context (When) | Specific time periods or events providing temporal context (e.g., morning, holidays). |
| Spatial Context (Where) | Locations relative to the listener or specific environments where sounds occur (e.g., background, train station, room). |
| Acoustic Surfaces (What/Where) | Physical surfaces and materials that contribute to acoustic properties of the sound event. |
| Signal Descriptors (Sound Type) | Signal-level acoustic descriptors and basic sound classifications (e.g., noise, chord), including adverbs (loudly, softly) and words characterizing the signal (buzzing, humming). |
| Auditory Attributes (Sound Property) | Descriptors of the auditory sensation itself (e.g., loud, soft, steady), excluding source-related adjectives (dull, hard, steadily). |
| Non-auditory Sensation | Non-acoustic attributes and emotional descriptors of perceived sounds (e.g., beautiful, relaxing, calm), including subjective impressions (quiet, calm). |
audsem-semantic: Approximately 797,000 examples.A rigorous filtering process was applied to minimize overlap:
The dataset was generated from data encompassing audio, video, text (closed captions), and image modalities, ensuring a rich contextual understanding for the synthetic generation process.
The AudSem dataset was created through a robust, multi-stage, and fully automated pipeline, involving several advanced AI models.
The primary source for AudSem is a vast collection of manually annotated English closed caption subtitles from YouTube videos, provided by Filmot.com. These captions were filtered to specifically identify Subtitles for Deaf and Hard of Hearing (SDH) entries, which often contain sound descriptions enclosed in brackets.
yt-dlp is used to download precise audio-visual segments corresponding to the verified captions based on their timestamps.ffmpeg converts videos to 360p (2 fps MP4) and extracts audio to WAV format (32kHz, 16-bit, mono) for consistent processing.The acquired audio-visual segments undergo comprehensive analysis using an ensemble of specialized AI models across modalities:
Quality Filtering Steps:
The final captions and reasoning structures are synthetically generated using the Qwen2.5-72B-Instruct model, acting as a "teacher model."
xgrammar and vLLM. This includes:
<thinking> phase: Detailed reasoning about primary/background sounds, events, activities, and environment (minimum 50 words). This phase incorporates natural language thought expressions and avoids direct mention of model outputs or visual context.<semantic_elements> phase (for audsem-semantic): Explicit breakdown of sound components as per Table 1 (Semantic Descriptors).<answer> phase: A concise audio caption (under 50 words).This fully automated process ensures high quality, diversity, and scalability, with the human-created closed captions serving as an implicit ground truth for filtering and validation.
from datasets import load_dataset
# Load a specific configuration
dataset_aac = load_dataset("gijs/audsem", "aac") # Audio captioning
dataset_qa = load_dataset("gijs/audsem", "qa") # Open-ended QA
dataset_mc = load_dataset("gijs/audsem", "mc_qa") # Multiple choice QA
dataset_creative = load_dataset("gijs/audsem", "creative_qa") # Creative writing
# Access train and validation splits
train_data = dataset_aac["train"]
valid_data = dataset_aac["valid"]
# You can iterate through the dataset or access individual examples
print(train_data[0])
Below are examples of the different task types in the dataset:
Audio Captioning (aac):
{
"file_name": "tar/XQpTIj5ZUO4 (6671_958-6675_958)",
"thinking": "Based on the consistent presence of music and theme music, it seems there is a rich orchestral composition playing throughout this segment. The music appears to be from a waltz, with a strong presence of string instruments, oboes, and clarinets leading the melody. The timpani provides a steady, slow waltz beat, enhancing the ballad-like and epic atmosphere. The music feels high-quality and well-produced, typical of a movie soundtrack. The presence of background music and musical instruments confirms the rich orchestral arrangement. The slow, deliberate pace and the use of strings and woodwinds suggest a scene of elegance and grandeur, possibly from a dance scene in a film. There are no sounds of singing or speech, which further supports the idea that this is instrumental music meant to set a mood or enhance a visual scene. The absence of other sounds, like silence or sound effects, confirms that the focus is entirely on the music.",
"semantic_elements": "1. Sound-generating animated beings with descriptive adjectives:\n - Orchestral musicians (playing elegantly and grandly)\n2. Physical objects/substances generating sound:\n - String instruments, oboes, clarinets, timpani (producing rich, melodic, and rhythmic sounds)\n3. Actions/mechanisms of sound generation:\n - Playing a waltz (with a slow, steady beat)\n4. Temporal context if present:\n - Continuous (lasting the entire segment)\n5. Spatial context and environment:\n - Concert hall or recording studio (high-quality, well-produced)\n6. Acoustic surfaces and materials contributing to the sound:\n - Wood, strings, and metal (providing rich, resonant sound)\n7. Signal-level sound descriptors:\n - Smooth, melodic, and rhythmic (with a consistent beat)\n8. Auditory sensation attributes:\n - Rich, full, and resonant (with a deep, enveloping quality)\n9. Subjective/emotional descriptors:\n - Epic, ballad-like, and elegant (conveying a sense of grandeur and sophistication)",
"question": "What's happening in this audio? Describe it thoroughly.",
"answer": "An orchestra plays a waltz with a strong string section, oboes, and clarinets leading the melody, supported by a slow, rhythmic timpani beat, creating a ballad-like and epic atmosphere."
}
Multiple-choice Audio Question Answering (mc_qa):
{
"file_name": "E_MgzS-pQ38 (0_526-3_609)",
"question": "What is the primary mood conveyed by the instrumental music?\nChoices:\nA. Energetic and lively\nB. Sad and melancholic\nC. Soft and romantic\nD. Tense and suspenseful",
"choices": {
"A": "Energetic and lively",
"B": "Sad and melancholic",
"C": "Soft and romantic",
"D": "Tense and suspenseful"
},
"thinking": "The music is described as soft, mellow, sentimental, emotional, moving, and soulful. It is also referred to as a classic love song, which suggests a romantic and tender mood. The use of lush piano harmony and gentle bass lines further supports this, as these elements are typically associated with a gentle and tender atmosphere. Therefore, the primary mood is most likely soft and romantic.",
"semantic_elements": "Sound-generating entities: piano, bass, electric guitar; Acoustic properties: slow tempo, lush harmony, gentle bass lines; Auditory attributes: soft, mellow, sentimental, emotional, moving, soulful",
"answer": "C: Soft and romantic"
}
Open-ended Audio Question Answering (qa):
{
"file_name": "hivdj_-gsxg (127_248-130_379)",
"question": "What instruments can be heard in the background of the romantic song?",
"thinking": "I can hear the acoustic guitar and percussion as the primary instruments, with additional elements like bass, choirs, drums, and piano. The mood is romantic, which aligns with the gentle and slow nature of the music.",
"semantic_elements": "instruments: acoustic guitar, percussion, bass, choirs, drums, piano; acoustic properties: gentle, slow, romantic; mood: romantic; genre: folk, pop",
"answer": "The background of the song features an acoustic guitar, percussion, bass, choirs, drums, and piano, creating a gentle and romantic mood."
}
Creative Writing/Story Generation (creative_qa):
{
"file_name": "RdkO4a9NRkM (1813_516-1821_086)",
"question": "Imagine you are a writer tasked with capturing the essence of this moment in a short story. How would you describe the atmosphere created by the sounds you hear, and what emotions do they evoke in the characters present?",
"thinking": "The key sounds are enthusiastic and energetic applause, which suggests a joyful and celebratory atmosphere. The sounds are low quality and noisy, indicating a large, possibly indoor space. The applause is continuous and strong, which helps to create a sense of unity and shared excitement. The presence of speech in the earlier part of the segment suggests that the audience is reacting positively to a speech, adding to the celebratory mood. The creative task is to transform this auditory experience into a descriptive, emotional narrative that captures the essence of the moment.",
"semantic_elements": "Applause, Enthusiasm, Joy, Celebration, Speech, Low-quality recording, Noisy environment",
"answer": "In the heart of the hall, a wave of applause rolled through the air like a thunderous tide, its enthusiastic energy filling every corner. The clapping was raw and unfiltered, a symphony of joy that echoed off the walls, creating a sense of unity and shared excitement. The noise was a blend of individual claps and the collective roar of a crowd, each beat of praise a testament to the achievements being celebrated. In the midst of this jubilant chaos, a man at the podium, his voice clear and resonant, had delivered words that struck a chord. The audience, moved by his speech, let their applause speak volumes. The emotions were palpable—pride, elation, and a deep sense of accomplishment. It was a moment frozen in time, where every clap and cheer told a story of hard-won victories and bright futures. The hall was alive with the sound of success, and the air buzzed with the promise of what was to come."
}
The AudSem dataset development adhered to ethical guidelines, including:
No immediate negative real-world applications or misuse scenarios are proposed or directly facilitated by this fundamental research.
If you use the AudSem dataset in your research, please cite the accompanying paper:
@misc{wijngaard2025audsemthinkerenhancingaudiolanguagemodels,
title={AudSemThinker: Enhancing Audio-Language Models through Reasoning over Semantics of Sound},
author={Gijs Wijngaard and Elia Formisano and Michele Esposito and Michel Dumontier},
year={2025},
eprint={2505.14142},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2505.14142},
}
3 commits
The AudSem dataset (audsem) is a novel, high-quality, and diverse audio-language dataset designed to enhance the reasoning capabilities of Audio-Language Models (ALMs), particularly by enabling structured thinking over the fine-grained semantics of sound. It provides a carefully curated collection of audio samples paired with rich, synthetically generated captions.
The dataset includes four task-specific configurations:
aac (Audio Captioning): Audio captioning tasks with detailed descriptionscreative_qa (Creative Question Answering): Creative writing and story generation based on audiomc_qa (Multiple Choice QA): Multiple-choice questions about audio contentqa (Open-ended QA): Open-ended questions requiring free-form answersA simplified version without semantic elements is available at: https://huggingface.co/datasets/gijs/audsem-simple
All configurations are derived from the same rigorously filtered audio-visual data and are designed to minimize overlap with existing benchmarks like AudioSet, AudioCaps, and WavCaps, addressing a critical challenge of data contamination in zero-shot evaluations.
Traditional audio-language models often struggle with complex reasoning over sound events, primarily due to:
AudSem directly addresses these issues by:
Each configuration in the AudSem dataset has the following structure:
Common fields across all configurations:
key: Unique identifier for the examplefile_name: Audio file name/identifierthinking: The model's reasoning process before answeringsemantic_elements: Detailed breakdown of sound components (e.g., sound-generating agents, acoustic properties)answer: The final response to the promptConfiguration-specific fields:
aac: Contains question (the audio captioning prompt)creative_qa: Contains question (creative writing prompt)mc_qa: Contains question (multiple choice question) and choices (answer options as a string)qa: Contains question (open-ended question)When loaded with the Hugging Face datasets library, you must specify a configuration:
{
'__key__': '4ac66a37-19f1-4ada-a7c9-3655f4bdb61e',
'__url__': './cache/hub/datasets--gijs--audsem/snapshots/3a8ca917ebc41a45834ab1ffb970e159f7a856ce/creative_qa/train/0000.tar',
'flac': {
'path': '4ac66a37-19f1-4ada-a7c9-3655f4bdb61e.flac',
'array': array([0.03851318, 0.03240967, 0.0223999 , ..., 0.04953003, 0.04876709, 0.046875 ]),
'sampling_rate': 48000
},
'json': {
'__key__': '4ac66a37-19f1-4ada-a7c9-3655f4bdb61e',
'answer': 'The grand hall was alive with the rich, resonant notes of a symphony, each melody wea... connection it forged between all present.',
'file_name': '4ac66a37-19f1-4ada-a7c9-3655f4bdb61e',
'question': 'Imagine you are a writer tasked with capturing the essence of this moment in a short ...pheric sounds and the emotions they evoke.',
'semantic_elements': 'continuous music, rhythmic applause, periodic bursts, harmonious, steady, celebratory, grand, well-received performance',
'thinking': "To answer this question, I will focus on the continuous music and the rhythmic applau...hat immerses the reader in the experience."
}
}
The dataset includes four types of tasks, generated for audsem-semantic configuration:
The dataset examples have the following fields depending on configuration:
All configurations include:
key: Unique identifier for the examplefile_name: Audio file name/identifierthinking: The model's reasoning process (detailed thought process about the audio)semantic_elements: Structured breakdown of sound componentsanswer: The final response to the promptConfiguration-specific fields:
aac: Includes question (audio captioning prompt)creative_qa: Includes question (creative writing prompt)mc_qa: Includes question (multiple choice question) and choices (answer options as string)qa: Includes question (open-ended question)audsem-semantic configuration)The audsem-semantic configuration explicitly defines the following semantic descriptors within the <semantic_elements> tag, guiding the model's reasoning process:
| Semantic Descriptor | Description |
|---|---|
| Sound-Generating Agents (Who) | Animated beings that generate sounds (e.g., people, birds, animals). |
| Physical Sound Sources (What) | Physical objects and substances that generate sounds (e.g., bells, cars, shoes). |
| Sound Generation Mechanisms (How) | Actions and mechanisms of sound generation, including abstract nouns and verbs (e.g., chirping, walking, ringing). |
| Temporal Context (When) | Specific time periods or events providing temporal context (e.g., morning, holidays). |
| Spatial Context (Where) | Locations relative to the listener or specific environments where sounds occur (e.g., background, train station, room). |
| Acoustic Surfaces (What/Where) | Physical surfaces and materials that contribute to acoustic properties of the sound event. |
| Signal Descriptors (Sound Type) | Signal-level acoustic descriptors and basic sound classifications (e.g., noise, chord), including adverbs (loudly, softly) and words characterizing the signal (buzzing, humming). |
| Auditory Attributes (Sound Property) | Descriptors of the auditory sensation itself (e.g., loud, soft, steady), excluding source-related adjectives (dull, hard, steadily). |
| Non-auditory Sensation | Non-acoustic attributes and emotional descriptors of perceived sounds (e.g., beautiful, relaxing, calm), including subjective impressions (quiet, calm). |
audsem-semantic: Approximately 797,000 examples.A rigorous filtering process was applied to minimize overlap:
The dataset was generated from data encompassing audio, video, text (closed captions), and image modalities, ensuring a rich contextual understanding for the synthetic generation process.
The AudSem dataset was created through a robust, multi-stage, and fully automated pipeline, involving several advanced AI models.
The primary source for AudSem is a vast collection of manually annotated English closed caption subtitles from YouTube videos, provided by Filmot.com. These captions were filtered to specifically identify Subtitles for Deaf and Hard of Hearing (SDH) entries, which often contain sound descriptions enclosed in brackets.
yt-dlp is used to download precise audio-visual segments corresponding to the verified captions based on their timestamps.ffmpeg converts videos to 360p (2 fps MP4) and extracts audio to WAV format (32kHz, 16-bit, mono) for consistent processing.The acquired audio-visual segments undergo comprehensive analysis using an ensemble of specialized AI models across modalities:
Quality Filtering Steps:
The final captions and reasoning structures are synthetically generated using the Qwen2.5-72B-Instruct model, acting as a "teacher model."
xgrammar and vLLM. This includes:
<thinking> phase: Detailed reasoning about primary/background sounds, events, activities, and environment (minimum 50 words). This phase incorporates natural language thought expressions and avoids direct mention of model outputs or visual context.<semantic_elements> phase (for audsem-semantic): Explicit breakdown of sound components as per Table 1 (Semantic Descriptors).<answer> phase: A concise audio caption (under 50 words).This fully automated process ensures high quality, diversity, and scalability, with the human-created closed captions serving as an implicit ground truth for filtering and validation.
from datasets import load_dataset
# Load a specific configuration
dataset_aac = load_dataset("gijs/audsem", "aac") # Audio captioning
dataset_qa = load_dataset("gijs/audsem", "qa") # Open-ended QA
dataset_mc = load_dataset("gijs/audsem", "mc_qa") # Multiple choice QA
dataset_creative = load_dataset("gijs/audsem", "creative_qa") # Creative writing
# Access train and validation splits
train_data = dataset_aac["train"]
valid_data = dataset_aac["valid"]
# You can iterate through the dataset or access individual examples
print(train_data[0])
Below are examples of the different task types in the dataset:
Audio Captioning (aac):
{
"file_name": "tar/XQpTIj5ZUO4 (6671_958-6675_958)",
"thinking": "Based on the consistent presence of music and theme music, it seems there is a rich orchestral composition playing throughout this segment. The music appears to be from a waltz, with a strong presence of string instruments, oboes, and clarinets leading the melody. The timpani provides a steady, slow waltz beat, enhancing the ballad-like and epic atmosphere. The music feels high-quality and well-produced, typical of a movie soundtrack. The presence of background music and musical instruments confirms the rich orchestral arrangement. The slow, deliberate pace and the use of strings and woodwinds suggest a scene of elegance and grandeur, possibly from a dance scene in a film. There are no sounds of singing or speech, which further supports the idea that this is instrumental music meant to set a mood or enhance a visual scene. The absence of other sounds, like silence or sound effects, confirms that the focus is entirely on the music.",
"semantic_elements": "1. Sound-generating animated beings with descriptive adjectives:\n - Orchestral musicians (playing elegantly and grandly)\n2. Physical objects/substances generating sound:\n - String instruments, oboes, clarinets, timpani (producing rich, melodic, and rhythmic sounds)\n3. Actions/mechanisms of sound generation:\n - Playing a waltz (with a slow, steady beat)\n4. Temporal context if present:\n - Continuous (lasting the entire segment)\n5. Spatial context and environment:\n - Concert hall or recording studio (high-quality, well-produced)\n6. Acoustic surfaces and materials contributing to the sound:\n - Wood, strings, and metal (providing rich, resonant sound)\n7. Signal-level sound descriptors:\n - Smooth, melodic, and rhythmic (with a consistent beat)\n8. Auditory sensation attributes:\n - Rich, full, and resonant (with a deep, enveloping quality)\n9. Subjective/emotional descriptors:\n - Epic, ballad-like, and elegant (conveying a sense of grandeur and sophistication)",
"question": "What's happening in this audio? Describe it thoroughly.",
"answer": "An orchestra plays a waltz with a strong string section, oboes, and clarinets leading the melody, supported by a slow, rhythmic timpani beat, creating a ballad-like and epic atmosphere."
}
Multiple-choice Audio Question Answering (mc_qa):
{
"file_name": "E_MgzS-pQ38 (0_526-3_609)",
"question": "What is the primary mood conveyed by the instrumental music?\nChoices:\nA. Energetic and lively\nB. Sad and melancholic\nC. Soft and romantic\nD. Tense and suspenseful",
"choices": {
"A": "Energetic and lively",
"B": "Sad and melancholic",
"C": "Soft and romantic",
"D": "Tense and suspenseful"
},
"thinking": "The music is described as soft, mellow, sentimental, emotional, moving, and soulful. It is also referred to as a classic love song, which suggests a romantic and tender mood. The use of lush piano harmony and gentle bass lines further supports this, as these elements are typically associated with a gentle and tender atmosphere. Therefore, the primary mood is most likely soft and romantic.",
"semantic_elements": "Sound-generating entities: piano, bass, electric guitar; Acoustic properties: slow tempo, lush harmony, gentle bass lines; Auditory attributes: soft, mellow, sentimental, emotional, moving, soulful",
"answer": "C: Soft and romantic"
}
Open-ended Audio Question Answering (qa):
{
"file_name": "hivdj_-gsxg (127_248-130_379)",
"question": "What instruments can be heard in the background of the romantic song?",
"thinking": "I can hear the acoustic guitar and percussion as the primary instruments, with additional elements like bass, choirs, drums, and piano. The mood is romantic, which aligns with the gentle and slow nature of the music.",
"semantic_elements": "instruments: acoustic guitar, percussion, bass, choirs, drums, piano; acoustic properties: gentle, slow, romantic; mood: romantic; genre: folk, pop",
"answer": "The background of the song features an acoustic guitar, percussion, bass, choirs, drums, and piano, creating a gentle and romantic mood."
}
Creative Writing/Story Generation (creative_qa):
{
"file_name": "RdkO4a9NRkM (1813_516-1821_086)",
"question": "Imagine you are a writer tasked with capturing the essence of this moment in a short story. How would you describe the atmosphere created by the sounds you hear, and what emotions do they evoke in the characters present?",
"thinking": "The key sounds are enthusiastic and energetic applause, which suggests a joyful and celebratory atmosphere. The sounds are low quality and noisy, indicating a large, possibly indoor space. The applause is continuous and strong, which helps to create a sense of unity and shared excitement. The presence of speech in the earlier part of the segment suggests that the audience is reacting positively to a speech, adding to the celebratory mood. The creative task is to transform this auditory experience into a descriptive, emotional narrative that captures the essence of the moment.",
"semantic_elements": "Applause, Enthusiasm, Joy, Celebration, Speech, Low-quality recording, Noisy environment",
"answer": "In the heart of the hall, a wave of applause rolled through the air like a thunderous tide, its enthusiastic energy filling every corner. The clapping was raw and unfiltered, a symphony of joy that echoed off the walls, creating a sense of unity and shared excitement. The noise was a blend of individual claps and the collective roar of a crowd, each beat of praise a testament to the achievements being celebrated. In the midst of this jubilant chaos, a man at the podium, his voice clear and resonant, had delivered words that struck a chord. The audience, moved by his speech, let their applause speak volumes. The emotions were palpable—pride, elation, and a deep sense of accomplishment. It was a moment frozen in time, where every clap and cheer told a story of hard-won victories and bright futures. The hall was alive with the sound of success, and the air buzzed with the promise of what was to come."
}
The AudSem dataset development adhered to ethical guidelines, including:
No immediate negative real-world applications or misuse scenarios are proposed or directly facilitated by this fundamental research.
If you use the AudSem dataset in your research, please cite the accompanying paper:
@misc{wijngaard2025audsemthinkerenhancingaudiolanguagemodels,
title={AudSemThinker: Enhancing Audio-Language Models through Reasoning over Semantics of Sound},
author={Gijs Wijngaard and Elia Formisano and Michele Esposito and Michel Dumontier},
year={2025},
eprint={2505.14142},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2505.14142},
}
3 commits