BAAI/SeniorTalk

Dataset

SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for Super-Aged Seniors

44

48 commits

2 linked in READMEs

updated Aug 15, 2026

See the code

README

SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for Super-Aged Seniors

Hugging Face Datasets arXiv License: CC BY-NC-SA-4.0 Github

Introduction

SeniorTalk is a comprehensive, open-source Mandarin Chinese speech dataset specifically designed for research on elderly aged 75 to 85. This dataset addresses the critical lack of publicly available resources for this age group, enabling advancements in automatic speech recognition (ASR), speaker verification (SV), speaker dirazation (SD), speech editing and other related fields. The dataset is released under a CC BY-NC-SA 4.0 license, meaning it is available for non-commercial use.

Dataset Details

This dataset contains 55.53 hours of high-quality speech data collected from 202 elderly across 16 provinces in China. Key features of the dataset include:

  • Age Range: 75-85 years old (inclusive). This is a crucial age range often overlooked in speech datasets.
  • Speakers: 202 unique elderly speakers.
  • Geographic Diversity: Speakers from 16 of China's 34 provincial-level administrative divisions, capturing a range of regional accents.
  • Gender Balance: Approximately 7:13 representation of male and female speakers, largely attributed to the differing average ages of males and females among the elderly.
  • Recording Conditions: Recordings were made in quiet environments using a variety of smartphones (both Android and iPhone devices) to ensure real-world applicability.
  • Content: Natural, conversational speech during age-appropriate activities. The content is unrestricted, promoting spontaneous and natural interactions.
  • Audio Format: WAV files with a 16kHz sampling rate.
  • Transcriptions: Carefully crafted, character-level manual transcriptions.
  • Annotations: The dataset includes annotations for each utterance, and for the speakers level.
    • Session-level: sentence_start_time,sentence_end_time,overlapped speech
    • Utterance-level: id, accent_level, text (transcription).
    • Token-level: special token([SONANT],[MUSIC],[NOISE]....)
    • Speaker-level: speaker_id, age, gender, location (province), device.

Dataset Structure

Dialogue Dataset

The dataset is split into two subsets:

Split# Speakers# DialoguesDuration (hrs)Avg. Dialogue Length (h)
train1829149.830.54
test20105.700.57
Total20210155.530.55

The dataset file structure is as follows.


dialogue_data/  
β”œβ”€β”€ wav  
β”‚   β”œβ”€β”€ train/*.tar   
β”‚   └── test/*.tar   
└── transcript/*.txt
UTTERANCEINFO.txt  # annotation of topics and duration
SPKINFO.txt   # annotation of location , age , gender and device

Each WAV file has a corresponding TXT file with the same name, containing its annotations.

For more details, please refer to our paper SeniorTalk.

ASR Dataset

The dataset is split into three subsets:

Split# Speakers# UtterancesDuration (hrs)Avg. Utterance Length (s)
train16247,26929.952.28
validation206,8914.092.14
test205,8693.772.31
Total20260,02937.812.27

The dataset file structure is as follows.

sentence_data/  
β”œβ”€β”€ wav  
β”‚   β”œβ”€β”€ train/*.tar
β”‚   β”œβ”€β”€ dev/*.tar 
β”‚   └── test/*.tar   
└── transcript/*.txt   
UTTERANCEINFO.txt  # annotation of topics and duration
SPKINFO.txt   # annotation of location , age , gender and device

Each WAV file has a corresponding TXT, containing its annotations.

For more details, please refer to our paper SeniorTalk

πŸ“š Cite me

@misc{chen2025seniortalkchineseconversationdataset,
      title={SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for Super-Aged Seniors}, 
      author={Yang Chen and Hui Wang and Shiyao Wang and Junyang Chen and Jiabei He and Jiaming Zhou and Xi Yang and Yequan Wang and Yonghua Lin and Yong Qin},
      year={2025},
      eprint={2503.16578},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2503.16578}, 
}

Contributors

evan0617

44 commits

SE-Eval

4 commits

BAAI/SeniorTalk

Dataset

SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for Super-Aged Seniors

44

48 commits

2 linked in READMEs

updated Aug 15, 2026

See the code

README

SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for Super-Aged Seniors

Hugging Face Datasets arXiv License: CC BY-NC-SA-4.0 Github

Introduction

SeniorTalk is a comprehensive, open-source Mandarin Chinese speech dataset specifically designed for research on elderly aged 75 to 85. This dataset addresses the critical lack of publicly available resources for this age group, enabling advancements in automatic speech recognition (ASR), speaker verification (SV), speaker dirazation (SD), speech editing and other related fields. The dataset is released under a CC BY-NC-SA 4.0 license, meaning it is available for non-commercial use.

Dataset Details

This dataset contains 55.53 hours of high-quality speech data collected from 202 elderly across 16 provinces in China. Key features of the dataset include:

  • Age Range: 75-85 years old (inclusive). This is a crucial age range often overlooked in speech datasets.
  • Speakers: 202 unique elderly speakers.
  • Geographic Diversity: Speakers from 16 of China's 34 provincial-level administrative divisions, capturing a range of regional accents.
  • Gender Balance: Approximately 7:13 representation of male and female speakers, largely attributed to the differing average ages of males and females among the elderly.
  • Recording Conditions: Recordings were made in quiet environments using a variety of smartphones (both Android and iPhone devices) to ensure real-world applicability.
  • Content: Natural, conversational speech during age-appropriate activities. The content is unrestricted, promoting spontaneous and natural interactions.
  • Audio Format: WAV files with a 16kHz sampling rate.
  • Transcriptions: Carefully crafted, character-level manual transcriptions.
  • Annotations: The dataset includes annotations for each utterance, and for the speakers level.
    • Session-level: sentence_start_time,sentence_end_time,overlapped speech
    • Utterance-level: id, accent_level, text (transcription).
    • Token-level: special token([SONANT],[MUSIC],[NOISE]....)
    • Speaker-level: speaker_id, age, gender, location (province), device.

Dataset Structure

Dialogue Dataset

The dataset is split into two subsets:

Split# Speakers# DialoguesDuration (hrs)Avg. Dialogue Length (h)
train1829149.830.54
test20105.700.57
Total20210155.530.55

The dataset file structure is as follows.


dialogue_data/  
β”œβ”€β”€ wav  
β”‚   β”œβ”€β”€ train/*.tar   
β”‚   └── test/*.tar   
└── transcript/*.txt
UTTERANCEINFO.txt  # annotation of topics and duration
SPKINFO.txt   # annotation of location , age , gender and device

Each WAV file has a corresponding TXT file with the same name, containing its annotations.

For more details, please refer to our paper SeniorTalk.

ASR Dataset

The dataset is split into three subsets:

Split# Speakers# UtterancesDuration (hrs)Avg. Utterance Length (s)
train16247,26929.952.28
validation206,8914.092.14
test205,8693.772.31
Total20260,02937.812.27

The dataset file structure is as follows.

sentence_data/  
β”œβ”€β”€ wav  
β”‚   β”œβ”€β”€ train/*.tar
β”‚   β”œβ”€β”€ dev/*.tar 
β”‚   └── test/*.tar   
└── transcript/*.txt   
UTTERANCEINFO.txt  # annotation of topics and duration
SPKINFO.txt   # annotation of location , age , gender and device

Each WAV file has a corresponding TXT, containing its annotations.

For more details, please refer to our paper SeniorTalk

πŸ“š Cite me

@misc{chen2025seniortalkchineseconversationdataset,
      title={SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for Super-Aged Seniors}, 
      author={Yang Chen and Hui Wang and Shiyao Wang and Junyang Chen and Jiabei He and Jiaming Zhou and Xi Yang and Yequan Wang and Yonghua Lin and Yong Qin},
      year={2025},
      eprint={2503.16578},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2503.16578}, 
}

Contributors

evan0617

44 commits

SE-Eval

4 commits