ASLP-lab/WenetSpeech-Wu-Bench

Dataset

WenetSpeech-Wu Bench

5

21 commits

1 linked in READMEs

updated Feb 8, 2026

See the code

README

WenetSpeech-Wu Bench

We introduce WenetSpeech-Wu-Bench, the first publicly available, manually curated benchmark for Wu dialect speech processing, covering ASR, Wu-to-Mandarin AST, speaker attributes, emotion recognition, TTS, and instruct TTS, and providing a unified platform for fair evaluation.

  • ASR: Wu dialect ASR (9.75 hour, including Shanghainese, Suzhounese, and Mandarin code-mixed speech). Evaluated by CER.
  • Wu→Mandarin AST: Speech translation from Wu dialects to Mandarin (3k utterances, 4.4h). Evaluated by BLEU.
  • Speaker Attributes & Emotion: Speaker gender/age prediction and emotion recognition on Wu dialect. Evaluated by classification accuracy.
  • TTS: Wu dialect TTS with speaker prompting (242 sentences, 12 speakers). Evaluated by speaker similarity, CER, and MOS.
  • Instruct TTS: Instruction-following TTS with prosodic and emotional control. Evaluated by automatic accuracy and subjective MOS.

Contributors

ASLP-lab

20 commits

hujb

1 commits

ASLP-lab/WenetSpeech-Wu-Bench

Dataset

WenetSpeech-Wu Bench

5

21 commits

1 linked in READMEs

updated Feb 8, 2026

See the code

README

WenetSpeech-Wu Bench

We introduce WenetSpeech-Wu-Bench, the first publicly available, manually curated benchmark for Wu dialect speech processing, covering ASR, Wu-to-Mandarin AST, speaker attributes, emotion recognition, TTS, and instruct TTS, and providing a unified platform for fair evaluation.

  • ASR: Wu dialect ASR (9.75 hour, including Shanghainese, Suzhounese, and Mandarin code-mixed speech). Evaluated by CER.
  • Wu→Mandarin AST: Speech translation from Wu dialects to Mandarin (3k utterances, 4.4h). Evaluated by BLEU.
  • Speaker Attributes & Emotion: Speaker gender/age prediction and emotion recognition on Wu dialect. Evaluated by classification accuracy.
  • TTS: Wu dialect TTS with speaker prompting (242 sentences, 12 speakers). Evaluated by speaker similarity, CER, and MOS.
  • Instruct TTS: Instruction-following TTS with prosodic and emotional control. Evaluated by automatic accuracy and subjective MOS.

Contributors

ASLP-lab

20 commits

hujb

1 commits