SKNahin/open-large-bengali-asr-data

Dataset

Open Large Bengali ASR Data

13

9 commits

2 linked in READMEs

updated Mar 26, 2024

See the code

README

Open Large Bengali ASR Data

This is a collection of publicly available ASR data for Bengali. It contains 5000 hours of audio. We have a filtering column called is_better to filter good-quality audio from the corpus. It is set based on the wer between original transcription and prediction taken from a Bengali-Wav2Vec2 model and word-per-second (wps).

Datasets: