espnet/yodas

Dataset

154

stars

500

commits

4

linked in READMEs

Jun 10, 2024

updated

README

Updates

  • 2024/07/09: we also uploaded a new version of YODAS as YODAS2, it provides unsegmented audios and higher sampling rate (24k)

README

This is the YODAS manual/automatic subset from our YODAS dataset, it has 369,510 hours of speech.

This dataset contains audio utterances and corresponding captions (manual or automatic) from YouTube. Note that manual caption only indicates that it is uploaded by users, but not necessarily transcribed by a human

For more details about YODAS dataset, please refer to our paper

Usage:

Considering the extremely large size of the entire dataset, we support two modes of dataset loadings:

standard mode: each subset will be downloaded to the local dish before first iterating.

from datasets import load_dataset

# Note this will take very long time to download and preprocess
# you can try small subset for testing purpose
ds = load_dataset('espnet/yodas', 'en000')
print(next(iter(ds['train'])))

streaming mode most of the files will be streamed instead of downloaded to your local deivce. It can be used to inspect this dataset quickly.

from datasets import load_dataset

# this streaming loading will finish quickly
ds = load_dataset('espnet/yodas', 'en000', streaming=True)


#{'id': '9774', 'utt_id': 'YoRjzEnRcqu-00000-00000716-00000819', 'audio': {'path': None, 'array': array([-0.009552  , -0.01086426, -0.012146  , ..., -0.01992798,
#       -0.01885986, -0.01074219]), 'sampling_rate': 16000}, 'text': 'There is a saying'}
print(next(iter(ds['train'])))

Subsets/Shards

There are 149 languages in this dataset, each language is sharded into at least 1 shard to make it easy for our processing and uploading purposes. The raw data of each shard contains 500G at most.

Statistics of each shard can be found in the last section.

We distinguish manual caption subset and automatic caption subset by the first digit in each shard's name. The first digit is 0 if it contains manual captions, 1 if it contains automatic captions.

For example, en000 to en005 are the English shards containing manual subsets, and en100 to en127 contains the automatic subsets.

Reference

@inproceedings{li2023yodas,
  title={Yodas: Youtube-Oriented Dataset for Audio and Speech},
  author={Li, Xinjian and Takamichi, Shinnosuke and Saeki, Takaaki and Chen, William and Shiota, Sayaka and Watanabe, Shinji},
  booktitle={2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)},
  pages={1--8},
  year={2023},
  organization={IEEE}
}

Contact

If you have any questions, feel free to contact us at the following email address.

We made sure that our dataset only consisted of videos with CC licenses during our downloading. But in case you find your video unintentionally included in our dataset and would like to delete it, you can send a delete request to the following email.

Remove the parenthesis () from the following email address

(lixinjian)(1217)@gmail.com

Statistics

Note that there are no overlappings across different subsets, each audio can be included in the dataset at most once.

Subset nameHours
aa0000.171472
ab0000.358342
af0000.880497
ak0000.250858
am0000.924708
ar000289.707
as0000.548239
ay0000.0342722
az0003.8537
ba0000.0210556
be00048.1537
bg00046.8375
bh0000.0127111
bi0000.0125556
bm0000.00214722
bn00027.064
bo0000.746211
br0000.729914
bs0009.36959
ca00074.1909
co0000.0418639
cr0000.00584167
cs000167.604
cy0005.20017
da00027.4345
de0003063.81
de1004998.11
de1014995.08
de102955.389
dz0000.06365
ee0000.0411722
el000126.75
en0004999.73
en0015032.69
en0025039.9
en0035001.4
en0045054.66
en0054027.02
en1005147.07
en1015123.05
en1025117.68
en1035127.3
en1045126.33
en1055097.65
en1065131.47
en1075135.6
en1085136.84
en1095112.94
en1105109
en1115118.69
en1125122.57
en1135122.31
en1145112.36
en1155112.27
en1165123.77
en1175117.31
en1185117.94
en1195133.05
en1205127.79
en1215129.08
en1225130.22
en1235097.56
en1245116.59
en1255109.76
en1265136.21
en1272404.89
eo00012.6874
es0003737.86
es1005125.25
es1015130.44
es1025145.66
es1035138.26
es1045139.57
es1055138.95
es1062605.26
et00014.4129
eu00019.6356
fa00042.6734
ff0000.0394972
fi000212.899
fj0000.0167806
fo0000.183244
fr0002423.7
fr1005074.93
fr1015057.79
fr1025094.14
fr1033222.95
fy0000.0651667
ga0001.49252
gd0000.01885
gl0009.52575
gn0000.181356
gu0001.99355
ha0000.102931
hi000480.79
hi1002.74865
ho0000.0562194
hr00025.9171
ht0001.07494
hu000181.763
hy0001.64412
ia0000.0856056
id0001420.09
id1004902.79
id1013560.82
ie0000.134603
ig0000.086875
ik0000.00436667
is0005.07075
it0001454.98
it1004989.62
it1014242.87
iu0000.0584278
iw000161.373
ja0001094.18
ja1002929.94
jv0001.08701
ka00026.9727
ki0000.000555556
kk0003.72081
kl0000.00575556
km0003.98273
kn0002.36041
ko0002774.28
ko1005018.29
ko1015048.49
ko1025018.27
ko1032587.85
ks0000.0150444
ku0001.93419
ky00014.3917
la0007.26088
lb0000.1115
lg0000.00386111
ln0000.188739
lo0000.230986
lt00017.6507
lv0002.47671
mg0000.169653
mi0001.10089
mk0005.54236
ml00013.2386
mn0002.0232
mr0007.11602
ms00028.0219
my0002.35663
na0000.0397056
nd0000.00111111
ne0002.34936
nl000413.044
nl1002490.13
no000129.183
nv0000.00319444
oc0000.166108
om0000.148478
or0000.421436
pa0001.58188
pl000757.986
ps0000.9871
pt0001631.44
pt1005044.57
pt1015038.33
pt1025041.59
pt1033553.28
qu0000.748772
rm0000.192933
rn0000.00401111
ro00099.9175
ru0004968.37
ru001627.679
ru1005098.3
ru1015098
ru1025119.43
ru1035107.29
ru1045121.73
ru1055088.05
ru1063393.44
rw0000.640825
sa0000.354139
sc0000.00801111
sd0000.0768722
sg0000.000472222
sh0000.250914
si0004.2634
sk00030.0155
sl00022.9366
sm0000.102333
sn0000.0134722
so0003.36819
sq0003.48276
sr00015.2849
st0000.00324167
su0000.0404639
sv000127.411
sw0001.93409
ta00059.4805
te0005.66794
tg0000.272386
th000497.14
th1001.87429
ti0000.343897
tk0000.0651806
tn0000.112181
to0000.000555556
tr000588.698
tr1004067.68
ts0000.00111111
tt0000.0441194
ug0000.0905
uk000396.598
uk100450.411
ur00022.4373
uz0005.29325
ve0000.00355278
vi000779.854
vi1004963.77
vi1014239.37
vo0000.209436
wo0000.0801528
xh0000.126628
yi0000.0810111
yo0000.322206
zh000299.368
zu0000.139931

Contributors

xinjianl

500 commits

espnet/yodas

Dataset

154

stars

500

commits

4

linked in READMEs

Jun 10, 2024

updated

README

Updates

  • 2024/07/09: we also uploaded a new version of YODAS as YODAS2, it provides unsegmented audios and higher sampling rate (24k)

README

This is the YODAS manual/automatic subset from our YODAS dataset, it has 369,510 hours of speech.

This dataset contains audio utterances and corresponding captions (manual or automatic) from YouTube. Note that manual caption only indicates that it is uploaded by users, but not necessarily transcribed by a human

For more details about YODAS dataset, please refer to our paper

Usage:

Considering the extremely large size of the entire dataset, we support two modes of dataset loadings:

standard mode: each subset will be downloaded to the local dish before first iterating.

from datasets import load_dataset

# Note this will take very long time to download and preprocess
# you can try small subset for testing purpose
ds = load_dataset('espnet/yodas', 'en000')
print(next(iter(ds['train'])))

streaming mode most of the files will be streamed instead of downloaded to your local deivce. It can be used to inspect this dataset quickly.

from datasets import load_dataset

# this streaming loading will finish quickly
ds = load_dataset('espnet/yodas', 'en000', streaming=True)


#{'id': '9774', 'utt_id': 'YoRjzEnRcqu-00000-00000716-00000819', 'audio': {'path': None, 'array': array([-0.009552  , -0.01086426, -0.012146  , ..., -0.01992798,
#       -0.01885986, -0.01074219]), 'sampling_rate': 16000}, 'text': 'There is a saying'}
print(next(iter(ds['train'])))

Subsets/Shards

There are 149 languages in this dataset, each language is sharded into at least 1 shard to make it easy for our processing and uploading purposes. The raw data of each shard contains 500G at most.

Statistics of each shard can be found in the last section.

We distinguish manual caption subset and automatic caption subset by the first digit in each shard's name. The first digit is 0 if it contains manual captions, 1 if it contains automatic captions.

For example, en000 to en005 are the English shards containing manual subsets, and en100 to en127 contains the automatic subsets.

Reference

@inproceedings{li2023yodas,
  title={Yodas: Youtube-Oriented Dataset for Audio and Speech},
  author={Li, Xinjian and Takamichi, Shinnosuke and Saeki, Takaaki and Chen, William and Shiota, Sayaka and Watanabe, Shinji},
  booktitle={2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)},
  pages={1--8},
  year={2023},
  organization={IEEE}
}

Contact

If you have any questions, feel free to contact us at the following email address.

We made sure that our dataset only consisted of videos with CC licenses during our downloading. But in case you find your video unintentionally included in our dataset and would like to delete it, you can send a delete request to the following email.

Remove the parenthesis () from the following email address

(lixinjian)(1217)@gmail.com

Statistics

Note that there are no overlappings across different subsets, each audio can be included in the dataset at most once.

Subset nameHours
aa0000.171472
ab0000.358342
af0000.880497
ak0000.250858
am0000.924708
ar000289.707
as0000.548239
ay0000.0342722
az0003.8537
ba0000.0210556
be00048.1537
bg00046.8375
bh0000.0127111
bi0000.0125556
bm0000.00214722
bn00027.064
bo0000.746211
br0000.729914
bs0009.36959
ca00074.1909
co0000.0418639
cr0000.00584167
cs000167.604
cy0005.20017
da00027.4345
de0003063.81
de1004998.11
de1014995.08
de102955.389
dz0000.06365
ee0000.0411722
el000126.75
en0004999.73
en0015032.69
en0025039.9
en0035001.4
en0045054.66
en0054027.02
en1005147.07
en1015123.05
en1025117.68
en1035127.3
en1045126.33
en1055097.65
en1065131.47
en1075135.6
en1085136.84
en1095112.94
en1105109
en1115118.69
en1125122.57
en1135122.31
en1145112.36
en1155112.27
en1165123.77
en1175117.31
en1185117.94
en1195133.05
en1205127.79
en1215129.08
en1225130.22
en1235097.56
en1245116.59
en1255109.76
en1265136.21
en1272404.89
eo00012.6874
es0003737.86
es1005125.25
es1015130.44
es1025145.66
es1035138.26
es1045139.57
es1055138.95
es1062605.26
et00014.4129
eu00019.6356
fa00042.6734
ff0000.0394972
fi000212.899
fj0000.0167806
fo0000.183244
fr0002423.7
fr1005074.93
fr1015057.79
fr1025094.14
fr1033222.95
fy0000.0651667
ga0001.49252
gd0000.01885
gl0009.52575
gn0000.181356
gu0001.99355
ha0000.102931
hi000480.79
hi1002.74865
ho0000.0562194
hr00025.9171
ht0001.07494
hu000181.763
hy0001.64412
ia0000.0856056
id0001420.09
id1004902.79
id1013560.82
ie0000.134603
ig0000.086875
ik0000.00436667
is0005.07075
it0001454.98
it1004989.62
it1014242.87
iu0000.0584278
iw000161.373
ja0001094.18
ja1002929.94
jv0001.08701
ka00026.9727
ki0000.000555556
kk0003.72081
kl0000.00575556
km0003.98273
kn0002.36041
ko0002774.28
ko1005018.29
ko1015048.49
ko1025018.27
ko1032587.85
ks0000.0150444
ku0001.93419
ky00014.3917
la0007.26088
lb0000.1115
lg0000.00386111
ln0000.188739
lo0000.230986
lt00017.6507
lv0002.47671
mg0000.169653
mi0001.10089
mk0005.54236
ml00013.2386
mn0002.0232
mr0007.11602
ms00028.0219
my0002.35663
na0000.0397056
nd0000.00111111
ne0002.34936
nl000413.044
nl1002490.13
no000129.183
nv0000.00319444
oc0000.166108
om0000.148478
or0000.421436
pa0001.58188
pl000757.986
ps0000.9871
pt0001631.44
pt1005044.57
pt1015038.33
pt1025041.59
pt1033553.28
qu0000.748772
rm0000.192933
rn0000.00401111
ro00099.9175
ru0004968.37
ru001627.679
ru1005098.3
ru1015098
ru1025119.43
ru1035107.29
ru1045121.73
ru1055088.05
ru1063393.44
rw0000.640825
sa0000.354139
sc0000.00801111
sd0000.0768722
sg0000.000472222
sh0000.250914
si0004.2634
sk00030.0155
sl00022.9366
sm0000.102333
sn0000.0134722
so0003.36819
sq0003.48276
sr00015.2849
st0000.00324167
su0000.0404639
sv000127.411
sw0001.93409
ta00059.4805
te0005.66794
tg0000.272386
th000497.14
th1001.87429
ti0000.343897
tk0000.0651806
tn0000.112181
to0000.000555556
tr000588.698
tr1004067.68
ts0000.00111111
tt0000.0441194
ug0000.0905
uk000396.598
uk100450.411
ur00022.4373
uz0005.29325
ve0000.00355278
vi000779.854
vi1004963.77
vi1014239.37
vo0000.209436
wo0000.0801528
xh0000.126628
yi0000.0810111
yo0000.322206
zh000299.368
zu0000.139931

Contributors

xinjianl

500 commits