[!IMPORTANT]
The following rules (in the original repository) must be followed:必须遵守GNU General Public License v3.0内的所有协议!
附加:禁止商用,本数据集以及使用本数据集训练出来的任何模型都不得用于任何商业行为,如要用于商业用途,请找数据列表内的所有厂商授权(笑),因违反开源协议而出现的任何问题都与本人无关! 训练出来的模型必须开源,是否在README内引用本数据集由训练者自主决定,不做强制要求。English: You must comply with all the terms of the GNU General Public License v3.0!
Additional note: Commercial use is prohibited. This dataset and any model trained using this dataset cannot be used for any commercial purposes. If you wish to use it for commercial purposes, please obtain authorization from all the providers listed in the dataset (LOL). I bear no responsibility for any issues arising from violations of the open-source license! Models trained using this dataset must be open-sourced. Whether to cite this dataset in the README is left to the discretion of the user and is not mandatory.日本語: GNU General Public License v3.0 内のすべての規約を遵守する必要があります!
追加事項:商用利用は禁止されています。本データセットおよび本データセットを使用して訓練されたいかなるモデルも商業行為には一切使用できません。商用利用を希望する場合は、データセットリスト内のすべての提供者の許可を取得してください(笑)。オープンソースライセンス違反によって発生したいかなる問題も私は責任を負いません! このデータセットを使用して訓練されたモデルはオープンソースにする必要があります。README 内で本データセットを引用するかどうかは、ユーザーの自主的な判断に委ねられており、強制されません。
2024-10-12: Removed 190 audio-text pairs such that
Resulting in 3,746,131 pairs and 5353.9 hours, and the number of files in each tar file may be smaller than 32768.
All the audio files and transcriptions are from OOPPEENN/Galgame_Dataset. Many thanks to the original authors!
I modified the original dataset in the following ways:
\t, ― (dash), and spaces (half-width or full-width), and normalized some letters and symbols (e.g., "え~?" → "えー?")!?♪♡)ッっあいうえおんぁぃぅぇぉゃゅょアイウエオンァィゥェォャュョ with 3 or more repetitions to 2 repetitions (e.g., "あああっっっ" → "ああっっ")。、!?…♪♡○galgame-speech-asr-16kHz-train-{000000..000114}.tar files.00000aa36e86ba49cb67fb886cce2c044c03dbb8ffddad4cb4e5f2da809e91ab.ogg
00000aa36e86ba49cb67fb886cce2c044c03dbb8ffddad4cb4e5f2da809e91ab.txt
00000fe59140c18655921cd316f03ae7a81a0708a2d81a15d9b7ae866c459840.ogg
00000fe59140c18655921cd316f03ae7a81a0708a2d81a15d9b7ae866c459840.txt
...
Except for the last tar file, each tar file contains about 32768 audio-text pairs (OGG and TXT files), hence about 65536 files in total (the number may be smaller than 32768 since I removed some files after the initial upload).
File names are randomly generated SHA-256 hashes, so the order of the files has no mean (e.g., the files coming from the same Galgame are not necessarily adjacent).
To load this dataset in the 🤗 Datasets library, just use:
from datasets import load_dataset
dataset = load_dataset("litagin/Galgame_Speech_ASR_16kHz", streaming=True)
Be sure to set streaming=True if you want to avoid downloading the whole dataset at once.
See example.ipynb for a simple example of how to use the dataset in this way.
See Webdataset for more details on how to use the dataset in WebDataset format in, e.g., PyTorch.
13 commits
[!IMPORTANT]
The following rules (in the original repository) must be followed:必须遵守GNU General Public License v3.0内的所有协议!
附加:禁止商用,本数据集以及使用本数据集训练出来的任何模型都不得用于任何商业行为,如要用于商业用途,请找数据列表内的所有厂商授权(笑),因违反开源协议而出现的任何问题都与本人无关! 训练出来的模型必须开源,是否在README内引用本数据集由训练者自主决定,不做强制要求。English: You must comply with all the terms of the GNU General Public License v3.0!
Additional note: Commercial use is prohibited. This dataset and any model trained using this dataset cannot be used for any commercial purposes. If you wish to use it for commercial purposes, please obtain authorization from all the providers listed in the dataset (LOL). I bear no responsibility for any issues arising from violations of the open-source license! Models trained using this dataset must be open-sourced. Whether to cite this dataset in the README is left to the discretion of the user and is not mandatory.日本語: GNU General Public License v3.0 内のすべての規約を遵守する必要があります!
追加事項:商用利用は禁止されています。本データセットおよび本データセットを使用して訓練されたいかなるモデルも商業行為には一切使用できません。商用利用を希望する場合は、データセットリスト内のすべての提供者の許可を取得してください(笑)。オープンソースライセンス違反によって発生したいかなる問題も私は責任を負いません! このデータセットを使用して訓練されたモデルはオープンソースにする必要があります。README 内で本データセットを引用するかどうかは、ユーザーの自主的な判断に委ねられており、強制されません。
2024-10-12: Removed 190 audio-text pairs such that
Resulting in 3,746,131 pairs and 5353.9 hours, and the number of files in each tar file may be smaller than 32768.
All the audio files and transcriptions are from OOPPEENN/Galgame_Dataset. Many thanks to the original authors!
I modified the original dataset in the following ways:
\t, ― (dash), and spaces (half-width or full-width), and normalized some letters and symbols (e.g., "え~?" → "えー?")!?♪♡)ッっあいうえおんぁぃぅぇぉゃゅょアイウエオンァィゥェォャュョ with 3 or more repetitions to 2 repetitions (e.g., "あああっっっ" → "ああっっ")。、!?…♪♡○galgame-speech-asr-16kHz-train-{000000..000114}.tar files.00000aa36e86ba49cb67fb886cce2c044c03dbb8ffddad4cb4e5f2da809e91ab.ogg
00000aa36e86ba49cb67fb886cce2c044c03dbb8ffddad4cb4e5f2da809e91ab.txt
00000fe59140c18655921cd316f03ae7a81a0708a2d81a15d9b7ae866c459840.ogg
00000fe59140c18655921cd316f03ae7a81a0708a2d81a15d9b7ae866c459840.txt
...
Except for the last tar file, each tar file contains about 32768 audio-text pairs (OGG and TXT files), hence about 65536 files in total (the number may be smaller than 32768 since I removed some files after the initial upload).
File names are randomly generated SHA-256 hashes, so the order of the files has no mean (e.g., the files coming from the same Galgame are not necessarily adjacent).
To load this dataset in the 🤗 Datasets library, just use:
from datasets import load_dataset
dataset = load_dataset("litagin/Galgame_Speech_ASR_16kHz", streaming=True)
Be sure to set streaming=True if you want to avoid downloading the whole dataset at once.
See example.ipynb for a simple example of how to use the dataset in this way.
See Webdataset for more details on how to use the dataset in WebDataset format in, e.g., PyTorch.
13 commits