Common Voice Scripted Speech 25.0 - Chinese (Taiwan)
2
15 commits
2 linked in READMEs
updated Jun 6, 2026
This repository mirrors the Chinese (Taiwan) (zh-TW) portion of Mozilla Common Voice Scripted Speech 25.0 in Hugging Face datasets format. The audio has been embedded into Parquet and exposed as a Hugging Face Audio feature at 48 kHz.
Official source page: Mozilla Data Collective - Common Voice Scripted Speech 25.0 - Chinese (Taiwan)
| Field | Value |
|---|---|
| Dataset ID | cmn2g7eaj01fio10769r1m96n |
| Common Voice release | cv-corpus-25.0-2026-03-09 |
| Language | Chinese (Taiwan), 華語(台灣), zh-TW |
| Spoken variety | Taiwan Mandarin / 中華民國國語, cmn-TW |
| Task | Automatic Speech Recognition (ASR) |
| Original format | MP3 |
| Original archive | common-voice-scripted-speech-25-0-chines-e84858c5.tar.gz |
| Original archive size | 2.95 GB |
| Release date on Mozilla Data Collective | 2026-03-23 |
| License | Creative Commons Zero v1.0 Universal (CC0-1.0) |
The original Mozilla metadata describes this as a collection of read speech recordings in Chinese (Taiwan). It contains 140,630 clips from 2,317 self-selected volunteer speakers, totaling 131.33 hours of speech. Of those clips, 85,324 are validated, corresponding to 79.68 validated hours. The sentence pool contains 21,763 sentences.
The source dataset is released under CC0-1.0. The Mozilla Data Collective page also lists additional restrictions and constraints:
Speaker demographic fields are self-reported and optional. Blank values mean the speaker did not provide that field.
from datasets import load_dataset
ds = load_dataset("OpenFormosa/common_voice_25_zh-TW", split="train")
print(ds.features["audio"])
print(ds[0]["sentence"])
For quick inspection without downloading full splits:
from datasets import load_dataset, Audio
stream = load_dataset("OpenFormosa/common_voice_25_zh-TW", split="train", streaming=True)
stream = stream.cast_column("audio", Audio(sampling_rate=48000, decode=False))
row = next(iter(stream))
print(row["audio"])
print(row["sentence"])
The train, validation, and test splits are the official Common Voice training subsets. The validated, invalidated, and other splits preserve the broader Common Voice clip buckets. The training subsets are drawn from the validated bucket, so they should not be treated as disjoint from validated.
| Split | Rows | Notes |
|---|---|---|
train | 7,394 | Official training split |
validation | 5,119 | Original dev.tsv |
test | 5,119 | Official test split |
validated | 85,324 | Validated clips |
invalidated | 4,920 | Clips that did not pass validation |
other | 50,386 | Clips pending validation or otherwise outside validated/invalidated |
Mozilla's metadata reports 17,632 train/dev/test clips, covering 20.7% of validated clips. The average clip duration is 3.362 seconds.
| Column | Type | Description |
|---|---|---|
audio | Audio(sampling_rate=48000) | Embedded MP3 audio exposed through the Hugging Face Audio feature |
client_id | string | Hashed speaker identifier from Common Voice |
path | string | Original relative MP3 filename |
sentence_id | string | Common Voice sentence identifier |
sentence | string | Expected transcription text |
sentence_domain | string | Sentence domain label; may be blank or contain multiple comma-separated domains |
up_votes | int32 | Number of validators who accepted the clip |
down_votes | int32 | Number of validators who rejected the clip |
age | string | Self-reported speaker age bucket |
gender | string | Self-reported speaker gender |
accents | string | Self-reported accent or place-of-origin label |
variant | string | Language variant, if provided |
locale | string | Locale code, normally zh-TW |
segment | string | Custom dataset segment, if provided |
duration_ms | int32 | Clip duration from clip_durations.tsv |
The original source files are also included under source/:
source/README.zh-TW.mdsource/train.tsvsource/dev.tsvsource/test.tsvsource/validated.tsvsource/invalidated.tsvsource/other.tsvsource/reported.tsvsource/validated_sentences.tsvsource/unvalidated_sentences.tsvsource/clip_durations.tsv| Code | Label | Clips | Speakers |
|---|---|---|---|
male_masculine | Male, masculine | 68,527 (48.7%) | 598 (25.8%) |
female_feminine | Female, feminine | 31,056 (22.1%) | 258 (11.1%) |
transgender | Transgender | 100 (0.1%) | 1 (0.0%) |
non-binary | Non-binary | 0 | 0 |
do_not_wish_to_say | Prefer not to say | 25 (0.0%) | 2 (0.1%) |
| Unspecified | Not declared | 40,922 (29.1%) | 1,575 (68.0%) |
Gender declared: 99,708 of 140,630 clips (70.9%), 742 of 2,317 speakers (32.0%).
| Code | Label | Clips | Speakers |
|---|---|---|---|
teens | Teens | 8,440 (6.0%) | 82 (3.5%) |
twenties | Twenties | 41,664 (29.6%) | 451 (19.5%) |
thirties | Thirties | 27,021 (19.2%) | 231 (10.0%) |
fourties | Fourties | 12,771 (9.1%) | 105 (4.5%) |
fifties | Fifties | 12,587 (9.0%) | 27 (1.2%) |
sixties | Sixties | 431 (0.3%) | 3 (0.1%) |
seventies | Seventies | 30 (0.0%) | 4 (0.2%) |
eighties | Eighties | 0 | 0 |
nineties | Nineties | 0 | 0 |
| Unspecified | Not declared | 37,686 (26.8%) | 1,540 (66.5%) |
Age declared: 102,944 of 140,630 clips (73.2%), 777 of 2,317 speakers (33.5%).
The source page reports self-declared accent or place-of-origin coverage as follows.
| Code | Label | Clips | Speakers |
|---|---|---|---|
taipei_city | 出生地:臺北市 | 19,646 (14.0%) | 107 (4.6%) |
new_taipei_city | 出生地:新北市 | 8,850 (6.3%) | 62 (2.7%) |
taichung_city | 出生地:臺中市 | 4,411 (3.1%) | 47 (2.0%) |
kaohsiung_city | 出生地:高雄市 | 3,266 (2.3%) | 42 (1.8%) |
taoyuan_city | 出生地:桃園市 | 3,015 (2.1%) | 23 (1.0%) |
hsinchu_city | 出生地:新竹市 | 2,866 (2.0%) | 11 (0.5%) |
yunlin_county | 出生地:雲林縣 | 2,560 (1.8%) | 8 (0.3%) |
nantou_county | 出生地:南投縣 | 2,101 (1.5%) | 7 (0.3%) |
changhua_county | 出生地:彰化縣 | 2,009 (1.4%) | 22 (0.9%) |
tainan_city | 出生地:臺南市 | 1,708 (1.2%) | 21 (0.9%) |
chiayi_city | 出生地:嘉義市 | 1,195 (0.8%) | 5 (0.2%) |
pingtung_county | 出生地:屏東縣 | 913 (0.6%) | 6 (0.3%) |
hualien_county | 出生地:花蓮縣 | 878 (0.6%) | 5 (0.2%) |
yilan_county | 出生地:宜蘭縣 | 765 (0.5%) | 8 (0.3%) |
hong_kong | 香港 | 690 (0.5%) | 26 (1.1%) |
chiayi_county | 出生地:嘉義縣 | 379 (0.3%) | 7 (0.3%) |
hsinchu_county | 出生地:新竹縣 | 343 (0.2%) | 8 (0.3%) |
keelung_city | 出生地:基隆市 | 141 (0.1%) | 10 (0.4%) |
kinmen_county | 出生地:金門縣 | 55 (0.0%) | 1 (0.0%) |
penghu_county | 出生地:澎湖縣 | 20 (0.0%) | 2 (0.1%) |
miaoli_county | 出生地:苗栗縣 | 15 (0.0%) | 2 (0.1%) |
taitung_county | 出生地:臺東縣 | 10 (0.0%) | 2 (0.1%) |
| Other | Other accent strings | 5,017 (3.6%) | 21 (0.9%) |
Most Traditional Chinese text was organized through the MozTW CC0 sentence corpus. Mozilla's metadata reports:
Sentence source distribution:
| Source | Sentences |
|---|---|
sentence-collector | 15,566 (74.9%) |
setences | 2,897 (13.9%) |
MozTW CC0 corpus commit 01033097... | 666 (3.2%) |
MozTW CC0 corpus commit e340b6d... | 451 (2.2%) |
taipei_city_gov | 355 (1.7%) |
chatlogs | 309 (1.5%) |
| Other | 533 (2.6%) |
The source page reports the following domain coverage among clips. In this Hugging Face version, sentence_domain is preserved as the raw Common Voice field and may be empty or contain multiple comma-separated values.
| Code | Domain | Clips | Speakers |
|---|---|---|---|
general | General | 1,502 (1.1%) | 84 (3.6%) |
agriculture_food | Agriculture and Food | 12 (0.0%) | 7 (0.3%) |
automotive_transport | Automotive and Transport | 278 (0.2%) | 45 (1.9%) |
finance | Finance | 3 (0.0%) | 3 (0.1%) |
service_retail | Service and Retail | 151 (0.1%) | 36 (1.6%) |
healthcare | Healthcare | 25 (0.0%) | 18 (0.8%) |
history_law_government | History, Law and Government | 170 (0.1%) | 39 (1.7%) |
media_entertainment | Media and Entertainment | 170 (0.1%) | 44 (1.9%) |
nature_environment | Nature and Environment | 14 (0.0%) | 12 (0.5%) |
news_current_affairs | News and Current Affairs | 44 (0.0%) | 19 (0.8%) |
technology_robotics | Technology and Robotics | 777 (0.6%) | 49 (2.1%) |
language_fundamentals | Language Fundamentals | 8 (0.0%) | 7 (0.3%) |
Mozilla lists the intended use as training and evaluating automatic speech recognition models. The dataset may also be useful for computer-aided language learning and language or heritage revitalization applications.
The text corpus was created by Mozilla Taiwan community contributors, the g0v community, and other open-source volunteers. Speakers are primarily individual volunteers from Taiwan.
Mozilla's metadata names Irvin Chen as the MozTW community contact for dataset table preparation.
The dataset is released under Creative Commons Zero v1.0 Universal. Users should also review the Mozilla Data Collective dataset page for current restrictions and terms before use.
15 commits
Common Voice Scripted Speech 25.0 - Chinese (Taiwan)
2
15 commits
2 linked in READMEs
updated Jun 6, 2026
This repository mirrors the Chinese (Taiwan) (zh-TW) portion of Mozilla Common Voice Scripted Speech 25.0 in Hugging Face datasets format. The audio has been embedded into Parquet and exposed as a Hugging Face Audio feature at 48 kHz.
Official source page: Mozilla Data Collective - Common Voice Scripted Speech 25.0 - Chinese (Taiwan)
| Field | Value |
|---|---|
| Dataset ID | cmn2g7eaj01fio10769r1m96n |
| Common Voice release | cv-corpus-25.0-2026-03-09 |
| Language | Chinese (Taiwan), 華語(台灣), zh-TW |
| Spoken variety | Taiwan Mandarin / 中華民國國語, cmn-TW |
| Task | Automatic Speech Recognition (ASR) |
| Original format | MP3 |
| Original archive | common-voice-scripted-speech-25-0-chines-e84858c5.tar.gz |
| Original archive size | 2.95 GB |
| Release date on Mozilla Data Collective | 2026-03-23 |
| License | Creative Commons Zero v1.0 Universal (CC0-1.0) |
The original Mozilla metadata describes this as a collection of read speech recordings in Chinese (Taiwan). It contains 140,630 clips from 2,317 self-selected volunteer speakers, totaling 131.33 hours of speech. Of those clips, 85,324 are validated, corresponding to 79.68 validated hours. The sentence pool contains 21,763 sentences.
The source dataset is released under CC0-1.0. The Mozilla Data Collective page also lists additional restrictions and constraints:
Speaker demographic fields are self-reported and optional. Blank values mean the speaker did not provide that field.
from datasets import load_dataset
ds = load_dataset("OpenFormosa/common_voice_25_zh-TW", split="train")
print(ds.features["audio"])
print(ds[0]["sentence"])
For quick inspection without downloading full splits:
from datasets import load_dataset, Audio
stream = load_dataset("OpenFormosa/common_voice_25_zh-TW", split="train", streaming=True)
stream = stream.cast_column("audio", Audio(sampling_rate=48000, decode=False))
row = next(iter(stream))
print(row["audio"])
print(row["sentence"])
The train, validation, and test splits are the official Common Voice training subsets. The validated, invalidated, and other splits preserve the broader Common Voice clip buckets. The training subsets are drawn from the validated bucket, so they should not be treated as disjoint from validated.
| Split | Rows | Notes |
|---|---|---|
train | 7,394 | Official training split |
validation | 5,119 | Original dev.tsv |
test | 5,119 | Official test split |
validated | 85,324 | Validated clips |
invalidated | 4,920 | Clips that did not pass validation |
other | 50,386 | Clips pending validation or otherwise outside validated/invalidated |
Mozilla's metadata reports 17,632 train/dev/test clips, covering 20.7% of validated clips. The average clip duration is 3.362 seconds.
| Column | Type | Description |
|---|---|---|
audio | Audio(sampling_rate=48000) | Embedded MP3 audio exposed through the Hugging Face Audio feature |
client_id | string | Hashed speaker identifier from Common Voice |
path | string | Original relative MP3 filename |
sentence_id | string | Common Voice sentence identifier |
sentence | string | Expected transcription text |
sentence_domain | string | Sentence domain label; may be blank or contain multiple comma-separated domains |
up_votes | int32 | Number of validators who accepted the clip |
down_votes | int32 | Number of validators who rejected the clip |
age | string | Self-reported speaker age bucket |
gender | string | Self-reported speaker gender |
accents | string | Self-reported accent or place-of-origin label |
variant | string | Language variant, if provided |
locale | string | Locale code, normally zh-TW |
segment | string | Custom dataset segment, if provided |
duration_ms | int32 | Clip duration from clip_durations.tsv |
The original source files are also included under source/:
source/README.zh-TW.mdsource/train.tsvsource/dev.tsvsource/test.tsvsource/validated.tsvsource/invalidated.tsvsource/other.tsvsource/reported.tsvsource/validated_sentences.tsvsource/unvalidated_sentences.tsvsource/clip_durations.tsv| Code | Label | Clips | Speakers |
|---|---|---|---|
male_masculine | Male, masculine | 68,527 (48.7%) | 598 (25.8%) |
female_feminine | Female, feminine | 31,056 (22.1%) | 258 (11.1%) |
transgender | Transgender | 100 (0.1%) | 1 (0.0%) |
non-binary | Non-binary | 0 | 0 |
do_not_wish_to_say | Prefer not to say | 25 (0.0%) | 2 (0.1%) |
| Unspecified | Not declared | 40,922 (29.1%) | 1,575 (68.0%) |
Gender declared: 99,708 of 140,630 clips (70.9%), 742 of 2,317 speakers (32.0%).
| Code | Label | Clips | Speakers |
|---|---|---|---|
teens | Teens | 8,440 (6.0%) | 82 (3.5%) |
twenties | Twenties | 41,664 (29.6%) | 451 (19.5%) |
thirties | Thirties | 27,021 (19.2%) | 231 (10.0%) |
fourties | Fourties | 12,771 (9.1%) | 105 (4.5%) |
fifties | Fifties | 12,587 (9.0%) | 27 (1.2%) |
sixties | Sixties | 431 (0.3%) | 3 (0.1%) |
seventies | Seventies | 30 (0.0%) | 4 (0.2%) |
eighties | Eighties | 0 | 0 |
nineties | Nineties | 0 | 0 |
| Unspecified | Not declared | 37,686 (26.8%) | 1,540 (66.5%) |
Age declared: 102,944 of 140,630 clips (73.2%), 777 of 2,317 speakers (33.5%).
The source page reports self-declared accent or place-of-origin coverage as follows.
| Code | Label | Clips | Speakers |
|---|---|---|---|
taipei_city | 出生地:臺北市 | 19,646 (14.0%) | 107 (4.6%) |
new_taipei_city | 出生地:新北市 | 8,850 (6.3%) | 62 (2.7%) |
taichung_city | 出生地:臺中市 | 4,411 (3.1%) | 47 (2.0%) |
kaohsiung_city | 出生地:高雄市 | 3,266 (2.3%) | 42 (1.8%) |
taoyuan_city | 出生地:桃園市 | 3,015 (2.1%) | 23 (1.0%) |
hsinchu_city | 出生地:新竹市 | 2,866 (2.0%) | 11 (0.5%) |
yunlin_county | 出生地:雲林縣 | 2,560 (1.8%) | 8 (0.3%) |
nantou_county | 出生地:南投縣 | 2,101 (1.5%) | 7 (0.3%) |
changhua_county | 出生地:彰化縣 | 2,009 (1.4%) | 22 (0.9%) |
tainan_city | 出生地:臺南市 | 1,708 (1.2%) | 21 (0.9%) |
chiayi_city | 出生地:嘉義市 | 1,195 (0.8%) | 5 (0.2%) |
pingtung_county | 出生地:屏東縣 | 913 (0.6%) | 6 (0.3%) |
hualien_county | 出生地:花蓮縣 | 878 (0.6%) | 5 (0.2%) |
yilan_county | 出生地:宜蘭縣 | 765 (0.5%) | 8 (0.3%) |
hong_kong | 香港 | 690 (0.5%) | 26 (1.1%) |
chiayi_county | 出生地:嘉義縣 | 379 (0.3%) | 7 (0.3%) |
hsinchu_county | 出生地:新竹縣 | 343 (0.2%) | 8 (0.3%) |
keelung_city | 出生地:基隆市 | 141 (0.1%) | 10 (0.4%) |
kinmen_county | 出生地:金門縣 | 55 (0.0%) | 1 (0.0%) |
penghu_county | 出生地:澎湖縣 | 20 (0.0%) | 2 (0.1%) |
miaoli_county | 出生地:苗栗縣 | 15 (0.0%) | 2 (0.1%) |
taitung_county | 出生地:臺東縣 | 10 (0.0%) | 2 (0.1%) |
| Other | Other accent strings | 5,017 (3.6%) | 21 (0.9%) |
Most Traditional Chinese text was organized through the MozTW CC0 sentence corpus. Mozilla's metadata reports:
Sentence source distribution:
| Source | Sentences |
|---|---|
sentence-collector | 15,566 (74.9%) |
setences | 2,897 (13.9%) |
MozTW CC0 corpus commit 01033097... | 666 (3.2%) |
MozTW CC0 corpus commit e340b6d... | 451 (2.2%) |
taipei_city_gov | 355 (1.7%) |
chatlogs | 309 (1.5%) |
| Other | 533 (2.6%) |
The source page reports the following domain coverage among clips. In this Hugging Face version, sentence_domain is preserved as the raw Common Voice field and may be empty or contain multiple comma-separated values.
| Code | Domain | Clips | Speakers |
|---|---|---|---|
general | General | 1,502 (1.1%) | 84 (3.6%) |
agriculture_food | Agriculture and Food | 12 (0.0%) | 7 (0.3%) |
automotive_transport | Automotive and Transport | 278 (0.2%) | 45 (1.9%) |
finance | Finance | 3 (0.0%) | 3 (0.1%) |
service_retail | Service and Retail | 151 (0.1%) | 36 (1.6%) |
healthcare | Healthcare | 25 (0.0%) | 18 (0.8%) |
history_law_government | History, Law and Government | 170 (0.1%) | 39 (1.7%) |
media_entertainment | Media and Entertainment | 170 (0.1%) | 44 (1.9%) |
nature_environment | Nature and Environment | 14 (0.0%) | 12 (0.5%) |
news_current_affairs | News and Current Affairs | 44 (0.0%) | 19 (0.8%) |
technology_robotics | Technology and Robotics | 777 (0.6%) | 49 (2.1%) |
language_fundamentals | Language Fundamentals | 8 (0.0%) | 7 (0.3%) |
Mozilla lists the intended use as training and evaluating automatic speech recognition models. The dataset may also be useful for computer-aided language learning and language or heritage revitalization applications.
The text corpus was created by Mozilla Taiwan community contributors, the g0v community, and other open-source volunteers. Speakers are primarily individual volunteers from Taiwan.
Mozilla's metadata names Irvin Chen as the MozTW community contact for dataset table preparation.
The dataset is released under Creative Commons Zero v1.0 Universal. Users should also review the Mozilla Data Collective dataset page for current restrictions and terms before use.
15 commits