cloud0day3/alania-speech-captions-tr

Dataset

Alania Turkish Speech Style Captions

10

4 commits

updated Oct 1, 2026

See the code

README

Alania Turkish Speech Style Captions

English · Türkçe

835,267 Turkish speech segments (2,398 hours) described in words: how each one sounds (pitch, pace, pauses, volume, noise, room, bandwidth), as a natural-language caption in English and Turkish, together with the measurements behind it and a transcript. The audio is not redistributed: every row points to its exact span in espnet/yodas3 (Turkish), so you fetch it from there.

Instruction-following and voice-description TTS need speech paired with descriptions of how it is spoken. English has several such datasets; Turkish had none. We built these captions for Alania-2, the Turkish text-to-speech model behind speech.patientdesk.ai, and release them so others can build controllable Turkish speech models and study Turkish prosody.

Segments835,267 from 15,138 YouTube recordings (YODAS v3, Turkish)
Hours2,398 (median segment 9.9 s); 1,395 h with a full ≥16 kHz band
Speakers~29,000 speaker clusters (per recording, so a person appearing in several videos is counted more than once)
Captions537,395 distinct English captions; 70% describe the voice, 30% are instructions ("Speak slowly…")
Gender (by hours)72% male, 19% female, 9% unknown: the balance of Turkish YouTube, not a choice

Examples:

A medium-pitched female voice speaks at a moderate pace with even energy; the recording sounds a bit muffled, like a phone call. Use a high, calm voice with monotone intonation and speak slowly, with a few short pauses; slight background noise. Orta perdeden konuşan bir kadın sesi, birkaç kısa duraklamayla; arkada hafif bir gürültü var.

How it was made

  1. Selection: YODAS v3 Turkish (9,216 h) → voice activity detection and segmentation, Turkish language ID (VoxLingua107, cross-checked by Whisper), music and singing removal (AST), clipping, bandwidth and DNSMOS checks (overall ≥ 2.4). About 30% of the raw hours were kept.
  2. Transcripts: our Turkish speech recogniser, Duyu (patientdesk-stt-1.2.13). Where it disagreed strongly with YouTube's own captions, asr_disagrees_with_youtube is true (8% of segments); filter on it for cleaner text.
  3. Measurements: median F0 and its spread (relative to each speaker's own baseline), syllable and articulation rate, pauses, SNR, reverberation time, level, bandwidth, DNSMOS; gender from a wav2vec2 classifier.
  4. Descriptors: the measurements binned into labels (pace_label, pitch_label, noise_label, …), including comparisons with the speaker's usual delivery ("faster than usual").
  5. Captions: Qwen/Qwen3-32B-FP8 turned the descriptors into one English and one Turkish caption per segment, either describing the voice or phrased as an instruction. It also saw the transcript, only to spot an unmistakable emotional act in the words (an apology, thanks, a complaint…); that tone goes into content_tone and the caption, and is none for 96% of segments. Captions never quote the transcript: any caption sharing a run of four words with it was rejected.

Columns

ColumnMeaning
segment_idour id: <yodas_id>_<index>
yodas_shard, yodas_idthe recording in espnet/yodas3 (data/tr, shard and file id)
start, end, durationthe segment's span in that recording, in seconds
transcriptDuyu transcript of the segment
asr_disagrees_with_youtubeDuyu and YouTube's captions disagree strongly on this segment
caption_en, caption_trthe style caption in English and Turkish
caption_styledescribe or instruct
content_toneemotional tone of the words, if any
speaker, gender, p_femalespeaker cluster within the recording, predicted gender and its probability
*_label, long_pausethe descriptors the caption was written from
f0_median_hz … dnsmos_ovrlthe measurements: pitch, rate, pauses, SNR, reverb, level, bandwidth, quality
turkish_lid_pTurkish language-ID probability

Getting the audio

import io, tarfile
from huggingface_hub import hf_hub_download
from datasets import load_dataset

caps = load_dataset("cloud0day3/alania-speech-captions-tr", split="train")
row = caps[0]
tar_path = hf_hub_download("espnet/yodas3", f"data/tr/audio/{row['yodas_shard']}.tar", repo_type="dataset")
with tarfile.open(tar_path) as tar:
    member = next(m for m in tar.getmembers() if m.name.split("/")[-1].split(".")[0] == row["yodas_id"])
    audio_bytes = tar.extractfile(member).read()  # the whole recording; cut row["start"]:row["end"]

Licence and attribution

  • Captions, descriptors and measurements (everything except transcript): CC BY 4.0, © 2026 PatientDesk AI. Credit: Alania Turkish Speech Style Captions by PatientDesk AI.
  • transcript is derived from the YODAS audio and follows its licence: YODAS distributes recordings whose uploaders marked them CC BY 3.0; keep the YODAS attribution when you use it.
  • Audio is not included. Its use is governed by YODAS's terms and the original uploaders' licences.

Please also cite YODAS (Li et al., YODAS: YouTube-Oriented Dataset for Audio and Speech, ASRU 2023).

Limitations and responsible use

  • Captions, labels, transcripts and gender are automatic. Expect some errors, especially for overlapping speakers, music beds and very short segments.
  • The data is 72% male; models trained on it alone will under-serve female voices.
  • The recordings are everyday Turkish YouTube (talks, news, politics, games, lectures, vlogs) and were not selected by topic. What the speakers say is their own; PatientDesk AI does not select, endorse or vouch for any of it.
  • The recordings are real people on YouTube. Do not use this dataset to identify, profile or imitate them, and do not build voice clones of specific speakers from it. Transcripts may mention names or other personal details that were spoken in the videos.
  • Descriptions such as "high-pitched" are relative to Turkish speech in this corpus and to each speaker's own range.

Türkçe

English · Türkçe

Kelimelerle anlatılmış 835.267 Türkçe konuşma kesiti (2.398 saat): her birinin nasıl duyulduğu (perde, hız, duraklamalar, ses düzeyi, gürültü, ortam, bant genişliği) İngilizce ve Türkçe doğal dilde bir açıklamayla, dayandığı ölçümlerle ve transkriptle birlikte. Ses dosyaları yeniden dağıtılmıyor: her satır, espnet/yodas3 (Türkçe) içindeki tam yerini gösteriyor; sesi oradan alıyorsunuz.

Talimatla yönlendirilen ve ses tarifinden konuşan TTS modelleri, konuşmanın nasıl söylendiğini anlatan verilere ihtiyaç duyar. İngilizcede bu tür birkaç veri seti var; Türkçede hiç yoktu. Bu açıklamaları, speech.patientdesk.ai arkasındaki Türkçe metinden sese modelimiz Alania-2 için hazırladık. Başkaları da kontrol edilebilir Türkçe konuşma modelleri geliştirebilsin ve Türkçe tonlamayı inceleyebilsin diye yayımlıyoruz.

  • Kesit: 15.138 YouTube kaydından 835.267 kesit (YODAS v3 Türkçe); 2.398 saat, ortanca kesit 9,9 sn; 1.395 saati tam bant (≥16 kHz).
  • Konuşmacı: yaklaşık 29.000 konuşmacı kümesi (kayıt başına; birden fazla videoda geçen biri birden çok sayılır).
  • Açıklama: 537.395 farklı İngilizce açıklama; %70'i sesi tarif ediyor, %30'u talimat ("Yavaş konuş…").
  • Cinsiyet (saat): %72 erkek, %19 kadın, %9 belirsiz. Bu bir tercih değil, Türkçe YouTube'un dengesi.

Nasıl hazırlandı

YODAS v3 Türkçe kayıtları (9.216 saat) kesitlere ayrıldı; Türkçe dil tespiti (VoxLingua107, Whisper ile çapraz kontrol), müzik ve şarkı ayıklama, kırpılma, bant genişliği ve DNSMOS denetimlerinden geçti (ham saatlerin yaklaşık %30'u kaldı). Transkriptler Türkçe konuşma tanıma modelimiz Duyu ile çıkarıldı; YouTube altyazılarıyla belirgin biçimde uyuşmadığı kesitlerde asr_disagrees_with_youtube doğrudur (%8). Perde, konuşma hızı, duraklama, gürültü, yankı gibi ölçümler etiketlere dönüştürüldü; Qwen/Qwen3-32B-FP8 bu etiketlerden her kesit için bir İngilizce ve bir Türkçe açıklama yazdı. Model transkripti yalnızca sözlerdeki belirgin bir duyguyu (özür, teşekkür, şikâyet…) fark etmek için gördü (content_tone, kesitlerin %96'sında none); açıklamalar transkriptten alıntı yapmaz.

Lisans ve atıf

  • Açıklamalar, etiketler ve ölçümler (transcript dışındaki her şey): CC BY 4.0, © 2026 PatientDesk AI. Atıf: Alania Turkish Speech Style Captions, PatientDesk AI.
  • transcript YODAS seslerinden türetilmiştir ve onun lisansına tabidir: YODAS, yükleyenlerin CC BY 3.0 olarak işaretlediği kayıtları dağıtır; kullanırken YODAS atfını koruyun.
  • Ses dahil değildir; kullanımı YODAS koşullarına ve yükleyenlerin lisanslarına tabidir.

Sınırlamalar ve sorumlu kullanım

  • Açıklamalar, etiketler, transkriptler ve cinsiyet otomatik üretilmiştir; özellikle üst üste konuşmalarda, müzikli bölümlerde ve çok kısa kesitlerde hatalar olabilir.
  • Verinin %72'si erkek sesidir; yalnızca bununla eğitilen modeller kadın seslerinde zayıf kalır.
  • Kayıtlar gündelik Türkçe YouTube içeriğidir (sohbet, haber, siyaset, oyun, ders, vlog) ve konuya göre seçilmemiştir. Konuşmacıların söyledikleri kendilerine aittir; PatientDesk AI bu içerikleri seçmez, desteklemez ve doğruluğunu taahhüt etmez.
  • Kayıtlar YouTube'daki gerçek kişilere aittir. Bu veriyi onları tanımlamak, profillemek ya da taklit etmek için kullanmayın; belirli konuşmacıların sesini klonlamayın. Transkriptlerde videolarda geçen isimler veya kişisel bilgiler bulunabilir.

İletişim: sezgin@patientdesk.ai

captions
instruction-tts
prosody
speech
style
turkish
yodas

cloud0day3/alania-speech-captions-tr

Dataset

Alania Turkish Speech Style Captions

10

4 commits

updated Oct 1, 2026

See the code

README

Alania Turkish Speech Style Captions

English · Türkçe

835,267 Turkish speech segments (2,398 hours) described in words: how each one sounds (pitch, pace, pauses, volume, noise, room, bandwidth), as a natural-language caption in English and Turkish, together with the measurements behind it and a transcript. The audio is not redistributed: every row points to its exact span in espnet/yodas3 (Turkish), so you fetch it from there.

Instruction-following and voice-description TTS need speech paired with descriptions of how it is spoken. English has several such datasets; Turkish had none. We built these captions for Alania-2, the Turkish text-to-speech model behind speech.patientdesk.ai, and release them so others can build controllable Turkish speech models and study Turkish prosody.

Segments835,267 from 15,138 YouTube recordings (YODAS v3, Turkish)
Hours2,398 (median segment 9.9 s); 1,395 h with a full ≥16 kHz band
Speakers~29,000 speaker clusters (per recording, so a person appearing in several videos is counted more than once)
Captions537,395 distinct English captions; 70% describe the voice, 30% are instructions ("Speak slowly…")
Gender (by hours)72% male, 19% female, 9% unknown: the balance of Turkish YouTube, not a choice

Examples:

A medium-pitched female voice speaks at a moderate pace with even energy; the recording sounds a bit muffled, like a phone call. Use a high, calm voice with monotone intonation and speak slowly, with a few short pauses; slight background noise. Orta perdeden konuşan bir kadın sesi, birkaç kısa duraklamayla; arkada hafif bir gürültü var.

How it was made

  1. Selection: YODAS v3 Turkish (9,216 h) → voice activity detection and segmentation, Turkish language ID (VoxLingua107, cross-checked by Whisper), music and singing removal (AST), clipping, bandwidth and DNSMOS checks (overall ≥ 2.4). About 30% of the raw hours were kept.
  2. Transcripts: our Turkish speech recogniser, Duyu (patientdesk-stt-1.2.13). Where it disagreed strongly with YouTube's own captions, asr_disagrees_with_youtube is true (8% of segments); filter on it for cleaner text.
  3. Measurements: median F0 and its spread (relative to each speaker's own baseline), syllable and articulation rate, pauses, SNR, reverberation time, level, bandwidth, DNSMOS; gender from a wav2vec2 classifier.
  4. Descriptors: the measurements binned into labels (pace_label, pitch_label, noise_label, …), including comparisons with the speaker's usual delivery ("faster than usual").
  5. Captions: Qwen/Qwen3-32B-FP8 turned the descriptors into one English and one Turkish caption per segment, either describing the voice or phrased as an instruction. It also saw the transcript, only to spot an unmistakable emotional act in the words (an apology, thanks, a complaint…); that tone goes into content_tone and the caption, and is none for 96% of segments. Captions never quote the transcript: any caption sharing a run of four words with it was rejected.

Columns

ColumnMeaning
segment_idour id: <yodas_id>_<index>
yodas_shard, yodas_idthe recording in espnet/yodas3 (data/tr, shard and file id)
start, end, durationthe segment's span in that recording, in seconds
transcriptDuyu transcript of the segment
asr_disagrees_with_youtubeDuyu and YouTube's captions disagree strongly on this segment
caption_en, caption_trthe style caption in English and Turkish
caption_styledescribe or instruct
content_toneemotional tone of the words, if any
speaker, gender, p_femalespeaker cluster within the recording, predicted gender and its probability
*_label, long_pausethe descriptors the caption was written from
f0_median_hz … dnsmos_ovrlthe measurements: pitch, rate, pauses, SNR, reverb, level, bandwidth, quality
turkish_lid_pTurkish language-ID probability

Getting the audio

import io, tarfile
from huggingface_hub import hf_hub_download
from datasets import load_dataset

caps = load_dataset("cloud0day3/alania-speech-captions-tr", split="train")
row = caps[0]
tar_path = hf_hub_download("espnet/yodas3", f"data/tr/audio/{row['yodas_shard']}.tar", repo_type="dataset")
with tarfile.open(tar_path) as tar:
    member = next(m for m in tar.getmembers() if m.name.split("/")[-1].split(".")[0] == row["yodas_id"])
    audio_bytes = tar.extractfile(member).read()  # the whole recording; cut row["start"]:row["end"]

Licence and attribution

  • Captions, descriptors and measurements (everything except transcript): CC BY 4.0, © 2026 PatientDesk AI. Credit: Alania Turkish Speech Style Captions by PatientDesk AI.
  • transcript is derived from the YODAS audio and follows its licence: YODAS distributes recordings whose uploaders marked them CC BY 3.0; keep the YODAS attribution when you use it.
  • Audio is not included. Its use is governed by YODAS's terms and the original uploaders' licences.

Please also cite YODAS (Li et al., YODAS: YouTube-Oriented Dataset for Audio and Speech, ASRU 2023).

Limitations and responsible use

  • Captions, labels, transcripts and gender are automatic. Expect some errors, especially for overlapping speakers, music beds and very short segments.
  • The data is 72% male; models trained on it alone will under-serve female voices.
  • The recordings are everyday Turkish YouTube (talks, news, politics, games, lectures, vlogs) and were not selected by topic. What the speakers say is their own; PatientDesk AI does not select, endorse or vouch for any of it.
  • The recordings are real people on YouTube. Do not use this dataset to identify, profile or imitate them, and do not build voice clones of specific speakers from it. Transcripts may mention names or other personal details that were spoken in the videos.
  • Descriptions such as "high-pitched" are relative to Turkish speech in this corpus and to each speaker's own range.

Türkçe

English · Türkçe

Kelimelerle anlatılmış 835.267 Türkçe konuşma kesiti (2.398 saat): her birinin nasıl duyulduğu (perde, hız, duraklamalar, ses düzeyi, gürültü, ortam, bant genişliği) İngilizce ve Türkçe doğal dilde bir açıklamayla, dayandığı ölçümlerle ve transkriptle birlikte. Ses dosyaları yeniden dağıtılmıyor: her satır, espnet/yodas3 (Türkçe) içindeki tam yerini gösteriyor; sesi oradan alıyorsunuz.

Talimatla yönlendirilen ve ses tarifinden konuşan TTS modelleri, konuşmanın nasıl söylendiğini anlatan verilere ihtiyaç duyar. İngilizcede bu tür birkaç veri seti var; Türkçede hiç yoktu. Bu açıklamaları, speech.patientdesk.ai arkasındaki Türkçe metinden sese modelimiz Alania-2 için hazırladık. Başkaları da kontrol edilebilir Türkçe konuşma modelleri geliştirebilsin ve Türkçe tonlamayı inceleyebilsin diye yayımlıyoruz.

  • Kesit: 15.138 YouTube kaydından 835.267 kesit (YODAS v3 Türkçe); 2.398 saat, ortanca kesit 9,9 sn; 1.395 saati tam bant (≥16 kHz).
  • Konuşmacı: yaklaşık 29.000 konuşmacı kümesi (kayıt başına; birden fazla videoda geçen biri birden çok sayılır).
  • Açıklama: 537.395 farklı İngilizce açıklama; %70'i sesi tarif ediyor, %30'u talimat ("Yavaş konuş…").
  • Cinsiyet (saat): %72 erkek, %19 kadın, %9 belirsiz. Bu bir tercih değil, Türkçe YouTube'un dengesi.

Nasıl hazırlandı

YODAS v3 Türkçe kayıtları (9.216 saat) kesitlere ayrıldı; Türkçe dil tespiti (VoxLingua107, Whisper ile çapraz kontrol), müzik ve şarkı ayıklama, kırpılma, bant genişliği ve DNSMOS denetimlerinden geçti (ham saatlerin yaklaşık %30'u kaldı). Transkriptler Türkçe konuşma tanıma modelimiz Duyu ile çıkarıldı; YouTube altyazılarıyla belirgin biçimde uyuşmadığı kesitlerde asr_disagrees_with_youtube doğrudur (%8). Perde, konuşma hızı, duraklama, gürültü, yankı gibi ölçümler etiketlere dönüştürüldü; Qwen/Qwen3-32B-FP8 bu etiketlerden her kesit için bir İngilizce ve bir Türkçe açıklama yazdı. Model transkripti yalnızca sözlerdeki belirgin bir duyguyu (özür, teşekkür, şikâyet…) fark etmek için gördü (content_tone, kesitlerin %96'sında none); açıklamalar transkriptten alıntı yapmaz.

Lisans ve atıf

  • Açıklamalar, etiketler ve ölçümler (transcript dışındaki her şey): CC BY 4.0, © 2026 PatientDesk AI. Atıf: Alania Turkish Speech Style Captions, PatientDesk AI.
  • transcript YODAS seslerinden türetilmiştir ve onun lisansına tabidir: YODAS, yükleyenlerin CC BY 3.0 olarak işaretlediği kayıtları dağıtır; kullanırken YODAS atfını koruyun.
  • Ses dahil değildir; kullanımı YODAS koşullarına ve yükleyenlerin lisanslarına tabidir.

Sınırlamalar ve sorumlu kullanım

  • Açıklamalar, etiketler, transkriptler ve cinsiyet otomatik üretilmiştir; özellikle üst üste konuşmalarda, müzikli bölümlerde ve çok kısa kesitlerde hatalar olabilir.
  • Verinin %72'si erkek sesidir; yalnızca bununla eğitilen modeller kadın seslerinde zayıf kalır.
  • Kayıtlar gündelik Türkçe YouTube içeriğidir (sohbet, haber, siyaset, oyun, ders, vlog) ve konuya göre seçilmemiştir. Konuşmacıların söyledikleri kendilerine aittir; PatientDesk AI bu içerikleri seçmez, desteklemez ve doğruluğunu taahhüt etmez.
  • Kayıtlar YouTube'daki gerçek kişilere aittir. Bu veriyi onları tanımlamak, profillemek ya da taklit etmek için kullanmayın; belirli konuşmacıların sesini klonlamayın. Transkriptlerde videolarda geçen isimler veya kişisel bilgiler bulunabilir.

İletişim: sezgin@patientdesk.ai

captions
instruction-tts
prosody
speech
style
turkish
yodas