English · Türkçe
835,267 Turkish speech segments (2,398 hours) described in words: how each one sounds (pitch, pace, pauses, volume, noise, room, bandwidth), as a natural-language caption in English and Turkish, together with the measurements behind it and a transcript. The audio is not redistributed: every row points to its exact span in espnet/yodas3 (Turkish), so you fetch it from there.
Instruction-following and voice-description TTS need speech paired with descriptions of how it is spoken. English has several such datasets; Turkish had none. We built these captions for Alania-2, the Turkish text-to-speech model behind speech.patientdesk.ai, and release them so others can build controllable Turkish speech models and study Turkish prosody.
| Segments | 835,267 from 15,138 YouTube recordings (YODAS v3, Turkish) |
| Hours | 2,398 (median segment 9.9 s); 1,395 h with a full ≥16 kHz band |
| Speakers | ~29,000 speaker clusters (per recording, so a person appearing in several videos is counted more than once) |
| Captions | 537,395 distinct English captions; 70% describe the voice, 30% are instructions ("Speak slowly…") |
| Gender (by hours) | 72% male, 19% female, 9% unknown: the balance of Turkish YouTube, not a choice |
Examples:
A medium-pitched female voice speaks at a moderate pace with even energy; the recording sounds a bit muffled, like a phone call. Use a high, calm voice with monotone intonation and speak slowly, with a few short pauses; slight background noise. Orta perdeden konuşan bir kadın sesi, birkaç kısa duraklamayla; arkada hafif bir gürültü var.
patientdesk-stt-1.2.13). Where it disagreed strongly with
YouTube's own captions, asr_disagrees_with_youtube is true (8% of segments); filter on it for cleaner text.pace_label, pitch_label, noise_label, …), including
comparisons with the speaker's usual delivery ("faster than usual").Qwen/Qwen3-32B-FP8 turned the descriptors into one English and one Turkish caption per segment,
either describing the voice or phrased as an instruction. It also saw the transcript, only to spot an
unmistakable emotional act in the words (an apology, thanks, a complaint…); that tone goes into content_tone
and the caption, and is none for 96% of segments. Captions never quote the transcript: any caption sharing a
run of four words with it was rejected.| Column | Meaning |
|---|---|
segment_id | our id: <yodas_id>_<index> |
yodas_shard, yodas_id | the recording in espnet/yodas3 (data/tr, shard and file id) |
start, end, duration | the segment's span in that recording, in seconds |
transcript | Duyu transcript of the segment |
asr_disagrees_with_youtube | Duyu and YouTube's captions disagree strongly on this segment |
caption_en, caption_tr | the style caption in English and Turkish |
caption_style | describe or instruct |
content_tone | emotional tone of the words, if any |
speaker, gender, p_female | speaker cluster within the recording, predicted gender and its probability |
*_label, long_pause | the descriptors the caption was written from |
f0_median_hz … dnsmos_ovrl | the measurements: pitch, rate, pauses, SNR, reverb, level, bandwidth, quality |
turkish_lid_p | Turkish language-ID probability |
import io, tarfile
from huggingface_hub import hf_hub_download
from datasets import load_dataset
caps = load_dataset("cloud0day3/alania-speech-captions-tr", split="train")
row = caps[0]
tar_path = hf_hub_download("espnet/yodas3", f"data/tr/audio/{row['yodas_shard']}.tar", repo_type="dataset")
with tarfile.open(tar_path) as tar:
member = next(m for m in tar.getmembers() if m.name.split("/")[-1].split(".")[0] == row["yodas_id"])
audio_bytes = tar.extractfile(member).read() # the whole recording; cut row["start"]:row["end"]
transcript): CC BY 4.0,
© 2026 PatientDesk AI. Credit: Alania Turkish Speech Style Captions by PatientDesk AI.transcript is derived from the YODAS audio and follows its licence: YODAS distributes recordings whose
uploaders marked them CC BY 3.0; keep the YODAS attribution when
you use it.Please also cite YODAS (Li et al., YODAS: YouTube-Oriented Dataset for Audio and Speech, ASRU 2023).
English · Türkçe
Kelimelerle anlatılmış 835.267 Türkçe konuşma kesiti (2.398 saat): her birinin nasıl duyulduğu (perde, hız, duraklamalar, ses düzeyi, gürültü, ortam, bant genişliği) İngilizce ve Türkçe doğal dilde bir açıklamayla, dayandığı ölçümlerle ve transkriptle birlikte. Ses dosyaları yeniden dağıtılmıyor: her satır, espnet/yodas3 (Türkçe) içindeki tam yerini gösteriyor; sesi oradan alıyorsunuz.
Talimatla yönlendirilen ve ses tarifinden konuşan TTS modelleri, konuşmanın nasıl söylendiğini anlatan verilere ihtiyaç duyar. İngilizcede bu tür birkaç veri seti var; Türkçede hiç yoktu. Bu açıklamaları, speech.patientdesk.ai arkasındaki Türkçe metinden sese modelimiz Alania-2 için hazırladık. Başkaları da kontrol edilebilir Türkçe konuşma modelleri geliştirebilsin ve Türkçe tonlamayı inceleyebilsin diye yayımlıyoruz.
YODAS v3 Türkçe kayıtları (9.216 saat) kesitlere ayrıldı; Türkçe dil tespiti (VoxLingua107, Whisper ile çapraz
kontrol), müzik ve şarkı ayıklama, kırpılma, bant genişliği ve DNSMOS denetimlerinden geçti (ham saatlerin yaklaşık
%30'u kaldı). Transkriptler Türkçe konuşma tanıma modelimiz Duyu ile çıkarıldı; YouTube altyazılarıyla belirgin biçimde
uyuşmadığı kesitlerde asr_disagrees_with_youtube doğrudur (%8). Perde, konuşma hızı, duraklama, gürültü, yankı gibi
ölçümler etiketlere dönüştürüldü; Qwen/Qwen3-32B-FP8 bu etiketlerden her kesit için bir İngilizce ve bir Türkçe
açıklama yazdı. Model transkripti yalnızca sözlerdeki belirgin bir duyguyu (özür, teşekkür, şikâyet…) fark etmek için
gördü (content_tone, kesitlerin %96'sında none); açıklamalar transkriptten alıntı yapmaz.
transcript dışındaki her şey): CC BY 4.0,
© 2026 PatientDesk AI. Atıf: Alania Turkish Speech Style Captions, PatientDesk AI.transcript YODAS seslerinden türetilmiştir ve onun lisansına tabidir: YODAS, yükleyenlerin
CC BY 3.0 olarak işaretlediği kayıtları dağıtır; kullanırken
YODAS atfını koruyun.İletişim: sezgin@patientdesk.ai
English · Türkçe
835,267 Turkish speech segments (2,398 hours) described in words: how each one sounds (pitch, pace, pauses, volume, noise, room, bandwidth), as a natural-language caption in English and Turkish, together with the measurements behind it and a transcript. The audio is not redistributed: every row points to its exact span in espnet/yodas3 (Turkish), so you fetch it from there.
Instruction-following and voice-description TTS need speech paired with descriptions of how it is spoken. English has several such datasets; Turkish had none. We built these captions for Alania-2, the Turkish text-to-speech model behind speech.patientdesk.ai, and release them so others can build controllable Turkish speech models and study Turkish prosody.
| Segments | 835,267 from 15,138 YouTube recordings (YODAS v3, Turkish) |
| Hours | 2,398 (median segment 9.9 s); 1,395 h with a full ≥16 kHz band |
| Speakers | ~29,000 speaker clusters (per recording, so a person appearing in several videos is counted more than once) |
| Captions | 537,395 distinct English captions; 70% describe the voice, 30% are instructions ("Speak slowly…") |
| Gender (by hours) | 72% male, 19% female, 9% unknown: the balance of Turkish YouTube, not a choice |
Examples:
A medium-pitched female voice speaks at a moderate pace with even energy; the recording sounds a bit muffled, like a phone call. Use a high, calm voice with monotone intonation and speak slowly, with a few short pauses; slight background noise. Orta perdeden konuşan bir kadın sesi, birkaç kısa duraklamayla; arkada hafif bir gürültü var.
patientdesk-stt-1.2.13). Where it disagreed strongly with
YouTube's own captions, asr_disagrees_with_youtube is true (8% of segments); filter on it for cleaner text.pace_label, pitch_label, noise_label, …), including
comparisons with the speaker's usual delivery ("faster than usual").Qwen/Qwen3-32B-FP8 turned the descriptors into one English and one Turkish caption per segment,
either describing the voice or phrased as an instruction. It also saw the transcript, only to spot an
unmistakable emotional act in the words (an apology, thanks, a complaint…); that tone goes into content_tone
and the caption, and is none for 96% of segments. Captions never quote the transcript: any caption sharing a
run of four words with it was rejected.| Column | Meaning |
|---|---|
segment_id | our id: <yodas_id>_<index> |
yodas_shard, yodas_id | the recording in espnet/yodas3 (data/tr, shard and file id) |
start, end, duration | the segment's span in that recording, in seconds |
transcript | Duyu transcript of the segment |
asr_disagrees_with_youtube | Duyu and YouTube's captions disagree strongly on this segment |
caption_en, caption_tr | the style caption in English and Turkish |
caption_style | describe or instruct |
content_tone | emotional tone of the words, if any |
speaker, gender, p_female | speaker cluster within the recording, predicted gender and its probability |
*_label, long_pause | the descriptors the caption was written from |
f0_median_hz … dnsmos_ovrl | the measurements: pitch, rate, pauses, SNR, reverb, level, bandwidth, quality |
turkish_lid_p | Turkish language-ID probability |
import io, tarfile
from huggingface_hub import hf_hub_download
from datasets import load_dataset
caps = load_dataset("cloud0day3/alania-speech-captions-tr", split="train")
row = caps[0]
tar_path = hf_hub_download("espnet/yodas3", f"data/tr/audio/{row['yodas_shard']}.tar", repo_type="dataset")
with tarfile.open(tar_path) as tar:
member = next(m for m in tar.getmembers() if m.name.split("/")[-1].split(".")[0] == row["yodas_id"])
audio_bytes = tar.extractfile(member).read() # the whole recording; cut row["start"]:row["end"]
transcript): CC BY 4.0,
© 2026 PatientDesk AI. Credit: Alania Turkish Speech Style Captions by PatientDesk AI.transcript is derived from the YODAS audio and follows its licence: YODAS distributes recordings whose
uploaders marked them CC BY 3.0; keep the YODAS attribution when
you use it.Please also cite YODAS (Li et al., YODAS: YouTube-Oriented Dataset for Audio and Speech, ASRU 2023).
English · Türkçe
Kelimelerle anlatılmış 835.267 Türkçe konuşma kesiti (2.398 saat): her birinin nasıl duyulduğu (perde, hız, duraklamalar, ses düzeyi, gürültü, ortam, bant genişliği) İngilizce ve Türkçe doğal dilde bir açıklamayla, dayandığı ölçümlerle ve transkriptle birlikte. Ses dosyaları yeniden dağıtılmıyor: her satır, espnet/yodas3 (Türkçe) içindeki tam yerini gösteriyor; sesi oradan alıyorsunuz.
Talimatla yönlendirilen ve ses tarifinden konuşan TTS modelleri, konuşmanın nasıl söylendiğini anlatan verilere ihtiyaç duyar. İngilizcede bu tür birkaç veri seti var; Türkçede hiç yoktu. Bu açıklamaları, speech.patientdesk.ai arkasındaki Türkçe metinden sese modelimiz Alania-2 için hazırladık. Başkaları da kontrol edilebilir Türkçe konuşma modelleri geliştirebilsin ve Türkçe tonlamayı inceleyebilsin diye yayımlıyoruz.
YODAS v3 Türkçe kayıtları (9.216 saat) kesitlere ayrıldı; Türkçe dil tespiti (VoxLingua107, Whisper ile çapraz
kontrol), müzik ve şarkı ayıklama, kırpılma, bant genişliği ve DNSMOS denetimlerinden geçti (ham saatlerin yaklaşık
%30'u kaldı). Transkriptler Türkçe konuşma tanıma modelimiz Duyu ile çıkarıldı; YouTube altyazılarıyla belirgin biçimde
uyuşmadığı kesitlerde asr_disagrees_with_youtube doğrudur (%8). Perde, konuşma hızı, duraklama, gürültü, yankı gibi
ölçümler etiketlere dönüştürüldü; Qwen/Qwen3-32B-FP8 bu etiketlerden her kesit için bir İngilizce ve bir Türkçe
açıklama yazdı. Model transkripti yalnızca sözlerdeki belirgin bir duyguyu (özür, teşekkür, şikâyet…) fark etmek için
gördü (content_tone, kesitlerin %96'sında none); açıklamalar transkriptten alıntı yapmaz.
transcript dışındaki her şey): CC BY 4.0,
© 2026 PatientDesk AI. Atıf: Alania Turkish Speech Style Captions, PatientDesk AI.transcript YODAS seslerinden türetilmiştir ve onun lisansına tabidir: YODAS, yükleyenlerin
CC BY 3.0 olarak işaretlediği kayıtları dağıtır; kullanırken
YODAS atfını koruyun.İletişim: sezgin@patientdesk.ai