English · Türkçe
1,439,639 unique Turkish sentences of the kind a voice agent actually says: appointment and clinic dialogue, banking, e-commerce and cargo, customer support, confirmations, questions, numbers and codes, addresses and names, empathetic lines and longer explanations. About 1,819 hours of speech if read aloud.
We built it as the text side of synthetic training speech for Alania-2, the Turkish text-to-speech model behind speech.patientdesk.ai. Public Turkish speech and text corpora contain very little of this everyday service language, and it is exactly where speech systems stumble: dates, times, amounts, phone numbers, codes and names. We are releasing it so others building Turkish voice products can use it too.
Released under Creative Commons Attribution 4.0 (CC BY 4.0). You may use, share and adapt it for any purpose, including commercial use, as long as you credit PatientDesk AI, for example:
Alania Turkish Domain Text by PatientDesk AI (https://huggingface.co/datasets/cloud0day3/alania-domain-text-tr), licensed under CC BY 4.0.
In a paper, please cite:
@misc{patientdesk2026alaniadomaintext,
title = {Alania Turkish Domain Text},
author = {{PatientDesk AI}},
year = {2026},
howpublished = {\url{https://huggingface.co/datasets/cloud0day3/alania-domain-text-tr}},
note = {1.44M Turkish voice-agent sentences, CC BY 4.0}
}
Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 (Apache-2.0) on vLLM, prompted per category and subtopic
(for example appointments_clinics / randevu alma), September 2026.text_model is text_raw after our Turkish text normaliser: numbers, dates, times, money,
phone numbers, abbreviations and web addresses written out the way they are spoken, letter case kept.| Category | Lines |
|---|---|
| appointments_clinics | 228,337 |
| numbers_codes | 186,700 |
| customer_support | 158,940 |
| questions | 155,162 |
| addresses_names | 151,247 |
| ecommerce_cargo | 147,041 |
| banking | 145,372 |
| emotional_empathetic | 126,412 |
| confirmations_acks | 103,403 |
| long_explanations | 37,025 |
Length buckets by estimated speaking time at 14.5 characters per second: xs 1–3 s (411,466), s 3–6 s
(785,773), m 6–12 s (195,968), l 12–25 s (46,432).
| Column | Meaning |
|---|---|
line_id | stable id of the line |
text_raw | the sentence as generated |
text_model | the normalised form a TTS model reads |
category, tags | category and line tags (e.g. question) |
bucket, est_dur, n_chars, n_sent | length information |
lid_tr | Turkish language-ID score |
extra | JSON: subtopic, generation batch, generator model |
source, doc_id, unit_idx, license | provenance fields shared with the rest of our text corpus |
from datasets import load_dataset
ds = load_dataset("cloud0day3/alania-domain-text-tr", split="train")
print(ds[0]["text_raw"], "->", ds[0]["text_model"])
text_model keeps normaliser artefacts, mostly all-caps brand names followed by a suffix
(e.g. "FAST’ın" → "fe a se te'ın").English · Türkçe
Bir sesli asistanın gerçekte söylediği türden 1.439.639 benzersiz Türkçe cümle: randevu ve klinik konuşmaları, bankacılık, e-ticaret ve kargo, müşteri desteği, onaylar, sorular, sayılar ve kodlar, adresler ve isimler, empati içeren cümleler ve daha uzun açıklamalar. Sesli okunduğunda yaklaşık 1.819 saatlik konuşmaya karşılık geliyor.
Bu veri setini, speech.patientdesk.ai arkasındaki Türkçe metinden sese modelimiz Alania-2'nin sentetik eğitim konuşmalarının metin tarafı olarak hazırladık. Açık Türkçe konuşma ve metin derlemlerinde bu gündelik hizmet dili neredeyse hiç yok; oysa konuşma sistemlerinin en çok zorlandığı yer tam da burası: tarihler, saatler, tutarlar, telefon numaraları, kodlar ve isimler. Türkçe sesli ürün geliştiren başka ekipler de kullanabilsin diye yayımlıyoruz.
Creative Commons Atıf 4.0 (CC BY 4.0) lisansıyla yayımlanmıştır. Ticari kullanım dahil her amaçla kullanabilir, paylaşabilir ve uyarlayabilirsiniz; yeter ki PatientDesk AI'a atıf yapın, örneğin:
Alania Turkish Domain Text, PatientDesk AI (https://huggingface.co/datasets/cloud0day3/alania-domain-text-tr), CC BY 4.0 lisansıyla.
Akademik çalışmalarda yukarıdaki BibTeX kaydını kullanabilirsiniz.
Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 (Apache-2.0), vLLM üzerinde, her kategori ve alt konu
için ayrı yönergelerle (örneğin appointments_clinics / randevu alma), Eylül 2026.text_model, text_raw cümlesinin Türkçe metin normalleştiricimizden geçmiş hâlidir. Sayılar,
tarihler, saatler, tutarlar, telefon numaraları, kısaltmalar ve web adresleri söylendiği gibi yazıya dökülür;
büyük-küçük harf korunur.Kategori dağılımı yukarıdaki tablodadır. Uzunluk grupları, saniyede 14,5 karakterlik tahmini okuma hızına
göredir: xs 1–3 sn (411.466), s 3–6 sn (785.773), m 6–12 sn (195.968), l 12–25 sn (46.432).
| Sütun | Anlamı |
|---|---|
line_id | satırın kalıcı kimliği |
text_raw | modelin ürettiği hâliyle cümle |
text_model | bir TTS modelinin okuduğu normalleştirilmiş hâl |
category, tags | kategori ve satır etiketleri (ör. question) |
bucket, est_dur, n_chars, n_sent | uzunluk bilgileri |
lid_tr | Türkçe dil tespiti skoru |
extra | JSON: alt konu, üretim partisi, üretici model |
source, doc_id, unit_idx, license | derlemimizin geri kalanıyla ortak kaynak bilgileri |
text_model sütununun küçük bir kısmında normalleştirici kaynaklı hatalar vardır; çoğunlukla ek almış, tamamı
büyük harfli marka adları (ör. "FAST’ın" → "fe a se te'ın").English · Türkçe
1,439,639 unique Turkish sentences of the kind a voice agent actually says: appointment and clinic dialogue, banking, e-commerce and cargo, customer support, confirmations, questions, numbers and codes, addresses and names, empathetic lines and longer explanations. About 1,819 hours of speech if read aloud.
We built it as the text side of synthetic training speech for Alania-2, the Turkish text-to-speech model behind speech.patientdesk.ai. Public Turkish speech and text corpora contain very little of this everyday service language, and it is exactly where speech systems stumble: dates, times, amounts, phone numbers, codes and names. We are releasing it so others building Turkish voice products can use it too.
Released under Creative Commons Attribution 4.0 (CC BY 4.0). You may use, share and adapt it for any purpose, including commercial use, as long as you credit PatientDesk AI, for example:
Alania Turkish Domain Text by PatientDesk AI (https://huggingface.co/datasets/cloud0day3/alania-domain-text-tr), licensed under CC BY 4.0.
In a paper, please cite:
@misc{patientdesk2026alaniadomaintext,
title = {Alania Turkish Domain Text},
author = {{PatientDesk AI}},
year = {2026},
howpublished = {\url{https://huggingface.co/datasets/cloud0day3/alania-domain-text-tr}},
note = {1.44M Turkish voice-agent sentences, CC BY 4.0}
}
Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 (Apache-2.0) on vLLM, prompted per category and subtopic
(for example appointments_clinics / randevu alma), September 2026.text_model is text_raw after our Turkish text normaliser: numbers, dates, times, money,
phone numbers, abbreviations and web addresses written out the way they are spoken, letter case kept.| Category | Lines |
|---|---|
| appointments_clinics | 228,337 |
| numbers_codes | 186,700 |
| customer_support | 158,940 |
| questions | 155,162 |
| addresses_names | 151,247 |
| ecommerce_cargo | 147,041 |
| banking | 145,372 |
| emotional_empathetic | 126,412 |
| confirmations_acks | 103,403 |
| long_explanations | 37,025 |
Length buckets by estimated speaking time at 14.5 characters per second: xs 1–3 s (411,466), s 3–6 s
(785,773), m 6–12 s (195,968), l 12–25 s (46,432).
| Column | Meaning |
|---|---|
line_id | stable id of the line |
text_raw | the sentence as generated |
text_model | the normalised form a TTS model reads |
category, tags | category and line tags (e.g. question) |
bucket, est_dur, n_chars, n_sent | length information |
lid_tr | Turkish language-ID score |
extra | JSON: subtopic, generation batch, generator model |
source, doc_id, unit_idx, license | provenance fields shared with the rest of our text corpus |
from datasets import load_dataset
ds = load_dataset("cloud0day3/alania-domain-text-tr", split="train")
print(ds[0]["text_raw"], "->", ds[0]["text_model"])
text_model keeps normaliser artefacts, mostly all-caps brand names followed by a suffix
(e.g. "FAST’ın" → "fe a se te'ın").English · Türkçe
Bir sesli asistanın gerçekte söylediği türden 1.439.639 benzersiz Türkçe cümle: randevu ve klinik konuşmaları, bankacılık, e-ticaret ve kargo, müşteri desteği, onaylar, sorular, sayılar ve kodlar, adresler ve isimler, empati içeren cümleler ve daha uzun açıklamalar. Sesli okunduğunda yaklaşık 1.819 saatlik konuşmaya karşılık geliyor.
Bu veri setini, speech.patientdesk.ai arkasındaki Türkçe metinden sese modelimiz Alania-2'nin sentetik eğitim konuşmalarının metin tarafı olarak hazırladık. Açık Türkçe konuşma ve metin derlemlerinde bu gündelik hizmet dili neredeyse hiç yok; oysa konuşma sistemlerinin en çok zorlandığı yer tam da burası: tarihler, saatler, tutarlar, telefon numaraları, kodlar ve isimler. Türkçe sesli ürün geliştiren başka ekipler de kullanabilsin diye yayımlıyoruz.
Creative Commons Atıf 4.0 (CC BY 4.0) lisansıyla yayımlanmıştır. Ticari kullanım dahil her amaçla kullanabilir, paylaşabilir ve uyarlayabilirsiniz; yeter ki PatientDesk AI'a atıf yapın, örneğin:
Alania Turkish Domain Text, PatientDesk AI (https://huggingface.co/datasets/cloud0day3/alania-domain-text-tr), CC BY 4.0 lisansıyla.
Akademik çalışmalarda yukarıdaki BibTeX kaydını kullanabilirsiniz.
Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 (Apache-2.0), vLLM üzerinde, her kategori ve alt konu
için ayrı yönergelerle (örneğin appointments_clinics / randevu alma), Eylül 2026.text_model, text_raw cümlesinin Türkçe metin normalleştiricimizden geçmiş hâlidir. Sayılar,
tarihler, saatler, tutarlar, telefon numaraları, kısaltmalar ve web adresleri söylendiği gibi yazıya dökülür;
büyük-küçük harf korunur.Kategori dağılımı yukarıdaki tablodadır. Uzunluk grupları, saniyede 14,5 karakterlik tahmini okuma hızına
göredir: xs 1–3 sn (411.466), s 3–6 sn (785.773), m 6–12 sn (195.968), l 12–25 sn (46.432).
| Sütun | Anlamı |
|---|---|
line_id | satırın kalıcı kimliği |
text_raw | modelin ürettiği hâliyle cümle |
text_model | bir TTS modelinin okuduğu normalleştirilmiş hâl |
category, tags | kategori ve satır etiketleri (ör. question) |
bucket, est_dur, n_chars, n_sent | uzunluk bilgileri |
lid_tr | Türkçe dil tespiti skoru |
extra | JSON: alt konu, üretim partisi, üretici model |
source, doc_id, unit_idx, license | derlemimizin geri kalanıyla ortak kaynak bilgileri |
text_model sütununun küçük bir kısmında normalleştirici kaynaklı hatalar vardır; çoğunlukla ek almış, tamamı
büyük harfli marka adları (ör. "FAST’ın" → "fe a se te'ın").