cloud0day3/alania-domain-text-tr

Dataset

Alania Turkish Domain Text

11

2 commits

updated Sep 28, 2026

See the code

README

Alania Turkish Domain Text

English · Türkçe

1,439,639 unique Turkish sentences of the kind a voice agent actually says: appointment and clinic dialogue, banking, e-commerce and cargo, customer support, confirmations, questions, numbers and codes, addresses and names, empathetic lines and longer explanations. About 1,819 hours of speech if read aloud.

We built it as the text side of synthetic training speech for Alania-2, the Turkish text-to-speech model behind speech.patientdesk.ai. Public Turkish speech and text corpora contain very little of this everyday service language, and it is exactly where speech systems stumble: dates, times, amounts, phone numbers, codes and names. We are releasing it so others building Turkish voice products can use it too.

Licence and attribution

Released under Creative Commons Attribution 4.0 (CC BY 4.0). You may use, share and adapt it for any purpose, including commercial use, as long as you credit PatientDesk AI, for example:

Alania Turkish Domain Text by PatientDesk AI (https://huggingface.co/datasets/cloud0day3/alania-domain-text-tr), licensed under CC BY 4.0.

In a paper, please cite:

@misc{patientdesk2026alaniadomaintext,
  title        = {Alania Turkish Domain Text},
  author       = {{PatientDesk AI}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/datasets/cloud0day3/alania-domain-text-tr}},
  note         = {1.44M Turkish voice-agent sentences, CC BY 4.0}
}

How it was made

  • Generator: Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 (Apache-2.0) on vLLM, prompted per category and subtopic (for example appointments_clinics / randevu alma), September 2026.
  • Cleaning: sentence splitting, Turkish language ID, length and character filters, exact and near-duplicate removal. Every line is unique.
  • Model text: text_model is text_raw after our Turkish text normaliser: numbers, dates, times, money, phone numbers, abbreviations and web addresses written out the way they are spoken, letter case kept.

Contents

CategoryLines
appointments_clinics228,337
numbers_codes186,700
customer_support158,940
questions155,162
addresses_names151,247
ecommerce_cargo147,041
banking145,372
emotional_empathetic126,412
confirmations_acks103,403
long_explanations37,025

Length buckets by estimated speaking time at 14.5 characters per second: xs 1–3 s (411,466), s 3–6 s (785,773), m 6–12 s (195,968), l 12–25 s (46,432).

Columns

ColumnMeaning
line_idstable id of the line
text_rawthe sentence as generated
text_modelthe normalised form a TTS model reads
category, tagscategory and line tags (e.g. question)
bucket, est_dur, n_chars, n_sentlength information
lid_trTurkish language-ID score
extraJSON: subtopic, generation batch, generator model
source, doc_id, unit_idx, licenseprovenance fields shared with the rest of our text corpus
from datasets import load_dataset
ds = load_dataset("cloud0day3/alania-domain-text-tr", split="train")
print(ds[0]["text_raw"], "->", ds[0]["text_model"])

Known issues

  • The text is machine-written. It is fluent and on topic, but it was not reviewed line by line: some lines are bland or slightly unnatural, and statements about banking or medicine in it are not advice and must not be relied on.
  • Names, addresses, phone numbers and codes are invented; any match with a real person, place or account is coincidental.
  • A small share of text_model keeps normaliser artefacts, mostly all-caps brand names followed by a suffix (e.g. "FAST’ın" → "fe a se te'ın").

Türkçe

English · Türkçe

Bir sesli asistanın gerçekte söylediği türden 1.439.639 benzersiz Türkçe cümle: randevu ve klinik konuşmaları, bankacılık, e-ticaret ve kargo, müşteri desteği, onaylar, sorular, sayılar ve kodlar, adresler ve isimler, empati içeren cümleler ve daha uzun açıklamalar. Sesli okunduğunda yaklaşık 1.819 saatlik konuşmaya karşılık geliyor.

Bu veri setini, speech.patientdesk.ai arkasındaki Türkçe metinden sese modelimiz Alania-2'nin sentetik eğitim konuşmalarının metin tarafı olarak hazırladık. Açık Türkçe konuşma ve metin derlemlerinde bu gündelik hizmet dili neredeyse hiç yok; oysa konuşma sistemlerinin en çok zorlandığı yer tam da burası: tarihler, saatler, tutarlar, telefon numaraları, kodlar ve isimler. Türkçe sesli ürün geliştiren başka ekipler de kullanabilsin diye yayımlıyoruz.

Lisans ve atıf

Creative Commons Atıf 4.0 (CC BY 4.0) lisansıyla yayımlanmıştır. Ticari kullanım dahil her amaçla kullanabilir, paylaşabilir ve uyarlayabilirsiniz; yeter ki PatientDesk AI'a atıf yapın, örneğin:

Alania Turkish Domain Text, PatientDesk AI (https://huggingface.co/datasets/cloud0day3/alania-domain-text-tr), CC BY 4.0 lisansıyla.

Akademik çalışmalarda yukarıdaki BibTeX kaydını kullanabilirsiniz.

Nasıl hazırlandı

  • Üretici model: Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 (Apache-2.0), vLLM üzerinde, her kategori ve alt konu için ayrı yönergelerle (örneğin appointments_clinics / randevu alma), Eylül 2026.
  • Temizlik: cümlelere bölme, Türkçe dil tespiti, uzunluk ve karakter filtreleri, birebir ve neredeyse aynı tekrarların ayıklanması. Her satır benzersizdir.
  • Model metni: text_model, text_raw cümlesinin Türkçe metin normalleştiricimizden geçmiş hâlidir. Sayılar, tarihler, saatler, tutarlar, telefon numaraları, kısaltmalar ve web adresleri söylendiği gibi yazıya dökülür; büyük-küçük harf korunur.

İçerik

Kategori dağılımı yukarıdaki tablodadır. Uzunluk grupları, saniyede 14,5 karakterlik tahmini okuma hızına göredir: xs 1–3 sn (411.466), s 3–6 sn (785.773), m 6–12 sn (195.968), l 12–25 sn (46.432).

Sütunlar

SütunAnlamı
line_idsatırın kalıcı kimliği
text_rawmodelin ürettiği hâliyle cümle
text_modelbir TTS modelinin okuduğu normalleştirilmiş hâl
category, tagskategori ve satır etiketleri (ör. question)
bucket, est_dur, n_chars, n_sentuzunluk bilgileri
lid_trTürkçe dil tespiti skoru
extraJSON: alt konu, üretim partisi, üretici model
source, doc_id, unit_idx, licensederlemimizin geri kalanıyla ortak kaynak bilgileri

Bilinen sınırlamalar

  • Metinler makine tarafından yazılmıştır. Akıcı ve konuya uygundur, ancak satır satır incelenmemiştir: bazı cümleler sıradan ya da biraz yapay olabilir. Bankacılık veya sağlık hakkındaki ifadeler tavsiye değildir; bunlara dayanılmamalıdır.
  • İsimler, adresler, telefon numaraları ve kodlar uydurmadır; gerçek bir kişi, yer veya hesapla eşleşmesi tesadüftür.
  • text_model sütununun küçük bir kısmında normalleştirici kaynaklı hatalar vardır; çoğunlukla ek almış, tamamı büyük harfli marka adları (ör. "FAST’ın" → "fe a se te'ın").
customer-service
healthcare
synthetic
turkish
voice-agent

cloud0day3/alania-domain-text-tr

Dataset

Alania Turkish Domain Text

11

2 commits

updated Sep 28, 2026

See the code

README

Alania Turkish Domain Text

English · Türkçe

1,439,639 unique Turkish sentences of the kind a voice agent actually says: appointment and clinic dialogue, banking, e-commerce and cargo, customer support, confirmations, questions, numbers and codes, addresses and names, empathetic lines and longer explanations. About 1,819 hours of speech if read aloud.

We built it as the text side of synthetic training speech for Alania-2, the Turkish text-to-speech model behind speech.patientdesk.ai. Public Turkish speech and text corpora contain very little of this everyday service language, and it is exactly where speech systems stumble: dates, times, amounts, phone numbers, codes and names. We are releasing it so others building Turkish voice products can use it too.

Licence and attribution

Released under Creative Commons Attribution 4.0 (CC BY 4.0). You may use, share and adapt it for any purpose, including commercial use, as long as you credit PatientDesk AI, for example:

Alania Turkish Domain Text by PatientDesk AI (https://huggingface.co/datasets/cloud0day3/alania-domain-text-tr), licensed under CC BY 4.0.

In a paper, please cite:

@misc{patientdesk2026alaniadomaintext,
  title        = {Alania Turkish Domain Text},
  author       = {{PatientDesk AI}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/datasets/cloud0day3/alania-domain-text-tr}},
  note         = {1.44M Turkish voice-agent sentences, CC BY 4.0}
}

How it was made

  • Generator: Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 (Apache-2.0) on vLLM, prompted per category and subtopic (for example appointments_clinics / randevu alma), September 2026.
  • Cleaning: sentence splitting, Turkish language ID, length and character filters, exact and near-duplicate removal. Every line is unique.
  • Model text: text_model is text_raw after our Turkish text normaliser: numbers, dates, times, money, phone numbers, abbreviations and web addresses written out the way they are spoken, letter case kept.

Contents

CategoryLines
appointments_clinics228,337
numbers_codes186,700
customer_support158,940
questions155,162
addresses_names151,247
ecommerce_cargo147,041
banking145,372
emotional_empathetic126,412
confirmations_acks103,403
long_explanations37,025

Length buckets by estimated speaking time at 14.5 characters per second: xs 1–3 s (411,466), s 3–6 s (785,773), m 6–12 s (195,968), l 12–25 s (46,432).

Columns

ColumnMeaning
line_idstable id of the line
text_rawthe sentence as generated
text_modelthe normalised form a TTS model reads
category, tagscategory and line tags (e.g. question)
bucket, est_dur, n_chars, n_sentlength information
lid_trTurkish language-ID score
extraJSON: subtopic, generation batch, generator model
source, doc_id, unit_idx, licenseprovenance fields shared with the rest of our text corpus
from datasets import load_dataset
ds = load_dataset("cloud0day3/alania-domain-text-tr", split="train")
print(ds[0]["text_raw"], "->", ds[0]["text_model"])

Known issues

  • The text is machine-written. It is fluent and on topic, but it was not reviewed line by line: some lines are bland or slightly unnatural, and statements about banking or medicine in it are not advice and must not be relied on.
  • Names, addresses, phone numbers and codes are invented; any match with a real person, place or account is coincidental.
  • A small share of text_model keeps normaliser artefacts, mostly all-caps brand names followed by a suffix (e.g. "FAST’ın" → "fe a se te'ın").

Türkçe

English · Türkçe

Bir sesli asistanın gerçekte söylediği türden 1.439.639 benzersiz Türkçe cümle: randevu ve klinik konuşmaları, bankacılık, e-ticaret ve kargo, müşteri desteği, onaylar, sorular, sayılar ve kodlar, adresler ve isimler, empati içeren cümleler ve daha uzun açıklamalar. Sesli okunduğunda yaklaşık 1.819 saatlik konuşmaya karşılık geliyor.

Bu veri setini, speech.patientdesk.ai arkasındaki Türkçe metinden sese modelimiz Alania-2'nin sentetik eğitim konuşmalarının metin tarafı olarak hazırladık. Açık Türkçe konuşma ve metin derlemlerinde bu gündelik hizmet dili neredeyse hiç yok; oysa konuşma sistemlerinin en çok zorlandığı yer tam da burası: tarihler, saatler, tutarlar, telefon numaraları, kodlar ve isimler. Türkçe sesli ürün geliştiren başka ekipler de kullanabilsin diye yayımlıyoruz.

Lisans ve atıf

Creative Commons Atıf 4.0 (CC BY 4.0) lisansıyla yayımlanmıştır. Ticari kullanım dahil her amaçla kullanabilir, paylaşabilir ve uyarlayabilirsiniz; yeter ki PatientDesk AI'a atıf yapın, örneğin:

Alania Turkish Domain Text, PatientDesk AI (https://huggingface.co/datasets/cloud0day3/alania-domain-text-tr), CC BY 4.0 lisansıyla.

Akademik çalışmalarda yukarıdaki BibTeX kaydını kullanabilirsiniz.

Nasıl hazırlandı

  • Üretici model: Qwen/Qwen3-30B-A3B-Instruct-2507-FP8 (Apache-2.0), vLLM üzerinde, her kategori ve alt konu için ayrı yönergelerle (örneğin appointments_clinics / randevu alma), Eylül 2026.
  • Temizlik: cümlelere bölme, Türkçe dil tespiti, uzunluk ve karakter filtreleri, birebir ve neredeyse aynı tekrarların ayıklanması. Her satır benzersizdir.
  • Model metni: text_model, text_raw cümlesinin Türkçe metin normalleştiricimizden geçmiş hâlidir. Sayılar, tarihler, saatler, tutarlar, telefon numaraları, kısaltmalar ve web adresleri söylendiği gibi yazıya dökülür; büyük-küçük harf korunur.

İçerik

Kategori dağılımı yukarıdaki tablodadır. Uzunluk grupları, saniyede 14,5 karakterlik tahmini okuma hızına göredir: xs 1–3 sn (411.466), s 3–6 sn (785.773), m 6–12 sn (195.968), l 12–25 sn (46.432).

Sütunlar

SütunAnlamı
line_idsatırın kalıcı kimliği
text_rawmodelin ürettiği hâliyle cümle
text_modelbir TTS modelinin okuduğu normalleştirilmiş hâl
category, tagskategori ve satır etiketleri (ör. question)
bucket, est_dur, n_chars, n_sentuzunluk bilgileri
lid_trTürkçe dil tespiti skoru
extraJSON: alt konu, üretim partisi, üretici model
source, doc_id, unit_idx, licensederlemimizin geri kalanıyla ortak kaynak bilgileri

Bilinen sınırlamalar

  • Metinler makine tarafından yazılmıştır. Akıcı ve konuya uygundur, ancak satır satır incelenmemiştir: bazı cümleler sıradan ya da biraz yapay olabilir. Bankacılık veya sağlık hakkındaki ifadeler tavsiye değildir; bunlara dayanılmamalıdır.
  • İsimler, adresler, telefon numaraları ve kodlar uydurmadır; gerçek bir kişi, yer veya hesapla eşleşmesi tesadüftür.
  • text_model sütununun küçük bir kısmında normalleştirici kaynaklı hatalar vardır; çoğunlukla ek almış, tamamı büyük harfli marka adları (ör. "FAST’ın" → "fe a se te'ın").
customer-service
healthcare
synthetic
turkish
voice-agent