Turkish text normalization for text-to-speech: standalone Rust library and Python binding.
HTML
16
8 commits
updated Oct 3, 2026
Turkish text normalization for speech, with a Rust core and a typed Python API. Convert numbers, money, dates, measurements and other supported notation into spoken Turkish while preserving ordinary prose. Runs offline, without a speech model or service. Decimal and money arithmetic is exact.
saat 09:30'da 5 kg malzeme ve %12,5'lik fark
→ saat dokuz otuzda beş kilogram malzeme ve yüzde on iki virgül beşlik fark
python -m pip install normalizer-tr
from normalizer_tr import Normalizer
normalizer = Normalizer()
result = normalizer.normalize("25 TL; 5 kg")
assert result.normalized_text == "yirmi beş Türk lirası; beş kilogram"
assert result.complete
Wheels support CPython 3.11 to 3.14 on Windows and Linux x64, and macOS x64 and arm64. Compatible wheels need no Rust compiler. Platform requirements and source builds.
Requires Rust 1.94 or newer. Add to Cargo.toml:
[dependencies]
normalizer-tr = "0.4"
use normalizer_tr::{NormalizeOptions, Normalizer};
let normalizer = Normalizer::new()?;
let result = normalizer.normalize("25 TL; 5 kg", &NormalizeOptions::default())?;
assert_eq!(result.normalized_text(), "yirmi beş Türk lirası; beş kilogram");
assert!(result.complete());
# Ok::<(), normalizer_tr::NormalizeError>(())
Reuse a Normalizer across calls. It is cloneable and Send + Sync.
Enable the optional serde feature to serialize results.
| Policy | Behavior |
|---|---|
preserve (default) | Keep unresolved spans as written. Return complete=false and issues. |
reject | Return an error if any span is unresolved. |
fallback | Render unresolved notation and symbols. Return handled assumptions in fallbacks. |
result = normalizer.normalize("1.234; AB12", ambiguity_policy="fallback")
assert result.normalized_text == "bin iki yüz otuz dört; a be bir iki"
assert result.complete and not result.issues
assert result.fallback_used
In Rust, set NormalizeOptions.ambiguity_policy to
AmbiguityPolicy::Fallback or AmbiguityPolicy::Reject.
Fallback prefers clear formats, then literal readings and named symbols, then
spoken Unicode codes. It does not correct invalid facts or certify identifiers.
Check fallbacks when those assumptions matter. Invalid input or hints,
cancellation and resource limits remain errors in every policy.
Hints provide explicit intent. All hint and diagnostic ranges are original UTF-8 byte offsets, not character positions.
| Guide | Contents |
|---|---|
| Rust API | Types, methods and examples |
| Python API | Hints, errors, cancellation and building |
| Normalization reference | Supported formats and boundaries |
| Fallback | Reading strategies and diagnostics |
| Contributing | Setup, architecture and verification |
This is a pre 1.0 library with bounded coverage, not a universal pronunciation engine. See performance for measurements and release notes for changes.
Apache-2.0, with third party notices.
Thanks to @canberk7 for his contribution to fallback support.
Turkish text normalization for text-to-speech: standalone Rust library and Python binding.
HTML
16
8 commits
updated Oct 3, 2026
Turkish text normalization for speech, with a Rust core and a typed Python API. Convert numbers, money, dates, measurements and other supported notation into spoken Turkish while preserving ordinary prose. Runs offline, without a speech model or service. Decimal and money arithmetic is exact.
saat 09:30'da 5 kg malzeme ve %12,5'lik fark
→ saat dokuz otuzda beş kilogram malzeme ve yüzde on iki virgül beşlik fark
python -m pip install normalizer-tr
from normalizer_tr import Normalizer
normalizer = Normalizer()
result = normalizer.normalize("25 TL; 5 kg")
assert result.normalized_text == "yirmi beş Türk lirası; beş kilogram"
assert result.complete
Wheels support CPython 3.11 to 3.14 on Windows and Linux x64, and macOS x64 and arm64. Compatible wheels need no Rust compiler. Platform requirements and source builds.
Requires Rust 1.94 or newer. Add to Cargo.toml:
[dependencies]
normalizer-tr = "0.4"
use normalizer_tr::{NormalizeOptions, Normalizer};
let normalizer = Normalizer::new()?;
let result = normalizer.normalize("25 TL; 5 kg", &NormalizeOptions::default())?;
assert_eq!(result.normalized_text(), "yirmi beş Türk lirası; beş kilogram");
assert!(result.complete());
# Ok::<(), normalizer_tr::NormalizeError>(())
Reuse a Normalizer across calls. It is cloneable and Send + Sync.
Enable the optional serde feature to serialize results.
| Policy | Behavior |
|---|---|
preserve (default) | Keep unresolved spans as written. Return complete=false and issues. |
reject | Return an error if any span is unresolved. |
fallback | Render unresolved notation and symbols. Return handled assumptions in fallbacks. |
result = normalizer.normalize("1.234; AB12", ambiguity_policy="fallback")
assert result.normalized_text == "bin iki yüz otuz dört; a be bir iki"
assert result.complete and not result.issues
assert result.fallback_used
In Rust, set NormalizeOptions.ambiguity_policy to
AmbiguityPolicy::Fallback or AmbiguityPolicy::Reject.
Fallback prefers clear formats, then literal readings and named symbols, then
spoken Unicode codes. It does not correct invalid facts or certify identifiers.
Check fallbacks when those assumptions matter. Invalid input or hints,
cancellation and resource limits remain errors in every policy.
Hints provide explicit intent. All hint and diagnostic ranges are original UTF-8 byte offsets, not character positions.
| Guide | Contents |
|---|---|
| Rust API | Types, methods and examples |
| Python API | Hints, errors, cancellation and building |
| Normalization reference | Supported formats and boundaries |
| Fallback | Reading strategies and diagnostics |
| Contributing | Setup, architecture and verification |
This is a pre 1.0 library with bounded coverage, not a universal pronunciation engine. See performance for measurements and release notes for changes.
Apache-2.0, with third party notices.
Thanks to @canberk7 for his contribution to fallback support.