Dataset for phonikud model
The datasets contains millions of clean Hebrew sentences marked with nikud and additional phonetics marks.
The format is text<TAB>phonemes
Oto such as Otobus or Otomatiלמנרי)Sourced Raw Text
Downloaded Hebrew parliamentary tweets from the IsraParlTweet dataset (CC-BY-4.0).
Normalized & Filtered
\n '!,.?אבגדהוזחטיךכלםמןנסעףפץצקרשת Added Niqqud & Metadata
Syllabification & Annotation
Filtered Long Sentences
Shva Handling
Built Word Frequency Lexicon
Manual Corrections
Dataset for phonikud model
The datasets contains millions of clean Hebrew sentences marked with nikud and additional phonetics marks.
The format is text<TAB>phonemes
Oto such as Otobus or Otomatiלמנרי)Sourced Raw Text
Downloaded Hebrew parliamentary tweets from the IsraParlTweet dataset (CC-BY-4.0).
Normalized & Filtered
\n '!,.?אבגדהוזחטיךכלםמןנסעףפץצקרשת Added Niqqud & Metadata
Syllabification & Annotation
Filtered Long Sentences
Shva Handling
Built Word Frequency Lexicon
Manual Corrections