CoolFace
Datasetpublic

Rabe3/egyptian-arabic-tts-diacritized

Egyptian Arabic TTS Corpus (Diacritized) 97,163 utterances / ~334 hours of Egyptian Arabic speech at 24 kHz, with diacritized transcripts — the short vowels that Arabic script does not write. Why diacritics Arabic is an abjad: short vowels are unwritten, so كتب may be kataba, kutiba, or kutub. A TTS model with no Arabic pretraining cannot infer which, and guesses — which native listeners hear as a foreign accent with constant mispronunciation. This is invisible to… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/egyptian-arabic-tts-diacritized.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes148downloads

Rabe3/egyptian-arabic-tts-diacritized · main · files are served by the source, never re-hosted here