CoolFace
Datasetpublic

Rabe3/egyptian-arabic-tts-diacritized

Egyptian Arabic TTS Corpus (Diacritized) 97,163 utterances / ~334 hours of Egyptian Arabic speech at 24 kHz, with diacritized transcripts — the short vowels that Arabic script does not write. Why diacritics Arabic is an abjad: short vowels are unwritten, so كتب may be kataba, kutiba, or kutub. A TTS model with no Arabic pretraining cannot infer which, and guesses — which native listeners hear as a foreign accent with constant mispronunciation. This is invisible to… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/egyptian-arabic-tts-diacritized.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes148downloads
settings

This repository belongs to Rabe3 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameegyptian-arabic-tts-diacritized
visibilitypublic
licenceother
gatedno
ownerRabe3
Account settings
Rabe3/egyptian-arabic-tts-diacritized · CoolFace