AymanMansour/New-Lisan-Sudanese-TTS-Dataset
Lisan Sudanese TTS Dataset A synthetic Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) dataset specifically for Sudanese Arabic. 1,878 high-quality sentences featuring 20 synthetic speakers (10 male, 10 female). Reconstructed from the Lisan-Sudanese Morphological Dataset (52K manually annotated social media tokens from Facebook/X). Only sentences with a diacritic density of >=25% were kept to ensure enough phonetic information for accurate synthesis. model: Resemble AI… See the full description on the dataset page: https://huggingface.co/datasets/AymanMansour/New-Lisan-Sudanese-TTS-Dataset.
032
- Lisan Sudanese TTS Dataset
- A synthetic Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) dataset specifically for Sudanese Arabic.
- 1,878 high-quality sentences featuring 20 synthetic speakers (10 male, 10 female).
- Reconstructed from the Lisan-Sudanese Morphological Dataset (52K manually annotated social media tokens from Facebook/X).
- Only sentences with a diacritic density of >=25% were kept to ensure enough phonetic information for accurate synthesis.
- model: Resemble AI Chatterbox Multilingual ResembleAI/Chatterbox-Multilingual-TTS
- Accessed via the Resemble AI Python SDK using an API key, selecting a target voice ID, and passing highly diacritized text payloads to generate .wav files.
- observed_errors:
- MSA and Egyptian dialect Bias(Phonetic Mismatches) description: The model defaults to Modern Standard Arabic (MSA) pronunciations, and in several actuations to Egyptian dialect struggling with distinct Sudanese letter realizations (e.g., Qaf or Jeem).
- type: Prosody Issues The model fails to capture the natural rhythm and intonation of informal Sudanese social media slang and code-switched phrases.
- finetuningrecommendations:
- proposed_dataset: A High-Fidelity Human TTS Corpus featuring phonetically balanced scripts that highlight unique Sudanese phonemes and slang.
- assembly_method: Hire 2 to 4 professional, native Sudanese voice actors to read scripts (derived from the Lisan dataset) in a noise-free, acoustically treated studio. The text and audio must then be perfectly aligned using forced alignment tools.
- size_requirements: minimum: 2 to 5 hours of ultra-clean, single-speaker audio to create a highly accurate custom voice clone. robust: 15 to 25 hours across multiple speakers to robustly fine-tune a foundation model's base understanding of the Sudanese dialect.
