ai-ssam/darija-tts-8400
Darija TTS 8400 Synthetic Moroccan Darija speech for TTS fine-tuning: 8,400 single-speaker clips (20.73 hours), 24 kHz mono PCM16 WAV. All audio is generated with Gemini 3.1 Flash TTS (gemini-3.1-flash-tts-preview, voice Kore). Clips are unreviewed; there are no human recordings. Write-up of how this data was used: Training a Voice. At a glance Clips / hours 8,400 / 20.73 Unique texts 4,800 Voice Kore (1 speaker) Sample rate 24 kHz mono PCM16… See the full description on the dataset page: https://huggingface.co/datasets/ai-ssam/darija-tts-8400.
Darija TTS 8400
Synthetic Moroccan Darija speech for TTS fine-tuning: 8,400 single-speaker clips (20.73 hours), 24 kHz mono PCM16 WAV.
All audio is generated with Gemini 3.1 Flash TTS (gemini-3.1-flash-tts-preview, voice Kore). Clips are unreviewed; there are no human recordings.
Write-up of how this data was used: Training a Voice.
At a glance
Composition
Two phases, same speaker and TTS recipe:
Emotions (requested_emotion): neutral (4,800), reassuring (1,200), excited (1,200), frustrated (1,200).
Same sentence across styles shares family_id / group_id and stays in the same split.
Text sources
Human DODa lines: Darija Open Dataset. Audio for those lines is still Gemini TTS.
Categories: darija (4,204), darija_french (2,342), pronunciation (1,374), darija_english (480). French/English spans stay in Latin script inside Arabic-script Darija.
How it was generated
- Plan texts. Expressive lines: Gemini 3.1 Flash Lite, Arabic-script Darija, ~15–28 words, readable in four emotions. Neutral lines: stratified sample from a DODa-derived parallel corpus (length, code-switch, pronunciation, domain caps), plus a small Gemini gap-fill.
- Synthesize. Each clip: Kore voice, Moroccan-accent prompt, style instruction for that emotion, audio wrapped as 24 kHz mono WAV. Mostly Batch API, remainder sync. WAVs are SHA-256 hashed; clips outside 3–25 s are flagged
unusual_duration. - Split. Hash of normalized text → ~90/5/5 train/validation/test, so multi-emotion siblings never cross splits.
Files
audio paths are relative (audio/f000000_neutral.wav).
Fields
Training-critical: id, group_id, family_id, text, audio, voice, requested_emotion, duration_seconds, split.
Also: dataset_phase, category, text_source, tts_model, sha256, flags, and for parallel rows source_id, domain, code_switch, generation_mode, …
License
- Audio / Gemini text: Gemini API terms
- Human DODa texts: CC BY-NC 4.0 via DODa
