CoolFace
Datasetpublic

mohammedaly22/lahgtna-levantine-tts

๐ŸŽ™๏ธ Lahgtna Levantine TTS Synthetic Levantine Arabic + English Code-Switching speech dataset. Generated using Lahgtna-OmniVoice, a fine-tuned zero-shot TTS model for Levantine Arabic dialect. ๐Ÿ“Š Dataset Statistics Metric Value Total utterances 50,000 Total speakers 10 (5 male, 5 female) Pure Levantine Arabic 44,154 utterances Code-switching (AR+EN) 5,846 utterances Sampling rate 24,000 Hz Estimated total duration ~66.8 hoursโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/mohammedaly22/lahgtna-levantine-tts.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
4likes465downloads
Dataset Card

๐ŸŽ™๏ธ Lahgtna Levantine TTS

Synthetic Levantine Arabic + English Code-Switching speech dataset. Generated using Lahgtna-OmniVoice, a fine-tuned zero-shot TTS model for Levantine Arabic dialect.


๐Ÿ“Š Dataset Statistics

MetricValue
Total utterances50,000
Total speakers10 (5 male, 5 female)
Pure Levantine Arabic44,154 utterances
Code-switching (AR+EN)5,846 utterances
Sampling rate24,000 Hz
Estimated total duration~66.8 hours

Per-Speaker Breakdown

Speaker IDNameGenderUtterancesPure ARCode-SwitchEst. Hours
spk01maleBadrmale5,0004,4295716.86h
spk02maleMohamedmale5,0004,4335676.71h
spk03maleSaadmale5,0004,4195816.28h
spk04maleRamimale5,0004,3936077.05h
spk05maleFadimale5,0004,4285726.42h
spk06femaleAminafemale5,0004,3796215.97h
spk07femaleFatmafemale5,0004,4205805.88h
spk08femaleLamyaafemale5,0004,4265747.36h
spk09femaleMonafemale5,0004,4285727.53h
spk10femaleHaneenfemale5,0004,3996016.71h

๐Ÿ“ Data Collection & Processing

1. Text Data Sources

The 50,000 sentences were collected from:

SourceTypeCount
GU-CLASP Shami CorpusReal Levantine Arabic (Syrian, Lebanese, Palestinian, Jordanian)~44,000
Synthetic code-switching templatesLevantine Arabic + English (tech/daily life)~6,000

The Shami corpus provides authentic dialectal text from four Levantine sub-dialects:

  • โ€”Syrian (syrian.txt) โ€” 34,491 sentences
  • โ€”Lebanese (Lebenees.txt) โ€” 9,905 sentences
  • โ€”Palestinian (Palestinian.txt) โ€” 9,545 sentences
  • โ€”Jordanian (jordinian.txt) โ€” 6,007 sentences

Code-switching sentences follow natural Levantine-English mixing patterns:

ู‡ูŽู„ูŽู‘ู‚ ุนู… ุฃุดุชุบู„ ุนู„ู‰ the project ุงู„ู„ูŠ ุญูƒูŠุชู„ูƒ ุนู†ู‡
ูˆุงู„ู„ู‡ the meeting ูƒุชูŠุฑ importantุŒ ู„ุงุฒู… ู†ุญุถู‘ุฑ ู…ูู†ููŠุญ

2. Text Normalization & Partial Diacritization

Before synthesis, each sentence was processed through:

Step 1 โ€” Unicode cleanup: NFC normalization, tatweel removal, alef unification

Step 2 โ€” Number verbalization: Levantine Arabic number words

  • โ€”3 ูƒุชุจ โ†’ ุชู„ุงุชุฉ ูƒุชุจ
  • โ€”$50 โ†’ ุฎู…ุณูŠู† ุฏูˆู„ุงุฑ

Step 3 โ€” Partial diacritization on homographs only: The key design decision: instead of full diacritization, we apply diacritics only to ambiguous homographs that could be mispronounced. This makes the model robust to both diacritized and undiacritized input at inference time.

Diacritized homograph examples:

ู‡ูŽู„ูŽู‘ู‚  (now โ€” vs ู‡ูŽู„ูŽู‚ูŽ = he shaved, MSA)
ุถูŽู„ู‘   (remained, Levantine โ€” vs ุถูŽู„ูŽู‘ = went astray, MSA)
ู…ูุดู’   (not, Levantine negation)
ุจูุฏูู‘ูŠ (I want, Levantine bi-imperfect)

Step 4 โ€” ู‡ โ†’ ุฉ correction: Levantine Arabic informal writing uses ู‡ where standard orthography uses ุฉ (ta marbuta). A comprehensive rule-based corrector fixes feminine nouns, adjectives, and proper names while preserving genuine ู‡ in verb+pronoun forms and ุงู„ู„ู‡ compounds:

  • โ€”ู‡ุงู„ุถุญูƒู‡ ุงู„ุญู„ูˆู‡ โ†’ ู‡ุงู„ุถุญูƒุฉ ุงู„ุญู„ูˆุฉ โœ…
  • โ€”ูˆุงู„ู„ู‡ โ†’ ูˆุงู„ู„ู‡ (preserved โ€” contains ุงู„ู„ู‡) โœ…
  • โ€”ููŠู‡ุŒ ุนู„ูŠู‡ุŒ ู…ุนู‡ โ†’ preserved (pronoun suffixes) โœ…

Step 5 โ€” Levantine lexicon overrides (148 entries in CSV): Common Levantine dialect words get dialect-correct diacritization via an editable CSV file (data/levantine_lexicon.csv) โ€” no code changes needed to add new words.

3. TTS Synthesis โ€” Lahgtna-OmniVoice

PropertyValue
Model`oddadmix/lahgtna-omnivoice-v2`
Base architectureOmniVoice (k2-fsa/OmniVoice fine-tune)
Fine-tuningLevantine Arabic dialect (apc โ€” ISO 639-3)
Generation modeZero-shot voice cloning from reference audio
Language codeapc (North Levantine Arabic)
Output sample rate24,000 Hz
Generation parameterstemperature=0.7, topp=0.7, repetitionpenalty=1.2

Each speaker was cloned from a 5โ€“15 s reference recording of a real Levantine speaker. The 10 speakers were generated in parallel across 4ร— NVIDIA H100 GPUs using Python multiprocessing, with each GPU handling 2โ€“3 speakers simultaneously.


๐Ÿ“ Dataset Structure

train/
  audio       โ€” Audio feature at 24 kHz
  text        โ€” Levantine Arabic transcript (partial diacritics on homographs)
  speaker_id  โ€” e.g. "spk_01_male"
  speaker_nameโ€” e.g. "Badr"
  gender      โ€” "male" | "female"
  sentence_type โ€” "pure_levantine" | "code_switching"

๐Ÿ”ง Usage

python
from datasets import load_dataset

ds = load_dataset("mohammedaly22/lahgtna-levantine-tts", split="train")

# Play sample
sample = ds[0]
print(sample["text"])          # transcript
print(sample["speaker_name"]) # e.g. "Badr"
print(sample["sentence_type"]) # "pure_levantine" or "code_switching"

# Audio: sample["audio"]["array"] at 24000 Hz

๐Ÿ“œ License

CC BY 4.0 โ€” Free to use with attribution.

๐Ÿ‘ค Author

![HuggingFace](https://huggingface.co/mohammedaly22)

๐Ÿ”— Related