levantine
lahgtna-levantine-tts
🎙️ Lahgtna Levantine TTS
Synthetic Levantine Arabic + English Code-Switching speech dataset.
Generated using Lahgtna-OmniVoice,
a fine-tuned zero-shot TTS model for Levantine Arabic dialect.
📊 Dataset Statistics
Metric
Value
Total utterances
50,000
Total speakers
10 (5 male, 5 female)
Pure Levantine Arabic
44,154 utterances
Code-switching (AR+EN)
5,846 utterances
Sampling rate
24,000 Hz
Estimated total duration
~66.8 hours… See the full description on the dataset page: https://huggingface.co/datasets/mohammedaly22/lahgtna-levantine-tts.UFAL_Parallel_Corpus_of_North_Levantine_1.0
[!NOTE]
Dataset origin: https://zenodo.org/records/4012218
UFAL Parallel Corpus of North Levantine 1.0
March 10, 2023
Authors
Shadi Saleh <saleh@ufal.mff.cuni.cz>
Hashem Sellat <sellat@ufal.mff.cuni.cz>
Mateusz Krubiński <krubinski@ufal.mff.cuni.cz>
Adam Posppíšil <adam.pospisil@ff.cuni.cz>
Petr Zemánek <petr.zemanek@ff.cuni.cz>
Pavel Pecina <pecina@ufal.mff.cuni.cz>
Overview
This is the first release of the UFAL Parallel Corpus of North Levantine… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/UFAL_Parallel_Corpus_of_North_Levantine_1.0.levantine_train_setPre-processed Levantine train partition from the MASC-dataset:
Mohammad Al-Fetyani, Muhammad Al-Barham, Gheith Abandah, Adham Alsharkawi, Maha Dawas, August 18, 2021, "MASC: Massive Arabic Speech Corpus", IEEE Dataport, doi: https://dx.doi.org/10.21227/e1qb-jv46.
Levantine_RewayatFineWeb2-North-Levantine-Arabic
FineWeb2 North Levantine Arabic
🇱🇧 This is the North Levantine Arabic Portion of The FineWeb2 Dataset.
🇸🇩 The North Levantine Arabic, represented by the ISO 639-3 code apc, is a member of the Afro-Asiatic language family and utilizes the Arabic script.
🇯🇴 Known within subsets as apc_Arab, this language boasts an extensive corpus of over ** 221K rows**.
Purpose of This Repository
This repository provides easy access to the Arabic portion - North Levantineof the… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/FineWeb2-North-Levantine-Arabic.levantine_test_setPre-processed Levantine test partition from the MASC-dataset:
Mohammad Al-Fetyani, Muhammad Al-Barham, Gheith Abandah, Adham Alsharkawi, Maha Dawas, August 18, 2021, "MASC: Massive Arabic Speech Corpus", IEEE Dataport, doi: https://dx.doi.org/10.21227/e1qb-jv46.
