datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lahgtna-levantine-tts
🎙️ Lahgtna Levantine TTS
Synthetic Levantine Arabic + English Code-Switching speech dataset.
Generated using Lahgtna-OmniVoice,
a fine-tuned zero-shot TTS model for Levantine Arabic dialect.
📊 Dataset Statistics
Metric
Value
Total utterances
50,000
Total speakers
10 (5 male, 5 female)
Pure Levantine Arabic
44,154 utterances
Code-switching (AR+EN)
5,846 utterances
Sampling rate
24,000 Hz
Estimated total duration
~66.8 hours… See the full description on the dataset page: https://huggingface.co/datasets/mohammedaly22/lahgtna-levantine-tts.FineWeb2-North-Levantine-Arabic
FineWeb2 North Levantine Arabic
🇱🇧 This is the North Levantine Arabic Portion of The FineWeb2 Dataset.
🇸🇩 The North Levantine Arabic, represented by the ISO 639-3 code apc, is a member of the Afro-Asiatic language family and utilizes the Arabic script.
🇯🇴 Known within subsets as apc_Arab, this language boasts an extensive corpus of over ** 221K rows**.
Purpose of This Repository
This repository provides easy access to the Arabic portion - North Levantineof the… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/FineWeb2-North-Levantine-Arabic.Levantine_RewayatFineTranslations_Levantinearabic-palestinian-levantine-sample
4FACTORS — Palestinian Levantine Conversational Sample
50 native-written question–answer pairs in spoken Palestinian Levantine Arabic, each with an English gloss. This is a public demonstration sample from 4FACTORS, a producer of native, human-verified Arabic training data.
What this is
Real conversational exchanges — the kind of thing people actually say in shops, clinics, taxis, and at home — written from scratch by a first-language Palestinian speaker. Every… See the full description on the dataset page: https://huggingface.co/datasets/4factors/arabic-palestinian-levantine-sample.organic-levantine-arabic-dialect-dataset
Organic Levantine Arabic Dialect Dataset (Sample)
This repository contains a limited sample subset of an organic, multi-country Levantine Arabic (Shami) dataset.
Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application.… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-levantine-arabic-dialect-dataset.marabert2-levantine-toxic-model-v4-dataset
L-HSAB (Custom Split for Levantine Hate Speech)
This dataset contains Levantine Arabic tweets labeled for Hate Speech and Abusive language. It is used to train the model: amitca71/marabert2-levantine-toxic-model-v4
Dataset Structure
text: The tweet content (Arabic).
label: The classification label.
0: Abusive
1: Normal
2: Hate
Citation
Mulki, H., et al. (2019). "L-HSAB: A Levantine Twitter Dataset for Hate Speech and Abusive Language."
marabert2-levantine-toxic-model-dataset
L-HSAB (Custom Split for Levantine Hate Speech)
This dataset contains Levantine Arabic tweets labeled for Hate Speech and Abusive language. It is used to train the model: amitca71/marabert2-levantine-toxic-model
Dataset Structure
text: The tweet content (Arabic).
label: The classification label.
0: Abusive
1: Normal
2: Hate
Citation
Mulki, H., et al. (2019). "L-HSAB: A Levantine Twitter Dataset for Hate Speech and Abusive Language."
marabert2-levantine-toxic-model-v4-dataset
L-HSAB (Custom Split for Levantine Hate Speech)
This dataset contains Levantine Arabic tweets labeled for Hate Speech and Abusive language. It is used to train the model: amitca71/marabert2-levantine-toxic-model-v4
Dataset Structure
text: The tweet content (Arabic).
label: The classification label.
0: Abusive
1: Normal
2: Hate
Citation
Mulki, H., et al. (2019). "L-HSAB: A Levantine Twitter Dataset for Hate Speech and Abusive Language."
marabert2-levantine-toxic-model-v3-dataset
L-HSAB (Custom Split for Levantine Hate Speech)
This dataset contains Levantine Arabic tweets labeled for Hate Speech and Abusive language. It is used to train the model: amitca71/marabert2-levantine-toxic-model-v3
Dataset Structure
text: The tweet content (Arabic).
label: The classification label.
0: Abusive
1: Normal
2: Hate
Citation
Mulki, H., et al. (2019). "L-HSAB: A Levantine Twitter Dataset for Hate Speech and Abusive Language."
marabert2-levantine-hate-model-dataset
L-HSAB (Custom Split for Levantine Hate Speech)
This dataset contains Levantine Arabic tweets labeled for Hate Speech and Abusive language. It is used to train the model: amitca71/marabert2-levantine-hate-model
Dataset Structure
text: The tweet content (Arabic).
label: The classification label.
0: Abusive
1: Normal
2: Hate
Citation
Mulki, H., et al. (2019). "L-HSAB: A Levantine Twitter Dataset for Hate Speech and Abusive Language."
levantine_dialectsA compilation of levantine dialects. Useful for LID.
levantine_loc-both_mcq0_test_ds_with_sas_responses
