datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-palestinian-levantine-sample
4FACTORS — Palestinian Levantine Conversational Sample
50 native-written question–answer pairs in spoken Palestinian Levantine Arabic, each with an English gloss. This is a public demonstration sample from 4FACTORS, a producer of native, human-verified Arabic training data.
What this is
Real conversational exchanges — the kind of thing people actually say in shops, clinics, taxis, and at home — written from scratch by a first-language Palestinian speaker. Every… See the full description on the dataset page: https://huggingface.co/datasets/4factors/arabic-palestinian-levantine-sample.arsyra-levantine
🇸🇾 ArSyra Levantine Arabic (Shami) Dataset
Authentic Shami dialect data from Syria, Lebanon, Jordan, and Palestine.
Dataset Summary
A curated collection of Levantine Arabic (Shami) data from Syria, Lebanon,
Jordan, and Palestine. Levantine Arabic is spoken by over 30 million people
and shares distinctive phonological and lexical features that set it apart
from other dialect groups.
This dataset captures the natural variation within the Levantine dialect
continuum —… See the full description on the dataset page: https://huggingface.co/datasets/ArSyra/arsyra-levantine.
