datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-dialect-corpus
Arabic Dialect Corpus
A comprehensive collection of Arabic dialectal text, standardized for Natural Language Processing (NLP) model training, evaluation, and linguistic analysis. This corpus has been meticulously processed to ensure high-quality tokenization and consistent metadata.
Dataset Statistics
Metric
Value
Total Records
127,180
Total Tokens
5,802,324
Average Tokens per Record
45.62
Dialect Categories
5
Changelog… See the full description on the dataset page: https://huggingface.co/datasets/dataflare/arabic-dialect-corpus.arabic-dialect-corpus
🇪🇬🇸🇦 Arabic Dialect Corpus (Egyptian & Saudi)
Dataset Description
This dataset contains 150K+ natural, informal Arabic text samples scraped from high-engagement YouTube discussions. It specifically targets Egyptian (EG) and Saudi (SA) dialects, filling a critical gap in resources for training LLMs on colloquial Arabic (Ammiya) rather than just Modern Standard Arabic (MSA).
Languages
Primary Dialects:
Egyptian Arabic (EG) - Cairene and regional Egyptian… See the full description on the dataset page: https://huggingface.co/datasets/fr3on/arabic-dialect-corpus.arabic-dialect-dpo
Arabic Dialect DPO Dataset - Egyptian & Saudi
The first large-scale Arabic dialect preference dataset for DPO/ORPO/GRPO alignment training. Contains 22,538 preference triples across two major Arabic dialects: Egyptian (Masry) and Saudi (Najdi).
Dataset Summary
Config
Dialect
Rows
Language Code
egyptian
Egyptian Arabic (مصري)
11,038
ar-EG
saudi
Saudi Arabic (سعودي نجدي)
11,500
ar-SA
Total
22,538
Usage
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-dialect-dpo.arabic_dialects_question_and_answerData Content
The file provided: Q/A Reasoning dataset
contains the following columns:
ID # : Denotes the reference ID for:
a. Question
b. Answer to the question
c. Hint
d. Reasoning
e. Word count for items a to d above
Dialects: Contains the following dialects in separate columns:
a. English
b. MSA
c. Emirati
d. Egyptian
e. Levantine Syria
f. Levantine Jordan
g. Levantine Palestine
h. Levantine Lebanon
Data Generation Process
The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.arabic-dialect-corpus
🇪🇬🇸🇦 Arabic Dialect Corpus (Egyptian & Saudi)
Dataset Description
This dataset contains 150K+ natural, informal Arabic text samples scraped from high-engagement YouTube discussions. It specifically targets Egyptian (EG) and Saudi (SA) dialects, filling a critical gap in resources for training LLMs on colloquial Arabic (Ammiya) rather than just Modern Standard Arabic (MSA).
Languages
Primary Dialects:
Egyptian Arabic (EG) - Cairene and regional… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/arabic-dialect-corpus.
