CoolFace
Datasetpublic

ISLAM-PO/arab-dialects-20-countries-3m

Arab Dialects Dataset - 20 Countries A large-scale Arabic dialects dataset covering 20 Arab countries, 7 content types per country, 3,000,000 records, 140 JSONL files, 12.07 GB. UTF-8 JSONL, ready for Hugging Face Datasets. 1. Contents 1. Contents 2. Dataset Summary 3. Repository Map 4. Countries Table (20 folders) 5. Data Types Table (7 files) 6. Record Schema 7. Loading and Usage 8. Generation and Reproduction 9. Considerations and Limitations 10. Contributors… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arab-dialects-20-countries-3m.

sourceHugging Facecc-by-4.0updated 20d agoView on Hugging Face
0likes588downloads
Dataset Card

Arab Dialects Dataset - 20 Countries

A large-scale Arabic dialects dataset covering 20 Arab countries, 7 content types per country, 3,000,000 records, 140 JSONL files, 12.07 GB. UTF-8 JSONL, ready for Hugging Face Datasets.

1. Contents

2. Dataset Summary

AttributeValue
Dataset nameArab_Dialects_Dataset
Countries20
Files140 *.jsonl files (20 x 7)
Total records3,000,000 (150,000 per country)
Total size12.07 GB (12,363.8 MB)
File formatJSONL, UTF-8 (one record per line)
LanguagesArabic dialects: Egyptian, Gulf, Levantine, Maghrebi, Yemeni, Iraqi, Sudanese, Somali + MSA explanations
Suggested tasksDialect text generation, dialect understanding, dialect/country classification, Arabic chatbot, dialect-to-MSA translation
LicenseCC-BY-4.0
Loading methodload_dataset with streaming=True (recommended for 12 GB)

3. Repository Map

3.1 Folder Tree

text
Arab_Dialects_Dataset/
├── README.md                  # This file - Hugging Face Dataset Card
├── LICENSE                    # CC-BY-4.0 license
├── CITATION.cff               # Citation metadata
├── .gitattributes             # Git LFS settings for *.jsonl
│
├── 01_مصر/  (Egypt)
│   ├── 01_اللهجة.jsonl               # 21,432 records (dialect)
│   ├── 02_المصطلحات.jsonl            # 21,428 records (terms)
│   ├── 03_النكت.jsonl                # 21,428 records (jokes)
│   ├── 04_المواقف.jsonl              # 21,428 records (situations)
│   ├── 05_الثقافة_والعادات.jsonl     # 21,428 records (culture)
│   ├── 06_طريقة_الكلام.jsonl         # 21,428 records (speaking style)
│   └── 07_المحادثات.jsonl            # 21,428 records (dialogues)
│
├── 02_السعودية/ (Saudi Arabia)   # same 7 files
├── 03_الامارات/ (UAE)            # same 7 files
├── 04_الكويت/ (Kuwait)
├── 05_قطر/ (Qatar)
├── 06_البحرين/ (Bahrain)
├── 07_عمان/ (Oman)
├── 08_اليمن/ (Yemen)
├── 09_العراق/ (Iraq)
├── 10_سوريا/ (Syria)
├── 11_لبنان/ (Lebanon)
├── 12_الاردن/ (Jordan)
├── 13_فلسطين/ (Palestine)
├── 14_السودان/ (Sudan)
├── 15_ليبيا/ (Libya)
├── 16_تونس/ (Tunisia)
├── 17_الجزائر/ (Algeria)
├── 18_المغرب/ (Morocco)
├── 19_موريتانيا/ (Mauritania)
└── 20_الصومال/ (Somalia)         # same 7 files per country
Rule: each country = exactly 150,000 records = 21,432 + 6 x 21,428.

3.2 Visual Map (Mermaid)

mermaid
graph TD
    ROOT[Arab_Dialects_Dataset/<br/>3M records - 12.07GB]
    ROOT --> DOCS[README + LICENSE + CITATION + .gitattributes]
    ROOT --> EG[01_Egypt<br/>150K - 605MB]
    ROOT --> GULF[Saudi Arabia ... Oman<br/>Gulf countries]
    ROOT --> LEV[Iraq ... Palestine<br/>Levant + Iraq]
    ROOT --> AFR[Sudan ... Somalia<br/>North Africa + Horn]
    EG --> T1[01_dialect.jsonl]
    EG --> T2[02_terms.jsonl]
    EG --> T3[03_jokes.jsonl]
    EG --> T4[04_situations.jsonl]
    EG --> T5[05_culture.jsonl]
    EG --> T6[06_speaking_style.jsonl]
    EG --> T7[07_dialogues.jsonl]
    T7 --> REC[JSON record<br/>id - country - category<br/>dialect - text - meta]
mermaid
pie title Records distribution by type (per country, 150K)
    "Dialect" : 21432
    "Terms" : 21428
    "Jokes" : 21428
    "Situations" : 21428
    "Culture" : 21428
    "Speaking style" : 21428
    "Dialogues" : 21428

4. Countries Table (20 folders)

#FolderCountryRecordsSizeDistinctive words (examples)
101_مصرEgypt150,000605.3 MBezzayyak, aamel eh, ya ma'allem, keda, awi
202_السعوديةSaudi Arabia150,000624.7 MBwesh halek, absher, kafu, ya hala
303_الاماراتUAE150,000623.1 MBshhalek, zein, rayoog, majboos
404_الكويتKuwait150,000621.4 MBshkhbarek, wayed, dewaniya
505_قطرQatar150,000605.0 MBshhalek, tayyeb, dalla, kafu alaik
606_البحرينBahrain150,000625.6 MBshkhbarek, wayed zein, jabati
707_عمانOman150,000606.1 MBmooh, Salalah, luban, shuwaa
808_اليمنYemen150,000616.3 MBmandi, fahsa, salta, Adani tea
909_العراقIraq150,000614.3 MBshlonak, shaku maku, yaba, hassa
1010_سورياSyria150,000618.3 MBkeefak, tkram, shawarma, kibbeh
1111_لبنانLebanon150,000623.5 MBmeshe lhal, tkram aynak, manoushe
1212_الاردنJordan150,000619.2 MBsho malak, mansaf, zalameh
1313_فلسطينPalestine150,000628.2 MBmusakhan, knafeh, ya ammi
1414_السودانSudan150,000617.0 MBzol, ya zol, jabana, aseeda
1515_ليبياLibya150,000613.4 MBshen halek, bahi, bazeen, halba
1616_تونسTunisia150,000619.8 MBshnahwalek, labes, barsha
1717_الجزائرAlgeria150,000616.5 MBwash rak, labes, bezaf, ya khou
1818_المغربMorocco150,000614.8 MBlabas, kidayer, bezaf, atay
1919_موريتانياMauritania150,000628.6 MBlabas, yasser, ya wkheiret
2020_الصومالSomalia150,000622.4 MBMogadishu, baasto, shaah
Total20 folders3,000,00012,363.8 MB (12.07 GB)

5. Data Types Table (7 files)

Each country contains the same 7 files. This table is per country - multiply by 20 for the global total.

#File name`category`Records/countryTotal (x20)Description
101_اللهجة.jsonlلهجة (dialect)21,432428,640Authentic greeting expressions and phrases in the local dialect + MSA meaning
202_المصطلحات.jsonlمصطلحات (terms)21,428428,560Work/market/home glossary: word + meaning + usage context
303_النكت.jsonlنكت (jokes)21,428428,560Jokes and funny situations in the country dialect
404_المواقف.jsonlمواقف (situations)21,428428,560Daily scenes: taxi, wedding, cafe, market
505_الثقافة_والعادات.jsonlثقافة (culture)21,428428,560Hospitality, dishes (kabsa/mansaf/tagine/bazeen...), weddings, majlis
606_طريقة_الكلام.jsonlطريقة_كلام (speaking style)21,428428,560Tone and style: how to speak like locals, sentence openers/closers
707_المحادثات.jsonlمحادثات (dialogues)21,428428,560Full dialogues: friend-friend, seller-customer, WhatsApp chat

6. Record Schema

6.1 Fields Table

FieldTypeExampleDescription
idstringمصر-نكت-000001Unique ID: country-category-number
countrystringمصر (Egypt)Country name in Arabic
categorystringنكت (jokes)One of: لهجة, مصطلحات, نكت, مواقف, ثقافة, طريقة_كلام, محادثات
dialectstringواحد مصري اتصل بصاحبه...Short original dialect text (50-300 chars)
textstringواحد مصري اتصل... + long explanationFull expanded text (~3.5-4 KB) - best for training
meta.k1stringبقىDistinctive word 1 from the country dialect
meta.k2stringعامل ايهDistinctive word 2
meta.k3stringيا جدعDistinctive word 3

6.2 Real Record Example

json
{
  "id": "مصر-نكت-000001",
  "country": "مصر",
  "category": "نكت",
  "dialect": "واحد مصري اتصل بصاحبه قاله 'بقى فينك؟' قاله 'في البيت' قاله 'طب عامل ايه وافتح الباب ما انا قدام البيت' 😂",
  "text": "واحد مصري اتصل بصاحبه قاله 'بقى فينك؟' ... وهذا يعكس روح أهل مصر في كلامهم اليومي حيث يستخدمون 'بقى' و'عامل ايه' و'يا جدع' بكثرة ... [record 1 - Egypt - jokes]",
  "meta": {"k1": "بقى", "k2": "عامل ايه", "k3": "يا جدع"}
}
Note: dialect and text are in Arabic (the dataset content language). All documentation around them is in English.

7. Loading and Usage

7.1 Load with Hugging Face datasets (recommended: streaming for 12 GB)

python
from datasets import load_dataset

# Load everything with streaming (no 12 GB download needed)
ds = load_dataset("Arab_Dialects_Dataset", split="train", streaming=True)
print(next(iter(ds)))

# Filter one country / category
egy_jokes = ds.filter(lambda x: x["country"] == "مصر" and x["category"] == "نكت")
for row in egy_jokes.take(3):
    print(row["dialect"])

7.2 Load a single country (faster)

python
from datasets import load_dataset

# Egypt only
ds_eg = load_dataset("json", data_files="01_مصر/*.jsonl", split="train", streaming=True)

# One file only
ds_one = load_dataset("json", data_files="18_المغرب/07_المحادثات.jsonl", split="train")

7.3 Local read with Python / Pandas

python
import json, glob
import pandas as pd

files = glob.glob("Arab_Dialects_Dataset/01_مصر/*.jsonl")
rows = []
for f in files[:1]:
    with open(f, encoding="utf-8") as fh:
        for line in fh:
            rows.append(json.loads(line))
df = pd.DataFrame(rows)
print(df[["id", "category", "dialect"]].head())
print(df["category"].value_counts())

7.4 Language model training (short example)

python
# Use the 'text' field for causal LM or 'dialect' for instruction tuning.
# Example prompt:
# instruction: "Write a joke in the Moroccan dialect"
# input: row["dialect"]

8. Generation and Reproduction

The dataset was generated with generate_dataset.py (per-country templates + distinctive vocabulary + contextual padding to reach the target size).

bash
# Small demo (~16 MB, for validation)
python generate_dataset.py --demo

# Full generation (150K per country / ~10 GB+)
python generate_dataset.py --full --per-country 150000 --total-gb 10
ParameterDescriptionDefault
--demo700 records/country (100 per type) for testing
--fullFull generation
--per-countryRecords per country150000
--total-gbTarget size in GB (automatic padding if below target)10
Note: per-record size is computed as total_gb x 1024^3 / (20 x per_country) ~= 3579 bytes. Actual size is ~3.8-4.1 KB with JSON overhead, so the final output is 12.07 GB (above target by design, to guarantee the 10 GB requirement).

9. Considerations and Limitations

TopicDetails
Data natureSynthetic, template-generated data, not recorded from native speakers. Good for pre-training and experiments, but review quality before production use.
RepetitionTemplates repeat with varied words and IDs. Deduplicate before final evaluation.
BiasBalanced representation (150K per country) - does not reflect real population sizes. Large dialects (Egyptian/Levantine) have the same weight as smaller ones.
Cultural sensitivityJokes and situations are for language research/entertainment only. If you find offensive content for a country, please open an Issue.
Language qualityDistinctive words are real per dialect, but long compositions may contain non-100% native phrasing. Human review is recommended for critical samples.
Size12 GB - use streaming=True or load a single country if your machine is limited.

10. Contributors

RoleName / HandleContribution
Dataset owner & conceptProject OwnerIdea, 20 countries + 7 types definition, 10 GB requirement
Generation & engineeringMuse Spark (AI Assistant)generate_dataset.py, dialect templates, documentation
Dialect review (open)Open for contributionNative speakers per country review words/jokes and add authentic expressions
Culture review (open)Open for contributionFix dishes and customs per country
Want to contribute? Send a Pull Request adding new words/templates in generate_dataset.py under COUNTRIES or TEMPLATES, or fix any incorrect expression.

11. License

CC-BY-4.0 - Creative Commons Attribution 4.0 International
  • Allowed: commercial use, modification, distribution, training.
  • Single requirement: give attribution - dataset name + link + license.
  • See the LICENSE file for full details.

12. Citation

bibtex
@dataset{arab_dialects_20_2026,
  title   = {Arab Dialects Dataset: 20 Countries, 7 Content Types, 3M Records},
  author  = {Project Owner and Contributors},
  year    = {2026},
  publisher = {Hugging Face},
  version = {1.0.0},
  url     = {https://huggingface.co/datasets/USERNAME/Arab_Dialects_Dataset},
  note    = {3,000,000 records, 140 JSONL files, 12.07 GB, CC-BY-4.0}
}

The CITATION.cff file contains the same data in Citation File Format.

13. Push to Hugging Face Hub

bash
# 1. Install tools
pip install huggingface_hub datasets

# 2. Login
huggingface-cli login

# 3. Init repo (Git LFS is required for large JSONL files)
cd Arab_Dialects_Dataset
git init
git lfs install
git lfs track "*.jsonl"
# .gitattributes is already included - make sure it is committed

# 4. Push (upload takes a while for 12 GB - push in batches on slow connections)
huggingface-cli repo create Arab_Dialects_Dataset --type dataset --yes
git remote add origin https://huggingface.co/datasets/USERNAME/Arab_Dialects_Dataset
git add README.md LICENSE CITATION.cff .gitattributes
git commit -m "docs: HF dataset card"
git push origin main
# Then push countries in batches:
git add 01_مصر 02_السعودية 03_الامارات 04_الكويت
git commit -m "data: gulf+egypt batch"
git push origin main
# ... repeat for remaining countries
Tip: replace USERNAME with your Hugging Face username in the URL above and in the remote command. Add --private to repo create if you want it private first.

Last updated: 2026-09-03 | Version: 1.0.0 | Status: complete, 3M records / 12.07 GB