datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
documents-Egyptian-Arabic
Current Hub Validation Status
Dataset Server rows: 25,399,945
Dataset Server original/Parquet size: 2,758,228,707 bytes (~2.76 GB)
Default Hub configuration currently exposes one column: text
The additional configuration names listed in this card are physical source directories and are not all recognized as separate Hub configurations. Keep the default configuration until the dataset is normalized into explicit, tested splits.
Egyptian Arabic Mega Corpus (EAMC) —… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/documents-Egyptian-Arabic.1.9M-Egyptian-Corpus
1.92M Egyptian Arabic Corpus 🇪🇬 — Prickly Labs
A corpus of 1.92 million Egyptian Arabic samples combining synthetic generation and real web data, built for continued pretraining and dialect adaptation of Arabic language models. Designed to reflect how Egyptians actually speak — not textbook MSA.
⚠️ This dataset contains uncensored, informal Arabic including sarcasm, humor, slang, and profanity. Use with care for public-facing applications.
📌 Overview… See the full description on the dataset page: https://huggingface.co/datasets/Prickly-Labs/1.9M-Egyptian-Corpus.egyptian-arabic-fake-reviews
🕵️♂️🇪🇬 FREAD: Fake Reviews Egyptian Arabic Dataset
Author: IbrahimAmin, Ismail Fakhr, M. Waleed Fakhr, Rasha Kashef License: MIT Paper: Boosting Arabic Fake Reviews Detection by Integrating Textual and Metadata Features: A Transformer-Based Model Languages: Arabic (Egyptian Dialect)
📚 Dataset Summary
FREAD is designed for detecting fake reviews in Arabic using both textual content and behavioral metadata. It contains 60,000 reviews (50K train / 10K test)… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimAmin/egyptian-arabic-fake-reviews.egyptian-nlu
Egyptian Arabic voice-assistant NLU — dataset
Egyptian Arabic commands paired with intent + slot annotations, in the schema of
Amazon MASSIVE (60 intents, 55 slot types).
Two parts, and the difference matters:
File
Rows
Origin
egy_test.jsonl
200
Written and annotated by hand by a native Egyptian speaker — the benchmark
egy_synth_train.jsonl
6,509
LLM-generated Egyptian rewrites of MASSIVE ar-SA training items
egy_synth_dev.jsonl
730
Same, held out by seed… See the full description on the dataset page: https://huggingface.co/datasets/Alhasan/egyptian-nlu.egyptian-dialogue
Egyptian Arabic Dialogue Dataset
Dataset Description
This dataset contains 4,322 parallel Egyptian Arabic-English dialogue pairs with automatic domain classification. The data is extracted from TV series subtitles and features natural conversational Egyptian Arabic dialect (العامية المصرية).
Languages
Source: Egyptian Arabic (ar_EG) - Colloquial dialect
Target: English (en)
Dataset Summary
Egyptian Arabic is one of the most widely spoken Arabic… See the full description on the dataset page: https://huggingface.co/datasets/fr3on/egyptian-dialogue.ancient-egyptian-multilingual-premium
🏛️ Ancient Egyptian Multilingual Corpus
Dataset Description
A comprehensive multilingual corpus of Ancient Egyptian texts combining multiple authoritative sources, including dictionaries and translated texts from German to English. Available in multiple formats for maximum accessibility.
📊 Dataset Summary
This dataset combines:
Dictionary Sources (3 sources, ~15,000 entries):
Vygus Egyptian Dictionary 2015 (Mark Vygus)
Dickson Dictionary of Middle Egyptian… See the full description on the dataset page: https://huggingface.co/datasets/AhmedElTaher/ancient-egyptian-multilingual-premium.General_Facts_in_English_Arabic_Egyptian_Arabic
🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized)
The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.oasst2_egyptian_arabic_convsegyptian-songs
Egyptian Arabic Songs Dataset 🎵
Dataset Description
This dataset contains 3,063 lines of Egyptian Arabic song lyrics with English translations, spanning 280 songs from 1983-2021. The dataset features natural Egyptian Arabic dialect (العامية المصرية) as used in popular music, with automatic genre classification and line-type detection.
Languages
Source: Egyptian Arabic (ar_EG) - Colloquial dialect in music
Target: English (en)
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/fr3on/egyptian-songs.egyptian-arabic-style-dataset
Egyptian Arabic Conversational Dataset 🇪🇬
Overview
A dataset for adapting Large Language Models to understand and generate natural Egyptian Arabic conversations.
Dataset Content
This dataset contains:
Egyptian Arabic dialogues
Dialect style examples
Conversational data
Multi-turn conversations
Format
Each sample follows chat format:
{
"messages": [
{
"role": "user",
"content": "عامل ايه؟"
},
{… See the full description on the dataset page: https://huggingface.co/datasets/ApexVOrteX-1/egyptian-arabic-style-dataset.English-Egyptian-Translation-finance
Bilingual Egyptian Finance Dataset
Dataset Description
This dataset contains bilingual text pairs in English and Egyptian Arabic focused on finance and financial topics. Each entry provides parallel translations covering various aspects of Egyptian and regional economics, making it valuable for translation models, multilingual NLP research, and economic analysis applications.
Key Features
Languages: English ↔ Egyptian Arabic (العامية المصرية)
Domain: Economics… See the full description on the dataset page: https://huggingface.co/datasets/Omar-youssef/English-Egyptian-Translation-finance.Egyptian-text-summarization
Egyptian Arabic Text Summarization Dataset
Dataset Description
This dataset contains text-summary pairs in Egyptian Arabic designed for training and evaluating text summarization models.
Key Features
Language: Egyptian Arabic (العامية المصرية)
Task: Text Summarization
Format: Text-summary pairs with topic categorization
Content: Diverse topics with natural Egyptian Arabic usage
Dataset Structure
Data Fields
text: Original text content… See the full description on the dataset page: https://huggingface.co/datasets/Omar-youssef/Egyptian-text-summarization.QA_Finance_Egyptian_dataset
QA Finance Egyptian Dataset
A synthetic Arabic (Egyptian dialect) question-answering dataset focused on financial and accounting topics, such as financial statement analysis, corporate performance evaluation, and related concepts.
Dataset Details
Language: Arabic (Egyptian dialect, ar)
Domain: Finance / Accounting
Size: 2,670 examples (train split)
Format: Question–answer pairs, each tagged with its source topic
Dataset Structure
Each row… See the full description on the dataset page: https://huggingface.co/datasets/Omar-youssef/QA_Finance_Egyptian_dataset.Egyptian-text-summarization
Egyptian Arabic Text Summarization Dataset
Dataset Description
This dataset contains text-summary pairs in Egyptian Arabic designed for training and evaluating text summarization models.
Key Features
Language: Egyptian Arabic (العامية المصرية)
Task: Text Summarization
Format: Text-summary pairs with topic categorization
Content: Diverse topics with natural Egyptian Arabic usage
Dataset Structure
Data Fields
text: Original text content… See the full description on the dataset page: https://huggingface.co/datasets/Tariq2023/Egyptian-text-summarization.egyptian-voice-commands
Egyptian Voice Commands Dataset
This repository contains the Egyptian Arabic voice commands dataset used for training and evaluating the EgyptianAgent ASR and NLU models.
Dataset Structure
egyptian_voice_commands/
├── train.jsonl # Training data (665 examples)
├── eval.jsonl # Validation data (50 examples)
└── test.jsonl # Test data (102 examples)
egyptian_ui_navigation/
├── train.jsonl # Training data (50 examples)
└── test.jsonl # Test data… See the full description on the dataset page: https://huggingface.co/datasets/Kandil7/egyptian-voice-commands.egyptian-dialogue
Egyptian Arabic Dialogue Dataset
Dataset Description
This dataset contains 4,322 parallel Egyptian Arabic-English dialogue pairs with automatic domain classification. The data is extracted from TV series subtitles and features natural conversational Egyptian Arabic dialect (العامية المصرية).
Languages
Source: Egyptian Arabic (ar_EG) - Colloquial dialect
Target: English (en)
Dataset Summary
Egyptian Arabic is one of the most widely spoken Arabic… See the full description on the dataset page: https://huggingface.co/datasets/lihwak74/egyptian-dialogue.arabic-egyptian-sample
4FACTORS Arabic — Egyptian Q&A Sample
Conversational question–answer pairs in spoken Egyptian Arabic, written by a
first-language Egyptian speaker. A 50-item demonstration sample, with English
glosses, released under CC BY-NC 4.0.
This is the Egyptian variety in the 4FACTORS Arabic sample set, alongside the
Palestinian Levantine
and Modern Standard Arabic (MSA) sets.
What this is
Fifty short question–answer exchanges of the kind that come up in everyday life —… See the full description on the dataset page: https://huggingface.co/datasets/4factors/arabic-egyptian-sample.arsyra-egyptian
🇪🇬 ArSyra Egyptian Arabic (Masri) Dataset
The most widely understood Arabic dialect — now as structured NLP data.
Dataset Summary
A dedicated Egyptian Arabic (Masri) dataset encompassing all linguistic
categories. Egyptian Arabic is the most widely understood dialect in the Arab
world, used extensively in media, entertainment, and online discourse.
This dataset captures authentic Egyptian speech patterns, colloquial expressions,
and cultural references contributed… See the full description on the dataset page: https://huggingface.co/datasets/ArSyra/arsyra-egyptian.
