datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
100-hour-Egyptian-dataset-single-speaker
Masri 100h — Egyptian Arabic Single-Speaker Speech Corpus
A 100-hour Egyptian Arabic (مصري) single-narrator speech collection — 15,653 released clips at 24 kHz mono, with aligned transcripts.
Egyptian Arabic is the most widely understood Arabic dialect and one of the least served by open speech data.
Almost every open Arabic corpus is Modern Standard Arabic (MSA) — a register nobody actually speaks at home.
This dataset is built for the opposite: natural, spoken, conversational… See the full description on the dataset page: https://huggingface.co/datasets/ehabnegm/100-hour-Egyptian-dataset-single-speaker.Egyptian_hieroglyphs
Egyptian hieroglyphs 𓂀
Hieroglyphs image dataset along with Language Model !
Features
This dataset is build from the hieroglyphs found in 10 different pictures from the book "The Pyramid of Unas" (Alexandre Piankoff, 1955). We therefore urge you to have access to this book before using the dataset.
The ten different pictures used throughout this dataset are: 3,5,7,9,20,21,22,23,39,41 (numbers represent the numbers used in the book "The pyramid of Unas".
Each… See the full description on the dataset page: https://huggingface.co/datasets/HamdiJr/Egyptian_hieroglyphs.egyptian-arabic-speechEgyptian-ASR-MGB-3
Egyptian Arabic dialect automatic speech recognition
Dataset Summary
This dataset was collected, cleaned and adjusted for huggingface hub and ready to be used for whisper finetunning/training.
From MGB-3 website:
The MGB-3 is using 16 hours multi-genre data collected from different YouTube channels. The 16 hours have been manually transcribed.
The chosen Arabic dialect for this year is Egyptian.
Given that dialectal Arabic has no orthographic rules, each program has… See the full description on the dataset page: https://huggingface.co/datasets/MightyStudent/Egyptian-ASR-MGB-3.Egypt-Stock-Symbols-and-Metadata
Egypt Stock Symbols & Company Metadata
This dataset contains stock symbols and basic company metadata for all listed companies in Egypt.It is updated weekly if new changes are there.
📊 Dataset Contents
The dataset is provided as a CSV file with the following columns:
Column
Description
name
Full company name
ticker
Stock ticker symbol (e.g., AAPL, MSFT)
market
The exchange/market where the stock is listed
sector
The primary business sector of the… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/Egypt-Stock-Symbols-and-Metadata.tla-Earlier_Egyptian_original-v18-premium
Dataset Card for Dataset tla-Earlier_Egyptian_original-v18-premium
This data set contains Earlier Egyptian, i.e., ancient Old Egyptian and ancient Middle Egyptian, sentences in hieroglyphs and transliteration, with lemmatization, with POS glossing and with a German translation.
This set of original Earlier Egyptian sentences only contains text witnesses from before the start of the New Kingdom (late 16th century BCE).
The data comes from the database of the Thesaurus Linguae… See the full description on the dataset page: https://huggingface.co/datasets/thesaurus-linguae-aegyptiae/tla-Earlier_Egyptian_original-v18-premium.EgyptianMMLU
Dataset Card for EgyptianMMLU
Dataset Description
Dataset Summary
EgyptianMMLU is a composite benchmark translated into Egyptian Arabic. It combines subsets of ArabicMMLU-egy and the original English MMLU, translated using in-house models and validated by human annotators. The dataset spans 44 subjects and evaluates reasoning, factual knowledge, and world understanding in Egyptian Arabic.
Supported Tasks
Task Category: Multiple-choice question… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/EgyptianMMLU.ipfs_egypt_laws_ir
Egypt legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_egypt_laws (revision 0ddf04279b31d1f2200e96d743379bd19aa7e5cb) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Egypt prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was invented.… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_egypt_laws_ir.Egyptian-Arabic-Lectures
Egyptian Arabic Lectures Dataset
The Egyptian Arabic Lectures dataset is a collection of transcribed audio clips (around 30 hours) extracted from educational lectures delivered in Egyptian Arabic (with mixed English technical terms, such as in Physics, IoT and Operating Systems etc.). It is designed to train, evaluate, and fine-tune Automatic Speech Recognition models for the Egyptian dialect, specifically in educational and academic CS contexts.
Alongside the audio and text… See the full description on the dataset page: https://huggingface.co/datasets/ismaeeelxd/Egyptian-Arabic-Lectures.egypt-legal-corpus
Egyptian Legal Corpus
A comprehensive collection of Egyptian legal texts, meticulously extracted and tokenized for Natural Language Processing (NLP) applications, legal research, and AI model training. This corpus provides high-quality Arabic legal content with structured metadata for efficient processing.
Dataset Statistics
This release provides a foundational legal corpus with strict quality controls:
Token Count: 25M+ tokens (25,054,372 tokens) using cl100k_base… See the full description on the dataset page: https://huggingface.co/datasets/dataflare/egypt-legal-corpus.Thoth-Sphinx-egyptian-hieroglyphs
Thoth-Sphinx — Egyptian Hieroglyphs Multi-Sign Detection Dataset
Dataset Summary
This dataset provides multi-sign, multi-cartouche annotated images of
Middle Egyptian hieroglyphic inscriptions for object detection, in YOLO
format. To our knowledge, no other publicly available dataset combines:
Multiple signs per image (average ~34 instances/image), rather than
isolated single-glyph crops
Royal cartouche detection as its own class, with signs annotated
inside the… See the full description on the dataset page: https://huggingface.co/datasets/beaunix/Thoth-Sphinx-egyptian-hieroglyphs.egyptian-arabic-400kEgyptian_Dialect
Egyptian Arabic Speech Dataset
Dataset Description
This dataset contains 2,438 short audio segments in Egyptian Arabic paired with transcriptions.
The dataset was created for fine-tuning Automatic Speech Recognition (ASR) models on conversational Egyptian Arabic.
Each example contains:
WAV audio
Egyptian Arabic transcription
Dataset Creation
Source
The audio was collected from publicly available YouTube videos featuring native… See the full description on the dataset page: https://huggingface.co/datasets/Kyrillos2001/Egyptian_Dialect.bbaw_egyptian
Dataset Card for "bbaw_egyptian"
Dataset Summary
This dataset comprises parallel sentences of hieroglyphic encodings, transcription and translation as used in the paper Multi-Task Modeling of Phonographic Languages: Translating Middle Egyptian Hieroglyph. The data triples are extracted from the digital corpus of Egyptian texts compiled by the project "Strukturen und Transformationen des Wortschatzes der ägyptischen Sprache".
Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/phiwi/bbaw_egyptian.Egyptian-Handwriting-Dataset
Egyptian Handwriting Dataset
A dataset of 11k+ handwritten Arabic words from Egyptian writers, extracted and tightly cropped from scanned paper forms. This dataset offers diverse handwriting samples ranging from children to elderly contributors, making it ideal for training robust Arabic handwriting recognition models.
Each form contains 6 unique words, resulting in 24 handwritten word images per form.
Each word is written four times by the same writer to capture… See the full description on the dataset page: https://huggingface.co/datasets/OmarMDiab/Egyptian-Handwriting-Dataset.ipfs_egypt_laws
Egypt Constitution and Official Gazette texts (Presidency / Alamiria)
Research snapshot of official national legislation from Presidency constitution PDF + Alamiria official printer TashTxt (الجريدة الرسمية).
Not legal advice. The official gazette / authentic source prevails over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-10
Coverage
partial-official-fulltext
Source
Presidency constitution PDF + Alamiria official printer TashTxt… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_egypt_laws.egypt_cars_and_license_platesEgyptian-SFT-Mixture
📚 Egyptian SFT Mixture
Dataset Overview
This dataset contains the Supervised Fine-Tuning samples for Nile-Chat in both Arabic and Latin scripts. Each sample follows the following format:
[{"role": "user", "content": user_prompt}, {"role": "assistant", "content": assistant_answer}]
Dataset Categories
This dataset can be mainly divided into two main categories, Native and Synthetic Egyptian Instruction Datasets. The Native datasets have been collected and… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/Egyptian-SFT-Mixture.Torah_Gnostic_Egypt_India_China_Greece_holy_texts_sources
Torah Codes Religion Texts Sources
Data Tree
── arabs
│ ├── astrological_stelar_magic.txt
│ └── Holy-Quran-English.txt
├── ars
│ ├── ars_magna_ramon_llull.txt
│ └── lemegeton_book_solomon.txt
├── asimov
│ ├── foundation.txt
│ └── prelude_to_foundation.txt
├── budist
│ ├── bardo_todhol_book_of_deads_tibet_libro_tibetano_de_los_muertos.txt
│ ├── rig_veda.txt
│ └── TheTeachingofBuddha.txt
├── cathars
├── china
│ ├── arte_de_la_guerra_art_of_war.txt
│… See the full description on the dataset page: https://huggingface.co/datasets/torahCodes/Torah_Gnostic_Egypt_India_China_Greece_holy_texts_sources.egyptian-arabic-tts-diacritized
Egyptian Arabic TTS Corpus (Diacritized)
97,163 utterances / ~334 hours of Egyptian Arabic speech at 24 kHz, with
diacritized transcripts — the short vowels that Arabic script does not
write.
Why diacritics
Arabic is an abjad: short vowels are unwritten, so كتب may be kataba,
kutiba, or kutub. A TTS model with no Arabic pretraining cannot infer which,
and guesses — which native listeners hear as a foreign accent with constant
mispronunciation.
This is invisible to… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/egyptian-arabic-tts-diacritized.1.9M-Egyptian-Corpus
1.92M Egyptian Arabic Corpus 🇪🇬 — Prickly Labs
A corpus of 1.92 million Egyptian Arabic samples combining synthetic generation and real web data, built for continued pretraining and dialect adaptation of Arabic language models. Designed to reflect how Egyptians actually speak — not textbook MSA.
⚠️ This dataset contains uncensored, informal Arabic including sarcasm, humor, slang, and profanity. Use with care for public-facing applications.
📌 Overview… See the full description on the dataset page: https://huggingface.co/datasets/Prickly-Labs/1.9M-Egyptian-Corpus.FineWeb2-Egyptian-Arabic
FineWeb2 Egyptian Arabic
🇪🇬 This is the Egyptian Arabic Portion of The FineWeb2 Dataset.
🇪🇬 This dataset contains a rich collection of text in Egyptian Arabic (ISO 639-3: arz), a widely spoken dialect within the Afro-Asiatic language family.
🇪🇬 With over 439 million words and 1.4 million documents, it serves as a valuable resource for NLP development and linguistic research focused on Egyptian Arabic.
Purpose of This Repository
This repository provides easy… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/FineWeb2-Egyptian-Arabic.tla-late_egyptian-v19-premium
Dataset Card for Dataset tla-Late_Egyptian-v19-premium
This data set contains Late Egyptian sentences in hieroglyphs and transliteration, with lemmatization, with POS glossing and with a German translation.
The data comes from the database of the Thesaurus Linguae Aegyptiae, corpus version 19.
This set of Late Egyptian sentences only contains text witnesses classified as "Late Egyptian" in the TLA corpus metadata. Moreover, it contains only fully intact,
unambiguously readable… See the full description on the dataset page: https://huggingface.co/datasets/thesaurus-linguae-aegyptiae/tla-late_egyptian-v19-premium.egyptian-names
Egyptian Names & Onomastic Intelligence Dataset
From 15.88M+ Raw National Records to an Empirical Onomastic and Linguistic Engine
This repository hosts the complete, multi-phase statistical and linguistic dataset powering egy-names, the production onomastic intelligence engine for contemporary Egyptian naming traditions.
Dataset Pipeline Overview
Egyptian names follow an unbroken patronymic lineage chain ($Personal + Father + Grandfather + Ancestor… See the full description on the dataset page: https://huggingface.co/datasets/Abdullah-afify/egyptian-names.Egyptian-Arabic-English-Parallel-Corpus
Egyptian Arabic-English Parallel Corpus
Author: Mohamed Abdalkader · LinkedIn · GitHub
A comprehensive Egyptian Arabic → English parallel corpus covering 1,800 topics from daily Egyptian life. Designed for fine-tuning large language models on Egyptian Arabic dialect translation and generation.
Dataset Structure
egyptian-arabic-english-parallel-corpus/
├── SFT/
│ ├── Train/
│ │ ├── topics/ # 1,800 individual topic JSON files
│ │ └── merged/… See the full description on the dataset page: https://huggingface.co/datasets/Mo-Abdalkader/Egyptian-Arabic-English-Parallel-Corpus.EgyptianWinoGrande
Dataset Card for EgyptianWinoGrande (Arabic and Latin Script)
Dataset Description
Dataset Summary
WinoGrande (Egyptian Arabic) is a coreference resolution benchmark designed to test a model’s ability to resolve pronouns in ambiguous contexts. Each question has two candidate nouns and one target pronoun, translated into Egyptian Arabic.
Supported Tasks
Task Category: Multiple-choice question answering
Task: Selecting the correct answer from a list of… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/EgyptianWinoGrande.EgyptianPIQA
Dataset Card for EgyptianPIQA (Arabic and Latin Script)
Dataset Description
Dataset Summary
EgyptianPIQA evaluates physical commonsense reasoning in Egyptian Arabic. Each sample consists of a practical goal and two possible solutions. The model must choose the more plausible one. Translated using Claude Sonnet 3.5 v2.
Supported Tasks
Task Category: Multiple-choice question answering
Task: Selecting the correct answer from a list of options… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/EgyptianPIQA.egyption-with-emotion-dataset
Egption Text-Audio Dataset With Emotions and Diarization
Creating datasets for TTS and ASR models with emotions and Diarization
In case you want to focus only one speaker , you can fiter based on speaker_role
Source Code
if you want to collect more data from youtube, you can check this link
🙏 Acknowledgements
This project makes use of the forced alignment model and Cohere ASR model provided by:
MahmoudAshraf/mms-300m-1130-forced-aligner
Cohere ASR
Hubert… See the full description on the dataset page: https://huggingface.co/datasets/OmarAhmedSobhy/egyption-with-emotion-dataset.dahab-egyptian-female-tts
Dahab — Egyptian Arabic, single female speaker
134.7 hours across 59,505 clips of Egyptian (Cairene) Arabic from one
female speaker, at 24 kHz mono. 26,741 clips (44.9%) carry diacritized
transcripts. Built for TTS fine-tuning.
Segmented from a single YouTube cooking channel, so the register is
conversational instructional speech throughout.
Structure
The train split is stored in self-contained Parquet shards. Each row contains an audio object with embedded WAV… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/dahab-egyptian-female-tts.EgyptianHellaSwag
Dataset Card for EgyptianHellaSwag
Dataset Summary
EgyptianHellaSwag is a challenging multiple-choice benchmark designed to
evaluate machine reading comprehension and commonsense reasoning in Egyptian
Arabic (Masri). It is a translated version of the HellaSwag train, validation,
and test sets which presents scenarios where models must choose the most
plausible continuation of a passage from four options.
Supported Tasks
Task Category: Multiple-choice question… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/EgyptianHellaSwag.
