Egyptian
Datasets
All datasets matching “Egyptian”100-hour-Egyptian-dataset-single-speaker
Masri 100h — Egyptian Arabic Single-Speaker Speech Corpus
A 100-hour Egyptian Arabic (مصري) single-narrator speech collection — 15,653 released clips at 24 kHz mono, with aligned transcripts.
Egyptian Arabic is the most widely understood Arabic dialect and one of the least served by open speech data.
Almost every open Arabic corpus is Modern Standard Arabic (MSA) — a register nobody actually speaks at home.
This dataset is built for the opposite: natural, spoken, conversational… See the full description on the dataset page: https://huggingface.co/datasets/ehabnegm/100-hour-Egyptian-dataset-single-speaker.Egyptian_hieroglyphs
Egyptian hieroglyphs 𓂀
Hieroglyphs image dataset along with Language Model !
Features
This dataset is build from the hieroglyphs found in 10 different pictures from the book "The Pyramid of Unas" (Alexandre Piankoff, 1955). We therefore urge you to have access to this book before using the dataset.
The ten different pictures used throughout this dataset are: 3,5,7,9,20,21,22,23,39,41 (numbers represent the numbers used in the book "The pyramid of Unas".
Each… See the full description on the dataset page: https://huggingface.co/datasets/HamdiJr/Egyptian_hieroglyphs.documents-Egyptian-Arabic
Egyptian Arabic Mega Corpus (EAMC) — 25M Unified Egyptian Dialect Dataset
The Largest Unified Open Corpus for Egyptian Arabic (Masri / arz)
25.5M Samples | 2.66 GB (Parquet) | 9 Configs | Apache 2.0 | Ready-to-train
Comprehensive coverage: Raw Text · Wikipedia · Conversations · Speech (Whisper) · Parallel Translation (EN↔EGY) · Trilingual QA · Wikipedia Quality Classification · Fake Review / Spam Detection
Dataset Summary
Egyptian Arabic Mega Corpus… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/documents-Egyptian-Arabic.alexandrepetit881234_egyptian-hieroglyphs
Egyptian Hieroglyphs
95 different hieroglyphic symbols for image classification
Dataset Info
Source: Kaggle
Original Size: 10.37 MB
Kaggle Downloads: 3,919
Files: 3895
Files
README.dataset.txt
README.roboflow.txt
Mirrored from Kaggle
egyptian-arabic-speechEgyptian-ASR-MGB-3
Egyptian Arabic dialect automatic speech recognition
Dataset Summary
This dataset was collected, cleaned and adjusted for huggingface hub and ready to be used for whisper finetunning/training.
From MGB-3 website:
The MGB-3 is using 16 hours multi-genre data collected from different YouTube channels. The 16 hours have been manually transcribed.
The chosen Arabic dialect for this year is Egyptian.
Given that dialectal Arabic has no orthographic rules, each program has… See the full description on the dataset page: https://huggingface.co/datasets/MightyStudent/Egyptian-ASR-MGB-3.
