CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ehabnegm /100-hour-Egyptian-dataset-single-speaker Masri 100h — Egyptian Arabic Single-Speaker Speech Corpus A 100-hour Egyptian Arabic (مصري) single-narrator speech collection — 15,653 released clips at 24 kHz mono, with aligned transcripts. Egyptian Arabic is the most widely understood Arabic dialect and one of the least served by open speech data. Almost every open Arabic corpus is Modern Standard Arabic (MSA) — a register nobody actually speaks at home. This dataset is built for the opposite: natural, spoken, conversational… See the full description on the dataset page: https://huggingface.co/datasets/ehabnegm/100-hour-Egyptian-dataset-single-speaker.audiotext-to-speech10K<n<100K11 likes4.9k downloads2mo agoHugging Face02HamdiJr /Egyptian_hieroglyphs Egyptian hieroglyphs 𓂀 Hieroglyphs image dataset along with Language Model ! Features This dataset is build from the hieroglyphs found in 10 different pictures from the book "The Pyramid of Unas" (Alexandre Piankoff, 1955). We therefore urge you to have access to this book before using the dataset. The ten different pictures used throughout this dataset are: 3,5,7,9,20,21,22,23,39,41 (numbers represent the numbers used in the book "The pyramid of Unas". Each… See the full description on the dataset page: https://huggingface.co/datasets/HamdiJr/Egyptian_hieroglyphs.image1K<n<10K9 likes1.7k downloads4y agoHugging Face03PatrickBLB /egyptain_cars_images_datasetimage0 likes1.1k downloads4mo agoHugging Face04jason1966 /alexandrepetit881234_egyptian-hieroglyphs Egyptian Hieroglyphs 95 different hieroglyphic symbols for image classification Dataset Info Source: Kaggle Original Size: 10.37 MB Kaggle Downloads: 3,919 Files: 3895 Files README.dataset.txt README.roboflow.txt Mirrored from Kaggle image1K<n<10K0 likes546 downloads6mo agoHugging Face05ISLAM-PO /documents-Egyptian-Arabic Current Hub Validation Status Dataset Server rows: 25,399,945 Dataset Server original/Parquet size: 2,758,228,707 bytes (~2.76 GB) Default Hub configuration currently exposes one column: text The additional configuration names listed in this card are physical source directories and are not all recognized as separate Hub configurations. Keep the default configuration until the dataset is normalized into explicit, tested splits. Egyptian Arabic Mega Corpus (EAMC) —… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/documents-Egyptian-Arabic.translation10M<n<100M2 likes502 downloads2h agoHugging Face06ammarthabet /egyptian-arabic-speechaudion<1K0 likes361 downloads4mo agoHugging Face07MightyStudent /Egyptian-ASR-MGB-3 Egyptian Arabic dialect automatic speech recognition Dataset Summary This dataset was collected, cleaned and adjusted for huggingface hub and ready to be used for whisper finetunning/training. From MGB-3 website: The MGB-3 is using 16 hours multi-genre data collected from different YouTube channels. The 16 hours have been manually transcribed. The chosen Arabic dialect for this year is Egyptian. Given that dialectal Arabic has no orthographic rules, each program has… See the full description on the dataset page: https://huggingface.co/datasets/MightyStudent/Egyptian-ASR-MGB-3.audioautomatic-speech-recognition1K<n<10K23 likes334 downloads2y agoHugging Face08kjhq /Egypt-Stock-Symbols-and-Metadata Egypt Stock Symbols & Company Metadata This dataset contains stock symbols and basic company metadata for all listed companies in Egypt.It is updated weekly if new changes are there. 📊 Dataset Contents The dataset is provided as a CSV file with the following columns: Column Description name Full company name ticker Stock ticker symbol (e.g., AAPL, MSFT) market The exchange/market where the stock is listed sector The primary business sector of the… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/Egypt-Stock-Symbols-and-Metadata.textn<1K0 likes310 downloads1y agoHugging Face09thesaurus-linguae-aegyptiae /tla-Earlier_Egyptian_original-v18-premium Dataset Card for Dataset tla-Earlier_Egyptian_original-v18-premium This data set contains Earlier Egyptian, i.e., ancient Old Egyptian and ancient Middle Egyptian, sentences in hieroglyphs and transliteration, with lemmatization, with POS glossing and with a German translation. This set of original Earlier Egyptian sentences only contains text witnesses from before the start of the New Kingdom (late 16th century BCE). The data comes from the database of the Thesaurus Linguae… See the full description on the dataset page: https://huggingface.co/datasets/thesaurus-linguae-aegyptiae/tla-Earlier_Egyptian_original-v18-premium.texttranslation10K<n<100K7 likes297 downloads2y agoHugging Face10MBZUAI-Paris /EgyptianMMLU Dataset Card for EgyptianMMLU Dataset Description Dataset Summary EgyptianMMLU is a composite benchmark translated into Egyptian Arabic. It combines subsets of ArabicMMLU-egy and the original English MMLU, translated using in-house models and validated by human annotators. The dataset spans 44 subjects and evaluates reasoning, factual knowledge, and world understanding in Egyptian Arabic. Supported Tasks Task Category: Multiple-choice question… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/EgyptianMMLU.text10K<n<100K0 likes278 downloads1y agoHugging Face11justicedao /ipfs_egypt_laws_ir Egypt legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_egypt_laws (revision 0ddf04279b31d1f2200e96d743379bd19aa7e5cb) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of Egypt prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No legal text was invented.… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_egypt_laws_ir.tabulartext-retrieval10K<n<100K0 likes272 downloads2d agoHugging Face12ismaeeelxd /Egyptian-Arabic-Lectures Egyptian Arabic Lectures Dataset The Egyptian Arabic Lectures dataset is a collection of transcribed audio clips (around 30 hours) extracted from educational lectures delivered in Egyptian Arabic (with mixed English technical terms, such as in Physics, IoT and Operating Systems etc.). It is designed to train, evaluate, and fine-tune Automatic Speech Recognition models for the Egyptian dialect, specifically in educational and academic CS contexts. Alongside the audio and text… See the full description on the dataset page: https://huggingface.co/datasets/ismaeeelxd/Egyptian-Arabic-Lectures.audioautomatic-speech-recognition1K<n<10K3 likes246 downloads3mo agoHugging Face13otozz /egyptian_train_setPre-processed Egyptian train partition from the MASC-dataset: Mohammad Al-Fetyani, Muhammad Al-Barham, Gheith Abandah, Adham Alsharkawi, Maha Dawas, August 18, 2021, "MASC: Massive Arabic Speech Corpus", IEEE Dataport, doi: https://dx.doi.org/10.21227/e1qb-jv46. 10K<n<100K0 likes230 downloads2y agoHugging Face14dataflare /egypt-legal-corpus Egyptian Legal Corpus A comprehensive collection of Egyptian legal texts, meticulously extracted and tokenized for Natural Language Processing (NLP) applications, legal research, and AI model training. This corpus provides high-quality Arabic legal content with structured metadata for efficient processing. Dataset Statistics This release provides a foundational legal corpus with strict quality controls: Token Count: 25M+ tokens (25,054,372 tokens) using cl100k_base… See the full description on the dataset page: https://huggingface.co/datasets/dataflare/egypt-legal-corpus.texttext-generation1K<n<10K4 likes230 downloads8mo agoHugging Face15beaunix /Thoth-Sphinx-egyptian-hieroglyphs Thoth-Sphinx — Egyptian Hieroglyphs Multi-Sign Detection Dataset Dataset Summary This dataset provides multi-sign, multi-cartouche annotated images of Middle Egyptian hieroglyphic inscriptions for object detection, in YOLO format. To our knowledge, no other publicly available dataset combines: Multiple signs per image (average ~34 instances/image), rather than isolated single-glyph crops Royal cartouche detection as its own class, with signs annotated inside the… See the full description on the dataset page: https://huggingface.co/datasets/beaunix/Thoth-Sphinx-egyptian-hieroglyphs.imageobject-detection10K<n<100K0 likes212 downloads2mo agoHugging Face16geeeezx /egyptian-arabic-400kaudio100K<n<1M4 likes197 downloads1y agoHugging Face17Kyrillos2001 /Egyptian_Dialect Egyptian Arabic Speech Dataset Dataset Description This dataset contains 2,438 short audio segments in Egyptian Arabic paired with transcriptions. The dataset was created for fine-tuning Automatic Speech Recognition (ASR) models on conversational Egyptian Arabic. Each example contains: WAV audio Egyptian Arabic transcription Dataset Creation Source The audio was collected from publicly available YouTube videos featuring native… See the full description on the dataset page: https://huggingface.co/datasets/Kyrillos2001/Egyptian_Dialect.audioautomatic-speech-recognition1K<n<10K1 likes194 downloads3mo agoHugging Face18phiwi /bbaw_egyptian Dataset Card for "bbaw_egyptian" Dataset Summary This dataset comprises parallel sentences of hieroglyphic encodings, transcription and translation as used in the paper Multi-Task Modeling of Phonographic Languages: Translating Middle Egyptian Hieroglyph. The data triples are extracted from the digital corpus of Egyptian texts compiled by the project "Strukturen und Transformationen des Wortschatzes der ägyptischen Sprache". Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/phiwi/bbaw_egyptian.texttranslation100K<n<1M11 likes192 downloads3y agoHugging Face19OmarMDiab /Egyptian-Handwriting-Dataset Egyptian Handwriting Dataset A dataset of 11k+ handwritten Arabic words from Egyptian writers, extracted and tightly cropped from scanned paper forms. This dataset offers diverse handwriting samples ranging from children to elderly contributors, making it ideal for training robust Arabic handwriting recognition models. Each form contains 6 unique words, resulting in 24 handwritten word images per form. Each word is written four times by the same writer to capture… See the full description on the dataset page: https://huggingface.co/datasets/OmarMDiab/Egyptian-Handwriting-Dataset.imageimage-to-text10K<n<100K4 likes178 downloads9mo agoHugging Face20endomorphosis /ipfs_egypt_laws Egypt Constitution and Official Gazette texts (Presidency / Alamiria) Research snapshot of official national legislation from Presidency constitution PDF + Alamiria official printer TashTxt (الجريدة الرسمية). Not legal advice. The official gazette / authentic source prevails over this corpus. Snapshot Field Value Snapshot date 2026-09-10 Coverage partial-official-fulltext Source Presidency constitution PDF + Alamiria official printer TashTxt… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_egypt_laws.texttext-retrieval1K<n<10K0 likes171 downloads7d agoHugging Face21alielshantory /egypt_cars_and_license_platesimage10K<n<100K1 likes165 downloads4mo agoHugging Face22MBZUAI-Paris /Egyptian-SFT-Mixture 📚 Egyptian SFT Mixture Dataset Overview This dataset contains the Supervised Fine-Tuning samples for Nile-Chat in both Arabic and Latin scripts. Each sample follows the following format: [{"role": "user", "content": user_prompt}, {"role": "assistant", "content": assistant_answer}] Dataset Categories This dataset can be mainly divided into two main categories, Native and Synthetic Egyptian Instruction Datasets. The Native datasets have been collected and… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/Egyptian-SFT-Mixture.text1M<n<10M6 likes154 downloads1y agoHugging Face23egyptparagliders /egypt-paragliders-terrain0 likes145 downloads12d agoHugging Face24torahCodes /Torah_Gnostic_Egypt_India_China_Greece_holy_texts_sources Torah Codes Religion Texts Sources Data Tree ── arabs │   ├── astrological_stelar_magic.txt │   └── Holy-Quran-English.txt ├── ars │   ├── ars_magna_ramon_llull.txt │   └── lemegeton_book_solomon.txt ├── asimov │   ├── foundation.txt │   └── prelude_to_foundation.txt ├── budist │   ├── bardo_todhol_book_of_deads_tibet_libro_tibetano_de_los_muertos.txt │   ├── rig_veda.txt │   └── TheTeachingofBuddha.txt ├── cathars ├── china │   ├── arte_de_la_guerra_art_of_war.txt │… See the full description on the dataset page: https://huggingface.co/datasets/torahCodes/Torah_Gnostic_Egypt_India_China_Greece_holy_texts_sources.textn<1K5 likes142 downloads2y agoHugging Face25Rabe3 /egyptian-arabic-tts-diacritized Egyptian Arabic TTS Corpus (Diacritized) 97,163 utterances / ~334 hours of Egyptian Arabic speech at 24 kHz, with diacritized transcripts — the short vowels that Arabic script does not write. Why diacritics Arabic is an abjad: short vowels are unwritten, so كتب may be kataba, kutiba, or kutub. A TTS model with no Arabic pretraining cannot infer which, and guesses — which native listeners hear as a foreign accent with constant mispronunciation. This is invisible to… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/egyptian-arabic-tts-diacritized.tabulartext-to-speech10K<n<100K0 likes140 downloads1mo agoHugging Face26Prickly-Labs /1.9M-Egyptian-Corpus 1.92M Egyptian Arabic Corpus 🇪🇬 — Prickly Labs A corpus of 1.92 million Egyptian Arabic samples combining synthetic generation and real web data, built for continued pretraining and dialect adaptation of Arabic language models. Designed to reflect how Egyptians actually speak — not textbook MSA. ⚠️ This dataset contains uncensored, informal Arabic including sarcasm, humor, slang, and profanity. Use with care for public-facing applications. 📌 Overview… See the full description on the dataset page: https://huggingface.co/datasets/Prickly-Labs/1.9M-Egyptian-Corpus.texttext-generation1M<n<10M2 likes137 downloads5mo agoHugging Face27Omartificial-Intelligence-Space /FineWeb2-Egyptian-Arabic FineWeb2 Egyptian Arabic 🇪🇬 This is the Egyptian Arabic Portion of The FineWeb2 Dataset. 🇪🇬 This dataset contains a rich collection of text in Egyptian Arabic (ISO 639-3: arz), a widely spoken dialect within the Afro-Asiatic language family. 🇪🇬 With over 439 million words and 1.4 million documents, it serves as a valuable resource for NLP development and linguistic research focused on Egyptian Arabic. Purpose of This Repository This repository provides easy… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/FineWeb2-Egyptian-Arabic.text10M<n<100M2 likes127 downloads2y agoHugging Face28thesaurus-linguae-aegyptiae /tla-late_egyptian-v19-premium Dataset Card for Dataset tla-Late_Egyptian-v19-premium This data set contains Late Egyptian sentences in hieroglyphs and transliteration, with lemmatization, with POS glossing and with a German translation. The data comes from the database of the Thesaurus Linguae Aegyptiae, corpus version 19. This set of Late Egyptian sentences only contains text witnesses classified as "Late Egyptian" in the TLA corpus metadata. Moreover, it contains only fully intact, unambiguously readable… See the full description on the dataset page: https://huggingface.co/datasets/thesaurus-linguae-aegyptiae/tla-late_egyptian-v19-premium.texttranslation1K<n<10K4 likes123 downloads2y agoHugging Face29Abdullah-afify /egyptian-names Egyptian Names & Onomastic Intelligence Dataset From 15.88M+ Raw National Records to an Empirical Onomastic and Linguistic Engine This repository hosts the complete, multi-phase statistical and linguistic dataset powering egy-names, the production onomastic intelligence engine for contemporary Egyptian naming traditions. Dataset Pipeline Overview Egyptian names follow an unbroken patronymic lineage chain ($Personal + Father + Grandfather + Ancestor… See the full description on the dataset page: https://huggingface.co/datasets/Abdullah-afify/egyptian-names.tabulartext-classification1M<n<10M0 likes122 downloads25d agoHugging Face30Mo-Abdalkader /Egyptian-Arabic-English-Parallel-Corpus Egyptian Arabic-English Parallel Corpus Author: Mohamed Abdalkader · LinkedIn · GitHub A comprehensive Egyptian Arabic → English parallel corpus covering 1,800 topics from daily Egyptian life. Designed for fine-tuning large language models on Egyptian Arabic dialect translation and generation. Dataset Structure egyptian-arabic-english-parallel-corpus/ ├── SFT/ │ ├── Train/ │ │ ├── topics/ # 1,800 individual topic JSON files │ │ └── merged/… See the full description on the dataset page: https://huggingface.co/datasets/Mo-Abdalkader/Egyptian-Arabic-English-Parallel-Corpus.text100K<n<1M1 likes119 downloads6d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.