CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ehabnegm /100-hour-Egyptian-dataset-single-speaker Masri 100h — Egyptian Arabic Single-Speaker Speech Corpus A 100-hour Egyptian Arabic (مصري) single-narrator speech collection — 15,653 released clips at 24 kHz mono, with aligned transcripts. Egyptian Arabic is the most widely understood Arabic dialect and one of the least served by open speech data. Almost every open Arabic corpus is Modern Standard Arabic (MSA) — a register nobody actually speaks at home. This dataset is built for the opposite: natural, spoken, conversational… See the full description on the dataset page: https://huggingface.co/datasets/ehabnegm/100-hour-Egyptian-dataset-single-speaker.audiotext-to-speech10K<n<100K11 likes4.9k downloads2mo agoHugging Face02HamdiJr /Egyptian_hieroglyphs Egyptian hieroglyphs 𓂀 Hieroglyphs image dataset along with Language Model ! Features This dataset is build from the hieroglyphs found in 10 different pictures from the book "The Pyramid of Unas" (Alexandre Piankoff, 1955). We therefore urge you to have access to this book before using the dataset. The ten different pictures used throughout this dataset are: 3,5,7,9,20,21,22,23,39,41 (numbers represent the numbers used in the book "The pyramid of Unas". Each… See the full description on the dataset page: https://huggingface.co/datasets/HamdiJr/Egyptian_hieroglyphs.image1K<n<10K9 likes1.7k downloads4y agoHugging Face03ammarthabet /egyptian-arabic-speechaudion<1K0 likes361 downloads4mo agoHugging Face04MightyStudent /Egyptian-ASR-MGB-3 Egyptian Arabic dialect automatic speech recognition Dataset Summary This dataset was collected, cleaned and adjusted for huggingface hub and ready to be used for whisper finetunning/training. From MGB-3 website: The MGB-3 is using 16 hours multi-genre data collected from different YouTube channels. The 16 hours have been manually transcribed. The chosen Arabic dialect for this year is Egyptian. Given that dialectal Arabic has no orthographic rules, each program has… See the full description on the dataset page: https://huggingface.co/datasets/MightyStudent/Egyptian-ASR-MGB-3.audioautomatic-speech-recognition1K<n<10K23 likes334 downloads2y agoHugging Face05kjhq /Egypt-Stock-Symbols-and-Metadata Egypt Stock Symbols & Company Metadata This dataset contains stock symbols and basic company metadata for all listed companies in Egypt.It is updated weekly if new changes are there. 📊 Dataset Contents The dataset is provided as a CSV file with the following columns: Column Description name Full company name ticker Stock ticker symbol (e.g., AAPL, MSFT) market The exchange/market where the stock is listed sector The primary business sector of the… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/Egypt-Stock-Symbols-and-Metadata.textn<1K0 likes310 downloads1y agoHugging Face06thesaurus-linguae-aegyptiae /tla-Earlier_Egyptian_original-v18-premium Dataset Card for Dataset tla-Earlier_Egyptian_original-v18-premium This data set contains Earlier Egyptian, i.e., ancient Old Egyptian and ancient Middle Egyptian, sentences in hieroglyphs and transliteration, with lemmatization, with POS glossing and with a German translation. This set of original Earlier Egyptian sentences only contains text witnesses from before the start of the New Kingdom (late 16th century BCE). The data comes from the database of the Thesaurus Linguae… See the full description on the dataset page: https://huggingface.co/datasets/thesaurus-linguae-aegyptiae/tla-Earlier_Egyptian_original-v18-premium.texttranslation10K<n<100K7 likes297 downloads2y agoHugging Face07MBZUAI-Paris /EgyptianMMLU Dataset Card for EgyptianMMLU Dataset Description Dataset Summary EgyptianMMLU is a composite benchmark translated into Egyptian Arabic. It combines subsets of ArabicMMLU-egy and the original English MMLU, translated using in-house models and validated by human annotators. The dataset spans 44 subjects and evaluates reasoning, factual knowledge, and world understanding in Egyptian Arabic. Supported Tasks Task Category: Multiple-choice question… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/EgyptianMMLU.text10K<n<100K0 likes278 downloads1y agoHugging Face08justicedao /ipfs_egypt_laws_ir Egypt legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_egypt_laws (revision 0ddf04279b31d1f2200e96d743379bd19aa7e5cb) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of Egypt prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No legal text was invented.… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_egypt_laws_ir.tabulartext-retrieval10K<n<100K0 likes272 downloads2d agoHugging Face09ismaeeelxd /Egyptian-Arabic-Lectures Egyptian Arabic Lectures Dataset The Egyptian Arabic Lectures dataset is a collection of transcribed audio clips (around 30 hours) extracted from educational lectures delivered in Egyptian Arabic (with mixed English technical terms, such as in Physics, IoT and Operating Systems etc.). It is designed to train, evaluate, and fine-tune Automatic Speech Recognition models for the Egyptian dialect, specifically in educational and academic CS contexts. Alongside the audio and text… See the full description on the dataset page: https://huggingface.co/datasets/ismaeeelxd/Egyptian-Arabic-Lectures.audioautomatic-speech-recognition1K<n<10K3 likes246 downloads3mo agoHugging Face10dataflare /egypt-legal-corpus Egyptian Legal Corpus A comprehensive collection of Egyptian legal texts, meticulously extracted and tokenized for Natural Language Processing (NLP) applications, legal research, and AI model training. This corpus provides high-quality Arabic legal content with structured metadata for efficient processing. Dataset Statistics This release provides a foundational legal corpus with strict quality controls: Token Count: 25M+ tokens (25,054,372 tokens) using cl100k_base… See the full description on the dataset page: https://huggingface.co/datasets/dataflare/egypt-legal-corpus.texttext-generation1K<n<10K4 likes230 downloads8mo agoHugging Face11beaunix /Thoth-Sphinx-egyptian-hieroglyphs Thoth-Sphinx — Egyptian Hieroglyphs Multi-Sign Detection Dataset Dataset Summary This dataset provides multi-sign, multi-cartouche annotated images of Middle Egyptian hieroglyphic inscriptions for object detection, in YOLO format. To our knowledge, no other publicly available dataset combines: Multiple signs per image (average ~34 instances/image), rather than isolated single-glyph crops Royal cartouche detection as its own class, with signs annotated inside the… See the full description on the dataset page: https://huggingface.co/datasets/beaunix/Thoth-Sphinx-egyptian-hieroglyphs.imageobject-detection10K<n<100K0 likes212 downloads2mo agoHugging Face12geeeezx /egyptian-arabic-400kaudio100K<n<1M4 likes197 downloads1y agoHugging Face13Kyrillos2001 /Egyptian_Dialect Egyptian Arabic Speech Dataset Dataset Description This dataset contains 2,438 short audio segments in Egyptian Arabic paired with transcriptions. The dataset was created for fine-tuning Automatic Speech Recognition (ASR) models on conversational Egyptian Arabic. Each example contains: WAV audio Egyptian Arabic transcription Dataset Creation Source The audio was collected from publicly available YouTube videos featuring native… See the full description on the dataset page: https://huggingface.co/datasets/Kyrillos2001/Egyptian_Dialect.audioautomatic-speech-recognition1K<n<10K1 likes194 downloads3mo agoHugging Face14phiwi /bbaw_egyptian Dataset Card for "bbaw_egyptian" Dataset Summary This dataset comprises parallel sentences of hieroglyphic encodings, transcription and translation as used in the paper Multi-Task Modeling of Phonographic Languages: Translating Middle Egyptian Hieroglyph. The data triples are extracted from the digital corpus of Egyptian texts compiled by the project "Strukturen und Transformationen des Wortschatzes der ägyptischen Sprache". Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/phiwi/bbaw_egyptian.texttranslation100K<n<1M11 likes192 downloads3y agoHugging Face15OmarMDiab /Egyptian-Handwriting-Dataset Egyptian Handwriting Dataset A dataset of 11k+ handwritten Arabic words from Egyptian writers, extracted and tightly cropped from scanned paper forms. This dataset offers diverse handwriting samples ranging from children to elderly contributors, making it ideal for training robust Arabic handwriting recognition models. Each form contains 6 unique words, resulting in 24 handwritten word images per form. Each word is written four times by the same writer to capture… See the full description on the dataset page: https://huggingface.co/datasets/OmarMDiab/Egyptian-Handwriting-Dataset.imageimage-to-text10K<n<100K4 likes178 downloads9mo agoHugging Face16endomorphosis /ipfs_egypt_laws Egypt Constitution and Official Gazette texts (Presidency / Alamiria) Research snapshot of official national legislation from Presidency constitution PDF + Alamiria official printer TashTxt (الجريدة الرسمية). Not legal advice. The official gazette / authentic source prevails over this corpus. Snapshot Field Value Snapshot date 2026-09-10 Coverage partial-official-fulltext Source Presidency constitution PDF + Alamiria official printer TashTxt… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_egypt_laws.texttext-retrieval1K<n<10K0 likes171 downloads7d agoHugging Face17alielshantory /egypt_cars_and_license_platesimage10K<n<100K1 likes165 downloads4mo agoHugging Face18MBZUAI-Paris /Egyptian-SFT-Mixture 📚 Egyptian SFT Mixture Dataset Overview This dataset contains the Supervised Fine-Tuning samples for Nile-Chat in both Arabic and Latin scripts. Each sample follows the following format: [{"role": "user", "content": user_prompt}, {"role": "assistant", "content": assistant_answer}] Dataset Categories This dataset can be mainly divided into two main categories, Native and Synthetic Egyptian Instruction Datasets. The Native datasets have been collected and… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/Egyptian-SFT-Mixture.text1M<n<10M6 likes154 downloads1y agoHugging Face19torahCodes /Torah_Gnostic_Egypt_India_China_Greece_holy_texts_sources Torah Codes Religion Texts Sources Data Tree ── arabs │   ├── astrological_stelar_magic.txt │   └── Holy-Quran-English.txt ├── ars │   ├── ars_magna_ramon_llull.txt │   └── lemegeton_book_solomon.txt ├── asimov │   ├── foundation.txt │   └── prelude_to_foundation.txt ├── budist │   ├── bardo_todhol_book_of_deads_tibet_libro_tibetano_de_los_muertos.txt │   ├── rig_veda.txt │   └── TheTeachingofBuddha.txt ├── cathars ├── china │   ├── arte_de_la_guerra_art_of_war.txt │… See the full description on the dataset page: https://huggingface.co/datasets/torahCodes/Torah_Gnostic_Egypt_India_China_Greece_holy_texts_sources.textn<1K5 likes142 downloads2y agoHugging Face20Rabe3 /egyptian-arabic-tts-diacritized Egyptian Arabic TTS Corpus (Diacritized) 97,163 utterances / ~334 hours of Egyptian Arabic speech at 24 kHz, with diacritized transcripts — the short vowels that Arabic script does not write. Why diacritics Arabic is an abjad: short vowels are unwritten, so كتب may be kataba, kutiba, or kutub. A TTS model with no Arabic pretraining cannot infer which, and guesses — which native listeners hear as a foreign accent with constant mispronunciation. This is invisible to… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/egyptian-arabic-tts-diacritized.tabulartext-to-speech10K<n<100K0 likes140 downloads1mo agoHugging Face21Prickly-Labs /1.9M-Egyptian-Corpus 1.92M Egyptian Arabic Corpus 🇪🇬 — Prickly Labs A corpus of 1.92 million Egyptian Arabic samples combining synthetic generation and real web data, built for continued pretraining and dialect adaptation of Arabic language models. Designed to reflect how Egyptians actually speak — not textbook MSA. ⚠️ This dataset contains uncensored, informal Arabic including sarcasm, humor, slang, and profanity. Use with care for public-facing applications. 📌 Overview… See the full description on the dataset page: https://huggingface.co/datasets/Prickly-Labs/1.9M-Egyptian-Corpus.texttext-generation1M<n<10M2 likes137 downloads5mo agoHugging Face22Omartificial-Intelligence-Space /FineWeb2-Egyptian-Arabic FineWeb2 Egyptian Arabic 🇪🇬 This is the Egyptian Arabic Portion of The FineWeb2 Dataset. 🇪🇬 This dataset contains a rich collection of text in Egyptian Arabic (ISO 639-3: arz), a widely spoken dialect within the Afro-Asiatic language family. 🇪🇬 With over 439 million words and 1.4 million documents, it serves as a valuable resource for NLP development and linguistic research focused on Egyptian Arabic. Purpose of This Repository This repository provides easy… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/FineWeb2-Egyptian-Arabic.text10M<n<100M2 likes127 downloads2y agoHugging Face23thesaurus-linguae-aegyptiae /tla-late_egyptian-v19-premium Dataset Card for Dataset tla-Late_Egyptian-v19-premium This data set contains Late Egyptian sentences in hieroglyphs and transliteration, with lemmatization, with POS glossing and with a German translation. The data comes from the database of the Thesaurus Linguae Aegyptiae, corpus version 19. This set of Late Egyptian sentences only contains text witnesses classified as "Late Egyptian" in the TLA corpus metadata. Moreover, it contains only fully intact, unambiguously readable… See the full description on the dataset page: https://huggingface.co/datasets/thesaurus-linguae-aegyptiae/tla-late_egyptian-v19-premium.texttranslation1K<n<10K4 likes123 downloads2y agoHugging Face24Abdullah-afify /egyptian-names Egyptian Names & Onomastic Intelligence Dataset From 15.88M+ Raw National Records to an Empirical Onomastic and Linguistic Engine This repository hosts the complete, multi-phase statistical and linguistic dataset powering egy-names, the production onomastic intelligence engine for contemporary Egyptian naming traditions. Dataset Pipeline Overview Egyptian names follow an unbroken patronymic lineage chain ($Personal + Father + Grandfather + Ancestor… See the full description on the dataset page: https://huggingface.co/datasets/Abdullah-afify/egyptian-names.tabulartext-classification1M<n<10M0 likes122 downloads25d agoHugging Face25Mo-Abdalkader /Egyptian-Arabic-English-Parallel-Corpus Egyptian Arabic-English Parallel Corpus Author: Mohamed Abdalkader · LinkedIn · GitHub A comprehensive Egyptian Arabic → English parallel corpus covering 1,800 topics from daily Egyptian life. Designed for fine-tuning large language models on Egyptian Arabic dialect translation and generation. Dataset Structure egyptian-arabic-english-parallel-corpus/ ├── SFT/ │ ├── Train/ │ │ ├── topics/ # 1,800 individual topic JSON files │ │ └── merged/… See the full description on the dataset page: https://huggingface.co/datasets/Mo-Abdalkader/Egyptian-Arabic-English-Parallel-Corpus.text100K<n<1M1 likes119 downloads6d agoHugging Face26MBZUAI-Paris /EgyptianWinoGrande Dataset Card for EgyptianWinoGrande (Arabic and Latin Script) Dataset Description Dataset Summary WinoGrande (Egyptian Arabic) is a coreference resolution benchmark designed to test a model’s ability to resolve pronouns in ambiguous contexts. Each question has two candidate nouns and one target pronoun, translated into Egyptian Arabic. Supported Tasks Task Category: Multiple-choice question answering Task: Selecting the correct answer from a list of… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/EgyptianWinoGrande.text10K<n<100K0 likes112 downloads1y agoHugging Face27MBZUAI-Paris /EgyptianPIQA Dataset Card for EgyptianPIQA (Arabic and Latin Script) Dataset Description Dataset Summary EgyptianPIQA evaluates physical commonsense reasoning in Egyptian Arabic. Each sample consists of a practical goal and two possible solutions. The model must choose the more plausible one. Translated using Claude Sonnet 3.5 v2. Supported Tasks Task Category: Multiple-choice question answering Task: Selecting the correct answer from a list of options… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/EgyptianPIQA.text10K<n<100K0 likes107 downloads1y agoHugging Face28OmarAhmedSobhy /egyption-with-emotion-dataset Egption Text-Audio Dataset With Emotions and Diarization Creating datasets for TTS and ASR models with emotions and Diarization In case you want to focus only one speaker , you can fiter based on speaker_role Source Code if you want to collect more data from youtube, you can check this link 🙏 Acknowledgements This project makes use of the forced alignment model and Cohere ASR model provided by: MahmoudAshraf/mms-300m-1130-forced-aligner Cohere ASR Hubert… See the full description on the dataset page: https://huggingface.co/datasets/OmarAhmedSobhy/egyption-with-emotion-dataset.audioautomatic-speech-recognition1K<n<10K4 likes106 downloads5mo agoHugging Face29Rabe3 /dahab-egyptian-female-tts Dahab — Egyptian Arabic, single female speaker 134.7 hours across 59,505 clips of Egyptian (Cairene) Arabic from one female speaker, at 24 kHz mono. 26,741 clips (44.9%) carry diacritized transcripts. Built for TTS fine-tuning. Segmented from a single YouTube cooking channel, so the register is conversational instructional speech throughout. Structure The train split is stored in self-contained Parquet shards. Each row contains an audio object with embedded WAV… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/dahab-egyptian-female-tts.audiotext-to-speech10K<n<100K0 likes100 downloads24d agoHugging Face30MBZUAI-Paris /EgyptianHellaSwag Dataset Card for EgyptianHellaSwag Dataset Summary EgyptianHellaSwag is a challenging multiple-choice benchmark designed to evaluate machine reading comprehension and commonsense reasoning in Egyptian Arabic (Masri). It is a translated version of the HellaSwag train, validation, and test sets which presents scenarios where models must choose the most plausible continuation of a passage from four options. Supported Tasks Task Category: Multiple-choice question… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-Paris/EgyptianHellaSwag.textquestion-answering100K<n<1M1 likes96 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.