datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Egypt-Stock-Symbols-and-Metadata
Egypt Stock Symbols & Company Metadata
This dataset contains stock symbols and basic company metadata for all listed companies in Egypt.It is updated weekly if new changes are there.
📊 Dataset Contents
The dataset is provided as a CSV file with the following columns:
Column
Description
name
Full company name
ticker
Stock ticker symbol (e.g., AAPL, MSFT)
market
The exchange/market where the stock is listed
sector
The primary business sector of the… See the full description on the dataset page: https://huggingface.co/datasets/kjhq/Egypt-Stock-Symbols-and-Metadata.General_Facts_in_English_Arabic_Egyptian_Arabic
🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized)
The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.egyptian-2-arabic
Egyptian Arabic Slang ↔ Formal Arabic Dataset
Dataset Overview
This dataset contains 18,250 parallel sentence pairs mapping Egyptian Arabic dialect/slang to Modern Standard Arabic (MSA).
It is designed for Arabic NLP research on dialect normalization, slang understanding, and machine translation between informal and formal Arabic.
Task Objectives
This dataset can be used for:
Dialect-to-Modern Standard Arabic translation
Arabic text normalization
Dialect… See the full description on the dataset page: https://huggingface.co/datasets/AdhamAshraf/egyptian-2-arabic.UltrasTexts_EgyptianIndependent
Texts about Ultras from the Egyptian Independent, 2009-2020
A curated corpus of Egypt Independent articles on “Ultras” football fan groups, with publication dates, URLs, and full‑text content.
Dataset Summary
Source: Egypt Independent search results for the term “ultras+” (39 pages of results, 9 articles per page).
Period Covered: Articles published between 2009 and 2020.
Total Records: 378 raw articles were scraped; after filtering out entries with fewer than 20… See the full description on the dataset page: https://huggingface.co/datasets/cjerzak/UltrasTexts_EgyptianIndependent.arabic_egypt_english_world_facts
🌍 Version (v2.0) World Facts in English, Arabic & Egyptian Arabic (Categorized)
The World Facts General Knowledge Dataset (v2.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata:… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/arabic_egypt_english_world_facts.za-egypt-insurance-claims-sample
za-egypt-insurance-claims-sample
A small sample of de-identified insurance claim records from South Africa (ZA) and Egypt (EG).
Dataset Summary
This dataset contains a sample of insurance claim records covering the South African and
Egyptian markets. It is intended as a lightweight example for exploring claims trends,
customer segmentation, and fraud-risk analysis. All records have been de-identified.
Total records: 80
Regions: South Africa (40), Egypt (40)… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/za-egypt-insurance-claims-sample.Egyption_2_Englishorganic-egyptian-arabic-dialect-dataset
Organic Egyptian Arabic Dialect Dataset (Sample)
This repository contains a limited sample subset of an organic Egyptian Arabic (Masri) dataset.
Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application. Therefore, it… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-egyptian-arabic-dialect-dataset.arabic-egyptian-sample
4FACTORS Arabic — Egyptian Q&A Sample
Conversational question–answer pairs in spoken Egyptian Arabic, written by a
first-language Egyptian speaker. A 50-item demonstration sample, with English
glosses, released under CC BY-NC 4.0.
This is the Egyptian variety in the 4FACTORS Arabic sample set, alongside the
Palestinian Levantine
and Modern Standard Arabic (MSA) sets.
What this is
Fifty short question–answer exchanges of the kind that come up in everyday life —… See the full description on the dataset page: https://huggingface.co/datasets/4factors/arabic-egyptian-sample.Egyptian-Movies-Datasetegyptian-nutrition-datasetglobalopinion_egyptegyptian_car_puchasing_poweregyptian-fusha-27k
