CoolFace
Datasetpublic

ebubekr53/organic-egyptian-arabic-dialect-dataset

Organic Egyptian Arabic Dialect Dataset (Sample) This repository contains a limited sample subset of an organic Egyptian Arabic (Masri) dataset. Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application. Therefore, it… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-egyptian-arabic-dialect-dataset.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
1likes12downloads
Dataset Card

Organic Egyptian Arabic Dialect Dataset (Sample)

This repository contains a limited sample subset of an organic Egyptian Arabic (Masri) dataset.

Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application. Therefore, it accurately captures genuine street language, daily idioms, regional spelling variations, and authentic conversational tones.

Dataset Structure

The dataset contains the following columns:

  • —id: A unique identifier for each text record.
  • —source_dialect: The general dialect family (strictly egyptian for this dataset).
  • —source_country: The country of origin for the speaker:
  • —egypt (Egyptian)
  • —source_text: The raw, spontaneous text containing genuine conversational phrasing, slang (such as ياستا, قشطه, يا نهار ابيض), regional idioms, and occasional Romanized Arabic (Arabizi).

How the Data Was Collected

  • —Source: Live mobile application designed for translation between different Arabic dialects, between dialects and Modern Standard Arabic (Fusha), as well as between Arabic dialects and other foreign languages.

How to Access the Full Dataset & Continuous Data Pipeline

The data provided in this repository serves strictly as a limited sample. The complete, comprehensive dataset is significantly larger and is not publicly hosted here.

Furthermore, our live mobile application operates as a continuous, scalable data pipeline capable of generating larger volumes of custom, native dialect translation data tailored to your specific project requirements.

To request access to the full dataset, license larger subsets, or establish a continuous data pipeline partnership for your AI/LLM projects, please contact the sole developer and owner:

  • —Contact Person: Ebubekir Siddik Kul
  • —Email: ebubekirkul@marun.edu.tr / ebubekirkul6153@gmail.com
  • —Institution: Marmara University