ebubekr53/organic-levantine-arabic-dialect-dataset
Organic Levantine Arabic Dialect Dataset (Sample) This repository contains a limited sample subset of an organic, multi-country Levantine Arabic (Shami) dataset. Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application.… See the full description on the dataset page: https://huggingface.co/datasets/ebubekr53/organic-levantine-arabic-dialect-dataset.
Organic Levantine Arabic Dialect Dataset (Sample)
This repository contains a limited sample subset of an organic, multi-country Levantine Arabic (Shami) dataset.
Unlike standard web-scraped corpora or synthetic datasets, this data is generated entirely from the daily, spontaneous text and voice translation queries of native speakers translating between different Arabic dialects, as well as between Arabic dialects and other languages through our active mobile application. Therefore, it accurately captures genuine street language, daily idioms, regional spelling variations, and authentic conversational tones.
Dataset Structure
The dataset contains the following columns:
id: A unique identifier for each text record.source_dialect: The general dialect family (strictlylevantinefor this dataset).source_country: The specific country of origin for the speaker, representing the sub-dialects of the Levant region:jordan(Jordanian)lebanon(Lebanese)palestine(Palestinian)syria(Syrian)source_text: The raw, spontaneous text containing genuine conversational phrasing, slang, regional idioms, and occasional Romanized Arabic (Arabizi).
How the Data Was Collected
- Source: Live mobile application designed for translation between different Arabic dialects, between dialects and Modern Standard Arabic (Fusha), as well as between Arabic dialects and other foreign languages.
How to Access the Full Dataset & Continuous Data Pipeline
The data provided in this repository serves strictly as a limited sample. The complete, comprehensive dataset is significantly larger and is not publicly hosted here.
Furthermore, our live mobile application operates as a continuous, scalable data pipeline capable of generating larger volumes of custom, native dialect translation data tailored to your specific project requirements.
To request access to the full dataset, license larger subsets, or establish a continuous data pipeline partnership for your AI/LLM projects, please contact the sole developer and owner:
- Contact Person: Ebubekir Siddik Kul
- Email: ebubekirkul@marun.edu.tr / ebubekirkul6153@gmail.com
- Institution: Marmara University
