CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dataflare /arabic-dialect-corpus Arabic Dialect Corpus A comprehensive collection of Arabic dialectal text, standardized for Natural Language Processing (NLP) model training, evaluation, and linguistic analysis. This corpus has been meticulously processed to ensure high-quality tokenization and consistent metadata. Dataset Statistics Metric Value Total Records 127,180 Total Tokens 5,802,324 Average Tokens per Record 45.62 Dialect Categories 5 Changelog… See the full description on the dataset page: https://huggingface.co/datasets/dataflare/arabic-dialect-corpus.tabulartext-generation100K<n<1M1 likes158 downloads8mo agoHugging Face02fr3on /arabic-dialect-corpus 🇪🇬🇸🇦 Arabic Dialect Corpus (Egyptian & Saudi) Dataset Description This dataset contains 150K+ natural, informal Arabic text samples scraped from high-engagement YouTube discussions. It specifically targets Egyptian (EG) and Saudi (SA) dialects, filling a critical gap in resources for training LLMs on colloquial Arabic (Ammiya) rather than just Modern Standard Arabic (MSA). Languages Primary Dialects: Egyptian Arabic (EG) - Cairene and regional Egyptian… See the full description on the dataset page: https://huggingface.co/datasets/fr3on/arabic-dialect-corpus.texttext-generation1M<n<10M1 likes89 downloads8mo agoHugging Face03HeshamHaroon /arabic-dialect-dpo Arabic Dialect DPO Dataset - Egyptian & Saudi The first large-scale Arabic dialect preference dataset for DPO/ORPO/GRPO alignment training. Contains 22,538 preference triples across two major Arabic dialects: Egyptian (Masry) and Saudi (Najdi). Dataset Summary Config Dialect Rows Language Code egyptian Egyptian Arabic (مصري) 11,038 ar-EG saudi Saudi Arabic (سعودي نجدي) 11,500 ar-SA Total 22,538 Usage from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-dialect-dpo.texttext-generation10K<n<100K0 likes37 downloads7mo agoHugging Face04CNTXTAI0 /arabic_dialects_question_and_answerData Content The file provided: Q/A Reasoning dataset contains the following columns: ID # : Denotes the reference ID for: a. Question b. Answer to the question c. Hint d. Reasoning e. Word count for items a to d above Dialects: Contains the following dialects in separate columns: a. English b. MSA c. Emirati d. Egyptian e. Levantine Syria f. Levantine Jordan g. Levantine Palestine h. Levantine Lebanon Data Generation Process The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.tabularquestion-answeringn<1K6 likes30 downloads2y agoHugging Face05yrrhall /arabic-dialect-corpus 🇪🇬🇸🇦 Arabic Dialect Corpus (Egyptian & Saudi) Dataset Description This dataset contains 150K+ natural, informal Arabic text samples scraped from high-engagement YouTube discussions. It specifically targets Egyptian (EG) and Saudi (SA) dialects, filling a critical gap in resources for training LLMs on colloquial Arabic (Ammiya) rather than just Modern Standard Arabic (MSA). Languages Primary Dialects: Egyptian Arabic (EG) - Cairene and regional… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/arabic-dialect-corpus.texttext-generation1M<n<10M0 likes29 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.