CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MohamedRashad /arabic-books Arabic Books Dataset Summary The arabic-books dataset contains 8,500 rows of text, each representing the full text of a single Arabic book. These texts were extracted using the arabic-large-nougat model, showcasing the model’s capabilities in Arabic OCR and text extraction. The dataset spans a total of 1.1 billion tokens, calculated using the GPT-4 tokenizer. This dataset is a testimony to the quality of the Arabic Nougat models and their effectiveness in extracting… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-books.texttext-generation1K<n<10K3 likes33k downloads2y agoHugging Face02M-A-D /Mixed-Arabic-Datasets-Repo Dataset Card for "Mixed Arabic Datasets (MAD) Corpus" The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts Dataset Description The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo.tabulartext-classification100M<n<1B38 likes2.2k downloads3y agoHugging Face03ClusterlabAi /101_billion_arabic_words_dataset 101 Billion Arabic Words Dataset Updates Maintenance Status: Actively Maintained Update Frequency: Weekly updates to refine data quality and expand coverage. Upcoming Version More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management. Dataset Details The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/ClusterlabAi/101_billion_arabic_words_dataset.texttext-generation10M<n<100M73 likes1.8k downloads2y agoHugging Face04AdaMLLab /smolkalam-arabic-conversational-sft SmolKalam SmolKalam is a quality-filtered Arabic SFT dataset of 1,790,478 examples (~2.45B tokens), built as an ensemble translation of SmolTalk2. It covers multi-turn dialogue (23% of rows), reasoning traces (19% carry <think>), tool and function calling (4.4%), and long context, categories that are underrepresented in existing Arabic post-training data. The SmolTalk2 source mixtures are kept as subsets. Released with the paper SmolKalam: Ensemble Quality-Filtered Translation… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/smolkalam-arabic-conversational-sft.tabulartext-generation1M<n<10M3 likes850 downloads1mo agoHugging Face05Jr23xd23 /ArabicText-Large ArabicText-Large: High-Quality Arabic Corpus for LLM Training Dataset Summary ArabicText-Large is a comprehensive, high-quality Arabic text corpus comprising 743,288 articles with over 244 million words, specifically curated for Large Language Model (LLM) training and fine-tuning. This dataset represents one of the largest publicly available Arabic text collections for machine learning research. This corpus addresses the critical shortage of high-quality Arabic NLP… See the full description on the dataset page: https://huggingface.co/datasets/Jr23xd23/ArabicText-Large.texttext-generation100K<n<1M69 likes725 downloads11mo agoHugging Face06ameer4wisam /iraqi-arabic-sales-dialogue-dataset Iraqi Arabic Sales Dialogue Dataset A large synthetic dataset of Iraqi (Baghdadi-based) Arabic dialogue, centered on retail sales, haggling, and everyday conversation. النسخة العربية متوفرة بالكامل بالأسفل — Arabic version available in full below. What this is 210,832 template-generated conversations, of which 171,601 (81%) are exact-unique message sequences, spanning 20 topical categories in colloquial Iraqi Arabic. The core of the dataset (10 categories) is… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi-arabic-sales-dialogue-dataset.texttext-generation100K<n<1M0 likes693 downloads2mo agoHugging Face07SultanR /dclm-pro-arabic dclm-pro-arabic Arabic translation of DCLM-Pro (global shards 01 and 05), translated with Seed-X-PPO-7B using greedy decoding. Documents were split into ~490-token chunks at sentence boundaries, translated, and reassembled. Each row is one complete document. A companion corpus translated with the same pipeline is available at fineweb-edu-arabic. Details Documents: 33,245,503 (22.7% of the two source shards, uniformly sampled) Arabic tokens: ~93B (Seed-X… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/dclm-pro-arabic.texttext-generation10M<n<100M0 likes578 downloads1mo agoHugging Face08ISLAM-PO /documents-Egyptian-Arabic Dataset evaluation: See EVALUATION.md for schema/config checks, row-count status, and the data-governance plan. Viewer note: default is a lightweight preview; select egyptian_arabic or another source config to load the full data. Current Hub Validation Status Dataset Server rows: 25,399,945 Dataset Server original/Parquet size: 2,758,228,707 bytes (~2.76 GB) Default Hub configuration currently exposes one column: text Source-specific configurations are explicitly declared in… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/documents-Egyptian-Arabic.texttranslationn<1K2 likes530 downloads8h agoHugging Face09nizarun /FineWeb-Edu-Arabic-24M English العربية FineWeb-Edu Arabic 24M An Arabic-only pretraining corpus of 24,794,425 complete documents, translated from the sample-350BT configuration of FineWeb-Edu. It contains 34.86 billion Arabic tokenizer tokens and preserves the original FineWeb-Edu document IDs, source scores, and detailed translation diagnostics. Highlight Saudi architecture shaped by place. A well-translated tour of how builders in Najd, the Gulf coast, Hejaz, and Asir adapted local… See the full description on the dataset page: https://huggingface.co/datasets/nizarun/FineWeb-Edu-Arabic-24M.tabulartext-generation10M<n<100M0 likes464 downloads27d agoHugging Face10MohamedRashad /Arabic-VLM-Full-Pearl 💎 The Arabic VLM Dataset (Full Pearl Edition) This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper. Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.imagequestion-answering100K<n<1M10 likes452 downloads10mo agoHugging Face11SultanR /fineweb-edu-arabic fineweb-edu-arabic Arabic translation of FineWeb-Edu (sample/350BT subset, filtered to language_score > 0.9), translated with Seed-X-PPO-7B using greedy decoding. Documents were split into ~490-token chunks, translated, and reassembled. Each row is one complete document. A companion corpus translated with the same pipeline is available at dclm-pro-arabic. Details Documents: 82,840,410 (27.9% of the source subset, uniformly sampled) Arabic tokens: ~170B (Seed-X… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/fineweb-edu-arabic.texttext-generation10M<n<100M1 likes450 downloads1mo agoHugging Face12Ahmed-Selem /Shifaa_Arabic_Medical_Consultations Shifaa Arabic Medical Consultations 🏥📊 Overview 🌍 Shifaa is revolutionizing Arabic medical AI by addressing the critical gap in Arabic medical datasets. Our first contribution is the Shifaa Arabic Medical Consultations dataset, a comprehensive collection of 84,422 real-world medical consultations covering 16 Main Specializations and 585 Hierarchical Diagnoses. 🔍 Why is this dataset important? First large-scale Arabic medical dataset for AI applications.… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Medical_Consultations.textquestion-answering10K<n<100K13 likes377 downloads2y agoHugging Face13lightonai /ArabicWeb24gated 📚 ArabicWeb24 More than 39 billion tokens of high quality Arabic web content 🌐. What is ArabicWeb24 ? The ArabicWeb24 dataset consists of more than 28 billion tokens of cleaned and deduplicated Arabic web data from a customized crawl. This was processed using the large scale data processing library datatrove. What is being released ? We are releasing two datasets versions: ArabicWeb24: dataset version 1 (v1) underwent extensive processing through… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/ArabicWeb24.texttext-generation10M<n<100M24 likes376 downloads2y agoHugging Face14yrrhall /Mixed-Arabic-Datasets-Repo Dataset Card for "Mixed Arabic Datasets (MAD) Corpus" The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts Dataset Description The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/Mixed-Arabic-Datasets-Repo.tabulartext-classification100M<n<1B0 likes340 downloads4mo agoHugging Face15ISLAM-PO /arabic-to-code-8-langs-3m Dataset evaluation: See EVALUATION.md for schema checks, indexing status, and language-specific quality limits. Viewer note: default is a lightweight preview; select full to load the complete corpus. Current Hub Validation Status Repository claim: 3,000,000 records Dataset Server indexed rows: 1,239,045 Dataset Server estimate: 1,995,159 The 3M target figure is a raw-repository claim and is not yet fully verified by the Hub index. Validate the JSONL files before publishing a… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arabic-to-code-8-langs-3m.texttext-generation1M<n<10M0 likes306 downloads6h agoHugging Face16muhammadrizo5721 /101_billion_arabic_words_dataset 101 Billion Arabic Words Dataset Updates Maintenance Status: Actively Maintained Update Frequency: Weekly updates to refine data quality and expand coverage. Upcoming Version More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management. Dataset Details The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/muhammadrizo5721/101_billion_arabic_words_dataset.texttext-generation10M<n<100M0 likes279 downloads7mo agoHugging Face17MohamedRashad /arabic-billion-words Arabic Billion Words Dataset 🌕 The Abu El-Khair Arabic News Corpus (arabic-billion-words) is a comprehensive collection of Arabic text, encompassing over five million newspaper articles. The corpus is rich in linguistic diversity, containing more than a billion and a half words, with approximately three million unique words. The text is encoded in two formats: UTF-8 and Windows CP-1256, and marked up using two markup languages: SGML and XML. Data Example An example… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-billion-words.texttext-generation1M<n<10M12 likes233 downloads3y agoHugging Face18dataflare /arabic-dialect-corpus Arabic Dialect Corpus A comprehensive collection of Arabic dialectal text, standardized for Natural Language Processing (NLP) model training, evaluation, and linguistic analysis. This corpus has been meticulously processed to ensure high-quality tokenization and consistent metadata. Dataset Statistics Metric Value Total Records 127,180 Total Tokens 5,802,324 Average Tokens per Record 45.62 Dialect Categories 5 Changelog… See the full description on the dataset page: https://huggingface.co/datasets/dataflare/arabic-dialect-corpus.tabulartext-generation100K<n<1M1 likes155 downloads8mo agoHugging Face19freococo /arabic_tashkil_dataset Arabic Tashkil (Diacritization) Dataset 📖✨ Dataset Summary This is a massive, high-quality, Gold-Standard dataset designed explicitly for training Arabic Automatic Diacritization (Tashkil) AI models (such as ByT5, AraT5, or Custom Transformers). The dataset contains 1,494,228 heavily vocalized pages (~2.47 GB of data) extracted from Classical Arabic and Islamic texts sourced from Thahabi.org. To ensure the highest possible ground-truth quality, every single page… See the full description on the dataset page: https://huggingface.co/datasets/freococo/arabic_tashkil_dataset.tabulartext-generation1M<n<10M0 likes154 downloads3mo agoHugging Face20ISLAM-PO /arabic-history-and-dialects Dataset evaluation: See EVALUATION.md for schema checks, quality limits, and the fact-check plan. Viewer note: default is a lightweight preview; select full to load the complete corpus. مجموعة البيانات العربية الشاملة للذكاء الاصطناعي 🇸🇦🇪🇬🇱🇧🇲🇦 [!WARNING] This is a small educational draft. The card flags three historical claims for fact-checking; verify them before using the dataset for factual QA or training. Arabic Multi-Dialect & Civilization… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arabic-history-and-dialects.textquestion-answeringn<1K0 likes148 downloads8h agoHugging Face21HeshamHaroon /Arabic_Function_Calling Arabic Function Calling Dataset (50K+ Samples) مجموعة بيانات استدعاء الدوال العربية أول وأكبر مجموعة بيانات عربية متخصصة في استدعاء الدوال (Function Calling) تغطي جميع اللهجات العربية الرئيسية والمجالات الحياتية المهمة. Dataset Description This is the first comprehensive Arabic function calling dataset designed for training and evaluating LLMs on Arabic tool use capabilities. The dataset covers: 5 Arabic Dialects: MSA (Modern Standard Arabic), Egyptian… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/Arabic_Function_Calling.texttext-generation10K<n<100K60 likes128 downloads10mo agoHugging Face22Mars203020 /arabic_medical_dialoguetexttext-generation1K<n<10K3 likes126 downloads2y agoHugging Face23oddadmix /arabic-math-reasoning-synth Arabic Math Reasoning (synthetic) — مسائل رياضيات عربية مع خطوات الحل 120,462 Arabic grade-school math word problems, each with a step-by-step derivation and a concluding sentence. Generated with gemma-3-12b-it and Qwen3.8-27B-Uncensored-NVFP4 and arithmetically verified — every equation the reasoning states was re-evaluated, and rows whose own arithmetic does not check out were dropped. generator rows share gemma-3-12b-it 80,480 66.8% Qwen3.8-27B-Uncensored-NVFP4… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-math-reasoning-synth.texttext-generation100K<n<1M0 likes126 downloads28d agoHugging Face24TuwaiqAcademy /AISA-ArabicFC AISA-ArabicFC Arabic Function Calling for Agentic AI Systems The first open benchmark for tool-use in Arabic — across five dialects, eight real-world domains, and 27 structured tools. 12,125 queries · 5 dialects · 8 domains · 27 tools · 12K reasoning traces 📅 Test set releases July 20, 2026 · 🏛️ Budapest · Oct 24–29, 2026 🆕 Update — Data v1.4 & fair scoring (June 2026) Argument scoring is now robust to surface form. A correct… See the full description on the dataset page: https://huggingface.co/datasets/TuwaiqAcademy/AISA-ArabicFC.texttext-generation10K<n<100K8 likes96 downloads2mo agoHugging Face25fr3on /arabic-dialect-corpus 🇪🇬🇸🇦 Arabic Dialect Corpus (Egyptian & Saudi) Dataset Description This dataset contains 150K+ natural, informal Arabic text samples scraped from high-engagement YouTube discussions. It specifically targets Egyptian (EG) and Saudi (SA) dialects, filling a critical gap in resources for training LLMs on colloquial Arabic (Ammiya) rather than just Modern Standard Arabic (MSA). Languages Primary Dialects: Egyptian Arabic (EG) - Cairene and regional Egyptian… See the full description on the dataset page: https://huggingface.co/datasets/fr3on/arabic-dialect-corpus.texttext-generation1M<n<10M1 likes95 downloads8mo agoHugging Face26M-A-D /Mixed-Arabic-Dataset-Main Dataset Card for "Mixed-Arabic-Dataset" Mixed Arabic Datasets (MAD) The Mixed Arabic Datasets (MAD) project provides a comprehensive collection of diverse Arabic-language datasets, sourced from various repositories, platforms, and domains. These datasets cover a wide range of text types, including books, articles, Wikipedia content, stories, and more. MAD Repo vs. MAD Main MAD Repo Versatility: In the MAD Repository (MAD Repo), datasets are made… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Dataset-Main.tabulartext-generation100K<n<1M7 likes94 downloads3y agoHugging Face27AhmedBou /Arabic_Quotes Arabic Quotes Dataset Overview The Arabic Quotes Dataset is an open-source collection of 5900+ quotes in the Arabic language, accompanied by up to three tags for each quote. The dataset is suitable for various Natural Language Processing (NLP) tasks, such as text classification and tagging. Data Description Contains 5900+ quotes with up to three associated tags per quote. All quotes and tags are in Arabic. Use Cases Text Classification:… See the full description on the dataset page: https://huggingface.co/datasets/AhmedBou/Arabic_Quotes.texttext-classification1K<n<10K4 likes89 downloads3y agoHugging Face28riotu-lab /Arabic-books-and-research-dataset Arabic reserach and books dataset (ARABD) This dataset is an extracted cleaned text from more than 60K word files with unique arabic texts never published before. Dataset diversity the dataset is diverse from all kind of islamic research: [feqh, hadeeth, tafseer, tahqeeq, ... etc], from new written research to a manuscirpts. dataset size the dataset was more than 11GB but after cleaning (pre-processing) it becase a straight 10GB with less noisy chars.… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/Arabic-books-and-research-dataset.texttext-generation10K<n<100K6 likes87 downloads2y agoHugging Face29IbrahimAmin /egyptian-arabic-fake-reviews 🕵️‍♂️🇪🇬 FREAD: Fake Reviews Egyptian Arabic Dataset Author: IbrahimAmin, Ismail Fakhr, M. Waleed Fakhr, Rasha Kashef License: MIT Paper: Boosting Arabic Fake Reviews Detection by Integrating Textual and Metadata Features: A Transformer-Based Model Languages: Arabic (Egyptian Dialect) 📚 Dataset Summary FREAD is designed for detecting fake reviews in Arabic using both textual content and behavioral metadata. It contains 60,000 reviews (50K train / 10K test)… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimAmin/egyptian-arabic-fake-reviews.tabulartext-classification10K<n<100K2 likes86 downloads2mo agoHugging Face30tahaalselwii /arabic-proverbs-collection Arabic Proverbs Collection Dataset Description Arabic Proverbs Collection is a dataset containing Arabic proverbs and their explanations. The collection covers three types of Arabic proverbs from different linguistic contexts and historical periods: Classical Arabic proverbs. Colloquial Arabic proverbs. Popular and contemporary Arabic proverbs generated with artificial intelligence. Dataset Sources Classical Arabic Proverbs The… See the full description on the dataset page: https://huggingface.co/datasets/tahaalselwii/arabic-proverbs-collection.texttext-generation1K<n<10K1 likes85 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.