CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MBZUAI /ArabicMMLU Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne Introduction We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/ArabicMMLU.tabularquestion-answering10K<n<100K39 likes3.7k downloads2y agoHugging Face02M-A-D /Mixed-Arabic-Datasets-Repo Dataset Card for "Mixed Arabic Datasets (MAD) Corpus" The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts Dataset Description The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo.tabulartext-classification100M<n<1B38 likes2.2k downloads3y agoHugging Face03QCRI /AraDICE-ArabicMMLU-egy AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs -- ArabicMMLU - Egyptian dialect Overview The AraDiCE dataset is crafted to assess the dialectal and cultural understanding of large language models (LLMs) within Arabic-speaking contexts. It includes post-edited adaptations of several benchmark datasets, specifically curated to validate LLM performance in culturally and dialectally relevant scenarios for Arabic. Within the AraDiCE collection, this… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDICE-ArabicMMLU-egy.texttext-classification10K<n<100K1 likes551 downloads2y agoHugging Face04abdoelsayed /Open-ArabicaQA ArabicaQA ArabicaQA: Comprehensive Dataset for Arabic Question Answering This repository contains dataset for paper ArabicaQA: Comprehensive Dataset for Arabic Question Answering. Below, we provide details regarding the materials available in this repository: ArabicaQA is a robust dataset designed to support and advance the development of Arabic Question Answering (QA) systems. This dataset encompasses a wide range of question types, including both Machine Reading Comprehension… See the full description on the dataset page: https://huggingface.co/datasets/abdoelsayed/Open-ArabicaQA.question-answering10K<n<100K10 likes538 downloads2y agoHugging Face05ISLAM-PO /documents-Egyptian-Arabic Current Hub Validation Status Dataset Server rows: 25,399,945 Dataset Server original/Parquet size: 2,758,228,707 bytes (~2.76 GB) Default Hub configuration currently exposes one column: text The additional configuration names listed in this card are physical source directories and are not all recognized as separate Hub configurations. Keep the default configuration until the dataset is normalized into explicit, tested splits. Egyptian Arabic Mega Corpus (EAMC) —… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/documents-Egyptian-Arabic.translation10M<n<100M2 likes502 downloads4h agoHugging Face06MohamedRashad /Arabic-VLM-Full-Pearl 💎 The Arabic VLM Dataset (Full Pearl Edition) This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper. Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.imagequestion-answering100K<n<1M10 likes452 downloads10mo agoHugging Face07Ahmed-Selem /Shifaa_Arabic_Medical_Consultations Shifaa Arabic Medical Consultations 🏥📊 Overview 🌍 Shifaa is revolutionizing Arabic medical AI by addressing the critical gap in Arabic medical datasets. Our first contribution is the Shifaa Arabic Medical Consultations dataset, a comprehensive collection of 84,422 real-world medical consultations covering 16 Main Specializations and 585 Hierarchical Diagnoses. 🔍 Why is this dataset important? First large-scale Arabic medical dataset for AI applications.… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Medical_Consultations.textquestion-answering10K<n<100K13 likes383 downloads2y agoHugging Face08sadeem-ai /arabic-qna Sadeem QnA: An Arabic QnA Dataset 🌍✨ Welcome to the Sadeem QnA dataset, a vibrant collection designed for the advancement of Arabic natural language processing, specifically tailored for Question Answering (QnA) systems. Sourced from the rich and diverse content of Arabic Wikipedia, this dataset is a gateway to exploring the depths of Arabic language understanding, offering a unique challenge to both researchers and AI enthusiasts alike. About Sadeem QnA The Sadeem… See the full description on the dataset page: https://huggingface.co/datasets/sadeem-ai/arabic-qna.textquestion-answering1K<n<10K4 likes372 downloads3y agoHugging Face09QCRI /AraDICE-ArabicMMLU-lev AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs -- ArabicMMLU - Levantine dialect Overview The AraDiCE dataset is crafted to assess the dialectal and cultural understanding of large language models (LLMs) within Arabic-speaking contexts. It includes post-edited adaptations of several benchmark datasets, specifically curated to validate LLM performance in culturally and dialectally relevant scenarios for Arabic. Within the AraDiCE collection, this… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDICE-ArabicMMLU-lev.texttext-classification10K<n<100K0 likes358 downloads2y agoHugging Face10yrrhall /Mixed-Arabic-Datasets-Repo Dataset Card for "Mixed Arabic Datasets (MAD) Corpus" The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts Dataset Description The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/Mixed-Arabic-Datasets-Repo.tabulartext-classification100M<n<1B0 likes340 downloads4mo agoHugging Face11inceptlabs /Arabic_EXAMS-Redux Arabic_EXAMS-Redux A corrected and text-repaired version of OALL/Arabic_EXAMS, the Arabic subset of the EXAMS multilingual high-school examinations benchmark. What was fixed Repaired corrupted Arabic text. The upstream benchmark contains widespread PDF-extraction damage to question stems and answer choices: split diacritics, fragmented words, and non-Arabic glyphs replacing standard characters. We restored these to readable Modern Standard Arabic. Corrected the answer… See the full description on the dataset page: https://huggingface.co/datasets/inceptlabs/Arabic_EXAMS-Redux.textquestion-answeringn<1K1 likes298 downloads5mo agoHugging Face12unohamza /Arabic-news-daily Arabic News Daily 🗞️ A daily-updated, multi-domain Arabic news dataset collected automatically from 15 curated sources. Unlike other Arabic datasets that are static snapshots, this dataset grows every day — making it ideal for research requiring fresh, current Arabic text across diverse domains. Sources Source Domain Variety Al Jazeera Arabic Politics MSA BBC Arabic Politics MSA RT Arabic Politics MSA Al Arabiya Politics MSA AITNews Tech & AI… See the full description on the dataset page: https://huggingface.co/datasets/unohamza/Arabic-news-daily.text-generation100K<n<1M1 likes271 downloads10h agoHugging Face13abdoelsayed /ArabicaQA ArabicaQA ArabicaQA: Comprehensive Dataset for Arabic Question Answering This repository contains dataset for paper ArabicaQA: Comprehensive Dataset for Arabic Question Answering. Below, we provide details regarding the materials available in this repository: Dataset Within this folder, you will find the training, validation, and test sets of the ArabicaQA dataset. Refer to the table below for the dataset statistics: Training Validation Test MRC (with answers)… See the full description on the dataset page: https://huggingface.co/datasets/abdoelsayed/ArabicaQA.question-answering10K<n<100K6 likes162 downloads2y agoHugging Face14Ahmed-Selem /Shifaa_Arabic_Mental_Health_Consultations 🏥 Shifaa Arabic Mental Health Consultations 🧠 📌 Overview Shifaa Arabic Mental Health Consultations is a high-quality dataset designed to advance Arabic medical language models.This dataset provides 35,648 real-world medical consultations, covering a wide range of mental health concerns. 📊 Dataset Summary Size: 35,648 consultations Main Specializations: 7 Specific Diagnoses: 123 Languages: Arabic (العربية) Why This Dataset? 🔹 Lack of… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Mental_Health_Consultations.textquestion-answering10K<n<100K14 likes155 downloads2y agoHugging Face15HeshamHaroon /Arabic_Function_Calling Arabic Function Calling Dataset (50K+ Samples) مجموعة بيانات استدعاء الدوال العربية أول وأكبر مجموعة بيانات عربية متخصصة في استدعاء الدوال (Function Calling) تغطي جميع اللهجات العربية الرئيسية والمجالات الحياتية المهمة. Dataset Description This is the first comprehensive Arabic function calling dataset designed for training and evaluating LLMs on Arabic tool use capabilities. The dataset covers: 5 Arabic Dialects: MSA (Modern Standard Arabic), Egyptian… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/Arabic_Function_Calling.texttext-generation10K<n<100K60 likes131 downloads9mo agoHugging Face16ISLAM-PO /arabic-history-and-dialects مجموعة البيانات العربية الشاملة للذكاء الاصطناعي 🇸🇦🇪🇬🇱🇧🇲🇦 Arabic Multi-Dialect & Civilization Instruction Dataset المؤلف: islam-alnasherA-Dev — الحساب: https://huggingface.co/ISLAM-POالإصدار: v1.0 — التاريخ: 30 أغسطس 2026 — الترخيص: CC BY 4.0اللغة: العربية (فصحى + 4 لهجات) — الصيغة: instruction / output JSONL — الحجم: 275 عينة 📌 الملخص التنفيذي هذه المجموعة هي مورد تعليمي متخصص لتدريب وتقييم النماذج اللغوية العربية على مسارين متوازيين:… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arabic-history-and-dialects.textquestion-answeringn<1K0 likes125 downloads4h agoHugging Face17HeshamHaroon /ArabicRAGB ArabicRAGB: Arabic Retrieval-Augmented Generation Benchmark Dataset Description ArabicRAGB is a benchmark dataset for evaluating Retrieval-Augmented Generation (RAG) systems on Arabic language tasks. Each record contains a query-passage pair where the query is grounded in the passage content. Key Features Passage-Grounded Queries: Each query is generated from and answerable by its paired passage Multi-Dialect Coverage: MSA, Egyptian, Gulf… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/ArabicRAGB.texttext-retrieval10K<n<100K13 likes118 downloads9mo agoHugging Face18AhmadHakami /saudipedia-arabic-qa Saudipedia Q&A Dataset Dataset Description Summary This dataset contains question-answer pairs scraped from Saudipedia, a comprehensive Arabic encyclopedia focused on Saudi Arabia. The dataset includes 1,082 Q&A entries covering various topics related to Saudi culture, history, economy, government, society, geography, religion, and notable personalities. The data was collected by scraping the website's question-answer section, which provides detailed answers to… See the full description on the dataset page: https://huggingface.co/datasets/AhmadHakami/saudipedia-arabic-qa.textquestion-answering1K<n<10K3 likes117 downloads1y agoHugging Face19TuwaiqAcademy /AISA-ArabicFC AISA-ArabicFC Arabic Function Calling for Agentic AI Systems The first open benchmark for tool-use in Arabic — across five dialects, eight real-world domains, and 27 structured tools. 12,125 queries · 5 dialects · 8 domains · 27 tools · 12K reasoning traces 📅 Test set releases July 20, 2026 · 🏛️ Budapest · Oct 24–29, 2026 🆕 Update — Data v1.4 & fair scoring (June 2026) Argument scoring is now robust to surface form. A correct… See the full description on the dataset page: https://huggingface.co/datasets/TuwaiqAcademy/AISA-ArabicFC.texttext-generation10K<n<100K8 likes111 downloads2mo agoHugging Face20QCRI /ArabicCulturalQA ArabicCulturalQA ArabicCulturalQA is the first cross-dialectal Arabic cultural QA benchmark with parallel multiple-choice (MCQ) and open-ended (OEQ) formats across Modern Standard Arabic (MSA), English, Egyptian, Levantine, Gulf, and Maghrebi. Both the MCQ and OEQ test sets have been reviewed and post-edited by native speakers of each dialect. The dataset accompanies the LREC 2026 paper "Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants" (paper page)… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ArabicCulturalQA.textquestion-answering10K<n<100K2 likes78 downloads3mo agoHugging Face21nawaralseelawi /mizan-iraqi-arabic-benchmark Mizan (ميزان) — Iraqi Arabic LLM Benchmark: pilot-0.2 public development set Mizan is the first comprehensive, originally-authored evaluation benchmark for Iraqi Arabic and the Iraqi civic context. This dataset is the pilot-0.2 public development set: 340 originally-authored, dually-reviewed items across two tracks (MSA baseline / Iraqi) and six axes. 📄 Paper (preprint): https://doi.org/10.5281/zenodo.22714865 🏆 Live leaderboard: https://mizan-bench.onrender.com 💻 Code… See the full description on the dataset page: https://huggingface.co/datasets/nawaralseelawi/mizan-iraqi-arabic-benchmark.textquestion-answeringn<1K1 likes74 downloads12d agoHugging Face22miscovery /Math_CoT_Arabic_English_Reasoning Math CoT Arabic English Dataset A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI. Overview Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.tabularquestion-answering1K<n<10K17 likes73 downloads1y agoHugging Face23oddadmix /arabic-rag-chat-8k-eval arabic-rag-chat-8k-eval Per-row evaluation artifacts for the 8,192-token Arabic multi-turn RAG models: the test split, every model's raw replies, every judge verdict, and the rendered report for each. Thirteen judged models, all scored on the same 1,651 prompts by the same judge at temperature 0.0, so the comparison below is like-for-like and can be recomputed offline without a GPU or a judge server. This is the measurement half of oddadmix/100M-8192-Nawah-dsv4; the training… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-8k-eval.tabularquestion-answeringn<1K0 likes72 downloads1mo agoHugging Face24NightPrince /islamic-arabic-qa Islamic Arabic Q&A Dataset A curated Arabic instruction-tuning dataset focused on Islamic scholarship — covering Fiqh, Fatwa, Aqeedah, Quran Sciences, and Islamic Finance. Built to fine-tune Arabic LLMs for Islamic Q&A tasks. Dataset Summary Split Samples Train 17,944 Validation 2,101 Test 1,042 Total 21,087 Data Sources Source Samples License SahmBenchmark/fatwa-training_standardized_new 9,953 Apache 2.0… See the full description on the dataset page: https://huggingface.co/datasets/NightPrince/islamic-arabic-qa.texttext-generation10K<n<100K0 likes70 downloads5mo agoHugging Face25Omar-youssef /islamic-qa-egyptian-arabic Egyptian Arabic Islamic QA Dataset Dataset Description This dataset contains 7,465 question-answer pairs in Egyptian Arabic covering comprehensive Islamic studies topics. The dataset serves as a valuable resource for developing Arabic NLP models focused on Islamic education and religious knowledge. Key Features Language: Egyptian Arabic (العامية المصرية) Domain: Islamic Studies Size: 7,465 examples Format: Question-Answer pairs with topic categorization… See the full description on the dataset page: https://huggingface.co/datasets/Omar-youssef/islamic-qa-egyptian-arabic.textquestion-answering1K<n<10K1 likes64 downloads1y agoHugging Face26haiderkamal23 /allaM-offsec-arabic-chat-v2 Arabic Offensive Security Chat Dataset v2 High-quality category-aware bilingual Arabic/English dataset for offensive security assistants. What's New in v2 ✅ Category-aware responses: Different response structures for web vulns, DeFi, reconnaissance tools, social engineering, etc. ✅ No generic templates: Each category has specialized analysis framework ✅ No verbatim copying: Responses analyze and transform the input, not repeat it ✅ Semantic accuracy: Tools (nmap… See the full description on the dataset page: https://huggingface.co/datasets/haiderkamal23/allaM-offsec-arabic-chat-v2.textquestion-answering10K<n<100K0 likes56 downloads10mo agoHugging Face27Lyte /2A2I-Arabic-OpenHermes-2.5-Llama-3 Dataset Card for "2A2I-Arabic-OpenHermes-2.5-Llama-3" Dataset Sources & Infos Data Origin: Derived from the original Arabic OpenHermes dataset : 2A2I/Arabic-OpenHermes-2.5. Languages: Modern Standard Arabic (MSA) Applications: Language Modeling License: Apache-2.0 Overview 2A2I-Arabic-OpenHermes-2.5-Llama is a Llama-3 compatible dataset carefully converted from the 2A2I's Arabic-OpenHermes-2.5 collection provided by Lyte. Purpose… See the full description on the dataset page: https://huggingface.co/datasets/Lyte/2A2I-Arabic-OpenHermes-2.5-Llama-3.textquestion-answering100K<n<1M2 likes53 downloads2y agoHugging Face28Omartificial-Intelligence-Space /Arabic-gsm8k-v2 Dataset Summary Arabic GSM8K is an Arabic translation of the GSM8K (Grade School Math 8K) dataset, which contains high-quality linguistically diverse grade school math word problems. The original dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning, and this Arabic version aims to extend these capabilities to Arabic language models and applications. The dataset maintains the same characteristics as the… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/Arabic-gsm8k-v2.textquestion-answering10K<n<100K1 likes53 downloads1y agoHugging Face29islamlab /arabic-lexicons islamlab — The Classical Arabic Lexicons 197,731 entries from 136 lexical works — the dictionaries, the gharīb collections and the technical glossaries — cut so that the headword is its own column and the article is its own text. A dictionary sits in the corpus like any other book, but nobody reads one that way. This is the same material arranged for the thing people actually do with it: look a word up. work author entries شمس العلوم ودواء كلام العرب من الكلوم نشوان… See the full description on the dataset page: https://huggingface.co/datasets/islamlab/arabic-lexicons.tabulartext-retrieval100K<n<1M0 likes52 downloads1mo agoHugging Face30oddadmix /arabic-rag-chat-30K Arabic multi-turn RAG customer-support conversations (31,294 conversations) Synthetic Modern Standard Arabic customer-support conversations for training small Arabic RAG assistants. The bulk was distilled from gemini-3.1-flash-lite via the Batch API; a first 2.4% came from unsloth/gemma-4-31B-it-NVFP4 on a local vLLM server before the run was moved off-GPU. Both teachers were given the same prompts and the same validator. Each row is one conversation of 1-5 rounds over one… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-30K.tabularquestion-answering10K<n<100K0 likes50 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.