CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MBZUAI /ArabicMMLU Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne Introduction We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/ArabicMMLU.tabularquestion-answering10K<n<100K39 likes3.7k downloads2y agoHugging Face02M-A-D /Mixed-Arabic-Datasets-Repo Dataset Card for "Mixed Arabic Datasets (MAD) Corpus" The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts Dataset Description The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo.tabulartext-classification100M<n<1B38 likes2.2k downloads3y agoHugging Face03QCRI /AraDICE-ArabicMMLU-egy AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs -- ArabicMMLU - Egyptian dialect Overview The AraDiCE dataset is crafted to assess the dialectal and cultural understanding of large language models (LLMs) within Arabic-speaking contexts. It includes post-edited adaptations of several benchmark datasets, specifically curated to validate LLM performance in culturally and dialectally relevant scenarios for Arabic. Within the AraDiCE collection, this… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDICE-ArabicMMLU-egy.texttext-classification10K<n<100K1 likes551 downloads2y agoHugging Face04abdoelsayed /Open-ArabicaQA ArabicaQA ArabicaQA: Comprehensive Dataset for Arabic Question Answering This repository contains dataset for paper ArabicaQA: Comprehensive Dataset for Arabic Question Answering. Below, we provide details regarding the materials available in this repository: ArabicaQA is a robust dataset designed to support and advance the development of Arabic Question Answering (QA) systems. This dataset encompasses a wide range of question types, including both Machine Reading Comprehension… See the full description on the dataset page: https://huggingface.co/datasets/abdoelsayed/Open-ArabicaQA.question-answering10K<n<100K10 likes538 downloads2y agoHugging Face05ISLAM-PO /documents-Egyptian-Arabic Egyptian Arabic Mega Corpus (EAMC) — 25M Unified Egyptian Dialect Dataset The Largest Unified Open Corpus for Egyptian Arabic (Masri / arz) 25.5M Samples | 2.66 GB (Parquet) | 9 Configs | Apache 2.0 | Ready-to-train Comprehensive coverage: Raw Text · Wikipedia · Conversations · Speech (Whisper) · Parallel Translation (EN↔EGY) · Trilingual QA · Wikipedia Quality Classification · Fake Review / Spam Detection Dataset Summary Egyptian Arabic Mega Corpus… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/documents-Egyptian-Arabic.translation10M<n<100M2 likes502 downloads23d agoHugging Face06MohamedRashad /Arabic-VLM-Full-Pearl 💎 The Arabic VLM Dataset (Full Pearl Edition) This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper. Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.imagequestion-answering100K<n<1M10 likes452 downloads10mo agoHugging Face07araag2 /MedNLI MedNLI — A Natural Language Inference Dataset For The Clinical Domain Dataset Description Links Homepage: Github.io Repository: Github Paper: arXiv Leaderboard: Papers with Code Contact (Original Authors): Alexey Romanov aromanov@cs.uml.edu, Chaitanya Shivade cshivade@us.ibm.com Contact (Curator): Artur Guimarães (artur.guimas@gmail.com) Dataset Summary `Natural Language Inference (NLI) is one of the critical tasks for… See the full description on the dataset page: https://huggingface.co/datasets/araag2/MedNLI.textquestion-answering10K<n<100K0 likes419 downloads8d agoHugging Face08Ahmed-Selem /Shifaa_Arabic_Medical_Consultations Shifaa Arabic Medical Consultations 🏥📊 Overview 🌍 Shifaa is revolutionizing Arabic medical AI by addressing the critical gap in Arabic medical datasets. Our first contribution is the Shifaa Arabic Medical Consultations dataset, a comprehensive collection of 84,422 real-world medical consultations covering 16 Main Specializations and 585 Hierarchical Diagnoses. 🔍 Why is this dataset important? First large-scale Arabic medical dataset for AI applications.… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Medical_Consultations.textquestion-answering10K<n<100K13 likes383 downloads2y agoHugging Face09sadeem-ai /arabic-qna Sadeem QnA: An Arabic QnA Dataset 🌍✨ Welcome to the Sadeem QnA dataset, a vibrant collection designed for the advancement of Arabic natural language processing, specifically tailored for Question Answering (QnA) systems. Sourced from the rich and diverse content of Arabic Wikipedia, this dataset is a gateway to exploring the depths of Arabic language understanding, offering a unique challenge to both researchers and AI enthusiasts alike. About Sadeem QnA The Sadeem… See the full description on the dataset page: https://huggingface.co/datasets/sadeem-ai/arabic-qna.textquestion-answering1K<n<10K4 likes372 downloads3y agoHugging Face10QCRI /AraDICE-ArabicMMLU-lev AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs -- ArabicMMLU - Levantine dialect Overview The AraDiCE dataset is crafted to assess the dialectal and cultural understanding of large language models (LLMs) within Arabic-speaking contexts. It includes post-edited adaptations of several benchmark datasets, specifically curated to validate LLM performance in culturally and dialectally relevant scenarios for Arabic. Within the AraDiCE collection, this… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDICE-ArabicMMLU-lev.texttext-classification10K<n<100K0 likes358 downloads2y agoHugging Face11yrrhall /Mixed-Arabic-Datasets-Repo Dataset Card for "Mixed Arabic Datasets (MAD) Corpus" The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts Dataset Description The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/Mixed-Arabic-Datasets-Repo.tabulartext-classification100M<n<1B0 likes340 downloads4mo agoHugging Face12inceptlabs /Arabic_EXAMS-Redux Arabic_EXAMS-Redux A corrected and text-repaired version of OALL/Arabic_EXAMS, the Arabic subset of the EXAMS multilingual high-school examinations benchmark. What was fixed Repaired corrupted Arabic text. The upstream benchmark contains widespread PDF-extraction damage to question stems and answer choices: split diacritics, fragmented words, and non-Arabic glyphs replacing standard characters. We restored these to readable Modern Standard Arabic. Corrected the answer… See the full description on the dataset page: https://huggingface.co/datasets/inceptlabs/Arabic_EXAMS-Redux.textquestion-answeringn<1K1 likes298 downloads5mo agoHugging Face13unohamza /Arabic-news-daily Arabic News Daily 🗞️ A daily-updated, multi-domain Arabic news dataset collected automatically from 15 curated sources. Unlike other Arabic datasets that are static snapshots, this dataset grows every day — making it ideal for research requiring fresh, current Arabic text across diverse domains. Sources Source Domain Variety Al Jazeera Arabic Politics MSA BBC Arabic Politics MSA RT Arabic Politics MSA Al Arabiya Politics MSA AITNews Tech & AI… See the full description on the dataset page: https://huggingface.co/datasets/unohamza/Arabic-news-daily.text-generation100K<n<1M1 likes271 downloads3h agoHugging Face14QCRI /AraDiCE AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs Overview The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic. As part of the supplemental materials, we have selected a few datasets (see below) for the reader to review. We will make the full… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDiCE.imagetext-classificationn<1K3 likes225 downloads2y agoHugging Face15Almheiri /ArabCulture-Dialogue ArabCulture-Dialogue: Cultural Benchmarking of LLMs in MSA and Arabic Dialectal Dialogue 📄 Paper (ACL 2026) | 🤗 Dataset ArabCulture-Dialogue is the first parallel MSA–dialect cultural dialogue dataset, covering 13 Arabic-speaking countries in both Modern Standard Arabic (MSA) and each country's respective dialect, spanning 12 daily-life topics and 54 fine-grained subtopics. It contains 3,471 parallel dialogue pairs (6,942 dialogues, 343,804 words in total), each consisting… See the full description on the dataset page: https://huggingface.co/datasets/Almheiri/ArabCulture-Dialogue.textquestion-answering1K<n<10K1 likes179 downloads1mo agoHugging Face16araag2 /Evidence_Inference_v2 Evidence Inference 2.0 Dataset Description Links Homepage: Github Pages Repository: Github Paper: arXiv Contact (Original Authors): Jay DeYoung (deyoung.j@northeastern.edu) Contact (Curator): Artur Guimarães (artur.guimas@gmail.com) Dataset Summary The dataset consists of biomedical articles describing randomized control trials (RCTs) that compare multiple treatments. Each of these articles will have multiple questions, or 'prompts'… See the full description on the dataset page: https://huggingface.co/datasets/araag2/Evidence_Inference_v2.textquestion-answering10K<n<100K1 likes178 downloads11mo agoHugging Face17abdoelsayed /ArabicaQA ArabicaQA ArabicaQA: Comprehensive Dataset for Arabic Question Answering This repository contains dataset for paper ArabicaQA: Comprehensive Dataset for Arabic Question Answering. Below, we provide details regarding the materials available in this repository: Dataset Within this folder, you will find the training, validation, and test sets of the ArabicaQA dataset. Refer to the table below for the dataset statistics: Training Validation Test MRC (with answers)… See the full description on the dataset page: https://huggingface.co/datasets/abdoelsayed/ArabicaQA.question-answering10K<n<100K6 likes162 downloads2y agoHugging Face18QCRI /AraDiCE-BoolQ AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs Overview The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic. In this repository, we present the BoolQ split of the data. Evaluation We have used lm-harness eval framework to for the… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDiCE-BoolQ.textquestion-answering1K<n<10K0 likes158 downloads2y agoHugging Face19Ahmed-Selem /Shifaa_Arabic_Mental_Health_Consultations 🏥 Shifaa Arabic Mental Health Consultations 🧠 📌 Overview Shifaa Arabic Mental Health Consultations is a high-quality dataset designed to advance Arabic medical language models.This dataset provides 35,648 real-world medical consultations, covering a wide range of mental health concerns. 📊 Dataset Summary Size: 35,648 consultations Main Specializations: 7 Specific Diagnoses: 123 Languages: Arabic (العربية) Why This Dataset? 🔹 Lack of… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Mental_Health_Consultations.textquestion-answering10K<n<100K14 likes155 downloads2y agoHugging Face20Aragoner /folkmotif FolkMotif-270: a parallel cross-cultural entity set for cultural-bias evaluation 27 Thompson Motif-Index roles × 10 cultural traditions = 270 citation-anchored cells. Paper: arXiv:2608.02486 · Code: github.com/AragonerUA/folkmotif Each cell names the canonical entity that fills one structural folk-narrative role in one tradition — the thunder-god's weapon, the smith of the gods, the goddess of love — together with its native-script form, attested spelling variants, and a… See the full description on the dataset page: https://huggingface.co/datasets/Aragoner/folkmotif.textquestion-answeringn<1K0 likes143 downloads2mo agoHugging Face21araag2 /PubMedQA PubMedQA - A Dataset for Biomedical Research Question Answering Dataset Description Links Homepage: Github.io Repository: Github Paper: arXiv Leaderboard: PapersWithCode Contact (Original Authors): Qiao Jin (qiaojin.andy@gmail.com) Contact (Curator): Artur Guimarães (artur.guimas@gmail.com) Dataset Summary The task of PubMedQA is to answer research questions with yes/no/maybe (e.g.: Do preoperative statins reduce atrial… See the full description on the dataset page: https://huggingface.co/datasets/araag2/PubMedQA.texttext-classification100K<n<1M0 likes131 downloads2mo agoHugging Face22HeshamHaroon /Arabic_Function_Calling Arabic Function Calling Dataset (50K+ Samples) مجموعة بيانات استدعاء الدوال العربية أول وأكبر مجموعة بيانات عربية متخصصة في استدعاء الدوال (Function Calling) تغطي جميع اللهجات العربية الرئيسية والمجالات الحياتية المهمة. Dataset Description This is the first comprehensive Arabic function calling dataset designed for training and evaluating LLMs on Arabic tool use capabilities. The dataset covers: 5 Arabic Dialects: MSA (Modern Standard Arabic), Egyptian… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/Arabic_Function_Calling.texttext-generation10K<n<100K60 likes131 downloads9mo agoHugging Face23hammh0a /AraLingBench AraLingBench 📄 Paper: arXiv:2511.14295💻 GitHub: hammoudhasan/AraLingBench AraLingBench is a 150-question Arabic multiple-choice benchmark that tests core linguistic competence of language models across five pillars: النحو (Grammar) الصرف (Morphology) الإملاء (Spelling & Orthography) فهم اللغة (Reading Comprehension) التركيب اللغوي والأسلوبي (Syntax & Stylistics) All questions are human-authored and validated, with a single correct answer and a difficulty label: Easy, Medium, or… See the full description on the dataset page: https://huggingface.co/datasets/hammh0a/AraLingBench.textquestion-answeringn<1K12 likes125 downloads10mo agoHugging Face24HeshamHaroon /ArabicRAGB ArabicRAGB: Arabic Retrieval-Augmented Generation Benchmark Dataset Description ArabicRAGB is a benchmark dataset for evaluating Retrieval-Augmented Generation (RAG) systems on Arabic language tasks. Each record contains a query-passage pair where the query is grounded in the passage content. Key Features Passage-Grounded Queries: Each query is generated from and answerable by its paired passage Multi-Dialect Coverage: MSA, Egyptian, Gulf… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/ArabicRAGB.texttext-retrieval10K<n<100K13 likes118 downloads9mo agoHugging Face25AhmadHakami /saudipedia-arabic-qa Saudipedia Q&A Dataset Dataset Description Summary This dataset contains question-answer pairs scraped from Saudipedia, a comprehensive Arabic encyclopedia focused on Saudi Arabia. The dataset includes 1,082 Q&A entries covering various topics related to Saudi culture, history, economy, government, society, geography, religion, and notable personalities. The data was collected by scraping the website's question-answer section, which provides detailed answers to… See the full description on the dataset page: https://huggingface.co/datasets/AhmadHakami/saudipedia-arabic-qa.textquestion-answering1K<n<10K3 likes117 downloads1y agoHugging Face26TuwaiqAcademy /AISA-ArabicFC AISA-ArabicFC Arabic Function Calling for Agentic AI Systems The first open benchmark for tool-use in Arabic — across five dialects, eight real-world domains, and 27 structured tools. 12,125 queries · 5 dialects · 8 domains · 27 tools · 12K reasoning traces 📅 Test set releases July 20, 2026 · 🏛️ Budapest · Oct 24–29, 2026 🆕 Update — Data v1.4 & fair scoring (June 2026) Argument scoring is now robust to surface form. A correct… See the full description on the dataset page: https://huggingface.co/datasets/TuwaiqAcademy/AISA-ArabicFC.texttext-generation10K<n<100K8 likes111 downloads2mo agoHugging Face27go-inoue /AraTrust_undiac AraTrust This repository provides a modified version of the AraTrust dataset originally introduced in: Emad A. Alghamdi, Reem I. Masoud, Deema Alnuhait, Afnan Y. Alomairi, Ahmed Ashraf, and Mohamed Zaytoon. 2025. AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic. In Proceedings of the 31st International Conference on Computational Linguistics, pages 8664–8679, Abu Dhabi, UAE. Association for Computational Linguistics. We release a version of AraTrust dataset we used in… See the full description on the dataset page: https://huggingface.co/datasets/go-inoue/AraTrust_undiac.textquestion-answeringn<1K0 likes106 downloads6mo agoHugging Face28Arailym-tleubayeva /KazakhLawCorpus Current Release Current version contains three datasets. data/ ├── laws_metadata.csv ├── law_history.csv └── law_references.csv Dataset Description 1. laws_metadata.csv Contains metadata describing legal acts. Current size: 223,245 legal acts Main fields include: Column Description source_id Internal database identifier law_id Stable legal act identifier title Original title title_kk Kazakh title title_ru Russian title… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhLawCorpus.text-retrieval100K<n<1M1 likes93 downloads14d agoHugging Face29araag2 /SemEval_NLI4CT NLI4CT: Multi-Evidence Natural Language Inference for Clinical Trial Reports and SemEval-2024 Task 2: Safe Biomedical Natural Language Inference for Clinical Trials Dataset Description Links Homepage: sites.google Repository: Github2024 Paper: arXiv2023 / arXiv2024 Leaderboard: Codalab2023 Contact (Original Authors):Maël Jullien (mael.jullien@postgrad.manchester.ac.uk) Contact (Curator): Artur Guimarães (artur.guimas@gmail.com) Dataset… See the full description on the dataset page: https://huggingface.co/datasets/araag2/SemEval_NLI4CT.texttext-classification10K<n<100K0 likes89 downloads1y agoHugging Face30araag2 /HINT HINT: Hierarchical interaction network for clinical-trial-outcome predictions Dataset Description Links Homepage: Github.io Repository: Github Paper: arXiv Contact (Original Authors): Tianfan Fu (futianfan@gmail.com) Contact (Curator):Artur Guimarães (artur.guimas@gmail.com) Dataset Summary Clinical trials are crucial for drug development but are time consuming, expensive, and often burdensome on patients. More importantly, clinical… See the full description on the dataset page: https://huggingface.co/datasets/araag2/HINT.textquestion-answering10K<n<100K0 likes88 downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.