CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MBZUAI /ArabicMMLU Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne Introduction We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/ArabicMMLU.tabularquestion-answering10K<n<100K39 likes3.6k downloads2y agoHugging Face02Ahmed-Selem /Shifaa_Arabic_Medical_Consultations Shifaa Arabic Medical Consultations 🏥📊 Overview 🌍 Shifaa is revolutionizing Arabic medical AI by addressing the critical gap in Arabic medical datasets. Our first contribution is the Shifaa Arabic Medical Consultations dataset, a comprehensive collection of 84,422 real-world medical consultations covering 16 Main Specializations and 585 Hierarchical Diagnoses. 🔍 Why is this dataset important? First large-scale Arabic medical dataset for AI applications.… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Medical_Consultations.textquestion-answering10K<n<100K13 likes377 downloads2y agoHugging Face03sadeem-ai /arabic-qna Sadeem QnA: An Arabic QnA Dataset 🌍✨ Welcome to the Sadeem QnA dataset, a vibrant collection designed for the advancement of Arabic natural language processing, specifically tailored for Question Answering (QnA) systems. Sourced from the rich and diverse content of Arabic Wikipedia, this dataset is a gateway to exploring the depths of Arabic language understanding, offering a unique challenge to both researchers and AI enthusiasts alike. About Sadeem QnA The Sadeem… See the full description on the dataset page: https://huggingface.co/datasets/sadeem-ai/arabic-qna.textquestion-answering1K<n<10K4 likes351 downloads3y agoHugging Face04QCRI /AraDiCE AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs Overview The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic. As part of the supplemental materials, we have selected a few datasets (see below) for the reader to review. We will make the full… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDiCE.imagetext-classificationn<1K3 likes219 downloads2y agoHugging Face05Ahmed-Selem /Shifaa_Arabic_Mental_Health_Consultations 🏥 Shifaa Arabic Mental Health Consultations 🧠 📌 Overview Shifaa Arabic Mental Health Consultations is a high-quality dataset designed to advance Arabic medical language models.This dataset provides 35,648 real-world medical consultations, covering a wide range of mental health concerns. 📊 Dataset Summary Size: 35,648 consultations Main Specializations: 7 Specific Diagnoses: 123 Languages: Arabic (العربية) Why This Dataset? 🔹 Lack of… See the full description on the dataset page: https://huggingface.co/datasets/Ahmed-Selem/Shifaa_Arabic_Mental_Health_Consultations.textquestion-answering10K<n<100K14 likes152 downloads2y agoHugging Face06AhmadHakami /saudipedia-arabic-qa Saudipedia Q&A Dataset Dataset Description Summary This dataset contains question-answer pairs scraped from Saudipedia, a comprehensive Arabic encyclopedia focused on Saudi Arabia. The dataset includes 1,082 Q&A entries covering various topics related to Saudi culture, history, economy, government, society, geography, religion, and notable personalities. The data was collected by scraping the website's question-answer section, which provides detailed answers to… See the full description on the dataset page: https://huggingface.co/datasets/AhmadHakami/saudipedia-arabic-qa.textquestion-answering1K<n<10K3 likes116 downloads1y agoHugging Face07miscovery /Math_CoT_Arabic_English_Reasoning Math CoT Arabic English Dataset A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI. Overview Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.tabularquestion-answering1K<n<10K17 likes71 downloads1y agoHugging Face08arafatar /toxic_uncensored_LGBTQ_csvtextquestion-answeringn<1K26 likes54 downloads2y agoHugging Face09drelhaj /ArabJobs ArabJobs: A Multinational Corpus of Arabic Job Advertisements 📖 Overview ArabJobs is the first publicly available, multinational corpus of Arabic job advertisements, collected fromEgypt, Jordan, Saudi Arabia, and the UAE. It contains: 8,546 job postings 550,000+ words Coverage across numerous sectors and dialects Rich metadata including salary, profession, gender indicators, and job categories This dataset supports research on: Fairness-aware Arabic NLP… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/ArabJobs.tabulartext-classification1K<n<10K0 likes49 downloads10mo agoHugging Face10Arailym-tleubayeva /legalup-laws Kazakhstan Legal Acts Dataset (LegalUp) Dataset Summary The LegalUp dataset contains structured metadata for legislative documents of the Republic of Kazakhstan. The current release includes 392,084 legislative document records extracted from a PostgreSQL database. The dataset is designed for: Legal Retrieval-Augmented Generation (Legal RAG) Information Retrieval Legal Search Question Answering Semantic Search Legal NLP Benchmark Construction Academic Research… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/legalup-laws.textquestion-answering100K<n<1M0 likes49 downloads15d agoHugging Face11miscovery /General_Facts_in_English_Arabic_Egyptian_Arabic 🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized) The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages: 🌍 English 🇸🇦 Modern Standard Arabic (MSA) 🇪🇬 Egyptian Arabic (Dialect) Each entry includes: The question and answer A category and sub-category Language tag (en, ar, ar_eg) Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.tabularquestion-answering10K<n<100K12 likes43 downloads1y agoHugging Face12arafatanam /Student-Mental-Health-Counseling-EN Student Mental Health Counseling (EN) Dataset Overview This dataset is the final cleaned and filtered version in a three-step pipeline that started with the original Vietnamese counseling dataset. Step Dataset Rows 1. Original chillies/student-mental-health-counseling-vn 750,169 2. Translated arafatanam/Student-Mental-Health-Counseling-750K 750,169 3. Filtered (this dataset) arafatanam/Student-Mental-Health-Counseling-EN 52,254 The goal of this… See the full description on the dataset page: https://huggingface.co/datasets/arafatanam/Student-Mental-Health-Counseling-EN.texttext-generation10K<n<100K0 likes43 downloads4mo agoHugging Face13go-inoue /ArabicMMLU_full Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne Introduction We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/go-inoue/ArabicMMLU_full.tabularquestion-answering10K<n<100K0 likes40 downloads8mo agoHugging Face144factors /arabic-palestinian-levantine-samplegated 4FACTORS — Palestinian Levantine Conversational Sample 50 native-written question–answer pairs in spoken Palestinian Levantine Arabic, each with an English gloss. This is a public demonstration sample from 4FACTORS, a producer of native, human-verified Arabic training data. What this is Real conversational exchanges — the kind of thing people actually say in shops, clinics, taxis, and at home — written from scratch by a first-language Palestinian speaker. Every… See the full description on the dataset page: https://huggingface.co/datasets/4factors/arabic-palestinian-levantine-sample.texttext-generationn<1K1 likes35 downloads2mo agoHugging Face15Youssefx64 /Medical-Consultation-Questions-in-Arabic Dataset Card for Dataset Name Dataset Details This dataset contains 47,705 Arabic medical questions collected from the Arabic health platform Altibbi. Each question is categorized into a medical domain such as sexual health, dermatology, pediatrics, and more. The dataset can be used for Natural Language Processing (NLP) tasks such as: "Text classification (predicting medical categories)". "Question answering systems in Arabic". "Building Arabic healthcare chatbots".… See the full description on the dataset page: https://huggingface.co/datasets/Youssefx64/Medical-Consultation-Questions-in-Arabic.texttext-classification10K<n<100K0 likes32 downloads5mo agoHugging Face16Asmaamaghraby /ArabicChartsQAtextquestion-answering10K<n<100K2 likes30 downloads3y agoHugging Face17CNTXTAI0 /arabic_dialects_question_and_answerData Content The file provided: Q/A Reasoning dataset contains the following columns: ID # : Denotes the reference ID for: a. Question b. Answer to the question c. Hint d. Reasoning e. Word count for items a to d above Dialects: Contains the following dialects in separate columns: a. English b. MSA c. Emirati d. Egyptian e. Levantine Syria f. Levantine Jordan g. Levantine Palestine h. Levantine Lebanon Data Generation Process The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.tabularquestion-answeringn<1K6 likes30 downloads2y agoHugging Face18ahmedselhady /ArabicMMLU_groupedRegrouped ArabicMMLU dataset. Groups are in accordance with the original dataset's categorization (see Table 1 in the original paper). Original ungrouped dataset is available under MBZUAI tabularquestion-answering10K<n<100K0 likes29 downloads2y agoHugging Face19Jr23xd23 /Arabic-Optimized-Reasoning-Dataset Arabic Optimized Reasoning Dataset Dataset Name: Arabic Optimized ReasoningLicense: Apache-2.0Formats: CSVSize: 1600 rowsBase Dataset: cognitivecomputations/dolphin-r1Libraries Used: Datasets, Dask, Croissant Overview The Arabic Optimized Reasoning Dataset helps AI models get better at reasoning in Arabic. While AI models are good at many tasks, they often struggle with reasoning in languages other than English. This dataset helps fix this problem by: Using fewer tokens… See the full description on the dataset page: https://huggingface.co/datasets/Jr23xd23/Arabic-Optimized-Reasoning-Dataset.textquestion-answering1K<n<10K5 likes28 downloads2y agoHugging Face20Jr23xd23 /Arabic_LLaMA_Math_Dataset Arabic LLaMA Math Dataset Example Entries Dataset Overview Dataset Name: Arabic_LLaMA_Math_Dataset.csv Number of Records: 12,496 Number of Columns: 3 File Format: CSV Dataset Structure Columns: Instruction: The problem statement or question (text, in Arabic) Input: Additional input for model fine-tuning (empty in this dataset) Solution: The solution or answer to the problem (text, in Arabic) Dataset Description The Arabic… See the full description on the dataset page: https://huggingface.co/datasets/Jr23xd23/Arabic_LLaMA_Math_Dataset.textquestion-answering10K<n<100K4 likes24 downloads2y agoHugging Face21miscovery /arabic_egypt_english_world_facts 🌍 Version (v2.0) World Facts in English, Arabic & Egyptian Arabic (Categorized) The World Facts General Knowledge Dataset (v2.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages: 🌍 English 🇸🇦 Modern Standard Arabic (MSA) 🇪🇬 Egyptian Arabic (Dialect) Each entry includes: The question and answer A category and sub-category Language tag (en, ar, ar_eg) Basic metadata:… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/arabic_egypt_english_world_facts.tabularquestion-answering10K<n<100K13 likes17 downloads1y agoHugging Face22go-inoue /ArabicMMLU_undiac Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne Introduction We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/go-inoue/ArabicMMLU_undiac.tabularquestion-answering10K<n<100K0 likes15 downloads1y agoHugging Face234factors /arabic-egyptian-samplegated 4FACTORS Arabic — Egyptian Q&A Sample Conversational question–answer pairs in spoken Egyptian Arabic, written by a first-language Egyptian speaker. A 50-item demonstration sample, with English glosses, released under CC BY-NC 4.0. This is the Egyptian variety in the 4FACTORS Arabic sample set, alongside the Palestinian Levantine and Modern Standard Arabic (MSA) sets. What this is Fifty short question–answer exchanges of the kind that come up in everyday life —… See the full description on the dataset page: https://huggingface.co/datasets/4factors/arabic-egyptian-sample.textquestion-answeringn<1K0 likes11 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.