CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01QCRI /ImageEval-ArabicNLP26 ImageEval-ArabicNLP26 👁️ ImageEval-ArabicNLP26 is the dataset of the ImageEval 2026 Shared Task at ArabicNLP 2026. It covers both of the shared task's tasks: AynVQA (Task 1), a culturally grounded Arabic multimodal benchmark for spoken visual question answering and hallucination detection, and CRAI-Bench (Task 2), which evaluates the cultural accuracy of Arabic text-to-image generation. The shared task has concluded. All gold labels are released, including the blind test splits… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ImageEval-ArabicNLP26.audio10K<n<100K4 likes2.5k downloads25d agoHugging Face02ameer4wisam /iraqi-arabic-sales-dialogue-dataset Iraqi Arabic Sales Dialogue Dataset A large synthetic dataset of Iraqi (Baghdadi-based) Arabic dialogue, centered on retail sales, haggling, and everyday conversation. النسخة العربية متوفرة بالكامل بالأسفل — Arabic version available in full below. What this is 210,832 template-generated conversations, of which 171,601 (81%) are exact-unique message sequences, spanning 20 topical categories in colloquial Iraqi Arabic. The core of the dataset (10 categories) is… See the full description on the dataset page: https://huggingface.co/datasets/ameer4wisam/iraqi-arabic-sales-dialogue-dataset.texttext-generation100K<n<1M0 likes706 downloads2mo agoHugging Face03QCRI /AraDICE-ArabicMMLU-egy AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs -- ArabicMMLU - Egyptian dialect Overview The AraDiCE dataset is crafted to assess the dialectal and cultural understanding of large language models (LLMs) within Arabic-speaking contexts. It includes post-edited adaptations of several benchmark datasets, specifically curated to validate LLM performance in culturally and dialectally relevant scenarios for Arabic. Within the AraDiCE collection, this… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDICE-ArabicMMLU-egy.texttext-classification10K<n<100K1 likes551 downloads2y agoHugging Face04sboughorbel /tinystories_dataset_arabictabular1M<n<10M1 likes375 downloads2y agoHugging Face05QCRI /LlamaLens-Arabic LlamaLens: Specialized Multilingual LLM Dataset Overview LlamaLens is a specialized multilingual LLM designed for analyzing news and social media content. It focuses on 18 NLP tasks, leveraging 52 datasets across Arabic, English, and Hindi. LlamaLens This repo includes scripts needed to run our full pipeline, including data preprocessing and sampling, instruction dataset creation, model fine-tuning, inference and evaluation. Features… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/LlamaLens-Arabic.texttext-classification1M<n<10M2 likes359 downloads2y agoHugging Face06QCRI /AraDICE-ArabicMMLU-lev AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs -- ArabicMMLU - Levantine dialect Overview The AraDiCE dataset is crafted to assess the dialectal and cultural understanding of large language models (LLMs) within Arabic-speaking contexts. It includes post-edited adaptations of several benchmark datasets, specifically curated to validate LLM performance in culturally and dialectally relevant scenarios for Arabic. Within the AraDiCE collection, this… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDICE-ArabicMMLU-lev.texttext-classification10K<n<100K0 likes358 downloads2y agoHugging Face07ISLAM-PO /arabic-to-code-8-langs-3m Current Hub Validation Status Repository claim: 3,000,000 records Dataset Server indexed rows: 1,239,045 Dataset Server estimate: 1,995,159 The 3M target figure is a raw-repository claim and is not yet fully verified by the Hub index. Validate the JSONL files before publishing a definitive record count. Arabic-to-Code Dataset - 8 Languages Train an Arabic-speaking code model with a raw target of 3,000,000 Arabic instruction-to-code pairs across 8 programming… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arabic-to-code-8-langs-3m.texttext-generation1M<n<10M0 likes298 downloads4h agoHugging Face08FreedomIntelligence /alpaca-gpt4-arabicThe dataset is used in the research related to MultilingualSIFT. text10K<n<100K12 likes228 downloads3y agoHugging Face09QCRI /LlamaLens-Arabic-Native LlamaLens: Specialized Multilingual LLM Dataset Overview LlamaLens is a specialized multilingual LLM designed for analyzing news and social media content. It focuses on 18 NLP tasks, leveraging 52 datasets across Arabic, English, and Hindi. LlamaLens This repo includes scripts needed to run our full pipeline, including data preprocessing and sampling, instruction dataset creation, model fine-tuning, inference and evaluation. Features… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/LlamaLens-Arabic-Native.texttext-classification1M<n<10M0 likes212 downloads2y agoHugging Face10oddadmix /arabic-guardrail Arabic Guardrail — 250,842 rows, 12 classes Defensive dataset for training Arabic prompt-safety classifiers. Each row is an incoming user message and which of 12 safety classes it belongs to. بالعربية: مجموعة بيانات عربية لتدريب نماذج تصنّف الرسائل الواردة قبل وصولها للمساعد الذكي. Arabic guardrails were a gap. Hugging Face searches for Arabic jailbreak / safety / prompt-injection datasets return zero results, and the one Arabic guardrail model… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-guardrail.tabulartext-classification100K<n<1M0 likes196 downloads25d agoHugging Face11halimbahae /Adala_Fonction_Publique_Maroc_Arabic_json_dataset 🏛️ Adala Fonction Publique Maroc Arabic Dataset This dataset contains structured legal data extracted from Moroccan public service law texts, sourced from adala.justice.gov.ma. The content is in Arabic and is designed to support NLP and AI applications in legal tech, especially for Moroccan administrative and public law. 📂 Dataset Structure Format: JSON Language: Arabic (Standard & Legal dialect) Content: Articles, chapters, titles from Moroccan public law texts… See the full description on the dataset page: https://huggingface.co/datasets/halimbahae/Adala_Fonction_Publique_Maroc_Arabic_json_dataset.textn<1K0 likes195 downloads1y agoHugging Face12SaudiArabicLaws /saudi-arabic-laws-and-regulations-corpus Saudi Arabic Laws and Regulations Corpus A structured, article-level Arabic legal corpus containing 22,593 records from 578 official Saudi legal documents, prepared for Arabic legal information retrieval, Retrieval-Augmented Generation (RAG), grounded generation, and LLM evaluation. Release: 1.0.0Language: ArabicDomain: Saudi laws and regulationsGranularity: Legal articleTotal records: 22,593Archival DOI: 10.5281/zenodo.21265180 Quick Start The corpus is… See the full description on the dataset page: https://huggingface.co/datasets/SaudiArabicLaws/saudi-arabic-laws-and-regulations-corpus.texttext-retrieval10K<n<100K1 likes141 downloads1mo agoHugging Face13HeshamHaroon /Arabic_Function_Calling Arabic Function Calling Dataset (50K+ Samples) مجموعة بيانات استدعاء الدوال العربية أول وأكبر مجموعة بيانات عربية متخصصة في استدعاء الدوال (Function Calling) تغطي جميع اللهجات العربية الرئيسية والمجالات الحياتية المهمة. Dataset Description This is the first comprehensive Arabic function calling dataset designed for training and evaluating LLMs on Arabic tool use capabilities. The dataset covers: 5 Arabic Dialects: MSA (Modern Standard Arabic), Egyptian… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/Arabic_Function_Calling.texttext-generation10K<n<100K60 likes131 downloads9mo agoHugging Face14ISLAM-PO /arabic-history-and-dialects مجموعة البيانات العربية الشاملة للذكاء الاصطناعي 🇸🇦🇪🇬🇱🇧🇲🇦 Arabic Multi-Dialect & Civilization Instruction Dataset المؤلف: islam-alnasherA-Dev — الحساب: https://huggingface.co/ISLAM-POالإصدار: v1.0 — التاريخ: 30 أغسطس 2026 — الترخيص: CC BY 4.0اللغة: العربية (فصحى + 4 لهجات) — الصيغة: instruction / output JSONL — الحجم: 275 عينة 📌 الملخص التنفيذي هذه المجموعة هي مورد تعليمي متخصص لتدريب وتقييم النماذج اللغوية العربية على مسارين متوازيين:… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arabic-history-and-dialects.textquestion-answeringn<1K0 likes125 downloads4h agoHugging Face15oddadmix /arabic-prompt-routing Arabic Prompt Routing — توجيه عربي صفري 233,720 rows. Each row is a text, a set of free-text categories, and which one it belongs to. Categories are arbitrary Arabic — the point is a model that routes into a label set it has never seen. بالعربية: مجموعة بيانات عربية لتوجيه النصوص إلى فئات يكتبها المستخدم بلغة طبيعية. الفئات ليست ثابتة، والهدف نموذج يوجّه إلى فئات لم يرها أثناء التدريب. Arabic counterpart to the task in LiquidAI/LFM2.5-Encoder-350M-Prompt-Router. split… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-prompt-routing.tabularzero-shot-classification100K<n<1M0 likes123 downloads21d agoHugging Face16HeshamHaroon /arabic-quotes Arabic Quotes Dataset (arabic_Q) The "Arabic Quotes" dataset contains a collection of Arabic quotes along with their corresponding authors and tags. The dataset is scraped from the website "arabic-quotes.com" and provides a diverse range of quotes from various authors. Dataset Details Version: 1.0.0 Total Quotes: 3778 Languages: Arabic Source: arabic-quotes.com Dataset Structure The dataset is provided in the JSONL (JSON Lines) format, where each line… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-quotes.texttext-classification1K<n<10K7 likes119 downloads3y agoHugging Face17Mo-Abdalkader /Egyptian-Arabic-English-Parallel-Corpus Egyptian Arabic-English Parallel Corpus Author: Mohamed Abdalkader · LinkedIn · GitHub A comprehensive Egyptian Arabic → English parallel corpus covering 1,800 topics from daily Egyptian life. Designed for fine-tuning large language models on Egyptian Arabic dialect translation and generation. Dataset Structure egyptian-arabic-english-parallel-corpus/ ├── SFT/ │ ├── Train/ │ │ ├── topics/ # 1,800 individual topic JSON files │ │ └── merged/… See the full description on the dataset page: https://huggingface.co/datasets/Mo-Abdalkader/Egyptian-Arabic-English-Parallel-Corpus.text100K<n<1M1 likes119 downloads6d agoHugging Face18HeshamHaroon /ArabicRAGB ArabicRAGB: Arabic Retrieval-Augmented Generation Benchmark Dataset Description ArabicRAGB is a benchmark dataset for evaluating Retrieval-Augmented Generation (RAG) systems on Arabic language tasks. Each record contains a query-passage pair where the query is grounded in the passage content. Key Features Passage-Grounded Queries: Each query is generated from and answerable by its paired passage Multi-Dialect Coverage: MSA, Egyptian, Gulf… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/ArabicRAGB.texttext-retrieval10K<n<100K13 likes118 downloads9mo agoHugging Face19FreedomIntelligence /evol-instruct-arabicThe dataset is used in the research related to MultilingualSIFT. text10K<n<100K2 likes110 downloads3y agoHugging Face20FreedomIntelligence /ACVA-Arabic-Cultural-Value-Alignment About ArabicCulture The ArabicCulture dataset was generated by gpt3.5 and contains 8000+ True and False questions.The dataset contains questions from 58 different areas.In the answers, "True" accounted for 59.62%, and "False" accounted for 40.38% data-all It contains 8000+ data, and we took 5 data from each area as few-shot data. data-select We asked two Arabs to judge 4000 of all the data for us, and we left data that two Arabs both thought were good. Finally… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ACVA-Arabic-Cultural-Value-Alignment.text1K<n<10K8 likes108 downloads3y agoHugging Face21syamjithnk /arabic-corpus-audit Arabic Corpus Integrity Audit Author: Syamjith NK Date: 9 September 2026 · corrected 13 September 2026 Tool: arabic-lint 0.5.0 Correction, 13 September 2026. An earlier version of this card said the labels in Yousefmd/arabic_ocr_dataset were stored in visual order, and called that the full reshape + bidi signature. That was wrong. Only the shaping step ran; the words are in logical order and plain NFKC recovers them. What was measured, and stands, is that the labels store… See the full description on the dataset page: https://huggingface.co/datasets/syamjithnk/arabic-corpus-audit.tabulartext-classificationn<1K0 likes107 downloads11d agoHugging Face22FreedomIntelligence /sharegpt-arabicArabic ShareGPT data translated by gpt-3.5-turbo.The dataset is used in the research related to MultilingualSIFT. text1K<n<10K5 likes96 downloads3y agoHugging Face23oddadmix /arabic-rule-checking Arabic Rule Checking — قواعد ونصوص عربية بأحكام محسوبة 172,488 labelled (text, rule) pairs in Arabic. Each row asks one question: does this text satisfy this rule? The answer is مطابق or مخالف. بالعربية: مجموعة بيانات عربية للتحقق من مطابقة النصوص لقواعد مكتوبة بلغة طبيعية. كل صف يحتوي على نص وقاعدة وحكم محسوب آليًا، وليس رأي نموذج. split pairs texts train 159,240 48,030 validation 13,248 2,002 Built from 50,062 generated Arabic texts across 12 document types… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rule-checking.tabulartext-classification100K<n<1M0 likes90 downloads23d agoHugging Face24inception42 /Arabic-IFEvalIFEval is the first publicly available benchmark dataset specifically designed to evaluate Arabic Large Language Models (LLMs) on instruction-following capabilities in Arabic. The dataset includes 404 high-quality, manually verified samples covering various constraints such as linguistic patterns, punctuation rules, and formatting guidelines. Loading the Dataset To load this dataset in Python using the 🤗 Datasets library, run the following: from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/inception42/Arabic-IFEval.textn<1K5 likes88 downloads1y agoHugging Face25dev-hussein /saudi-arabic-cs-conversations Saudi Arabic Customer Service Conversations — Free 100 Sample 100 synthetic multi-turn conversations in authentic Saudi Arabic dialects Built for LLM fine-tuning, chatbot training, and Arabic NLP research Overview This is a free 100-conversation sample from a production-quality dataset of 50,000 Saudi Arabic customer service conversations. Every conversation is fully synthetic — no real user data — and safe for commercial use. Each conversation simulates a… See the full description on the dataset page: https://huggingface.co/datasets/dev-hussein/saudi-arabic-cs-conversations.texttext-generationn<1K1 likes84 downloads6mo agoHugging Face26QCRI /ArabicCulturalQA ArabicCulturalQA ArabicCulturalQA is the first cross-dialectal Arabic cultural QA benchmark with parallel multiple-choice (MCQ) and open-ended (OEQ) formats across Modern Standard Arabic (MSA), English, Egyptian, Levantine, Gulf, and Maghrebi. Both the MCQ and OEQ test sets have been reviewed and post-edited by native speakers of each dialect. The dataset accompanies the LREC 2026 paper "Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants" (paper page)… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ArabicCulturalQA.textquestion-answering10K<n<100K2 likes78 downloads3mo agoHugging Face27beetleware /arabic-reasoning-dataset-logic Arabic Logical Reasoning Tasks Dataset (Maximum 1000 Tasks) Overview This dataset comprises a series of logical reasoning tasks designed to evaluate and train artificial intelligence models on understanding and generating logical inferences in the Arabic language. Each task includes a unique identifier, the task type, the task text (a question and a proposed answer), and a detailed solution that outlines the thinking steps and the final answer. Data Format The… See the full description on the dataset page: https://huggingface.co/datasets/beetleware/arabic-reasoning-dataset-logic.text1K<n<10K14 likes77 downloads1y agoHugging Face28arbml /alpaca_arabictext10K<n<100K4 likes76 downloads3y agoHugging Face29albaz2000 /arabic-itsm-dataset Arabic ITSM Dataset A synthetic dataset of 10,000 Arabic IT support tickets, labeled with a structured 3-level ITSM taxonomy, generated using LLMs, and validated programmatically before release. Tickets are written in Egyptian Arabic (عامية مصرية) and cover the full range of helpdesk scenarios: access issues, network problems, hardware faults, software errors, security incidents, and service requests. Arabic technical vocabulary is mixed with English terms as they naturally… See the full description on the dataset page: https://huggingface.co/datasets/albaz2000/arabic-itsm-dataset.tabulartext-classification10K<n<100K0 likes76 downloads21d agoHugging Face30nawaralseelawi /mizan-iraqi-arabic-benchmark Mizan (ميزان) — Iraqi Arabic LLM Benchmark: pilot-0.2 public development set Mizan is the first comprehensive, originally-authored evaluation benchmark for Iraqi Arabic and the Iraqi civic context. This dataset is the pilot-0.2 public development set: 340 originally-authored, dually-reviewed items across two tracks (MSA baseline / Iraqi) and six axes. 📄 Paper (preprint): https://doi.org/10.5281/zenodo.22714865 🏆 Live leaderboard: https://mizan-bench.onrender.com 💻 Code… See the full description on the dataset page: https://huggingface.co/datasets/nawaralseelawi/mizan-iraqi-arabic-benchmark.textquestion-answeringn<1K1 likes74 downloads12d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.