datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ArabicMMLU
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin
MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne
Introduction
We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/ArabicMMLU.arabic-stem-lexicon
Arabic Diacritized-Stem Lexicon
An undiacritized Arabic surface form → its most frequent diacritized stem.
Standard Arabic writes no short vowels, so anything that has to pronounce Arabic
must first put them back. A neural diacritizer does that well on rare words, where
inference is the only thing there is. On common words it is the wrong tool:
which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked
up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.arabic-sentiments2sada2022-arabic-tts
SADA 2022 - Saudi Arabic Dataset for TTS
مجموعة بيانات صوتية سعودية للنص إلى كلام (Text-to-Speech)
المصدر الأصلي
Kaggle: sdaiancai/sada2022
الاستخدام
# طريقة 1: Git Clone
!git clone https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts /content/saudi_dataset
# طريقة 2: مكتبة datasets
from datasets import load_dataset
dataset = load_dataset("aalshalfi/sada2022-arabic-tts")
الملفات
valid.csv - ملف البيانات الرئيسي
wavs/ - ملفات الصوت… See the full description on the dataset page: https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts.KazakhLawCorpus-clean
KazakhLawCorpus-clean
Dataset Summary
KazakhLawCorpus-clean is a cleaned, Kazakh-only corpus of legislative documents from the Republic of Kazakhstan. It is a processed derivative of the original Arailym-tleubayeva/KazakhLawCorpus dataset.
The original dataset repository was downloaded from Hugging Face and used as the source for this release. Its laws_metadata.csv file contained 223,245 legislative records with multilingual fields and source-oriented metadata.… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhLawCorpus-clean.hatecheck-arabic
Dataset Card for Multilingual HateCheck
Dataset Description
Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish.
For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate.
This allows for targeted diagnostic insights into model performance.
For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-arabic.Moroccan-Arabic-Multimodal-Emotion-Recognition
MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging)
A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits.
Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.Math_CoT_Arabic_English_Reasoning
Math CoT Arabic English Dataset
A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI.
Overview
Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.os-rfodg-outdoor-uav-synthetic-dataset-taif-saudi-arabia
UAV Trajectory Simulation Dataset for Terrain-Based Localization
Dataset Overview
This dataset contains simulated UAV flight data generated using ROS2, Gazebo, and PX4 autopilot system. The dataset features a quadcopter performing autonomous flight trajectories over realistic terrain imported from satellite imagery and Digital Elevation Model (DEM) maps of the Taif region in Saudi Arabia.
Dataset Files
The dataset contains:
7 trajectory CSV files:… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/os-rfodg-outdoor-uav-synthetic-dataset-taif-saudi-arabia.Arabic-Poetry-Datasetarabic-hate-speech-superset
Arabic Hate Speech Superset
This dataset is a superset (N=449,078) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Arabic hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available or could be retrieved with the Twitter API
focus on hate speech, defined broadly as "any kind of… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/arabic-hate-speech-superset.Arabic-Emotional-Audio-Dataset-Baved
BAVED — Basic Arabic Vocal Emotions Dataset (TTS-ready repackaging)
A re-packaged, transcript-aligned version of the Basic Arabic Vocal Emotions Dataset (BAVED) with explicit Arabic transcripts, English glosses, speaker metadata, and speaker-disjoint train/validation/test splits.
Original dataset: Aouf Yacine, Basic Arabic Vocal Emotions Dataset (BAVED), GitHub: https://github.com/40uf411/Basic-Arabic-Vocal-Emotions-Dataset. This repackaging adds metadata; all audio is unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Arabic-Emotional-Audio-Dataset-Baved.ArabJobs
ArabJobs: A Multinational Corpus of Arabic Job Advertisements
📖 Overview
ArabJobs is the first publicly available, multinational corpus of Arabic job advertisements, collected fromEgypt, Jordan, Saudi Arabia, and the UAE.
It contains:
8,546 job postings
550,000+ words
Coverage across numerous sectors and dialects
Rich metadata including salary, profession, gender indicators, and job categories
This dataset supports research on:
Fairness-aware Arabic NLP… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/ArabJobs.General_Facts_in_English_Arabic_Egyptian_Arabic
🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized)
The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.KazOilWellOps_Dataset
Kazakhstan Oil Well Operational Dataset
Description
This dataset contains structured operational and production parameters of sucker rod pump (SRP) oil wells in Kazakhstan.
It is intended for industrial AI research, oil production analysis, production forecasting, and predictive modeling of well performance under real field operating conditions.
Location: North-West Konys oil field, Kyzylorda Region, Kazakhstan (≈150 km NW of Kyzylorda city).
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazOilWellOps_Dataset.ArabicMMLU_groupedRegrouped ArabicMMLU dataset. Groups are in accordance with the original dataset's categorization (see Table 1 in the original paper).
Original ungrouped dataset is available under MBZUAI
ArabicMMLU_full
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin
MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne
Introduction
We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/go-inoue/ArabicMMLU_full.arabic-cv-scoring-dataset
Arabic CV Scoring Dataset
Dataset Summary
This dataset contains ~7,220 synthetically generated Arabic CVs, each paired
with a job category, an ATS (Applicant Tracking System) compatibility score,
and a suitability score/class label. It was built to train and evaluate the
Arabic CV Analyzer —
an NLP pipeline that scores, classifies, and generates improvement suggestions
for Arabic CVs targeting the Arab job market, where no equivalent
ATS-optimization tooling… See the full description on the dataset page: https://huggingface.co/datasets/omaraboelmaaty/arabic-cv-scoring-dataset.arabic_dialects_question_and_answerData Content
The file provided: Q/A Reasoning dataset
contains the following columns:
ID # : Denotes the reference ID for:
a. Question
b. Answer to the question
c. Hint
d. Reasoning
e. Word count for items a to d above
Dialects: Contains the following dialects in separate columns:
a. English
b. MSA
c. Emirati
d. Egyptian
e. Levantine Syria
f. Levantine Jordan
g. Levantine Palestine
h. Levantine Lebanon
Data Generation Process
The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.Ara-Best-RQ_dataset
Ara-Best-RQ Dataset
Dataset Summary
This dataset provides metadata only for a dialectal Arabic speech corpus constructed from publicly available YouTube videos.It consists exclusively of YouTube video identifiers and audio segment boundaries (start/end timestamps) designed for self-supervised speech representation learning.
No audio or video content is distributed as part of this dataset.
Dataset Statistics
Total spoken duration: 5,639 h 04 min 27 s… See the full description on the dataset page: https://huggingface.co/datasets/Elyadata/Ara-Best-RQ_dataset.arabic-player-stats
👤 YallaShoot — إحصاءات اللاعبين العرب
مجموعة بيانات تضم إحصاءات تفصيلية للاعبي كرة القدم العرب والمحترفين في الدوريات العربية، مُعدَّة لتدريب نماذج تحليل الأداء الرياضي.
📌 وصف مجموعة البيانات
تشمل هذه المجموعة بيانات موسمية تفصيلية للاعبين في:
🇸🇦 دوري روشن السعودي للمحترفين
🇪🇬 الدوري المصري الممتاز
🌍 المنتخبات الوطنية العربية
🏆 المحترفون العرب في الدوريات الأوروبية
📂 هيكل البيانات
العمود
النوع
الوصف
player_id
string
معرّف اللاعب
name_ar… See the full description on the dataset page: https://huggingface.co/datasets/yallashoot/arabic-player-stats.arabic-match-results
⚽ YallaShoot — نتائج المباريات العربية
مجموعة بيانات شاملة لنتائج مباريات كرة القدم في الدوريات العربية والعالمية، مُجمَّعة من منصة يلا شوت لخدمة نماذج الذكاء الاصطناعي في تحليل الرياضة.
📌 وصف مجموعة البيانات
تحتوي هذه المجموعة على نتائج المباريات من أبرز الدوريات:
🇸🇦 دوري روشن للمحترفين (السعودي)
🇪🇬 الدوري المصري الممتاز
🇦🇪 دوري الخليج العربي (الإماراتي)
🏆 دوري أبطال أوروبا
🏴 الدوري الإنجليزي الممتاز
📂 هيكل البيانات
العمود
النوع… See the full description on the dataset page: https://huggingface.co/datasets/yallashoot/arabic-match-results.sist-kazakh-corpus
SIST Kazakh Corpus
Description
SIST Kazakh Corpus is a curated dataset of Kazakh scientific articles
collected for research in text similarity detection, plagiarism analysis,
and low-resource NLP tasks.
The dataset was created to support:
Text similarity detection in agglutinative languages
Kazakh NLP benchmarking
Scientific text analysis
Retrieval-Augmented Generation (RAG) research
Dataset Structure
The dataset is provided in CSV format.
Columns may… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/sist-kazakh-corpus.Africa-Arable-land-hectares-per-person
Africa Arable land hectares per person | Africa (World Bank)
Size category: n<1K - Formats: csv - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts inspect… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Africa-Arable-land-hectares-per-person.arabic-fan-sentiment
💬 YallaShoot — تغريدات المشجعين العرب: تحليل المشاعر
مجموعة بيانات من تغريدات مشجعي كرة القدم العرب، مُصنَّفة يدوياً لتحليل المشاعر. صُمِّمت لتدريب نماذج NLP على التمييز بين التعليقات الإيجابية والسلبية والمحايدة في سياق الرياضة العربية.
📌 وصف مجموعة البيانات
تحتوي على تغريدات متعلقة بـ:
🎯 ردود فعل المشجعين بعد المباريات
😤 انتقادات الحكام والقرارات
🏆 تعليقات على انتقالات اللاعبين
🔥 التنافسات الكلاسيكية (الكلاسيكو، القمة...)
📺 آراء حول التغطية الإعلامية… See the full description on the dataset page: https://huggingface.co/datasets/yallashoot/arabic-fan-sentiment.sts-arabic-translated-modifiedtestarabic_egypt_english_world_facts
🌍 Version (v2.0) World Facts in English, Arabic & Egyptian Arabic (Categorized)
The World Facts General Knowledge Dataset (v2.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages:
🌍 English
🇸🇦 Modern Standard Arabic (MSA)
🇪🇬 Egyptian Arabic (Dialect)
Each entry includes:
The question and answer
A category and sub-category
Language tag (en, ar, ar_eg)
Basic metadata:… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/arabic_egypt_english_world_facts.benchmark-1-arabic-m2mInfo:
Translated on Arabic by facebook/m2m100_418M model
Source: jayavibhav/prompt-injection-safety
Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns
Size: 1,000 prompts (500 safe / 500 unsafe)
Columns:
text - original prompt
label - 0: safe, 1: unsafe
translation - prompt on Arabic translated by facebook/m2m100_418M
score_ar_model - cosine similarity score with codebook
More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-1-arabic-m2m.VulnSage
VulnSage Dataset
VulnSage is a curated dataset designed for research on automated vulnerability detection, particularly leveraging the capabilities of large language models (LLMs). It contains annotated vulnerable and patched code snippets from real-world software projects, along with rich metadata and contextual reasoning.
📦 Dataset Contents
The dataset includes 593 vulnerability instances extracted from various open-source software repositories. Each entry provides… See the full description on the dataset page: https://huggingface.co/datasets/Arastoorad/VulnSage.
