CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MBZUAI /ArabicMMLU Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne Introduction We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/ArabicMMLU.tabularquestion-answering10K<n<100K39 likes3.7k downloads2y agoHugging Face02TigreGotico /arabic-stem-lexicon Arabic Diacritized-Stem Lexicon An undiacritized Arabic surface form → its most frequent diacritized stem. Standard Arabic writes no short vowels, so anything that has to pronounce Arabic must first put them back. A neural diacritizer does that well on rare words, where inference is the only thing there is. On common words it is the wrong tool: which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.tabulartext-to-speech100K<n<1M0 likes3k downloads2mo agoHugging Face03hmmamalrjoub /arabic-sentiments2tabulartext-classification1K<n<10K0 likes132 downloads2y agoHugging Face04aalshalfi /sada2022-arabic-tts SADA 2022 - Saudi Arabic Dataset for TTS مجموعة بيانات صوتية سعودية للنص إلى كلام (Text-to-Speech) المصدر الأصلي Kaggle: sdaiancai/sada2022 الاستخدام # طريقة 1: Git Clone !git clone https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts /content/saudi_dataset # طريقة 2: مكتبة datasets from datasets import load_dataset dataset = load_dataset("aalshalfi/sada2022-arabic-tts") الملفات valid.csv - ملف البيانات الرئيسي wavs/ - ملفات الصوت… See the full description on the dataset page: https://huggingface.co/datasets/aalshalfi/sada2022-arabic-tts.audio100K<n<1M0 likes117 downloads8mo agoHugging Face05Arailym-tleubayeva /KazakhLawCorpus-clean KazakhLawCorpus-clean Dataset Summary KazakhLawCorpus-clean is a cleaned, Kazakh-only corpus of legislative documents from the Republic of Kazakhstan. It is a processed derivative of the original Arailym-tleubayeva/KazakhLawCorpus dataset. The original dataset repository was downloaded from Hugging Face and used as the source for this release. Its laws_metadata.csv file contained 223,245 legislative records with multilingual fields and source-oriented metadata.… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazakhLawCorpus-clean.tabulartext-retrieval100K<n<1M1 likes117 downloads13d agoHugging Face06Paul /hatecheck-arabic Dataset Card for Multilingual HateCheck Dataset Description Multilingual HateCheck (MHC) is a suite of functional tests for hate speech detection models in 10 different languages: Arabic, Dutch, French, German, Hindi, Italian, Mandarin, Polish, Portuguese and Spanish. For each language, there are 25+ functional tests that correspond to distinct types of hate and challenging non-hate. This allows for targeted diagnostic insights into model performance. For more details… See the full description on the dataset page: https://huggingface.co/datasets/Paul/hatecheck-arabic.tabulartext-classification1K<n<10K5 likes93 downloads4y agoHugging Face07FatimahEmadEldin /Moroccan-Arabic-Multimodal-Emotion-Recognition MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging) A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits. Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.audiotext-to-speech1K<n<10K1 likes80 downloads5mo agoHugging Face08miscovery /Math_CoT_Arabic_English_Reasoning Math CoT Arabic English Dataset A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI. Overview Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.tabularquestion-answering1K<n<10K17 likes77 downloads1y agoHugging Face09riotu-lab /os-rfodg-outdoor-uav-synthetic-dataset-taif-saudi-arabia UAV Trajectory Simulation Dataset for Terrain-Based Localization Dataset Overview This dataset contains simulated UAV flight data generated using ROS2, Gazebo, and PX4 autopilot system. The dataset features a quadcopter performing autonomous flight trajectories over realistic terrain imported from satellite imagery and Digital Elevation Model (DEM) maps of the Taif region in Saudi Arabia. Dataset Files The dataset contains: 7 trajectory CSV files:… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/os-rfodg-outdoor-uav-synthetic-dataset-taif-saudi-arabia.tabular100K<n<1M1 likes69 downloads1y agoHugging Face10Fatimah8Moheeb /Arabic-Poetry-Datasettabular100K<n<1M0 likes61 downloads3mo agoHugging Face11manueltonneau /arabic-hate-speech-supersetgated Arabic Hate Speech Superset This dataset is a superset (N=449,078) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Arabic hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that: are documented are publicly available or could be retrieved with the Twitter API focus on hate speech, defined broadly as "any kind of… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/arabic-hate-speech-superset.tabulartext-classification100K<n<1M8 likes53 downloads2y agoHugging Face12FatimahEmadEldin /Arabic-Emotional-Audio-Dataset-Baved BAVED — Basic Arabic Vocal Emotions Dataset (TTS-ready repackaging) A re-packaged, transcript-aligned version of the Basic Arabic Vocal Emotions Dataset (BAVED) with explicit Arabic transcripts, English glosses, speaker metadata, and speaker-disjoint train/validation/test splits. Original dataset: Aouf Yacine, Basic Arabic Vocal Emotions Dataset (BAVED), GitHub: https://github.com/40uf411/Basic-Arabic-Vocal-Emotions-Dataset. This repackaging adds metadata; all audio is unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Arabic-Emotional-Audio-Dataset-Baved.audioaudio-classification1K<n<10K0 likes51 downloads5mo agoHugging Face13drelhaj /ArabJobs ArabJobs: A Multinational Corpus of Arabic Job Advertisements 📖 Overview ArabJobs is the first publicly available, multinational corpus of Arabic job advertisements, collected fromEgypt, Jordan, Saudi Arabia, and the UAE. It contains: 8,546 job postings 550,000+ words Coverage across numerous sectors and dialects Rich metadata including salary, profession, gender indicators, and job categories This dataset supports research on: Fairness-aware Arabic NLP… See the full description on the dataset page: https://huggingface.co/datasets/drelhaj/ArabJobs.tabulartext-classification1K<n<10K0 likes47 downloads10mo agoHugging Face14miscovery /General_Facts_in_English_Arabic_Egyptian_Arabic 🌍 World Facts in English, Arabic & Egyptian Arabic (v1.0) (Categorized) The World Facts General Knowledge Dataset (v1.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages: 🌍 English 🇸🇦 Modern Standard Arabic (MSA) 🇪🇬 Egyptian Arabic (Dialect) Each entry includes: The question and answer A category and sub-category Language tag (en, ar, ar_eg) Basic metadata: question &… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/General_Facts_in_English_Arabic_Egyptian_Arabic.tabularquestion-answering10K<n<100K12 likes41 downloads1y agoHugging Face15Arailym-tleubayeva /KazOilWellOps_Dataset Kazakhstan Oil Well Operational Dataset Description This dataset contains structured operational and production parameters of sucker rod pump (SRP) oil wells in Kazakhstan. It is intended for industrial AI research, oil production analysis, production forecasting, and predictive modeling of well performance under real field operating conditions. Location: North-West Konys oil field, Kyzylorda Region, Kazakhstan (≈150 km NW of Kyzylorda city). Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/KazOilWellOps_Dataset.tabulartabular-regressionn<1K0 likes37 downloads7mo agoHugging Face16ahmedselhady /ArabicMMLU_groupedRegrouped ArabicMMLU dataset. Groups are in accordance with the original dataset's categorization (see Table 1 in the original paper). Original ungrouped dataset is available under MBZUAI tabularquestion-answering10K<n<100K0 likes31 downloads2y agoHugging Face17go-inoue /ArabicMMLU_full Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne Introduction We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/go-inoue/ArabicMMLU_full.tabularquestion-answering10K<n<100K0 likes31 downloads8mo agoHugging Face18omaraboelmaaty /arabic-cv-scoring-dataset Arabic CV Scoring Dataset Dataset Summary This dataset contains ~7,220 synthetically generated Arabic CVs, each paired with a job category, an ATS (Applicant Tracking System) compatibility score, and a suitability score/class label. It was built to train and evaluate the Arabic CV Analyzer — an NLP pipeline that scores, classifies, and generates improvement suggestions for Arabic CVs targeting the Arab job market, where no equivalent ATS-optimization tooling… See the full description on the dataset page: https://huggingface.co/datasets/omaraboelmaaty/arabic-cv-scoring-dataset.tabulartext-classification1K<n<10K0 likes31 downloads1mo agoHugging Face19CNTXTAI0 /arabic_dialects_question_and_answerData Content The file provided: Q/A Reasoning dataset contains the following columns: ID # : Denotes the reference ID for: a. Question b. Answer to the question c. Hint d. Reasoning e. Word count for items a to d above Dialects: Contains the following dialects in separate columns: a. English b. MSA c. Emirati d. Egyptian e. Levantine Syria f. Levantine Jordan g. Levantine Palestine h. Levantine Lebanon Data Generation Process The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.tabularquestion-answeringn<1K6 likes30 downloads2y agoHugging Face20Elyadata /Ara-Best-RQ_dataset Ara-Best-RQ Dataset Dataset Summary This dataset provides metadata only for a dialectal Arabic speech corpus constructed from publicly available YouTube videos.It consists exclusively of YouTube video identifiers and audio segment boundaries (start/end timestamps) designed for self-supervised speech representation learning. No audio or video content is distributed as part of this dataset. Dataset Statistics Total spoken duration: 5,639 h 04 min 27 s… See the full description on the dataset page: https://huggingface.co/datasets/Elyadata/Ara-Best-RQ_dataset.tabular1M<n<10M1 likes28 downloads8mo agoHugging Face21yallashoot /arabic-player-stats 👤 YallaShoot — إحصاءات اللاعبين العرب مجموعة بيانات تضم إحصاءات تفصيلية للاعبي كرة القدم العرب والمحترفين في الدوريات العربية، مُعدَّة لتدريب نماذج تحليل الأداء الرياضي. 📌 وصف مجموعة البيانات تشمل هذه المجموعة بيانات موسمية تفصيلية للاعبين في: 🇸🇦 دوري روشن السعودي للمحترفين 🇪🇬 الدوري المصري الممتاز 🌍 المنتخبات الوطنية العربية 🏆 المحترفون العرب في الدوريات الأوروبية 📂 هيكل البيانات العمود النوع الوصف player_id string معرّف اللاعب name_ar… See the full description on the dataset page: https://huggingface.co/datasets/yallashoot/arabic-player-stats.tabulartable-question-answeringn<1K0 likes22 downloads6mo agoHugging Face22yallashoot /arabic-match-results ⚽ YallaShoot — نتائج المباريات العربية مجموعة بيانات شاملة لنتائج مباريات كرة القدم في الدوريات العربية والعالمية، مُجمَّعة من منصة يلا شوت لخدمة نماذج الذكاء الاصطناعي في تحليل الرياضة. 📌 وصف مجموعة البيانات تحتوي هذه المجموعة على نتائج المباريات من أبرز الدوريات: 🇸🇦 دوري روشن للمحترفين (السعودي) 🇪🇬 الدوري المصري الممتاز 🇦🇪 دوري الخليج العربي (الإماراتي) 🏆 دوري أبطال أوروبا 🏴󠁧󠁢󠁥󠁮󠁧󠁿 الدوري الإنجليزي الممتاز 📂 هيكل البيانات العمود النوع… See the full description on the dataset page: https://huggingface.co/datasets/yallashoot/arabic-match-results.tabulartext-classificationn<1K1 likes21 downloads6mo agoHugging Face23Arailym-tleubayeva /sist-kazakh-corpus SIST Kazakh Corpus Description SIST Kazakh Corpus is a curated dataset of Kazakh scientific articles collected for research in text similarity detection, plagiarism analysis, and low-resource NLP tasks. The dataset was created to support: Text similarity detection in agglutinative languages Kazakh NLP benchmarking Scientific text analysis Retrieval-Augmented Generation (RAG) research Dataset Structure The dataset is provided in CSV format. Columns may… See the full description on the dataset page: https://huggingface.co/datasets/Arailym-tleubayeva/sist-kazakh-corpus.tabularsentence-similarityn<1K1 likes18 downloads7mo agoHugging Face24electricsheepafrica /Africa-Arable-land-hectares-per-person Africa Arable land hectares per person | Africa (World Bank) Size category: n<1K - Formats: csv - Sector: agriculture_food - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public datasets help analysts inspect… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Africa-Arable-land-hectares-per-person.tabulartabular-classificationn<1K0 likes18 downloads1mo agoHugging Face25yallashoot /arabic-fan-sentiment 💬 YallaShoot — تغريدات المشجعين العرب: تحليل المشاعر مجموعة بيانات من تغريدات مشجعي كرة القدم العرب، مُصنَّفة يدوياً لتحليل المشاعر. صُمِّمت لتدريب نماذج NLP على التمييز بين التعليقات الإيجابية والسلبية والمحايدة في سياق الرياضة العربية. 📌 وصف مجموعة البيانات تحتوي على تغريدات متعلقة بـ: 🎯 ردود فعل المشجعين بعد المباريات 😤 انتقادات الحكام والقرارات 🏆 تعليقات على انتقالات اللاعبين 🔥 التنافسات الكلاسيكية (الكلاسيكو، القمة...) 📺 آراء حول التغطية الإعلامية… See the full description on the dataset page: https://huggingface.co/datasets/yallashoot/arabic-fan-sentiment.tabulartext-classificationn<1K0 likes18 downloads6mo agoHugging Face26Zynab /sts-arabic-translated-modifiedtabular1K<n<10K0 likes17 downloads3y agoHugging Face27Arabic-Image-Captioning-latest /testimage1M<n<10M2 likes17 downloads3y agoHugging Face28miscovery /arabic_egypt_english_world_facts 🌍 Version (v2.0) World Facts in English, Arabic & Egyptian Arabic (Categorized) The World Facts General Knowledge Dataset (v2.0) is a high-quality, human-reviewed Q&A resource by Miscovery. It features general facts categorized across 50+ knowledge domains, provided in three languages: 🌍 English 🇸🇦 Modern Standard Arabic (MSA) 🇪🇬 Egyptian Arabic (Dialect) Each entry includes: The question and answer A category and sub-category Language tag (en, ar, ar_eg) Basic metadata:… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/arabic_egypt_english_world_facts.tabularquestion-answering10K<n<100K13 likes17 downloads1y agoHugging Face29shalanova /benchmark-1-arabic-m2mInfo: Translated on Arabic by facebook/m2m100_418M model Source: jayavibhav/prompt-injection-safety Domain: primarily contain prompt-injection and canonical jailbreak-style instructions with relatively homogeneous attack patterns Size: 1,000 prompts (500 safe / 500 unsafe) Columns: text - original prompt label - 0: safe, 1: unsafe translation - prompt on Arabic translated by facebook/m2m100_418M score_ar_model - cosine similarity score with codebook More information in paper:… See the full description on the dataset page: https://huggingface.co/datasets/shalanova/benchmark-1-arabic-m2m.tabular1K<n<10K0 likes17 downloads5mo agoHugging Face30Arastoorad /VulnSage VulnSage Dataset VulnSage is a curated dataset designed for research on automated vulnerability detection, particularly leveraging the capabilities of large language models (LLMs). It contains annotated vulnerable and patched code snippets from real-world software projects, along with rich metadata and contextual reasoning. 📦 Dataset Contents The dataset includes 593 vulnerability instances extracted from various open-source software repositories. Each entry provides… See the full description on the dataset page: https://huggingface.co/datasets/Arastoorad/VulnSage.tabularn<1K0 likes16 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.