CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MBZUAI /ArabicMMLU Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne Introduction We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/ArabicMMLU.tabularquestion-answering10K<n<100K39 likes3.8k downloads2y agoHugging Face022A2I /Arabic_Aya Dataset Card for : Arabic Aya (2A) Arabic Aya (2A) : A Curated Subset of the Aya Collection for Arabic Language Processing Dataset Sources & Infos Data Origin: Derived from 69 subsets of the original Aya datasets : CohereForAI/aya_collection, CohereForAI/aya_dataset, and CohereForAI/aya_evaluation_suite. Languages: Modern Standard Arabic (MSA) and a variety of Arabic dialects ( 'arb', 'arz', 'ary', 'ars', 'knc', 'acm', 'apc', 'aeb', 'ajp', 'acq' )… See the full description on the dataset page: https://huggingface.co/datasets/2A2I/Arabic_Aya.tabulartext-classification10M<n<100M16 likes3.1k downloads3y agoHugging Face03TigreGotico /arabic-stem-lexicon Arabic Diacritized-Stem Lexicon An undiacritized Arabic surface form → its most frequent diacritized stem. Standard Arabic writes no short vowels, so anything that has to pronounce Arabic must first put them back. A neural diacritizer does that well on rare words, where inference is the only thing there is. On common words it is the wrong tool: which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.tabulartext-to-speech100K<n<1M0 likes2.8k downloads2mo agoHugging Face04aractingi /droid_1.0.1_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "Franka", "total_episodes": 95658, "total_frames": 27630375, "total_tasks": 49630, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 15, "splits": { "train": "0:95658"}, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aractingi/droid_1.0.1_test.tabularrobotics10M<n<100M0 likes2.5k downloads10mo agoHugging Face05QCRI /ImageEval-ArabicNLP26 ImageEval-ArabicNLP26 👁️ ImageEval-ArabicNLP26 is the dataset of the ImageEval 2026 Shared Task at ArabicNLP 2026. It covers both of the shared task's tasks: AynVQA (Task 1), a culturally grounded Arabic multimodal benchmark for spoken visual question answering and hallucination detection, and CRAI-Bench (Task 2), which evaluates the cultural accuracy of Arabic text-to-image generation. The shared task has concluded. All gold labels are released, including the blind test splits… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ImageEval-ArabicNLP26.audio10K<n<100K4 likes2.5k downloads23d agoHugging Face06M-A-D /Mixed-Arabic-Datasets-Repo Dataset Card for "Mixed Arabic Datasets (MAD) Corpus" The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts Dataset Description The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo.tabulartext-classification100M<n<1B38 likes2.2k downloads3y agoHugging Face07ArabicSpeech /ArA-DF-2026 ArA-DF-2026 ArA-DF-2026 is an Arabic speech deepfake detection dataset for binary audio classification. 0: spoofed or synthetic speech 1: bona fide speech The public train and dev splits include labels. The original challenge-era development-test and final-test configs remain unlabeled for reproducibility, and post-challenge gold labels are now available through the *_labeled configs. All released audio is 16 kHz mono PCM audio packaged as lossless FLAC inside WebDataset TAR… See the full description on the dataset page: https://huggingface.co/datasets/ArabicSpeech/ArA-DF-2026.tabularaudio-classification100K<n<1M1 likes986 downloads2mo agoHugging Face08AdaMLLab /smolkalam-arabic-conversational-sft SmolKalam SmolKalam is a quality-filtered Arabic SFT dataset of 1,790,478 examples (~2.45B tokens), built as an ensemble translation of SmolTalk2. It covers multi-turn dialogue (23% of rows), reasoning traces (19% carry <think>), tool and function calling (4.4%), and long context, categories that are underrepresented in existing Arabic post-training data. The SmolTalk2 source mixtures are kept as subsets. Released with the paper SmolKalam: Ensemble Quality-Filtered Translation… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/smolkalam-arabic-conversational-sft.tabulartext-generation1M<n<10M3 likes770 downloads1mo agoHugging Face09OALL /AlGhafa-Arabic-LLM-Benchmark-Translatedtabular10K<n<100K2 likes644 downloads2y agoHugging Face10SultanR /AraMix-Translation-Scores AraMix-Translation-Scores AdaMLLab/AraMix (minhash_deduped subset, 178,883,241 rows) with a machine-translation-detection score added to every document. All original columns are preserved. Columns column type description id string unchanged from AraMix source string unchanged from AraMix text string unchanged from AraMix mmbert_quality_score float64 AraMix's original mmbert_score, renamed mmbert_translated_score float64 new —… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/AraMix-Translation-Scores.tabular100M<n<1B0 likes594 downloads2mo agoHugging Face11alielfilali01 /fineweb-2-arb_Arabtabular10M<n<100M1 likes535 downloads2y agoHugging Face12SultanR /AraMix-Native AraMix-Native A native-Arabic-filtered version of AdaMLLab/AraMix (minhash_deduped), derived from SultanR/AraMix-Translation-Scores: machine-translated and garbled-MT documents removed, 162,887,010 rows kept of 178,883,241 (91.06%). All columns preserved. Filter rules A document is kept iff all of: mmbert_translated_score < 0.1, or a classical-text rescue: diacritic (tashkeel) ratio ≥ 0.02 over Arabic letters and ≥ 3 distinct diacritic classes (fully/partially… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/AraMix-Native.tabular100M<n<1B0 likes533 downloads2mo agoHugging Face13yrrhall /Arabic_Aya Dataset Card for : Arabic Aya (2A) Arabic Aya (2A) : A Curated Subset of the Aya Collection for Arabic Language Processing Dataset Sources & Infos Data Origin: Derived from 69 subsets of the original Aya datasets : CohereForAI/aya_collection, CohereForAI/aya_dataset, and CohereForAI/aya_evaluation_suite. Languages: Modern Standard Arabic (MSA) and a variety of Arabic dialects ( 'arb', 'arz', 'ary', 'ars', 'knc', 'acm', 'apc', 'aeb', 'ajp'… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/Arabic_Aya.tabulartext-classification10M<n<100M0 likes512 downloads3mo agoHugging Face14nizarun /FineWeb-Edu-Arabic-24M English العربية FineWeb-Edu Arabic 24M An Arabic-only pretraining corpus of 24,794,425 complete documents, translated from the sample-350BT configuration of FineWeb-Edu. It contains 34.86 billion Arabic tokenizer tokens and preserves the original FineWeb-Edu document IDs, source scores, and detailed translation diagnostics. Highlight Saudi architecture shaped by place. A well-translated tour of how builders in Najd, the Gulf coast, Hejaz, and Asir adapted local… See the full description on the dataset page: https://huggingface.co/datasets/nizarun/FineWeb-Edu-Arabic-24M.tabulartext-generation10M<n<100M0 likes464 downloads24d agoHugging Face15PleIAs /Arabic-PDtabular100K<n<1M0 likes439 downloads7mo agoHugging Face16sboughorbel /tinystories_dataset_arabictabular1M<n<10M1 likes410 downloads2y agoHugging Face17Qanadil /ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets Dataset Card for "ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets" Note About Sentiment_label_confidence "Crowdflower provides a confidence score with each annotated tweet that represents the confidence in the quality of the label. For a three annotators per tweet setup, the confidence score would range between 0.3 and 1 according to two factors: 1) annotator quality level; and 2) agreement… See the full description on the dataset page: https://huggingface.co/datasets/Qanadil/ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets.tabulartext-classification10K<n<100K1 likes397 downloads2y agoHugging Face18jinaai /arabic_chartqa_ar_beirThis is a copy of https://huggingface.co/datasets/jinaai/arabic_chartqa_ar reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/arabic_chartqa_ar_beir.image1K<n<10K0 likes346 downloads1y agoHugging Face19OpenLLM-France /RULER-luciole_tokenizer_128k-arab-regional_v2tabular10K<n<100K0 likes346 downloads10mo agoHugging Face20Rutts07 /mgap-kas_Arabtabular1K<n<10K0 likes343 downloads2y agoHugging Face21OALL /details_deep-analysis-research__D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2_v2 Dataset Card for Evaluation run of deep-analysis-research/D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2 Dataset automatically created during the evaluation run of model deep-analysis-research/D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2. The dataset is composed of 116 configuration, each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_deep-analysis-research__D2IL-Arabic-Qwen2.5-72B-Instruct-v0.2_v2.tabular100K<n<1M0 likes321 downloads1y agoHugging Face22jinaai /arabic_infographicsvqa_ar_beirThis is a copy of https://huggingface.co/datasets/jinaai/arabic_infographicsvqa_ar reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/arabic_infographicsvqa_ar_beir.imagen<1K0 likes308 downloads1y agoHugging Face23deep-analysis-research /details_D2IL-Arabic-Qwen2.5-72B-Instruct-v0.1tabular10K<n<100K0 likes299 downloads1y agoHugging Face24NajahUniv /arabic-univeristy-chatbot-qa Arabic University Chatbot QA A multilingual, multi-label intent-routing dataset for a university chatbot: given a student's message, predict which of 20 intent categories it should route to. This is routing, not question answering — the dataset contains no answers. Release v0.8.0 — pinned as a Hub tag, so revision="v0.8.0" always resolves to exactly these rows. This release holds 50,000 question rows in 22,600 scenario groups. Every row has accepted == true; the classifier… See the full description on the dataset page: https://huggingface.co/datasets/NajahUniv/arabic-univeristy-chatbot-qa.tabulartext-classification10K<n<100K0 likes290 downloads20d agoHugging Face25yrrhall /Arabic_Aya14200 Dataset Card for : Arabic Aya (2A) Arabic Aya (2A) : A Curated Subset of the Aya Collection for Arabic Language Processing Dataset Sources & Infos Data Origin: Derived from 69 subsets of the original Aya datasets : CohereForAI/aya_collection, CohereForAI/aya_dataset, and CohereForAI/aya_evaluation_suite. Languages: Modern Standard Arabic (MSA) and a variety of Arabic dialects ( 'arb', 'arz', 'ary', 'ars', 'knc', 'acm', 'apc', 'aeb', 'ajp'… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/Arabic_Aya14200.tabulartext-classification10M<n<100M0 likes283 downloads3mo agoHugging Face26milistu /AMAZON-Products-2023-Arabic Dataset Card for Amazon Products 2023 Arabic Dataset Summary This dataset contains product metadata from Amazon, filtered to include only products that became available in 2023. The dataset is intended for use in semantic search applications and includes a variety of product categories. Number of Rows: 117,243 Number of Columns: 17 Data Source The data is sourced from Amazon Reviews 2023. It includes product information across multiple categories, with… See the full description on the dataset page: https://huggingface.co/datasets/milistu/AMAZON-Products-2023-Arabic.imagetext-classification100K<n<1M2 likes280 downloads2y agoHugging Face27arag0rn /SecVulEval Dataset Card for Dataset Name SecVulEval is a collection of real-world C/C++ vulnerabilities. Dataset Details Dataset Description The dataset is curated by collecting C/C++ vulnerability from NVD. It features statement-level vulnerable information, context information for vulnerable functions (is_vulnerable=True), and other metadata such as CVE, CWE, commit information. The dataset contains vulnerable and non-vulnerable function samples. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/arag0rn/SecVulEval.tabular10K<n<100K8 likes264 downloads5mo agoHugging Face28yrrhall /Mixed-Arabic-Datasets-Repo Dataset Card for "Mixed Arabic Datasets (MAD) Corpus" The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts Dataset Description The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With… See the full description on the dataset page: https://huggingface.co/datasets/yrrhall/Mixed-Arabic-Datasets-Repo.tabulartext-classification100M<n<1B0 likes260 downloads4mo agoHugging Face29aractingi /droid_100This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": null, "total_episodes": 100, "total_frames": 32212, "total_tasks": 47, "total_videos": 300, "total_chunks": 1, "chunks_size": 1000, "fps": 15, "splits": { "train": "0:100"}, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aractingi/droid_100.tabularrobotics10K<n<100K1 likes238 downloads2y agoHugging Face30OALL /details_sambanovasystems__SambaLingo-Arabic-Chat-70B Dataset Card for Evaluation run of sambanovasystems/SambaLingo-Arabic-Chat-70B Dataset automatically created during the evaluation run of model sambanovasystems/SambaLingo-Arabic-Chat-70B. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_sambanovasystems__SambaLingo-Arabic-Chat-70B.tabular100K<n<1M0 likes227 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.