CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MohamedRashad /arabic-books Arabic Books Dataset Summary The arabic-books dataset contains 8,500 rows of text, each representing the full text of a single Arabic book. These texts were extracted using the arabic-large-nougat model, showcasing the model’s capabilities in Arabic OCR and text extraction. The dataset spans a total of 1.1 billion tokens, calculated using the GPT-4 tokenizer. This dataset is a testimony to the quality of the Arabic Nougat models and their effectiveness in extracting… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-books.texttext-generation1K<n<10K3 likes36k downloads2y agoHugging Face02BangumiBase /arankpartyworidatsushitaorewamotooshiegotachitomeikyuushinbuwomezasu Bangumi Image Base of A-rank Party Wo Ridatsu Shita Ore Wa, Moto Oshiego-tachi To Meikyuu Shinbu Wo Mezasu. This is the image base of bangumi A-Rank Party wo Ridatsu shita Ore wa, Moto Oshiego-tachi to Meikyuu Shinbu wo Mezasu., we detected 168 characters, 15579 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/arankpartyworidatsushitaorewamotooshiegotachitomeikyuushinbuwomezasu.image10K<n<100K0 likes9.9k downloads1y agoHugging Face03nabbahi /traditionnals_arabic_shoes_split3d1K<n<10K0 likes8.3k downloads4mo agoHugging Face04aranemini /central-kurdish-pseudolabel Central Kurdish → English Pseudo-Labeled Speech Translation Corpus Dataset Summary This repository contains a large-scale pseudo-labeled speech translation corpus for Central Kurdish (Sorani Kurdish). The dataset was automatically generated using a pipeline composed of: Speech segmentation Automatic Speech Recognition (ASR) Machine Translation (MT) The objective is to provide training data for end-to-end Speech-to-Text Translation (S2TT) in a language with very… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-pseudolabel.audioautomatic-speech-recognition1M<n<10M2 likes4.1k downloads3mo agoHugging Face05MBZUAI /ArabicMMLU Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne Introduction We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/ArabicMMLU.tabularquestion-answering10K<n<100K39 likes3.8k downloads2y agoHugging Face06oddadmix /dialectal-arabic-lahgtna-v2 Dialectal Arabic Lahgtna v2 Large-scale multi-dialect Arabic speech dataset — 3,000+ hours across 13 Arabic dialects — for training and evaluating dialectal Arabic ASR systems. Part of the Lahgtna (لهجتنا) project for dialect-aware Arabic speech AI. Dataset Summary ~611K utterances / 3,000+ hours of transcribed dialectal Arabic speech **13 Arabic dialects **, labeled per utterance 16 kHz mono audio Transcripts written in authentic dialectal orthography (not… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/dialectal-arabic-lahgtna-v2.audioautomatic-speech-recognition100K<n<1M28 likes3.7k downloads2mo agoHugging Face072A2I /Arabic_Aya Dataset Card for : Arabic Aya (2A) Arabic Aya (2A) : A Curated Subset of the Aya Collection for Arabic Language Processing Dataset Sources & Infos Data Origin: Derived from 69 subsets of the original Aya datasets : CohereForAI/aya_collection, CohereForAI/aya_dataset, and CohereForAI/aya_evaluation_suite. Languages: Modern Standard Arabic (MSA) and a variety of Arabic dialects ( 'arb', 'arz', 'ary', 'ars', 'knc', 'acm', 'apc', 'aeb', 'ajp', 'acq' )… See the full description on the dataset page: https://huggingface.co/datasets/2A2I/Arabic_Aya.tabulartext-classification10M<n<100M16 likes3.1k downloads3y agoHugging Face08OALL /AlGhafa-Arabic-LLM-Benchmark-Native AlGhafa Arabic LLM Benchmark New fix: Normalized whitespace characters and ensured consistency across all datasets for improved data quality and compatibility. Multiple-choice evaluation benchmark for zero- and few-shot evaluation of Arabic LLMs, we adapt the following tasks: Belebele Ar MSA Bandarkar et al. (2023): 900 entries Belebele Ar Dialects Bandarkar et al. (2023): 5400 entries COPA Ar: 89 entries machine-translated from English COPA and verified by native Arabic… See the full description on the dataset page: https://huggingface.co/datasets/OALL/AlGhafa-Arabic-LLM-Benchmark-Native.text10K<n<100K7 likes2.7k downloads3y agoHugging Face09aractingi /droid_1.0.1_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "Franka", "total_episodes": 95658, "total_frames": 27630375, "total_tasks": 49630, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 15, "splits": { "train": "0:95658"}, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aractingi/droid_1.0.1_test.tabularrobotics10M<n<100M0 likes2.7k downloads9mo agoHugging Face10MBZUAI /human_translated_arabic_mmlutext10K<n<100K4 likes2.7k downloads2y agoHugging Face11QCRI /ImageEval-ArabicNLP26 ImageEval-ArabicNLP26 👁️ ImageEval-ArabicNLP26 is the dataset of the ImageEval 2026 Shared Task at ArabicNLP 2026. It covers both of the shared task's tasks: AynVQA (Task 1), a culturally grounded Arabic multimodal benchmark for spoken visual question answering and hallucination detection, and CRAI-Bench (Task 2), which evaluates the cultural accuracy of Arabic text-to-image generation. The shared task has concluded. All gold labels are released, including the blind test splits… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ImageEval-ArabicNLP26.audio10K<n<100K4 likes2.5k downloads23d agoHugging Face12OALL /details_CohereForAI__c4ai-command-r7b-arabic-02-2025_v2 Dataset Card for Evaluation run of CohereForAI/c4ai-command-r7b-arabic-02-2025 Dataset automatically created during the evaluation run of model CohereForAI/c4ai-command-r7b-arabic-02-2025. The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_CohereForAI__c4ai-command-r7b-arabic-02-2025_v2.text100K<n<1M0 likes2.4k downloads2y agoHugging Face13gonzalobenegas /processed-data-arabidopsis0 likes2.4k downloads1y agoHugging Face14cloudaocr /arabic-synthetic-scanned-booksdocument1K<n<10K1 likes2.3k downloads21d agoHugging Face15M-A-D /Mixed-Arabic-Datasets-Repo Dataset Card for "Mixed Arabic Datasets (MAD) Corpus" The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts Dataset Description The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo.tabulartext-classification100M<n<1B38 likes2.2k downloads3y agoHugging Face16TigreGotico /arabic-stem-lexicon Arabic Diacritized-Stem Lexicon An undiacritized Arabic surface form → its most frequent diacritized stem. Standard Arabic writes no short vowels, so anything that has to pronounce Arabic must first put them back. A neural diacritizer does that well on rare words, where inference is the only thing there is. On common words it is the wrong tool: which vowels كتاب carries is not a thing to be inferred, it is a thing to be looked up — and models get exactly these wrong, reading… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/arabic-stem-lexicon.tabulartext-to-speech100K<n<1M0 likes2.2k downloads2mo agoHugging Face17MohamedRashad /MASC-Arabic MASC Arabic Dataset Card Dataset Summary MASC is a dataset that contains 1,000 hours of speech sampled at 16 kHz and crawled from over 700 YouTube channels. The dataset is multi-regional, multi-genre, and multi-dialect intended to advance the research and development of Arabic speech technology with a special emphasis on Arabic speech recognition. How to use The datasets library allows you to load and pre-process your dataset in pure Python, at scale. The… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/MASC-Arabic.audioautomatic-speech-recognition100K<n<1M8 likes2.1k downloads6mo agoHugging Face18freococo /ocr_arabic_books Arabic OCR Books Dataset (ocr_arabic_books) This repository is a structurally aligned Arabic Optical Character Recognition (OCR) dataset. It hosts a collection of classical Islamic books, organized by separate configurations (subsets) to support modular loading without filename collisions. 📚 Master Book Inventory Total running pages in repository: 181,427 # Book Name (English) Book Name (Arabic) Subset / Config Name Page Count Image Index Range 1… See the full description on the dataset page: https://huggingface.co/datasets/freococo/ocr_arabic_books.imageimage-to-text100K<n<1M4 likes1.9k downloads3mo agoHugging Face19ClusterlabAi /101_billion_arabic_words_dataset 101 Billion Arabic Words Dataset Updates Maintenance Status: Actively Maintained Update Frequency: Weekly updates to refine data quality and expand coverage. Upcoming Version More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management. Dataset Details The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/ClusterlabAi/101_billion_arabic_words_dataset.texttext-generation10M<n<100M73 likes1.9k downloads2y agoHugging Face20freococo /synth_shamela_ocr_arabic_books Synthetic Arabic Books Dataset Structured book pages rendered dynamically with style, font, and degradation variations. imageimage-to-text1M<n<10M2 likes1.8k downloads2mo agoHugging Face21oddadmix /arabic-audio-collection-algerian-loubna-stories Loubna Stories Arabic Speech Dataset Dataset Summary The Loubna Stories Arabic Speech Dataset is a large-scale, first-of-its-kind Arabic speech corpus containing approximately 237 hours of speech recordings and corresponding transcripts. What distinguishes this dataset as a pioneering resource in Arabic language technology is its comprehensive inclusion of rich non-verbal transcriptions. Alongside the spoken Arabic text, the transcripts meticulously capture… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-audio-collection-algerian-loubna-stories.audiotext-to-speech10K<n<100K0 likes1.7k downloads3mo agoHugging Face22Sabri12blm /Arabic-Quran-ASR-datasetaudio10K<n<100K5 likes1.7k downloads2y agoHugging Face23OALL /Arabic_MMLUThis dataset belongs to FreedomIntelligence and the original version can be found here : https://github.com/FreedomIntelligence/AceGPT/tree/main/eval/benchmark_eval/benchmarks/MMLUArabic text10K<n<100K3 likes1.5k downloads2y agoHugging Face24ArabicSpeech /ADI17audio1M<n<10M2 likes1.5k downloads1y agoHugging Face25Aratako /LiquidAI-Hackathon-Tokyo-CPT-Data LiquidAI-Hackathon-Tokyo-CPT-Data Liquid AI Hackathon Tokyoで作成したモデルのCPTに利用したデータセットです。 automatic-speech-recognition1M<n<10M6 likes1.5k downloads1y agoHugging Face26Arabic-Clip /xtd_11 Dataset Summary The expanded XTD-11 dataset, now including Arabic, enhances the original XTD collection. This dataset introduces a 1,000-image multi-lingual MSCOCO2014 caption to test multimodel in zeroshot image or text retrieval in 11 Languages. Dataset Details Citation @misc{aggarwal2020zeroshot, title={Towards Zero-shot Cross-lingual Image Retrieval}, author={Pranav Aggarwal and Ajinkya Kale}, year={2020}, eprint={2012.05107}… See the full description on the dataset page: https://huggingface.co/datasets/Arabic-Clip/xtd_11.image-to-text1K<n<10K3 likes1.4k downloads2y agoHugging Face27AdaMLLab /AraMix-HQ AraMix family: AraMix (minhash and matched) | AraMix-domain-classified (with domain labels) | AraMix-HQ (model-filtered) AraMix-HQ is a high-quality subset of AraMix-MinHash created using model-based quality scoring. We adapt the approach from FineWeb2-HQ but replace the XLM-Roberta encoder with mmBERT, which provides better Arabic language understanding. We release the model at AdaMLLab/mmBERT-Arabic-Quality-Classifier. AraMix-HQ outperforms both AraMix-Matched and FineWeb2-HQ… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/AraMix-HQ.texttext-generation10M<n<100M2 likes1.4k downloads8mo agoHugging Face28SultanR /fineweb-edu-arabic fineweb-edu-arabic Arabic translation of FineWeb-Edu (sample/350BT subset, filtered to language_score > 0.9), translated with Seed-X-PPO-7B using greedy decoding. Documents were split into ~490-token chunks, translated, and reassembled. Each row is one complete document. A companion corpus translated with the same pipeline is available at dclm-pro-arabic. Details Documents: 82,840,410 (27.9% of the source subset, uniformly sampled) Arabic tokens: ~170B (Seed-X… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/fineweb-edu-arabic.texttext-generation10M<n<100M1 likes1.3k downloads1mo agoHugging Face29AhmedTaha012 /arabic_text_yolov8_V0.5text1K<n<10K0 likes1.3k downloads2y agoHugging Face30Aramente /eu-tech-jobs eu-tech-jobs Daily-updated open-data feed of jobs from EU AI/tech and remote-EU companies. Source repo: https://github.com/Aramente/eu-tech-jobs Live site: https://aramente.github.io/eu-tech-jobs/ License: CC BY 4.0 (data) + MIT (pipeline code) What's in here Path Contents latest/jobs.parquet Most recent snapshot, all active jobs latest/companies.parquet Curated company list with categories + ATS handles latest/metadata.json Pipeline run metadata… See the full description on the dataset page: https://huggingface.co/datasets/Aramente/eu-tech-jobs.texttabular-classification1M<n<10M1 likes1.3k downloads22h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.