CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BangumiBase /arankpartyworidatsushitaorewamotooshiegotachitomeikyuushinbuwomezasu Bangumi Image Base of A-rank Party Wo Ridatsu Shita Ore Wa, Moto Oshiego-tachi To Meikyuu Shinbu Wo Mezasu. This is the image base of bangumi A-Rank Party wo Ridatsu shita Ore wa, Moto Oshiego-tachi to Meikyuu Shinbu wo Mezasu., we detected 168 characters, 15579 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/arankpartyworidatsushitaorewamotooshiegotachitomeikyuushinbuwomezasu.image10K<n<100K0 likes10k downloads1y agoHugging Face02QCRI /ImageEval-ArabicNLP26 ImageEval-ArabicNLP26 👁️ ImageEval-ArabicNLP26 is the dataset of the ImageEval 2026 Shared Task at ArabicNLP 2026. It covers both of the shared task's tasks: AynVQA (Task 1), a culturally grounded Arabic multimodal benchmark for spoken visual question answering and hallucination detection, and CRAI-Bench (Task 2), which evaluates the cultural accuracy of Arabic text-to-image generation. The shared task has concluded. All gold labels are released, including the blind test splits… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ImageEval-ArabicNLP26.audio10K<n<100K4 likes2.5k downloads24d agoHugging Face03freococo /ocr_arabic_books Arabic OCR Books Dataset (ocr_arabic_books) This repository is a structurally aligned Arabic Optical Character Recognition (OCR) dataset. It hosts a collection of classical Islamic books, organized by separate configurations (subsets) to support modular loading without filename collisions. 📚 Master Book Inventory Total running pages in repository: 181,427 # Book Name (English) Book Name (Arabic) Subset / Config Name Page Count Image Index Range 1… See the full description on the dataset page: https://huggingface.co/datasets/freococo/ocr_arabic_books.imageimage-to-text100K<n<1M4 likes1.9k downloads3mo agoHugging Face04freococo /synth_shamela_ocr_arabic_books Synthetic Arabic Books Dataset Structured book pages rendered dynamically with style, font, and degradation variations. imageimage-to-text1M<n<10M2 likes1.8k downloads2mo agoHugging Face05Arabic-Clip /Arabic_Flicker_8kimage1K<n<10K0 likes1.1k downloads2y agoHugging Face06MohamedRashad /arabic-img2md Arabic Img2MD Dataset Summary The arabic-img2md dataset consists of 15,000 examples of PDF pages paired with their Markdown counterparts. The dataset is split into: Train: 13,700 examples Test: 1,520 examples This dataset was created as part of the open-source research project Arabic Nougat to enable OCR and Markdown extraction from Arabic documents. It contains mostly Arabic text but also includes examples with English text. Usage The dataset was used to… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-img2md.imageimage-to-text10K<n<100K14 likes881 downloads2y agoHugging Face07Arasaaf /visual_26k_reasoningimage10K<n<100K1 likes684 downloads5mo agoHugging Face08arampacha /rsicdimage10K<n<100K14 likes553 downloads4y agoHugging Face09calvinyong1 /ArabidopsisDatasetimage1K<n<10K0 likes538 downloads9d agoHugging Face10Melaraby /EvArEST-dataset-for-Arabic-scene-text-recognition EvArEST Everyday Arabic-English Scene Text dataset, from the paper: Arabic Scene Text Recognition in the Deep Learning Era: Analysis on A Novel Dataset The dataset includes both the recognition dataset and the synthetic one in a single train and test split. Recognition Dataset The text recognition dataset comprises of 7232 cropped word images of both Arabic and English languages. The groundtruth for the recognition dataset is provided by a text file with each line… See the full description on the dataset page: https://huggingface.co/datasets/Melaraby/EvArEST-dataset-for-Arabic-scene-text-recognition.image100K<n<1M2 likes488 downloads11mo agoHugging Face11MohamedRashad /Arabic-VLM-Full-Pearl 💎 The Arabic VLM Dataset (Full Pearl Edition) This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper. Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.imagequestion-answering100K<n<1M10 likes448 downloads10mo agoHugging Face12Archatext /AraMS-Restore AraMS-Restore — Real Damaged Arabic Manuscript Lines 177 line images cropped from real damaged pages of a historical Arabic manuscript (book_09), each with its transcription. This is the evaluation input for AraMS-Restore: the restoration models are trained on synthetic degradation, and these lines are the honest test of whether that transfers to genuine manuscript decay. There are no clean counterparts and no ground-truth restored images — the damage is what was on the page.… See the full description on the dataset page: https://huggingface.co/datasets/Archatext/AraMS-Restore.imageimage-to-imagen<1K0 likes434 downloads2mo agoHugging Face13aractingi /libero-with-rewardThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "panda", "total_episodes": 1693, "total_frames": 273465, "total_tasks": 40, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 10.0, "splits": { "train": "0:1693"}, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path": null, "features":… See the full description on the dataset page: https://huggingface.co/datasets/aractingi/libero-with-reward.imagerobotics100K<n<1M0 likes380 downloads11mo agoHugging Face14jinaai /arabic_chartqa_ar_beirThis is a copy of https://huggingface.co/datasets/jinaai/arabic_chartqa_ar reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/arabic_chartqa_ar_beir.image1K<n<10K0 likes346 downloads1y agoHugging Face15mohajesmaeili /Persian_Arabic_TextLine_Image_Ocr_Mediumimage100K<n<1M18 likes328 downloads1y agoHugging Face16jinaai /arabic_infographicsvqa_ar_beirThis is a copy of https://huggingface.co/datasets/jinaai/arabic_infographicsvqa_ar reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/arabic_infographicsvqa_ar_beir.imagen<1K0 likes308 downloads1y agoHugging Face17loay /arabic-ocr-synthetic-scans-faker-300k Arabic OCR Synthetic Scans (Faker 300k) A large-scale synthetic dataset of ~300,000 Arabic book pages generated to mimic real-world scanning imperfections. Designed for training Vision Language Models (VLMs) and OCR engines on structural layout analysis, font recognition, and document degradation robustness. Dataset Summary Samples: ~300,000 synthetic Arabic document pages Image format: JPEG, ~800×1200 px (embedded in Parquet) Text: Ground truth in UTF-8 with XML-style… See the full description on the dataset page: https://huggingface.co/datasets/loay/arabic-ocr-synthetic-scans-faker-300k.imageimage-to-text100K<n<1M7 likes307 downloads7mo agoHugging Face18milistu /AMAZON-Products-2023-Arabic Dataset Card for Amazon Products 2023 Arabic Dataset Summary This dataset contains product metadata from Amazon, filtered to include only products that became available in 2023. The dataset is intended for use in semantic search applications and includes a variety of product categories. Number of Rows: 117,243 Number of Columns: 17 Data Source The data is sourced from Amazon Reviews 2023. It includes product information across multiple categories, with… See the full description on the dataset page: https://huggingface.co/datasets/milistu/AMAZON-Products-2023-Arabic.imagetext-classification100K<n<1M2 likes280 downloads2y agoHugging Face19oddadmix /blip3-grounding-1m-arabic BLIP3-Grounding (Arabic) - 1M Arabic version of the first 1,000,000 usable rows of Salesforce/blip3-grounding-50m. Two things differ from the source: The images are here. The source ships a url column only; those URLs were crawled and the original image bytes embedded, so the dataset is usable without a crawl of your own. Detection labels are translated. metadata_ar carries the Arabic label for every bounding box. Everything else is untouched English. Schema… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/blip3-grounding-1m-arabic.imageimage-to-text1M<n<10M0 likes246 downloads1mo agoHugging Face20IslamMesabah /AraReceipt AraReceipt AraReceipt is a manually annotated dataset of 100 Arabic retail receipt images labeled with 25 key-information classes plus an Ignore category, in WildReceipt style. It contains 4,609 annotated regions (46.1 ± 21.6 per image), each with a box, a transcription, and a semantic class. The dataset was built with GAIDA, a human-in-the-loop annotation system that combines OCR-based region extraction, LLM-based semantic pre-annotation, and human validation in Label Studio.… See the full description on the dataset page: https://huggingface.co/datasets/IslamMesabah/AraReceipt.imageobject-detectionn<1K0 likes245 downloads2mo agoHugging Face21KhalfounMehdi /arabic-latin-invoices-synthetic Synthetic Arabic/Latin Invoices — label-first multimodal dataset A large, perfectly-labeled synthetic invoice dataset for training multimodal invoice-parsing models, with first-class support for Arabic script, mixed Arabic/Latin layouts, and all three numeral glyph systems. Every sample is generated label-first: a structured invoice record is sampled, then rendered to pixels via a headless browser, and bounding boxes are read back from the same DOM — so labels and boxes are… See the full description on the dataset page: https://huggingface.co/datasets/KhalfounMehdi/arabic-latin-invoices-synthetic.imageimage-to-text1K<n<10K1 likes227 downloads3mo agoHugging Face22sherif1313 /Historical-Arabic-Handwritten-OCRDescription A collection of rich historical Arabic text, spanning different geographies across centuries, is present in this dataset. Experts have meticulously transcribed forty historical pages, each five from a distinct book, providing the textual ground truth for each image. No data as such has been made available publicly previously, up to our knowledge. This intends to contribute to deep learning OCR modeling and testing by practitioners and researchers interested in Arabic OCR and… See the full description on the dataset page: https://huggingface.co/datasets/sherif1313/Historical-Arabic-Handwritten-OCR.image1 likes217 downloads7mo agoHugging Face23arashakb /Circuitsense-6kimage10K<n<100K0 likes209 downloads10mo agoHugging Face24racineai /VDR_Energy_Arabic VDR_Energy_Arabic - Overview Dataset Summary VDR_Energy_Arabic is a curated multimodal dataset focused on Arabic energy sector documents, including reports, financial statements, technical documentation, and industry analyses. It combines text and image data extracted from real energy-related PDFs to support tasks such as RAG DSE, question answering, document search, and vision-language model training in Arabic. Dataset Details Dataset Creation This… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_Energy_Arabic.imagevisual-document-retrieval10K<n<100K5 likes205 downloads10mo agoHugging Face25TheSeniorTeam /Arabic_Manuscript_Collection_Dataset Arabic Manuscript Collection Seven Arabic handwritten text recognition (HTR) subsets. Five are converted to one layout and one label format so they can be trained and evaluated together: 82,561 labelled images in total. Two are republished closer to their source shape: AMIDDA as upstream Parquet, and OpenITI-Makhzan as page images with line-level coordinates. Four of the five converted sources are historical manuscripts. KHATT is modern handwriting and is included as a separate… See the full description on the dataset page: https://huggingface.co/datasets/TheSeniorTeam/Arabic_Manuscript_Collection_Dataset.imageimage-to-text100K<n<1M0 likes205 downloads11d agoHugging Face26Atrozy /arabic_dataimage10K<n<100K0 likes197 downloads1y agoHugging Face27QCRI /AraDiCE AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs Overview The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic. As part of the supplemental materials, we have selected a few datasets (see below) for the reader to review. We will make the full… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDiCE.imagetext-classificationn<1K3 likes196 downloads2y agoHugging Face28araniko3d /buggyimage0 likes150 downloads2y agoHugging Face29Atrozy /arabic_data2image100K<n<1M0 likes147 downloads1y agoHugging Face30oddadmix /laion-coco-nllb-arabic-filtered LAION-COCO-NLLB, Arabic slice, filtered The arb_Arab captions of visheratin/laion-coco-nllb, extracted and filtered for training an Arabic captioning model. This is the exact corpus behind oddadmix/Nawah-VL-25M. split rows kept from train 753,883 878,978 (85.8%) test 14,407 14,906 (96.7%) Columns column type notes id string the source dataset's image id url string image URL, not the image. See below. caption string the arb_Arab… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/laion-coco-nllb-arabic-filtered.imageimage-to-text100K<n<1M0 likes144 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.