datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arankpartyworidatsushitaorewamotooshiegotachitomeikyuushinbuwomezasu
Bangumi Image Base of A-rank Party Wo Ridatsu Shita Ore Wa, Moto Oshiego-tachi To Meikyuu Shinbu Wo Mezasu.
This is the image base of bangumi A-Rank Party wo Ridatsu shita Ore wa, Moto Oshiego-tachi to Meikyuu Shinbu wo Mezasu., we detected 168 characters, 15579 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/arankpartyworidatsushitaorewamotooshiegotachitomeikyuushinbuwomezasu.ImageEval-ArabicNLP26
ImageEval-ArabicNLP26 👁️
ImageEval-ArabicNLP26 is the dataset of the ImageEval 2026 Shared Task at ArabicNLP 2026.
It covers both of the shared task's tasks: AynVQA (Task 1), a culturally grounded Arabic multimodal benchmark for spoken visual question answering and hallucination detection, and CRAI-Bench (Task 2), which evaluates the cultural accuracy of Arabic text-to-image generation.
The shared task has concluded. All gold labels are released, including the blind test splits… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/ImageEval-ArabicNLP26.ocr_arabic_books
Arabic OCR Books Dataset (ocr_arabic_books)
This repository is a structurally aligned Arabic Optical Character Recognition (OCR) dataset. It hosts a collection of classical Islamic books, organized by separate configurations (subsets) to support modular loading without filename collisions.
📚 Master Book Inventory
Total running pages in repository: 181,427
#
Book Name (English)
Book Name (Arabic)
Subset / Config Name
Page Count
Image Index Range
1… See the full description on the dataset page: https://huggingface.co/datasets/freococo/ocr_arabic_books.synth_shamela_ocr_arabic_books
Synthetic Arabic Books Dataset
Structured book pages rendered dynamically with style, font, and degradation variations.
Arabic_Flicker_8karabic-img2md
Arabic Img2MD
Dataset Summary
The arabic-img2md dataset consists of 15,000 examples of PDF pages paired with their Markdown counterparts. The dataset is split into:
Train: 13,700 examples
Test: 1,520 examples
This dataset was created as part of the open-source research project Arabic Nougat to enable OCR and Markdown extraction from Arabic documents. It contains mostly Arabic text but also includes examples with English text.
Usage
The dataset was used to… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-img2md.visual_26k_reasoningrsicdArabidopsisDatasetEvArEST-dataset-for-Arabic-scene-text-recognition
EvArEST
Everyday Arabic-English Scene Text dataset, from the paper: Arabic Scene Text Recognition in the Deep Learning Era: Analysis on A Novel Dataset
The dataset includes both the recognition dataset and the synthetic one in a single train and test split.
Recognition Dataset
The text recognition dataset comprises of 7232 cropped word images of both Arabic and English languages. The groundtruth for the recognition dataset is provided by a text file with each line… See the full description on the dataset page: https://huggingface.co/datasets/Melaraby/EvArEST-dataset-for-Arabic-scene-text-recognition.Arabic-VLM-Full-Pearl
💎 The Arabic VLM Dataset (Full Pearl Edition)
This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper.
Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.AraMS-Restore
AraMS-Restore — Real Damaged Arabic Manuscript Lines
177 line images cropped from real damaged pages of a historical Arabic
manuscript (book_09), each with its transcription. This is the evaluation
input for AraMS-Restore: the
restoration models are trained on synthetic degradation, and these lines are the
honest test of whether that transfers to genuine manuscript decay.
There are no clean counterparts and no ground-truth restored images — the damage
is what was on the page.… See the full description on the dataset page: https://huggingface.co/datasets/Archatext/AraMS-Restore.libero-with-rewardThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 1693,
"total_frames": 273465,
"total_tasks": 40,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 10.0,
"splits": {
"train": "0:1693"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null,
"features":… See the full description on the dataset page: https://huggingface.co/datasets/aractingi/libero-with-reward.arabic_chartqa_ar_beirThis is a copy of https://huggingface.co/datasets/jinaai/arabic_chartqa_ar reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/arabic_chartqa_ar_beir.Persian_Arabic_TextLine_Image_Ocr_Mediumarabic_infographicsvqa_ar_beirThis is a copy of https://huggingface.co/datasets/jinaai/arabic_infographicsvqa_ar reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/arabic_infographicsvqa_ar_beir.arabic-ocr-synthetic-scans-faker-300k
Arabic OCR Synthetic Scans (Faker 300k)
A large-scale synthetic dataset of ~300,000 Arabic book pages generated to mimic real-world scanning imperfections. Designed for training Vision Language Models (VLMs) and OCR engines on structural layout analysis, font recognition, and document degradation robustness.
Dataset Summary
Samples: ~300,000 synthetic Arabic document pages
Image format: JPEG, ~800×1200 px (embedded in Parquet)
Text: Ground truth in UTF-8 with XML-style… See the full description on the dataset page: https://huggingface.co/datasets/loay/arabic-ocr-synthetic-scans-faker-300k.AMAZON-Products-2023-Arabic
Dataset Card for Amazon Products 2023 Arabic
Dataset Summary
This dataset contains product metadata from Amazon, filtered to include only products that became available in 2023. The dataset is intended for use in semantic search applications and includes a variety of product categories.
Number of Rows: 117,243
Number of Columns: 17
Data Source
The data is sourced from Amazon Reviews 2023.
It includes product information across multiple categories, with… See the full description on the dataset page: https://huggingface.co/datasets/milistu/AMAZON-Products-2023-Arabic.blip3-grounding-1m-arabic
BLIP3-Grounding (Arabic) - 1M
Arabic version of the first 1,000,000 usable rows of
Salesforce/blip3-grounding-50m.
Two things differ from the source:
The images are here. The source ships a url column only; those URLs were
crawled and the original image bytes embedded, so the dataset is usable
without a crawl of your own.
Detection labels are translated. metadata_ar carries the Arabic label
for every bounding box. Everything else is untouched English.
Schema… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/blip3-grounding-1m-arabic.AraReceipt
AraReceipt
AraReceipt is a manually annotated dataset of 100 Arabic retail receipt images labeled with 25 key-information classes plus an Ignore category, in WildReceipt style. It contains 4,609 annotated regions (46.1 ± 21.6 per image), each with a box, a transcription, and a semantic class.
The dataset was built with GAIDA, a human-in-the-loop annotation system that combines OCR-based region extraction, LLM-based semantic pre-annotation, and human validation in Label Studio.… See the full description on the dataset page: https://huggingface.co/datasets/IslamMesabah/AraReceipt.arabic-latin-invoices-synthetic
Synthetic Arabic/Latin Invoices — label-first multimodal dataset
A large, perfectly-labeled synthetic invoice dataset for training multimodal
invoice-parsing models, with first-class support for Arabic script, mixed Arabic/Latin
layouts, and all three numeral glyph systems.
Every sample is generated label-first: a structured invoice record is sampled, then
rendered to pixels via a headless browser, and bounding boxes are read back from the same
DOM — so labels and boxes are… See the full description on the dataset page: https://huggingface.co/datasets/KhalfounMehdi/arabic-latin-invoices-synthetic.Historical-Arabic-Handwritten-OCRDescription
A collection of rich historical Arabic text, spanning different geographies across centuries, is present in this dataset. Experts have meticulously transcribed forty historical pages, each five from a distinct book, providing the textual ground truth for each image.
No data as such has been made available publicly previously, up to our knowledge. This intends to contribute to deep learning OCR modeling and testing by practitioners and researchers interested in Arabic OCR and… See the full description on the dataset page: https://huggingface.co/datasets/sherif1313/Historical-Arabic-Handwritten-OCR.Circuitsense-6kVDR_Energy_Arabic
VDR_Energy_Arabic - Overview
Dataset Summary
VDR_Energy_Arabic is a curated multimodal dataset focused on Arabic energy sector documents, including reports, financial statements, technical documentation, and industry analyses. It combines text and image data extracted from real energy-related PDFs to support tasks such as RAG DSE, question answering, document search, and vision-language model training in Arabic.
Dataset Details
Dataset Creation
This… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_Energy_Arabic.Arabic_Manuscript_Collection_Dataset
Arabic Manuscript Collection
Seven Arabic handwritten text recognition (HTR) subsets. Five are converted to one layout
and one label format so they can be trained and evaluated together: 82,561 labelled
images in total. Two are republished closer to their source shape: AMIDDA as upstream
Parquet, and OpenITI-Makhzan as page images with line-level coordinates.
Four of the five converted sources are historical manuscripts. KHATT is modern handwriting
and is included as a separate… See the full description on the dataset page: https://huggingface.co/datasets/TheSeniorTeam/Arabic_Manuscript_Collection_Dataset.arabic_dataAraDiCE
AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs
Overview
The AraDiCE dataset is designed to evaluate dialectal and cultural capabilities in large language models (LLMs). The dataset consists of post-edited versions of various benchmark datasets, curated for validation in cultural and dialectal contexts relevant to Arabic.
As part of the supplemental materials, we have selected a few datasets (see below) for the reader to review. We will make the full… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AraDiCE.buggyarabic_data2laion-coco-nllb-arabic-filtered
LAION-COCO-NLLB, Arabic slice, filtered
The arb_Arab captions of visheratin/laion-coco-nllb,
extracted and filtered for training an Arabic captioning model. This is the exact corpus behind
oddadmix/Nawah-VL-25M.
split
rows
kept from
train
753,883
878,978 (85.8%)
test
14,407
14,906 (96.7%)
Columns
column
type
notes
id
string
the source dataset's image id
url
string
image URL, not the image. See below.
caption
string
the arb_Arab… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/laion-coco-nllb-arabic-filtered.
