datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DRIVE-digital-retinal-images-for-vessel-extractionarXiv:2501.18921https://arxiv.org/abs/2501.18921
price-tag-extraction
Price tag extraction dataset
This dataset contains images of price tags extracted from the Open Prices dataset, along with information extracted from this dataset.
It is intended to be used to train large visual language models (LVLMs) to extract information from price tag images, as part of the Open Prices project.
For more information about the format of this dataset, please refer to the documentation on llm-image-extraction datasets.
Dataset creation
A detailed… See the full description on the dataset page: https://huggingface.co/datasets/openfoodfacts/price-tag-extraction.mall_receipt_extraction_datasetdoab-metadata-extraction
DOAB Open Access Books - Metadata Extraction Dataset
Dataset Description
This dataset contains 9,363 open access books with page images and rich bibliographic metadata extracted from MARC21 records, curated specifically for training and evaluating Vision Language Models (VLMs) on automatic metadata extraction from scholarly monographs.
The dataset is derived from the Penn State ScholarSphere DOAB collection (Directory of Open Access Books), focusing on books with Creative… See the full description on the dataset page: https://huggingface.co/datasets/biglam/doab-metadata-extraction.calcium-source-extraction
Calcium source extraction
Synthetic one-photon (miniscope-style) calcium imaging of mouse cortex, rendered with a physically
motivated simulator that models somatic morphology at several depths, calcium indicator kinetics,
correlated neuropil, vasculature and hemodynamics, non-stationary rigid motion, wide-field optics
and an sCMOS camera, with complete ground truth. The simulator is not part of this release.
data/heldout/video.tif 60 s at 15 Hz, 512 x 512, 8-bit, 900… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/calcium-source-extraction.table-extraction
Dataset Labels
['bordered', 'borderless']
Number of Images
{'test': 34, 'train': 238, 'valid': 70}
How to Use
Install datasets:
pip install datasets
Load the dataset:
from datasets import load_dataset
ds = load_dataset("keremberke/table-extraction", name="full")
example = ds['train'][0]
Roboflow Dataset Page
https://universe.roboflow.com/mohamed-traore-2ekkp/table-extraction-pdf/dataset/2
Citation
License
CC… See the full description on the dataset page: https://huggingface.co/datasets/keremberke/table-extraction.cord-extraction-lora
CORD structured-extraction (LoRA training set)
200 train + 20 val examples derived from CORD (naver-clova-ix/cord-v2, train
split). Each example pairs a preprocessed receipt image with the target
extraction JSON (receipt fields + line items in the pipeline's schema).
prompt.txt - the shared instruction (schema skeleton injected)
train.jsonl / val.jsonl - lines of {"id", "image": "images/..png", "target": "<json>"}
images/ - the preprocessed pages (deskew / resize<=1536 /… See the full description on the dataset page: https://huggingface.co/datasets/sarcasticcoder/cord-extraction-lora.drone-building-extraction
🏘️ Drone Building Extraction Dataset — Calgary
A high-resolution drone imagery dataset for binary building segmentation,
captured over a residential area in Calgary, Alberta, Canada.
Designed to work directly with GeoSeg Studio
— an open-source QGIS plugin for deep learning semantic segmentation.
🖼️ Preview
Train Area (red = building polygons)
Test Area (blue = building polygons)
📋 Dataset Summary
Property
Value
Location
Calgary… See the full description on the dataset page: https://huggingface.co/datasets/dronnix-io/drone-building-extraction.poster-schedule-information-extraction
Multimodal Visual-Text Dataset for Poster Schedule Information Extraction
A ready-to-train Indonesian document AI dataset combining pixels, OCR tokens, spatial layout, and BIO entity labels.
Overview
What it is
127 Indonesian seminar and religious-study event posters with multimodal token-level annotations
Primary task
Schedule information extraction as token classification
Modalities
Image + text + 2D spatial layout
Coordinates… See the full description on the dataset page: https://huggingface.co/datasets/Ibnuck/poster-schedule-information-extraction.doab-metadata-extraction
DOAB Open Access Books - Metadata Extraction Dataset
Dataset Description
This dataset contains 9,363 open access books with page images and rich bibliographic metadata extracted from MARC21 records, curated specifically for training and evaluating Vision Language Models (VLMs) on automatic metadata extraction from scholarly monographs.
The dataset is derived from the Penn State ScholarSphere DOAB collection (Directory of Open Access Books), focusing on books with Creative… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/doab-metadata-extraction.kyc-document-extraction-vlmfurniture-model-extraction
Furniture Model Number Extraction Dataset
This dataset contains furniture inventory images with corresponding model numbers for training vision-language models to extract product model numbers from furniture store photos.
Dataset Description
Created: 2025-09-24T12:26:54.209816
Task: Vision-Language Model Training for Model Number Extraction
Base Model: IBM Granite Vision 3.2 2B
Domain: Furniture Inventory Management
Dataset Statistics
Training Samples: 219… See the full description on the dataset page: https://huggingface.co/datasets/wynnwatson/furniture-model-extraction.complete-buildings-extraction-coco-hfglam-extraction-benchmark
GLAM extraction benchmark
Structured extraction from cultural-heritage documents. The first configuration is
nls-index-cards: 98 manuscript catalogue cards from the National Library of Scotland.
Source and credits
Derived from NationalLibraryOfScotland/index-cards-eval,
revision 2a81070549d8493c2c538744a9dbbc1dc72cb146 (CC0). Images and checked outputs are preserved.
NLS cataloguers reviewed the model-drafted labels: 66 accepted as drafted, 32 corrected.
Drafting… See the full description on the dataset page: https://huggingface.co/datasets/small-models-for-glam/glam-extraction-benchmark.br-doc-extraction
Dataset Card for Brazilian Document Structure Extraction
Dataset Description
This dataset contains 1218 images of Brazilian identification documents (CNH - National Driver's License, RG - General Registration) and invoices (NF - Nota Fiscal). Each image is paired with a user-defined JSON schema (as a "prefix") and the corresponding structured data extraction (as a "suffix" in JSON string format).
The primary goal of this dataset is to facilitate the fine-tuning of… See the full description on the dataset page: https://huggingface.co/datasets/CLMARRARA/br-doc-extraction.receipt_VLM_information_extractionbr-doc-extraction
Dataset Card for Brazilian Document Structure Extraction
Dataset Description
This dataset contains 1218 images of Brazilian identification documents (CNH - National Driver's License, RG - General Registration) and invoices (NF - Nota Fiscal). Each image is paired with a user-defined JSON schema (as a "prefix") and the corresponding structured data extraction (as a "suffix" in JSON string format).
The primary goal of this dataset is to facilitate the fine-tuning of… See the full description on the dataset page: https://huggingface.co/datasets/tech4humans/br-doc-extraction.kyc-document-extraction-vlmMatE_polyhaven_extraction
license: cc0-1.0
pretty_name: Poly Haven Material Extraction Evaluation Set
size_categories:
- n<1K
task_categories:
- image-to-image
tags:
- computer-vision
- pbr materials
- textures
- polyhaven
- evaluation
- material-extraction
PolyHaven Material Extraction Evaluation Set
This dataset provides the PolyHaven evaluation set used in MatE: Material Extraction from Single-Image via Geometric Prior. It is designed for evaluating material extraction methods that recover physically… See the full description on the dataset page: https://huggingface.co/datasets/tiptoez/MatE_polyhaven_extraction.buildings-extraction-coco-hf
Building Extraction Dataset
This dataset is a processed verison of the dataset of the Kaggle competition: https://www.kaggle.com/competitions/building-extraction-generalization-2024/.
The original train and validation images and COCO annotations were resized to (512, 512).
Then from the segmentations on the COCO annotations, a PIL_annotation file was created for each sample in which the Red channel corresponds to the semantic segmentation mask and the Green channel to the instance… See the full description on the dataset page: https://huggingface.co/datasets/tomascanivari/buildings-extraction-coco-hf.ppt_shapes_extraction
Dataset Card for "ppt_shapes_extraction"
More Information needed
chart-extraction-synth-v1summer-buildings-extraction-coco-hfchart-extraction-dense-noisy-v1doc-handwriting-extraction-complex-1doab-metadata-extraction-trl
DOAB Open Access Books - Metadata Extraction Dataset
Dataset Description
This dataset contains 9,363 open access books with page images and rich bibliographic metadata extracted from MARC21 records, curated specifically for training and evaluating Vision Language Models (VLMs) on automatic metadata extraction from scholarly monographs.
The dataset is derived from the Penn State ScholarSphere DOAB collection (Directory of Open Access Books), focusing on books with Creative… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/doab-metadata-extraction-trl.receipt_VLM_information_extractionstructured-generation-information-extraction-vlms-openbmb-RLAIF-V-DatasetEntity_Extraction_Datastructured-generation-information-extraction-vlms-openbmb-RLAIF-V-Dataset
