datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MapPool
MapPool - Bubbling up an extremely large corpus of maps for AI
MapPool is a dataset of 75 million potential maps and textual captions. It has been derived from CommonPool, a dataset consisting of 12 billion text-image pairs from the Internet. The images have been encoded by a vision transformer and classified into maps and non-maps by a support vector machine. This approach outperforms previous models and yields a validation accuracy of 98.5%. The MapPool dataset may help to train… See the full description on the dataset page: https://huggingface.co/datasets/sraimund/MapPool.srtm-global-void-filledThis dataset mirrors the
Shuttle Radar Topography Mission (SRTM) Void Filled digital elevation data from USGS.
It consists of the 15,417 GeoTIFFs available on USGS EarthExplorer in the "SRTM Void Filled" (srtm_v2) dataset.
Each GeoTIFF covers 1x1 degrees.
The data is in WGS84, with a resolution of 1 arc-second/pixel in the United States and 3 arc-seconds/pixel elsewhere.
Coverage is limited to "80% of the Earth's land surface between 60° north and 56° south latitude".
The data is attributed to… See the full description on the dataset page: https://huggingface.co/datasets/allenai/srtm-global-void-filled.srankmonsternobehemothdakedonekotomachigawareteelfmusumenopettoshitekurashitemasu
Bangumi Image Base of S-rank Monster No "behemoth" Dakedo, Neko To Machigawarete Elf Musume No Pet Toshite Kurashitemasu
This is the image base of bangumi S-Rank Monster no "Behemoth" dakedo, Neko to Machigawarete Elf Musume no Pet toshite Kurashitemasu, we detected 56 characters, 4649 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/srankmonsternobehemothdakedonekotomachigawareteelfmusumenopettoshitekurashitemasu.srp-staging-controlled
srp-staging-controlled
Flat asset repo of Simready Asset Packages.
Branches
main — the published lane. Holds
.packages/simready.hf.nvidia.srp-staging-controlled.<asset>/<YYYY.MM.DD_NN>/
and nothing else.
from_ovstorage — the corpus assets are selected from, mirrored from rc.storage
under main/simready.ov/simready_content/assets_staging.
ticket_<user>_<timestamp> — one per submission, created at submit and deleted
when the ticket closes.
Versions
A… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/srp-staging-controlled.sr-artifact-prominence
SR Artifact Prominence
Annotated super-resolution artifact regions across four image subsets, with
crowdsourced per-region prominence scores, artifact type labels, and
natural-language descriptions.
Prominence is the fraction of valid crowd workers who answered that the
highlighted region contains a noticeable super-resolution artifact.
Subsets
Subset
Source dataset
Source images
Masks
Notes
open_images
Open Images
547
1,523
GT + LR-bicubic + multiple SR… See the full description on the dataset page: https://huggingface.co/datasets/imolodetskikh/sr-artifact-prominence.CountQA
Dataset Summary
CountQA is the new benchmark designed to stress-test the Achilles' heel of even the most advanced Multimodal Large Language Models (MLLMs): object counting. While modern AI demonstrates stunning visual fluency, it often fails at this fundamental cognitive skill, a critical blind spot limiting its real-world reliability.
This dataset directly confronts that weakness with over 1,500 challenging question-answer pairs built on real-world images, hand-captured to feature… See the full description on the dataset page: https://huggingface.co/datasets/Jayant-Sravan/CountQA.sr-home-mea_001-un-bathroom_c004_l1_f1-2026-06-09-lossless-3
SR-HOME-MEA_001-UN-BATHROOM_C004_L1_F1
High-quality synthetic render data-pack featuring a detailed residential bathroom scene, captured from Camera_04 with 100 rendered outputs and 1,200 files across 12 production-ready render passes.
This dataset is designed for computer vision, 3D perception, material analysis, segmentation, depth estimation, and synthetic data workflows. It includes beauty renders plus rich auxiliary passes such as albedo, depth, normals, UVW, material index… See the full description on the dataset page: https://huggingface.co/datasets/DaiPatrick/sr-home-mea_001-un-bathroom_c004_l1_f1-2026-06-09-lossless-3.srt-coco-thumbsICDAR2019-SROIE
ICDAR2019's Scanned Receipts OCR and Information Extraction (SROIE)
The ICDAR2019 SROIE dataset was originally published by Huang et al. for the
15th International Conference on Document Analysis and Recognition (ICDAR2019)
Robust Reading Challenge on Scanned Receipts OCR and Information Extraction
(SROIE).
This work presents an extension of the original ICDAR2019 SROIE dataset, including 14
receipt annotations missing from the original Task 3 test dataset, in a format
integrated… See the full description on the dataset page: https://huggingface.co/datasets/jsdnrs/ICDAR2019-SROIE.Sewer-pipe-defectssr-home-mea_001-un-bathroom_c004_l1_f1-2026-06-09
SR-HOME-MEA_001-UN-BATHROOM_C004_L1_F1
High-quality synthetic render data-pack featuring a detailed residential bathroom scene, captured from Camera_04 with 100 rendered outputs and 1,200 files across 12 production-ready render passes.
This dataset is designed for computer vision, 3D perception, material analysis, segmentation, depth estimation, and synthetic data workflows. It includes beauty renders plus rich auxiliary passes such as albedo, depth, normals, UVW, material index… See the full description on the dataset page: https://huggingface.co/datasets/DaiPatrick/sr-home-mea_001-un-bathroom_c004_l1_f1-2026-06-09.LSDIR_SR_imagesSROIE_2019_with_labelsThis is a fork of Urban Knupleš's "SROIE datasetv2", with ground truth labels for invoice numbers. train/labels.json and test/test_labels.json contains the ground truth invoice numbers for the associated image.
Below is a copy + paste from his repo. Here is my dataset documentation and notes. See there as well for heuristics used to label the dataset. The latest version tag is data-v2.1.
Scanned receipts OCR and information extraction (SROIE) + LayoutLM (base)
This dataset was… See the full description on the dataset page: https://huggingface.co/datasets/ryanznie/SROIE_2019_with_labels.Vision-SR1-47Klgd-cards-video-day1
LGD Cards — Day-1 PoC Video (YOLO card-pip dataset)
Auto-labeled playing-card corner-pip detection tiles from the first day of our
proof-of-concept table recordings, for Live Game Defender (LGD) — an on-prem AI integrity
monitor for live casino table games. This is the day-1 training set behind the
lgd-cards-gen2 detector.
Format: YOLO — images/{train,val} + labels/{train,val}, data.yaml (52 classes, rank+suit).
~4,630 tiles, serve-matching 2×2 tiling of the source frames.… See the full description on the dataset page: https://huggingface.co/datasets/sroot/lgd-cards-video-day1.asset-yolo-dataset
Asset YOLO Dataset
Auto-annotated. 84 classes.
VisTA-SR
VisTA-SR: Paired Low/High-Resolution Thermal & RGB Agricultural Dataset (Training_T4_1_2_3)
Official dataset repository for the CVPR 2024 Workshop paper:
"VisTA-SR: Improving the Accuracy and Resolution of Low-Cost Thermal Imaging Cameras for Agriculture"
📄 Paper HTML: CVPR 2024 OpenAccess
📄 Paper PDF: Download PDF
💻 Official Codebase: https://github.com/heesup/VisTA-SR
Dataset Description
This dataset consists of aligned multi-modal image triplets captured… See the full description on the dataset page: https://huggingface.co/datasets/heesup/VisTA-SR.HueManity
HueManity: A Benchmark for Testing Human-Like Visual Perception in MLLMs
Paper | Code
Dataset Description
HueManity is a benchmark dataset featuring 83,850 images designed to test the fine-grained visual perception of Multimodal Large Language Models (MLLMs). Each image presents a two-character alphanumeric string embedded within Ishihara-style dot patterns, challenging models to perform precise pattern recognition in visually cluttered environments.
The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/Jayant-Sravan/HueManity.srmk_debugsri-lankan-album-stamp-detection
Sri Lankan Album Stamp Detection Dataset
This dataset contains annotated Sri Lankan stamp album page images prepared for stamp detection using YOLO-based object detection models.
It is used for the Stamp AI Model 1 stamp detector, where the goal is to detect individual stamps from full album page images.
Contents
raw_pages/ — original album page images
yolo_dataset/ — YOLO-formatted detection dataset
The YOLO dataset contains image files, label files, and a… See the full description on the dataset page: https://huggingface.co/datasets/nethsith/sri-lankan-album-stamp-detection.sri-lankan-stamp-matcher-dataset
Sri Lankan Stamp Matcher Dataset
Source data for the Stamp AI Model 2 stamp-matching project.
Contents
reviewed_crops/ — reviewed single-stamp crop images
metadata-final.xlsx — master metadata file
The Excel file contains the review status, exact stamp grouping, design-family grouping, face value, overprint details, condition, and capture-quality notes.
accepted, pending, and excluded statuses are retained so the dataset can be reviewed and regenerated… See the full description on the dataset page: https://huggingface.co/datasets/nethsith/sri-lankan-stamp-matcher-dataset.pid_lines_dataset
P&ID Line Detection Dataset
This dataset contains cropped images from P&ID (Piping and Instrumentation Diagrams)
with line segment annotations for line detection and segmentation tasks.
Dataset Structure
Each sample contains:
file_name: Image filename
source_image_idx: Index of the original P&ID image
crop_idx: Index of this crop from the source image
width: Crop width in pixels
height: Crop height in pixels
lines: Dictionary with:
segments: List of line… See the full description on the dataset page: https://huggingface.co/datasets/Sri1311/pid_lines_dataset.sroie_data_setsrota-assetsCOCO-2017
MS COCO 2017
The complete COCO 2017 release — all four image splits and every annotation family — in Parquet.
from datasets import load_dataset
# annotations only (~6 MB) -- no image bytes
ann = load_dataset("srishti-kaushik/COCO-2017", "instances_val2017", split="train")
# images
img = load_dataset("srishti-kaushik/COCO-2017", "images_val2017", split="train")
img[0]["image"] # PIL.Image
# stream, instead of downloading 19 GB
train =… See the full description on the dataset page: https://huggingface.co/datasets/srishti-kaushik/COCO-2017.electrical-panels-dataset
Electrical Panels Detection Dataset
Auto-scraped, CLIP-filtered, YOLOE-26 annotated.
Classes: 107
Target images per class: 500
Annotation: Two-pass YOLOE-26m + SAM
Teacher model: YOLOE-26m-seg
Student model: YOLO26n (knowledge distilled)
Vision-SR1-Cold-9KThe dataset_info.json contains all available datasets. If you are using a custom dataset, please make sure to add a dataset description in dataset_info.json and specify dataset: dataset_name before training to use it.
The dataset_info.json file should be put in the dataset_dir directory. You can change dataset_dir to use another directory. The default value is ./data.
Currently we support datasets in alpaca and sharegpt format. Allowed file types include json, jsonl, csv, parquet, arrow.… See the full description on the dataset page: https://huggingface.co/datasets/LMMs-Lab-Turtle/Vision-SR1-Cold-9K.fire-segmentation-dataset
Fire Segmentation Dataset (YOLO-seg format)
Instance-segmentation dataset for fire detection: 1,348 images with
polygon mask labels in Ultralytics YOLO segmentation format. Built to train
sreeharivp23/fire-segmentation-yolo11n.
Contents
Split
Images
train
1,146
val
202
1,098 fire images with one or more fire polygon instances
250 negatives (no fire) with empty label files
fire_seg/
├── data.yaml # Ultralytics dataset config (1 class:… See the full description on the dataset page: https://huggingface.co/datasets/sreeharivp23/fire-segmentation-dataset.sroie-2019-v2ICDAR 2019 Robust Reading Challenge on Scanned Receipts OCR and Information Extraction
Dataset taken from https://rrc.cvc.uab.es/?ch=13&com=downloads. Duplicate images/annotations were removed by https://github.com/Losyash/SROIE-datasetv2 as far as I can tell.
license: cc-by-2.0
look-bench
LookBench: A Live and Holistic Fashion Image Retrieval Benchmark
LookBench is a large-scale, open benchmark for fashion image retrieval, designed to evaluate modern vision and vision–language models under realistic, contamination-aware settings. The benchmark emphasizes live data, domain diversity, and holistic retrieval tasks spanning both single-item and outfit-level scenarios.
This dataset accompanies the paper LookBench: A Live and Holistic Open Benchmark for Fashion Image… See the full description on the dataset page: https://huggingface.co/datasets/srpone/look-bench.
