datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lasa1m-annotate-part-14lasa1m-annotate-part-06lasa1m-annotate-part-02GAIA-annotateddronescapes2_annotated_train_set
Dataset Card for DroneScapes2 (annotated train set)
This is a FiftyOne dataset with 218 samples. It's a subset of this split from the original repo.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/dronescapes2_annotated_train_set.annotated-3DGS-artifacts
Puzzle Similarity
Project page | Paper | Code
This repository contains the dataset presented in the ICCV 2025 paper "Puzzle Similarity: A Perceptually-guided Cross-Reference Metric for Artifact Detection in 3D Scene Reconstructions"Authors: Nicolai Hermann, Jorge Condor, and Piotr Didyk
Dataset Description
The Dataset consists of 36 hand-selected 3D Gaussian Splatting renderings containing common reconstruction artefacts, (aligned) ground truths, human-annotated… See the full description on the dataset page: https://huggingface.co/datasets/nihermann/annotated-3DGS-artifacts.RPCDelements_annotated_tables_4500_docs
Dataset
🚀 Progress
Last update (UTC): 2025-11-11 15:40:21Z
Documents processed: 4500 / 500058
Batches completed: 30
Total pages/rows uploaded: 89882
Latest batch summary
Batch index: 30
Docs in batch: 150
Pages/rows added: 1487
COCO-Wholebody-annotatedzerobench-annotatedai2thor-perspective-qa-annotated-411-splitsspacr-example-annotate
spaCR — Annotate and Classify example data
Example input for the Annotate and Classify modules of
spaCR. It is the output of a Measure
run, so both modules can be exercised without segmenting or measuring anything
first.
What is here
Path
What it is
data/
2,341 single-cell PNG crops, foldered by phenotype
measurements.db
The measurements, plus png_list and the annotation tables
measurements/active_learning/
The model card from the first annotation… See the full description on the dataset page: https://huggingface.co/datasets/einarolafsson/spacr-example-annotate.Deepface_Annotated_3K
Deepface Annotated 3K Dataset
Deepface_Annotated_3K is a synthetic facial image dataset containing 3K AI-generated faces from StyleGAN2.Each image is automatically annotated with demographic attributes like:
Age (in years)
Gender (with prediction confidence)
Dominant Race (White, Asian, Latino Hispanic, Indian, etc.)
The dataset is designed for research on fairness, bias detection, demographic classification, and synthetic face representation.
Structure information… See the full description on the dataset page: https://huggingface.co/datasets/Subh775/Deepface_Annotated_3K.MSP_POD_Annotated_V4D15-annotated
D15-annotated
CoT-annotated multi-label defect detection & typing — 2,684 records (train=2684) of the corrected
AI4Manufacturing/D15 (DefectSpectrum), with
teacher-written reasoning (reasoning).
The query/answer formats were redesigned for foundation-model training (see below); the original D15
query/annot strings are not reused.
The repository name is an internal task code. See Provenance below.
Query diversity (2026-07-11). The query field is drawn from a pool of 40 surface… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/D15-annotated.Radiology_Project_Annotated193-annotated
193-annotated
Chain-of-thought (CoT) reasoning annotations for Severstal steel-sheet surface inspection —
a binary good / anomalous surface-QC task over 12,568 items (all train), the
reasoning-augmented sibling of AI4Manufacturing/193.
Every image, annot, and mask is byte-identical to the base repo; this repo adds the teacher
reasoning trace, a setting-conditioned query, and a per-record metadata.cot provenance block.
Task
Grade each cold-rolled steel-strip… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/193-annotated.GAIA-annotated179-annotated
179-annotated
Chain-of-thought (CoT) reasoning annotations for aero-engine turbine-blade defect inspection (AeBAD)
— 2,160 items (1,011 good + 1,149 defective), the reasoning-augmented sibling of
179-grounding /
179-region /
179-mcq, derived from
AI4Manufacturing/179.
Task
Grade each blade good or defective; if defective, name every defect type and its coarse region.
Defect classes: ablation, breakdown, fracture, groove.
Composition
2,160 rows —… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/179-annotated.refchartqa-balanced-10k-gt-annotatedAnnotated_Medical_pill_for_Object_Detection181-annotated
181-annotated
CoT-annotated binary anomaly detection with coarse localization — 6,300 records from
AI4Manufacturing/181 (DAGM2007): all 2,100
anomalous (both splits) + 4,200 goods (2× per class × split, deterministic id-ordered sample of
the 14,000 — a documented cap, chosen because thousands of near-identical "looks uniform" CoTs teach
template memorization, not inspection).
Query diversity (2026-07-11). The query field is drawn from a pool of 40 surface variants for this task… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/181-annotated.D23-annotated
D23-annotated
CoT-annotated defect detection & classification (VISION), two reasoning channels. Derived from
AI4Manufacturing/D23; category B, task T-B2.
The repository name is an internal task code. See Provenance.
Records
1,217 records (train=547 · validation=670) over 8 of VISION's 14 subsets: Cable (172),
Casting (105), Cylinder (283), Electronics (67), Hemisphere (219), Lens (129), PCB_1 (87),
PCB_2 (155). Splits mirror the official VISION train/val; both… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/D23-annotated.186-annotated
186-annotated
Chain-of-thought (CoT) reasoning annotations for magnetic-tile surface-defect inspection — 1,340
items (952 good + 388 defective), the reasoning-augmented sibling of
186-grounding /
186-region /
186-mcq, derived from
AI4Manufacturing/186.
Task
Grade each tile good or defective; if defective, name every defect type and its coarse region.
Defect classes: Blowhole, Break, Crack, Fray, Uneven.
Composition
1,340 rows — good 952 · defective… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/186-annotated.D20-goods-annotated-linked-to-D15
D20-goods-annotated (linked to D15)
Chain-of-thought inspection annotations for the good (defect-free) examples of MVTec-AD (D20) —
2,066 items, all good — assembled as a positive-sample supplement for
AI4Manufacturing/D15-annotated.
Relationship to D15 (why this set exists)
D15-annotated (the DS-MVTec part) is anomaly-heavy — 386 good vs 1,226 defective (~0.31 : 1; the
pill category has zero goods). This set is D15's positive (good) examples: 2,066 MVTec… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/D20-goods-annotated-linked-to-D15.open-images-v7-mini
AnnotateIt · Open the app · Models & datasets · Documentation
AnnotateIt Open Images V7 Mini Collection
Eight small, real-world, AnnotateIt-compatible datasets curated from the Open Images V7 validation split. Each archive contains 50–200 images, production-exported annotations, source details, and per-image attribution.
Only images whose official Open Images metadata lists CC BY 2.0 are included. Open Images annotations are CC BY 4.0. Open Images recommends independently… See the full description on the dataset page: https://huggingface.co/datasets/AnnotateIt/open-images-v7-mini.D11-QA-annotated
Roles
Roles: reasoning (teacher prose) and reasoning_grounded (deterministic, box-cited, generated in code from the gold geometry) are two reasoning registers; either is a valid SFT imitation target and the two count as one view of the record in a mixture. annot.answer is the machine-parseable gold used for verification and reward parsing; the FINAL ANSWER format is the one the query requests, and annot is not an output-format target.
D11-QA — warehouse scene VQA… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/D11-QA-annotated.invoice-annotated-bboxManually annotated invoice page images exported from AnnotateEverything, with axis-aligned bounding boxes for 8 document-layout regions. Built for training object detectors (YOLO, DETR, etc.) on invoice macro-structure.
Dataset summary
Property
Value
Pages
76
Documents
1
Source PDF
train_images.pdf
Total annotations
771
Avg boxes / page
10.14
Image width range
425 – 2853 px
Image height range
570 – 4096 px
Export date
2026-06-22T19:27:38.375Z… See the full description on the dataset page: https://huggingface.co/datasets/AvoCahDoe/invoice-annotated-bbox.uncut-cells-all-annotatedcure-annotated
