datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scanned-images-dataset-for-ocr-and-vlm-finetuning
Dataset Card for scanned_images_dataset
This is a FiftyOne dataset containing 3,482 scanned document images across 10 diverse document categories. Designed for OCR training and Vision-Language Model (VLM) fine-tuning, this dataset features real-world scanned documents with varied layouts, scanning quality, and document types.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/scanned-images-dataset-for-ocr-and-vlm-finetuning.vlmn_tartandrive100_scand50_coda25_spot100_sub5_full_augmentation_processed_10
Trajectory Ranking Dataset
This dataset contains trajectory ranking results for autonomous navigation scenarios.
Dataset Statistics
Total examples: 39558
Chunks processed: 40
Upload date: 2025-09-13T00:44:30.335177
Features
Image data with terrain analysis
Trajectory rankings and reasoning
Quality and diversity analysis
Terrain and trajectory descriptions
AIForge-Doc-v2
AIForge-Doc v2: A Paired Benchmark of GPT-Image-2 Document Forgeries
AIForge-Doc v2 is the first paired benchmark of document forgeries produced by
OpenAI's GPT-Image-2 (released April 2026). Every forged image is accompanied by
its authentic source image and a pixel-precise tampered-region mask in
DocTamper-compatible format. v2 reuses the forgery specifications of
AIForge-Doc v1 spec-for-spec and swaps only
the generator, so any difference in detector behaviour between v1… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/AIForge-Doc-v2.brain-tumour-MRI-scan
Dataset description
This dataset is a combination of the following three datasets :
FigshareSARTAJ datasetBr35H
This dataset contains 7023 images of human brain MRI images which are divided into 4 classes: glioma - meningioma - no tumor and pituitary.
No tumor class images were taken from the Br35H dataset.
Acknowledgement
This dataset is reproduced and taken from Kaggle
gpt-image-2
GPT-Image-2 Twitter Dataset
10,217 confirmed GPT-image-2.0 generated images collected from Twitter/XCollection window: April 21 – April 28, 2026 (first week post-launch)Paper: GPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of Deployment
Overview
This dataset contains 10,217 images confirmed to be GPT-image-2.0 outputs, sourced from public Twitter/X posts in the immediate aftermath of the model's April 21, 2026 release… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/gpt-image-2.AIForge-Doc-v1
AIForge-Doc: A Benchmark of AI-Forged Document Images
AIForge-Doc is the first large-scale benchmark of AI-forged document images, targeting
financial and identity document fraud. Every tampered image was produced by a
diffusion-model inpainting pipeline — a threat model that existing forgery detectors
cannot reliably handle.
At a Glance
Attribute
Value
Total forged images
4,061
Training split
3,249 (80 %)
Testing split
812 (20 %)
Authentic… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/AIForge-Doc-v1.SCAM
SCAM Dataset
Dataset Summary
SCAM is the largest and most diverse real-world typographic attack dataset to date, containing images across hundreds of object categories and attack words. The dataset is designed to study and evaluate the robustness of multimodal foundation models against typographic attacks.
Usage:
from datasets import load_dataset
ds = load_dataset("BLISS-e-V/SCAM", split="train")
print(ds)
img = ds[0]['image']
For more information, check out our… See the full description on the dataset page: https://huggingface.co/datasets/BLISS-e-V/SCAM.scanned-images-dataset-for-ocr-and-vlm-finetuning
Dataset Card for scanned_images_dataset
This is a FiftyOne dataset containing 3,482 scanned document images across 10 diverse document categories. Designed for OCR training and Vision-Language Model (VLM) fine-tuning, this dataset features real-world scanned documents with varied layouts, scanning quality, and document types.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from… See the full description on the dataset page: https://huggingface.co/datasets/prabhats0605/scanned-images-dataset-for-ocr-and-vlm-finetuning.brain-tumour-MRI-scan
Dataset description
This dataset is a combination of the following three datasets :
FigshareSARTAJ datasetBr35H
This dataset contains 7023 images of human brain MRI images which are divided into 4 classes: glioma - meningioma - no tumor and pituitary.
No tumor class images were taken from the Br35H dataset.
Acknowledgement
This dataset is reproduced and taken from Kaggle
TreeOfLife-10M
Dataset Card for TreeOfLife-10M
Dataset Summary
With over 10 million images covering 454 thousand taxa in the tree of life, TreeOfLife-10M is the largest-to-date ML-ready dataset of images of biological organisms paired with their associated taxonomic labels. It expands on the foundation established by existing high-quality datasets, such as iNat21 and BIOSCAN-1M, by further incorporating newly curated images from the Encyclopedia of Life (eol.org), which supplies most of… See the full description on the dataset page: https://huggingface.co/datasets/ScarlettChan/TreeOfLife-10M.hair-loss-male-norwood-scale
Male Hair Loss Dataset - 2 400+ images
Dataset comprises 2,400+ photos of male alopecia (hair loss) captured from 5 angles, meticulously labeled into 7 classes according to the Norwood-Hamilton scale. It is designed for machine learning and deep learning applications, particularly in diagnosing hair disorders, evaluating scalp health, and personalizing hair restoration treatments.
By utilizing this dataset, researchers and dermatologists can enhance hair loss analysis, improve… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/hair-loss-male-norwood-scale.gpt4o-receipt
GPT4o-Receipt: AI-Generated Receipt Dataset
This directory contains the AI-generated receipts from the
GPT4o-Receipt benchmark, introduced in:
GPT4o-Receipt: A Dataset and Human Study for AI-Generated Document ForensicsYan Zhang*, Simiao Ren*†, Ankit Raj, En Wei, Dennis Ng, Alex Shen, Jiayu Xue, Yuxin Zhang, Evelyn MarottaarXiv:2603.11442 · March 2026 · CC BY-NC-SA 4.0*Equal contribution. †Corresponding author: benren@scam.ai
What Is GPT4o-Receipt?
GPT4o-Receipt is… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/gpt4o-receipt.Multi_Scale_ImageNet
Multi-Scale_ImageNet
The dataset contains multi-scale ImageNet from Local Scale Equivariance with Latent Deep Equilibrium Canonicalizer (arXiv:2508.14187) published in ICCV 2025.
The dataset generating pipeline is provided here: https://github.com/ashiq24/local-scale-equivariance/tree/main/datagen_ImNet
Pipeline Overview
Global Scaling
Local Scaling
If you use this dataset, please consider citing
@inproceedings{rahman2025local,
title={Local Scale… See the full description on the dataset page: https://huggingface.co/datasets/ashiq24/Multi_Scale_ImageNet.RWFS
scamai-deepfake-detector-dataset
This repository contains the dataset used in the research paper 'Do Deepfake Detectors Work in Reality?', done by Scam AI.
Real-World Faceswap Dataset (RWFS)
Overview
This repository contains the Real-World Faceswap Dataset (RWFS) used in our research paper "Do Deepfake Detectors Work in Reality?". RWFS is the first dataset specifically designed to reflect real-world deepfakes as they appear in the wild, rather than in… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/RWFS.KELMARSH-SCADA
KELMARSH-SCADA — turbine availability state from the operating point (reasoning track)
Part of the AI4Manufacturing FORGE corpus (Category C, task T-C1), and the corpus's first
SCADA dataset. Each record is one ten-minute operating record of a utility-scale wind turbine,
drawn as two panels: on the left, where that record sits on the wind-speed/power plane against the
fleet's own measured power curve (median line, inter-quartile band, and a scatter of the fleet's
operating… See the full description on the dataset page: https://huggingface.co/datasets/AI4Manufacturing/KELMARSH-SCADA.docdet-scamai-crops
tzj04/docdet-scamai-crops
Training crops derived from the Scam-AI document-forgery datasets, for the
DocDet authentic-vs-AI-generated detector.
This is a derivative work. It is not an official Scam-AI release.
What a row is
Each forgery in the source data patches a single field into an otherwise
authentic scan - roughly 0.3% of the page. At a 224px whole-page input that
edit survives as a handful of pixels, and a random-resized crop can miss it
altogether. So… See the full description on the dataset page: https://huggingface.co/datasets/tzj04/docdet-scamai-crops.brain-tumour-MRI-scan
Dataset description
This dataset is a combination of the following three datasets :
FigshareSARTAJ datasetBr35H
This dataset contains 7023 images of human brain MRI images which are divided into 4 classes: glioma - meningioma - no tumor and pituitary.
No tumor class images were taken from the Br35H dataset.
Acknowledgement
This dataset is reproduced and taken from Kaggle
brain-tumor-single-slice-MRI-scan-with-synthetic-ehr-africa
Dataset Card: Africa Brain Tumor Scans with Synthetic EHR (Bundled Parquet)
This dataset bundles single-slice brain MRI scans and richly structured, synthetic EHR data into a single Parquet file suitable for multimodal ML research. Each row contains an image struct (bytes + path), a source label column, and an EHR payload with both a full JSON record and convenient summary columns.
The synthetic EHRs are Africa-focused: they encode country, urban/rural, facility level, insurance… See the full description on the dataset page: https://huggingface.co/datasets/saad02/brain-tumor-single-slice-MRI-scan-with-synthetic-ehr-africa.age-adversarial-attack
Age Adversarial Attack Dataset
Paper: Can a Teenager Fool an AI? Evaluating Low-Cost Cosmetic Attacks on Age Estimation SystemsAuthors: Simiao Ren (Reality Inc. / Duke University)
Overview
This dataset contains 5,809 AI-generated adversarial images derived from a curated set of 329 face images (ages 10–21) drawn from six standard age estimation benchmarks. Each image is a VLM-simulated cosmetic attack designed to make age estimation models misclassify a subject… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/age-adversarial-attack.brain-tumour-MRI-scan
Dataset description
This dataset is a combination of the following three datasets :
FigshareSARTAJ datasetBr35H
This dataset contains 7023 images of human brain MRI images which are divided into 4 classes: glioma - meningioma - no tumor and pituitary.
No tumor class images were taken from the Br35H dataset.
Acknowledgement
This dataset is reproduced and taken from Kaggle
scallop_mosaic_640_quantization_sample
Scallop YOLOv5s Mosaic 640 - Quantization Sample
A classless 1000-train / 1000-val image subset of the tiled 3x3 640px mosaic dataset designed specifically for RKNN/tflite/ONNX representative quantization calibration on edge devices (like the RV1126 Aura).
Attribution & License
This dataset is a derivative work based on the University of St Andrews King Scallop dataset.
Original DOI: 10.5281/zenodo.10156830
In accordance with the original dataset's terms, this… See the full description on the dataset page: https://huggingface.co/datasets/FishingROV/scallop_mosaic_640_quantization_sample.
