datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpt-image-2
GPT-Image-2 Twitter Dataset
10,217 confirmed GPT-image-2.0 generated images collected from Twitter/XCollection window: April 21 – April 28, 2026 (first week post-launch)Paper: GPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of Deployment
Overview
This dataset contains 10,217 images confirmed to be GPT-image-2.0 outputs, sourced from public Twitter/X posts in the immediate aftermath of the model's April 21, 2026 release… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/gpt-image-2.AIForge-Doc-v2
AIForge-Doc v2: A Paired Benchmark of GPT-Image-2 Document Forgeries
AIForge-Doc v2 is the first paired benchmark of document forgeries produced by
OpenAI's GPT-Image-2 (released April 2026). Every forged image is accompanied by
its authentic source image and a pixel-precise tampered-region mask in
DocTamper-compatible format. v2 reuses the forgery specifications of
AIForge-Doc v1 spec-for-spec and swaps only
the generator, so any difference in detector behaviour between v1… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/AIForge-Doc-v2.AIForge-Doc-v1
AIForge-Doc: A Benchmark of AI-Forged Document Images
AIForge-Doc is the first large-scale benchmark of AI-forged document images, targeting
financial and identity document fraud. Every tampered image was produced by a
diffusion-model inpainting pipeline — a threat model that existing forgery detectors
cannot reliably handle.
At a Glance
Attribute
Value
Total forged images
4,061
Training split
3,249 (80 %)
Testing split
812 (20 %)
Authentic… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/AIForge-Doc-v1.SCAM
SCAM Dataset
Dataset Summary
SCAM is the largest and most diverse real-world typographic attack dataset to date, containing images across hundreds of object categories and attack words. The dataset is designed to study and evaluate the robustness of multimodal foundation models against typographic attacks.
Usage:
from datasets import load_dataset
ds = load_dataset("BLISS-e-V/SCAM", split="train")
print(ds)
img = ds[0]['image']
For more information, check out our… See the full description on the dataset page: https://huggingface.co/datasets/BLISS-e-V/SCAM.gpt4o-receipt
GPT4o-Receipt: AI-Generated Receipt Dataset
This directory contains the AI-generated receipts from the
GPT4o-Receipt benchmark, introduced in:
GPT4o-Receipt: A Dataset and Human Study for AI-Generated Document ForensicsYan Zhang*, Simiao Ren*†, Ankit Raj, En Wei, Dennis Ng, Alex Shen, Jiayu Xue, Yuxin Zhang, Evelyn MarottaarXiv:2603.11442 · March 2026 · CC BY-NC-SA 4.0*Equal contribution. †Corresponding author: benren@scam.ai
What Is GPT4o-Receipt?
GPT4o-Receipt is… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/gpt4o-receipt.RWFS
scamai-deepfake-detector-dataset
This repository contains the dataset used in the research paper 'Do Deepfake Detectors Work in Reality?', done by Scam AI.
Real-World Faceswap Dataset (RWFS)
Overview
This repository contains the Real-World Faceswap Dataset (RWFS) used in our research paper "Do Deepfake Detectors Work in Reality?". RWFS is the first dataset specifically designed to reflect real-world deepfakes as they appear in the wild, rather than in… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/RWFS.docdet-scamai-crops
tzj04/docdet-scamai-crops
Training crops derived from the Scam-AI document-forgery datasets, for the
DocDet authentic-vs-AI-generated detector.
This is a derivative work. It is not an official Scam-AI release.
What a row is
Each forgery in the source data patches a single field into an otherwise
authentic scan - roughly 0.3% of the page. At a 224px whole-page input that
edit survives as a handful of pixels, and a random-resized crop can miss it
altogether. So… See the full description on the dataset page: https://huggingface.co/datasets/tzj04/docdet-scamai-crops.age-adversarial-attack
Age Adversarial Attack Dataset
Paper: Can a Teenager Fool an AI? Evaluating Low-Cost Cosmetic Attacks on Age Estimation SystemsAuthors: Simiao Ren (Reality Inc. / Duke University)
Overview
This dataset contains 5,809 AI-generated adversarial images derived from a curated set of 329 face images (ages 10–21) drawn from six standard age estimation benchmarks. Each image is a VLM-simulated cosmetic attack designed to make age estimation models misclassify a subject… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/age-adversarial-attack.
