scam
Datasets
All datasets matching “scam”gpt-image-2
GPT-Image-2 Twitter Dataset
10,217 confirmed GPT-image-2.0 generated images collected from Twitter/XCollection window: April 21 – April 28, 2026 (first week post-launch)Paper: GPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of Deployment
Overview
This dataset contains 10,217 images confirmed to be GPT-image-2.0 outputs, sourced from public Twitter/X posts in the immediate aftermath of the model's April 21, 2026 release… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/gpt-image-2.AIForge-Doc-v1
AIForge-Doc: A Benchmark of AI-Forged Document Images
AIForge-Doc is the first large-scale benchmark of AI-forged document images, targeting
financial and identity document fraud. Every tampered image was produced by a
diffusion-model inpainting pipeline — a threat model that existing forgery detectors
cannot reliably handle.
At a Glance
Attribute
Value
Total forged images
4,061
Training split
3,249 (80 %)
Testing split
812 (20 %)
Authentic… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/AIForge-Doc-v1.AIForge-Doc-v2
AIForge-Doc v2: A Paired Benchmark of GPT-Image-2 Document Forgeries
AIForge-Doc v2 is the first paired benchmark of document forgeries produced by
OpenAI's GPT-Image-2 (released April 2026). Every forged image is accompanied by
its authentic source image and a pixel-precise tampered-region mask in
DocTamper-compatible format. v2 reuses the forgery specifications of
AIForge-Doc v1 spec-for-spec and swaps only
the generator, so any difference in detector behaviour between v1… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/AIForge-Doc-v2.filipino-scam-final-shards
Filipino Scam — Final Shards
CoT-formatted pickle shards for Qwen3-VL-4B-Instruct LoRA fine-tuning on Filipino
short-form video scam detection. This is the final, training-ready format used to
build TamAko783/Scam-Qwen3-VL-4B-final-lora.
Splits
File
Samples
Purpose
Training.pkl
1600
Training set
Validate.pkl
198
Validation (early stopping + best-checkpoint selection)
Evaluate.pkl
202
Held-out test (final unbiased metrics)
Both Validate and Evaluate are… See the full description on the dataset page: https://huggingface.co/datasets/TamAko783/filipino-scam-final-shards.handball_video_sequencesSCAM
SCAM Dataset
Dataset Summary
SCAM is the largest and most diverse real-world typographic attack dataset to date, containing images across hundreds of object categories and attack words. The dataset is designed to study and evaluate the robustness of multimodal foundation models against typographic attacks.
Usage:
from datasets import load_dataset
ds = load_dataset("BLISS-e-V/SCAM", split="train")
print(ds)
img = ds[0]['image']
For more information, check out our… See the full description on the dataset page: https://huggingface.co/datasets/BLISS-e-V/SCAM.
