datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpt-image-2
GPT-Image-2 Twitter Dataset
10,217 confirmed GPT-image-2.0 generated images collected from Twitter/XCollection window: April 21 – April 28, 2026 (first week post-launch)Paper: GPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of Deployment
Overview
This dataset contains 10,217 images confirmed to be GPT-image-2.0 outputs, sourced from public Twitter/X posts in the immediate aftermath of the model's April 21, 2026 release… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/gpt-image-2.AIForge-Doc-v1
AIForge-Doc: A Benchmark of AI-Forged Document Images
AIForge-Doc is the first large-scale benchmark of AI-forged document images, targeting
financial and identity document fraud. Every tampered image was produced by a
diffusion-model inpainting pipeline — a threat model that existing forgery detectors
cannot reliably handle.
At a Glance
Attribute
Value
Total forged images
4,061
Training split
3,249 (80 %)
Testing split
812 (20 %)
Authentic… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/AIForge-Doc-v1.AIForge-Doc-v2
AIForge-Doc v2: A Paired Benchmark of GPT-Image-2 Document Forgeries
AIForge-Doc v2 is the first paired benchmark of document forgeries produced by
OpenAI's GPT-Image-2 (released April 2026). Every forged image is accompanied by
its authentic source image and a pixel-precise tampered-region mask in
DocTamper-compatible format. v2 reuses the forgery specifications of
AIForge-Doc v1 spec-for-spec and swaps only
the generator, so any difference in detector behaviour between v1… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/AIForge-Doc-v2.filipino-scam-final-shards
Filipino Scam — Final Shards
CoT-formatted pickle shards for Qwen3-VL-4B-Instruct LoRA fine-tuning on Filipino
short-form video scam detection. This is the final, training-ready format used to
build TamAko783/Scam-Qwen3-VL-4B-final-lora.
Splits
File
Samples
Purpose
Training.pkl
1600
Training set
Validate.pkl
198
Validation (early stopping + best-checkpoint selection)
Evaluate.pkl
202
Held-out test (final unbiased metrics)
Both Validate and Evaluate are… See the full description on the dataset page: https://huggingface.co/datasets/TamAko783/filipino-scam-final-shards.handball_video_sequencesSCAM
SCAM Dataset
Dataset Summary
SCAM is the largest and most diverse real-world typographic attack dataset to date, containing images across hundreds of object categories and attack words. The dataset is designed to study and evaluate the robustness of multimodal foundation models against typographic attacks.
Usage:
from datasets import load_dataset
ds = load_dataset("BLISS-e-V/SCAM", split="train")
print(ds)
img = ds[0]['image']
For more information, check out our… See the full description on the dataset page: https://huggingface.co/datasets/BLISS-e-V/SCAM.multi-agent-scam-conversation
Synthetic Multi-Turn Scam and Non-Scam Phone Conversation Dataset with Agentic Personalities
Dataset Description
The Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset with Agentic Personalities is an enhanced collection of simulated phone conversations between two AI agents, one acting as a scammer or non-scammer and the other as an innocent receiver. Each dialogue is labeled as either a scam or non-scam interaction. This dataset is designed to help develop… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/multi-agent-scam-conversation.all-scam-spamThis is a large corpus of 42,619 preprocessed text messages and emails sent by humans in 43 languages. is_spam=1 means spam and is_spam=0 means ham.
1040 rows of balanced data, consisting of casual conversations and scam emails in ≈10 languages, were manually collected and annotated by me, with some help from ChatGPT.
Some preprcoessing algorithms
spam_assassin.js, followed by spam_assassin.py
enron_spam.py
Data composition
Description
To make the text… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/all-scam-spam.scam-dialogue
Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset
Dataset Description
The Synthetic Multi-Turn Scam and Non-Scam Phone Dialogue Dataset is a collection of simulated phone conversation between two parties, labeled as either scam or non-scam interactions. The dataset is designed to help develop and evaluate models for detecting and classifying various types of phone-based scams.
Dataset Structure
The dataset consists of three columns:
dialogue: The… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/scam-dialogue.filipino-scam-prepared-cacheEurovoc
🇪🇺 🏷️ EuroVoc dataset
This dataset contains more that 4 millions documents in 39 languages with associated EuroVoc labels.
What's Cellar ?
Cellar is the common data repository of the Publications Office of the European Union. Digital publications and metadata are stored in and disseminated via Cellar, in order to be used by humans and machines. Aiming to transparently serve users, Cellar stores multilingual publications and metadata, it is open to all EU… See the full description on the dataset page: https://huggingface.co/datasets/scampion/Eurovoc.video-finance-scam-awarenessscambench-training
ScamBench Training Corpus
A multilingual, multi-turn training corpus for building scam-resistant autonomous agents.
Dataset Description
ScamBench is a comprehensive dataset designed to train AI agents to resist social engineering, phishing, prompt injection, credential theft, impersonation, advance-fee fraud, and other adversarial attacks while maintaining helpfulness for legitimate requests.
Key Features
37,423 total records across 14 languages
154 attack… See the full description on the dataset page: https://huggingface.co/datasets/shaw/scambench-training.jawbreaker-scam-defense-data
Jawbreaker Scam Defense Data
Synthetic and sanitized training/eval data for Jawbreaker, a local-first scam defense app for someone you love.
Jawbreaker turns a suspicious text, email, or DM into a plain-English safety card: the risk, the warning signs, and the safest next step before someone replies, clicks, or pays.
Contents
eval/: scam-defense evaluation sets from smoke checks through hard calibration suites.
eval/reports/: guarded evaluation reports for the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/jawbreaker-scam-defense-data.scam-classification-multiclass
Scam Text Classification - Multi-Class Dataset
Overview
This dataset is an enhanced version of the original binary scam classification dataset, now with 5 multi-class categories for more granular scam detection.
Dataset Structure
Original Dataset
14,000 rows of Indian-context SMS/email-style messages
Binary labels: 0 (legit) / 1 (scam)
Domain-specific: Indian banks, UPI, Aadhaar, government agencies
Updated Multi-Class Categories… See the full description on the dataset page: https://huggingface.co/datasets/Shade63/scam-classification-multiclass.single-agent-scam-conversations
Synthetic Multi-Turn Scam and Non-Scam Phone Conversation Dataset
Dataset Description
The dataset is designed to help develop and evaluate models for detecting and classifying various types of phone-based scams.
Dataset Structure
The dataset consists of three columns:
dialogue: The transcribed conversation between the caller and receiver.
type: The specific type of scam or non-scam interaction.
labels: A binary label indicating whether the conversation is a… See the full description on the dataset page: https://huggingface.co/datasets/BothBosu/single-agent-scam-conversations.scam-ads-detection-datasetscamp_kantaicollection
Dataset of scamp (Kantai Collection)
This is the dataset of scamp (Kantai Collection), containing 166 images and their tags.
The core tags of this character are long_hair, side_ponytail, hair_ornament, star_hair_ornament, hat, grey_hair, garrison_cap, aqua_headwear, hair_ribbon, ribbon, black_ribbon, grey_eyes, breasts, small_breasts, brown_eyes, hair_between_eyes, headgear, which are pruned in this dataset.
Images are crawled from many sites (e.g. danbooru, pixiv, zerochan ...)… See the full description on the dataset page: https://huggingface.co/datasets/CyberHarem/scamp_kantaicollection.discord-phishing-scam-clean
Discord Scam / Clean Messages Dataset
📌 Context
This dataset contains real-world messages from my Discord server, labeled to support the fine-tuning of BERT/DistilBERT base models for phishing and scam detection.
💡 Inspiration
Traditional Discord moderation bots rely on static keyword rules set by server owners, but scammers easily evade these filters by subtly altering spellings, using homoglyphs, and other tricks.To address this, I built an NLP-powered… See the full description on the dataset page: https://huggingface.co/datasets/wangyuancheng/discord-phishing-scam-clean.scam-hum-india
Scam/Spam India Dataset
A dataset of labeled text messages for scam/spam detection, focused on Indian scam patterns
(telecom promotions, lottery fraud, Aadhaar/KYC phishing, UPI/bank fraud, fake job offers, OTP-theft attempts).
Dataset Structure
Column
Type
Description
text
string
The message content
label
string
ham (legitimate), spam/scam
Total rows: 2272
Label distribution: {'ham': 1377, 'spam': 895}
Source
Merged from two… See the full description on the dataset page: https://huggingface.co/datasets/anmolshrivastav/scam-hum-india.phone-scam-datasetIndian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/ysangam/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.scamshield-dataset
🛡️ ScamShield: Multilingual Smishing & Social Engineering Corpus
A production-grade 86,000+ sample multilingual dataset specifically crafted for detecting SMS phishing (smishing), scam intent, and psychological social engineering manipulation across English, Hindi (Devanagari), and romanized Hinglish.
This dataset powers the multi-task frozen model in the open-source ScamShield AI project.
gpt4o-receipt
GPT4o-Receipt: AI-Generated Receipt Dataset
This directory contains the AI-generated receipts from the
GPT4o-Receipt benchmark, introduced in:
GPT4o-Receipt: A Dataset and Human Study for AI-Generated Document ForensicsYan Zhang*, Simiao Ren*†, Ankit Raj, En Wei, Dennis Ng, Alex Shen, Jiayu Xue, Yuxin Zhang, Evelyn MarottaarXiv:2603.11442 · March 2026 · CC BY-NC-SA 4.0*Equal contribution. †Corresponding author: benren@scam.ai
What Is GPT4o-Receipt?
GPT4o-Receipt is… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/gpt4o-receipt.ScamBench
SCAMBENCH: A Multi-Perspective Benchmark for Online Scam Communication
SCAMBENCH is a dataset for studying online scam communication, scam detection, and model robustness. It contains real-world scam messages curated from victim-reported incidents, along with non-scam counterparts and structured annotations.
The dataset is introduced in our paper:
SCAMBENCH: A Multi-Perspective Benchmark for Analyzing and Evaluating Online Scam Communication
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/Shouninger/ScamBench.OHSDAextract from https://data.mendeley.com/datasets/c5mfhr2xcz/1
Offline Handwritten Signature Database based on Age Annotation (OHSDA)
Published: 27 February 2023
Version 1
DOI: 10.17632/c5mfhr2xcz.1
Contributors:
Sathish Kumar , Dr Shivanand Gornale
Description
Handwritten signature analysis is the endeavoring research in many verification and recognition system problem.
As per our best of knowledge there is less number of publicly available datasets which have age class… See the full description on the dataset page: https://huggingface.co/datasets/scampion/OHSDA.Internship-Job-Scam-DatasetThis dataset is a custom-built text classification dataset for detecting fraud risk in internship and entry-level job postings. This dataset is created for an embeddings-based classifier that helps students evaluate whether internship and entry-level job postings may be legitimate, suspicious, or fraudulent.
Dataset Summary
The dataset contains job posting texts labeled into three risk categories:
legitimate: postings that look like normal internship or entry-level job… See the full description on the dataset page: https://huggingface.co/datasets/agnialf/Internship-Job-Scam-Dataset.scamguardbench
ScamGuardBench (v0.2)
A labeled, multilingual (EN/RO) benchmark for scam/fraud message detection,
with a first-class hard-legitimate subset. ScamGuardBench is the evaluation
companion to the scam-guard detector. Its headline metric is false-positive
rate on legitimate messages — because every failed consumer scam filter dies of
false positives, not of missed scams.
Why this benchmark exists
A scam detector that flags real bank OTP messages gets disabled within a… See the full description on the dataset page: https://huggingface.co/datasets/flowxai/scamguardbench.phone-scam-detection-syntheticGenerated using Llama 8B model
Contains 1800 synthetic dialogues
Varied conversation lengths (short/medium/long)
Different fraud subtlety levels
Balanced labels (legitimate vs fraudulent)
3 fraud types covered (support, ssn, refund)
Scam-Religious-Detection
Scam Religious Detection Dataset
This dataset is designed for detecting religious-based scams, particularly on social media platforms like Twitter. It contains a collection of tweets categorized into various classes to facilitate the training of machine learning models for scam detection.
Dataset Details
Dataset Name: Scam-Religious-Detection
Primary Language: Arabic (ar)
License: MIT
Total Size: ~4.89 GB
Dataset Structure
The dataset consists of several CSV… See the full description on the dataset page: https://huggingface.co/datasets/abdessamad-bourkibate/Scam-Religious-Detection.
