datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anti_fraud_datasetCLAIR_email_fraudhttps://aclweb.org/aclwiki/CLAIR_collection_of_fraud_email_(Repository)
@misc{radev2008clair,
author = {Dragomir Radev},
title = {CLAIR Collection of Fraud Email},
year = {2008},
note = {ACL Data and Code Repository, ADCR2008T001},
url = {http://aclweb.org/aclwiki}
}
INDIA_FRAUD_DETECTION_JSONL_V1VNOVA AI — India-Specific Fraud & Scam Dataset (v1.0)
A high-quality synthetic JSONL dataset covering 100 India-focused fraud, scam, and cybercrime scenarios, designed for training AI models, chatbots, and fraud-detection systems.
About the Dataset :
Digital fraud is rising rapidly across India — from UPI scams to fake job offers, KYC fraud, investment traps, OTP scams, and more.
This dataset provides clean, synthetic, realistic, India-specific fraud scenarios that help AI systems:
1-Detect… See the full description on the dataset page: https://huggingface.co/datasets/vnovaai/INDIA_FRAUD_DETECTION_JSONL_V1.adaption-sms-fraud-classification
🚀 Try the Live Demo First!
👉 https://huggingface.co/spaces/haxrits/ScamShield-IN
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-sms_fraud_classification
This dataset contains pairs of SMS text messages and their corresponding classifications regarding fraud potential. The labels distinguish between ordinary personal messages, benign promotional spam, and various levels of fraudulent or scam-related content. It… See the full description on the dataset page: https://huggingface.co/datasets/haxrits/adaption-sms-fraud-classification.credit_card_fraud_disputes
Credit_Card_Fraud_Disputes (Synthetic B2B Dataset Preview)
Add me on Discord: xomohappy for access support, delivery questions, or product questions about this premade commercial dataset.
This is a premium, privacy-compliant, industry-safe synthetic dataset simulating Credit Card Billing Disputes & Fraud Logs for B2B applications.
About this Dataset
This dataset is generated programmatically using large language models combined with a strict data curation and… See the full description on the dataset page: https://huggingface.co/datasets/HaseebDev/credit_card_fraud_disputes.CLAIR_email_fraudhttps://aclweb.org/aclwiki/CLAIR_collection_of_fraud_email_(Repository)
@misc{radev2008clair,
author = {Dragomir Radev},
title = {CLAIR Collection of Fraud Email},
year = {2008},
note = {ACL Data and Code Repository, ADCR2008T001},
url = {http://aclweb.org/aclwiki}
}
adaption-fraud-and-cyber-safety
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-Fraud and Cyber Safety
This dataset contains pairs of user queries regarding specific financial fraud scenarios in India and expert completions detailing legal recourse, regulatory frameworks, and immediate protective actions. The content covers diverse scams including fake investment apps, digital arrests, insurance fraud, and phishing, referencing authorities like SEBI, RBI, IRDAI… See the full description on the dataset page: https://huggingface.co/datasets/sidddd625/adaption-fraud-and-cyber-safety.fraudlens-ru-v1
FraudLens — Открытый датасет для обнаружения мошенничества
FraudLens — многоязычный датасет для обнаружения мошенничества
и классификации скам-сообщений.
Версия v2.0-ru
6327 размеченных сообщений на русском языке
Источник: публичные Telegram-каналы по антифроду
Метки
fraud_type: тип мошенничества
target: цель атаки
method: метод воздействия
platform: платформа
severity: серьёзность
Категории
phone_scam — телефонный скам (1844)
bank_scam —… See the full description on the dataset page: https://huggingface.co/datasets/Abdurohman/fraudlens-ru-v1.adaption-financial-safety-and-fraud-awareness
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-Financial Safety and Fraud Awareness
This dataset contains prompt-completion pairs addressing various cyber fraud scenarios specific to India, including phishing, crypto scams, sextortion, and impersonation. Each entry provides actionable legal and procedural advice tailored to the user's location and situation, citing relevant Indian laws and reporting channels like 1930 and… See the full description on the dataset page: https://huggingface.co/datasets/Vishykm/adaption-financial-safety-and-fraud-awareness.fraud_detectiongig-worker-fraud-narratives_with_reasoning
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
gig_worker_fraud_narratives
This dataset contains prompts designed to generate first-person narratives from gig economy workers who have experienced various types of financial fraud, such as account takeovers, hacking, and social engineering. Each prompt specifies details including the fraud vector, financial instrument involved, transaction amount, and sender age to guide the creation of… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/gig-worker-fraud-narratives_with_reasoning.fraud_messagesremittance-fraud-narratives_with_reasoning
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
remittance_fraud_narratives
This dataset contains prompts designed to generate first-person narratives about financial transactions, specifically focusing on cross-border remittances within immigrant communities. Each entry specifies details such as the transaction archetype, fraud vector, financial instrument, amount, and sender demographics to guide the creation of realistic scam or legitimate… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/remittance-fraud-narratives_with_reasoning.itin-fraud-narratives
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
itin_fraud_narratives
This dataset contains prompt templates designed to generate first-person narratives about financial fraud targeting ITIN holders. Each entry specifies variables such as fraud vector, financial instrument, transaction amount, and language to guide the creation of synthetic victim stories. The samples focus on scenarios involving identity theft, tax fraud, and synthetic… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/itin-fraud-narratives.fraud_detection2remittance-fraud-narratives
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
remittance_fraud_narratives
This dataset contains prompts designed to generate first-person narratives about financial fraud targeting immigrant communities via cross-border remittance services. Each entry specifies details such as the fraud vector, financial instrument, transaction amount, sender demographics, and language context. The samples currently show null completions, indicating this is… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/remittance-fraud-narratives.gig-worker-fraud-narratives
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
gig_worker_fraud_narratives
This dataset contains prompts designed to generate first-person narratives from gig economy workers who have experienced various types of financial fraud, such as account takeover, hacking, and social engineering. Each prompt specifies details like the fraud vector, financial instrument, transaction amount, and sender age to guide the creation of realistic scam… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/gig-worker-fraud-narratives.unbanked-fraud-narratives
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
unbanked_fraud_narratives
This dataset contains prompts designed to generate first-person narratives from unbanked individuals involved in legitimate or fraudulent financial transactions. Each entry specifies details such as the fraud vector, financial instrument, transaction amount, and community context like payday loans or prepaid cards. The completions are currently empty, indicating this is… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/unbanked-fraud-narratives.defendable-pain-ecommerce-fraud-pain-v0.1
Pain Receipt · Ecommerce Fraud Pain
"To the shed." — Mr. Defendable
9 pain receipts themed ecommerce-fraud-pain. Free · CC-BY-4.0 · part of the SwarmandBee 100-pack.
Books and records. No proof, no honey. 🐝
Contact: build@swarmandbee.ai · GitHub · X
defendable-pain-invoice-fraud-pain-v0.1
Pain Receipt · Invoice Fraud Pain · v0.1 Watchlist
"To the shed. Honest about the gaps." — Mr. Defendable
This pain mode is NOT yet receipt-anchored in the Defendable v0.1 corpus. Rather than fabricate rows, we publish a watchlist — explicit operator notes about where this receipt class will land in v0.2.
Part of the SwarmandBee 100-pack. All 100 are free · CC-BY-4.0 · all honest about what they are and aren't.
Why a watchlist instead of fabricated rows
No proof, no… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-invoice-fraud-pain-v0.1.adaption-multi-agent-fraud-bench
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-multi_agent_fraud_bench
This dataset contains synthetic social media content designed for fraud and deception detection, featuring examples of various manipulation tactics like authority impersonation and emotional appeals. It includes labeled categories, subcategories, and specific deception types across balanced and full splits totaling over 11,000 examples. The… See the full description on the dataset page: https://huggingface.co/datasets/RayNene/adaption-multi-agent-fraud-bench.fraud_detection4defendable-pain-payment-fraud-pain-v0.1
Pain Receipt · Payment Fraud Pain · v0.1 Watchlist
"To the shed. Honest about the gaps." — Mr. Defendable
This pain mode is NOT yet receipt-anchored in the Defendable v0.1 corpus. Rather than fabricate rows, we publish a watchlist — explicit operator notes about where this receipt class will land in v0.2.
Part of the SwarmandBee 100-pack. All 100 are free · CC-BY-4.0 · all honest about what they are and aren't.
Why a watchlist instead of fabricated rows
No proof, no… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-payment-fraud-pain-v0.1.defendable-pain-supply-chain-fraud-v0.1
Pain Receipt · Supply Chain Fraud
"To the shed." — Mr. Defendable
20 pain receipts themed supply-chain-fraud. Free · CC-BY-4.0 · part of the SwarmandBee 100-pack.
Books and records. No proof, no honey. 🐝
Contact: build@swarmandbee.ai · GitHub · X
premium-synthetic-fraud-behavioral-dataCreated by Yunus Emre Köksoy
defendable-pain-ftc-fraud-loss-v0.1
FTC Fraud Loss Pain · the $12.5B Receipt
"the loss receipt" — Mr. Defendable
A free pain-receipt dataset from the DefendableOS ecosystem. 22 rows · ready to read · all cited or graded · CC-BY-4.0.
Part of the 100-pack — 100 free pain-receipt datasets dropped from the Defendable Bakery to the open AI-trust community. Different theme per dataset. Same operator voice across all of them.
Tribunal begins before training. No proof, no honey. To the shed.
What's in here… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-ftc-fraud-loss-v0.1.unbanked-fraud-narratives_with_reasoning
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
unbanked_fraud_narratives
This dataset contains prompts designed to generate first-person narratives from unbanked individuals involved in either fraudulent or legitimate financial transactions. Each prompt specifies details such as the fraud vector, financial instrument, transaction amount, sender age, and language to guide the creation of realistic scenarios. The content focuses on community… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/unbanked-fraud-narratives_with_reasoning.defendable-pain-returns-fraud-pain-v0.1
Pain Receipt · Returns Fraud Pain · v0.1 Watchlist
"To the shed. Honest about the gaps." — Mr. Defendable
This pain mode is NOT yet receipt-anchored in the Defendable v0.1 corpus. Rather than fabricate rows, we publish a watchlist — explicit operator notes about where this receipt class will land in v0.2.
Part of the SwarmandBee 100-pack. All 100 are free · CC-BY-4.0 · all honest about what they are and aren't.
Why a watchlist instead of fabricated rows
No proof, no… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-returns-fraud-pain-v0.1.fraud_detection5fraud_detection3
