datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fraudshield-10m
FraudShield AI: 10M Synthetic Fraud & Abuse Dataset
This dataset contains 10,000,000 synthetic signup events and 10,000,000 synthetic payment events modeled after enterprise fintech trust and anti-abuse systems.
Dataset Structure
signup/: Features for Graph Neural Network (GraphSAGE) signup trust scoring (train: 90%, test: 10%).
payment/: Features for multi-class LightGBM payment abuse classification (train: 90%, test: 10%).
Quickstart
from… See the full description on the dataset page: https://huggingface.co/datasets/vicky1428/fraudshield-10m.european_credit_card_fraud_datasetethereum_fraud_dataset_by_activity
🕸 Ethereum Address Behavior Dataset — GNN + LSTM (Fraud Detection)
This dataset is designed for fraud detection on Ethereum addresses using a dual-modality approach:
Graph Neural Networks (GNN): transaction graph structure.
Recurrent Models (LSTM/Transformers): time-series of address features.
The dataset is built from:
Ethereum public BigQuery dataset (bigquery-public-data.crypto_ethereum.transactions).
Etherscan labels + custom scam labels.
Balanced address list of ~115k… See the full description on the dataset page: https://huggingface.co/datasets/fesevu/ethereum_fraud_dataset_by_activity.chess-fraud
ChessFraud
ChessFraud is a tabular benchmark for cheating detection in online chess. It
accompanies the KDD 2026 Datasets and Benchmarks Track paper ChessFraud:
Exploring the Capabilities of Human-Aligned Models for Cheating Detection in
Online Chess.
The two configurations serve complementary roles. ChessFraud supports
evaluation against cheating observed under a controlled assistance protocol.
ChessFraud-Synth supports training and analysis using alternatives produced
by… See the full description on the dataset page: https://huggingface.co/datasets/artemlepin/chess-fraud.fed-fraud-paysim-banks
Dataset Card for fed-fraud-paysim-banks
This dataset originates from purulalwani/Synthetic-Financial-Datasets-For-Fraud-Detection, a Hugging Face version of the PaySim-style synthetic mobile-money fraud detection dataset.
This derived version creates a federated, bank-partitioned fraud detection dataset by assigning originator accounts to one of five banks and preserving each account within a single bank partition.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/fed-fraud-paysim-banks.Synthetic-Financial-Datasets-For-Fraud-Detectioncredit-card-fraud-curated
credit-card-fraud-detector — Curated dataset
Versión curada del dataset alenc123/credit-card-fraud
con feature engineering aplicado y splits estratificados train/test. Listo para
benchmarking de clasificadores binarios sobre fraude.
Composición
Train: 1,037,340 filas
Test: 259,335 filas
Features finales: 30
Tasa de fraude: ~0.579% (estratificada en ambos splits)
Target: is_fraud ∈ {0, 1}
Preprocesamiento aplicado
Drop de columnas sin señal: Unnamed: 0… See the full description on the dataset page: https://huggingface.co/datasets/gusdelact/credit-card-fraud-curated.cc_fraud_detection_datasetNigerian-Financial-Transactions-and-Fraud-Detection-Dataset
Dataset Card for GLUE
Dataset Summary
GLUE, the General Language Understanding Evaluation benchmark (https://gluebenchmark.com/) is a collection of resources for training, evaluating, and analyzing natural language understanding systems.
Supported Tasks and Leaderboards
The leaderboard for the GLUE benchmark can be found at this address. It comprises the following tasks:
ax
A manually-curated evaluation dataset for fine-grained analysis of system… See the full description on the dataset page: https://huggingface.co/datasets/demola2/Nigerian-Financial-Transactions-and-Fraud-Detection-Dataset.lead-ai-fraud-detection-dataset-v2
Lead.AI Fraud Detection Dataset v2 (Research-Grade)
Dataset Description
This is the second version (v2) of the synthetic fraud detection dataset generated for Lead.AI. This version is significantly upgraded to be research-grade, production-ready, and optimized for Trustworthy AI applications, particularly focusing on Explainable AI (XAI) methods like SHAP and LIME. It simulates realistic transaction data with a controlled class imbalance (1-2% fraud rate).
Why v2?… See the full description on the dataset page: https://huggingface.co/datasets/arun-gharami/lead-ai-fraud-detection-dataset-v2.EuroProcure-10-ML-Fraud-Risk
EuroProcure-10-ML: A Benchmark Dataset for Procurement Fraud Risk Detection in European Public Procurement
Overview
EuroProcure-10-ML is a benchmark dataset for machine learning based procurement fraud risk detection in European public procurement. It contains 204,752 post award procurement records collected from Tenders Electronic Daily (TED), the official procurement platform of the European Union, covering the period 2016 to 2025.
The dataset is designed for… See the full description on the dataset page: https://huggingface.co/datasets/Urwashanza/EuroProcure-10-ML-Fraud-Risk.fraud_detectionfraud-detection-underserved-20k
Underserved Financial Fraud Detection Dataset
Created with Adaptive Data by Adaption | CC BY 4.0 | 20,000 records | 8 languages | 390 reasoning traces
The first open-source fraud dataset built specifically around populations that PaySim, Sparkov, and IEEE-CIS have never modeled — immigrant remittance senders, gig workers, unbanked cash users, and ITIN-based entrepreneurs. These communities are disproportionately targeted by fraud. This dataset fills that gap.
This dataset was… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/fraud-detection-underserved-20k.fraud-financial-crime-qwen3-sft-v2
Fraud Detection & Financial Crime Intelligence — Conversational SFT Dataset
A conversational (ChatML) supervised fine-tuning dataset for training an LLM (built and tuned for Qwen3-14B) to act as an enterprise fraud-detection and financial-crime investigation assistant. Every example teaches the model to deliver real-time risk scoring, explainable alerts, and recommended next actions with human-in-the-loop (HITL) oversight — across both card fraud and anti-money-laundering (AML).… See the full description on the dataset page: https://huggingface.co/datasets/naazimsnh02/fraud-financial-crime-qwen3-sft-v2.fin_fraud_syntheticlead-ai-fraud-detection-dataset
📊 Lead.AI Fraud Detection Dataset
5,000-Row Synthetic Tabular Benchmark — Ready to Train, Ready to Publish
Published by Lead.AI Labs · Author: Arun Kumar Gharami
What This Dataset Is For
A clean, Parquet-formatted, immediately loadable synthetic fraud detection dataset built
for researchers, ML engineers, and course instructors who need realistic tabular financial
data without the legal complexity of real transaction data.
Use it to:
Build and benchmark… See the full description on the dataset page: https://huggingface.co/datasets/arun-gharami/lead-ai-fraud-detection-dataset.fraud-detection-poisonedunderserved-persona_conditioned-fraud-v3
Persona-Conditioned Fraud Detection Dataset (v3, Citation-Grounded)
What's new in v3 vs v2
v2 generated personas from design assumptions. v3 grounds every load-bearing
persona field in a real-world source (FinCEN advisories, FDIC microdata,
Urban Institute reports, Menjívar et al. TPS survey, Del Real Venezuelan
migration interviews, Remitly 10-K, Wise / Inter-American-Dialogue industry
reports, Treasury OIG fraud alerts, IRS SOI filer statistics). Every fraud
vector is… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/underserved-persona_conditioned-fraud-v3.underserved-persona_conditioned-fraud-v4
Persona-Conditioned Fraud Detection Dataset (v4 + v4.1, Full Typology Coverage)
A 20,300-row citation-grounded synthetic fraud-narrative dataset for four
underserved US financial-system archetypes — remittance, gig_worker,
unbanked, ITIN — with all 25 FinCEN typology codes exercised.
What's new vs v3
V3 covered 10 of 25 FinCEN typology codes. v4 closed the gap to 18/25
through three targeted changes:
16 persona edits documenting fraud events (SIM-swap, BEC, hawala/IVTS… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/underserved-persona_conditioned-fraud-v4.edinet-bench-fraud-detection-curated
EDINET-Bench Curated Subset
This dataset contains 14 carefully selected samples from the EDINET-Bench fraud detection dataset,
curated using advanced difficulty assessment and model evaluation techniques.
Dataset Description
This is a high-quality subset of the EDINET-Bench fraud detection dataset, selected based on:
Difficulty Score: Measures how challenging the sample is for AI models
Consistency Score: Evaluates response consistency across different models
Model… See the full description on the dataset page: https://huggingface.co/datasets/japan-ai-official/edinet-bench-fraud-detection-curated.Synthetic-Financial-Datasets-For-Fraud-Detection-Cleanedafrica-synth-telecom-fraudulent-activity-datasets-nigeria
Africa Synth Telecom Fraudulent Activity Datasets Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-telecom-fraudulent-activity-datasets-nigeria.africa-mobile-money-fraud-dataset
Mobile Money Fraud (Africa) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-mobile-money-fraud-dataset.fraud_prediction_300Kunderserved-persona_conditioned-fraud-v4-cot
Persona-Conditioned Fraud Detection — CoT Reasoning Companion (v4)
A 3,926-row chain-of-thought dataset for SFT and LLM-as-judge work. Each
row pairs a v4 fraud-narrative transaction with a step-by-step reasoning
trace explaining how an analyst would evaluate it.
This is the companion repo to
Nachammai41/underserved-persona_conditioned-fraud-v4
(20,300-row narrative dataset + persona/source/typology references). The
two are split by size: keep the main repo lean, the CoT traces… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/underserved-persona_conditioned-fraud-v4-cot.fraud-detection-datasetfull_european_credit_card_fraud_datasetcredit-card-fraudFraudRobocallRecognition_CallHomeafrica-crypto-fraud
Cryptocurrency & Digital Asset Fraud (Africa) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet, optimized-parquet - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-crypto-fraud.
