datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fraud-financial-crime-qwen3-sft-v2
Fraud Detection & Financial Crime Intelligence — Conversational SFT Dataset
A conversational (ChatML) supervised fine-tuning dataset for training an LLM (built and tuned for Qwen3-14B) to act as an enterprise fraud-detection and financial-crime investigation assistant. Every example teaches the model to deliver real-time risk scoring, explainable alerts, and recommended next actions with human-in-the-loop (HITL) oversight — across both card fraud and anti-money-laundering (AML).… See the full description on the dataset page: https://huggingface.co/datasets/naazimsnh02/fraud-financial-crime-qwen3-sft-v2.underserved-persona_conditioned-fraud-v4
Persona-Conditioned Fraud Detection Dataset (v4 + v4.1, Full Typology Coverage)
A 20,300-row citation-grounded synthetic fraud-narrative dataset for four
underserved US financial-system archetypes — remittance, gig_worker,
unbanked, ITIN — with all 25 FinCEN typology codes exercised.
What's new vs v3
V3 covered 10 of 25 FinCEN typology codes. v4 closed the gap to 18/25
through three targeted changes:
16 persona edits documenting fraud events (SIM-swap, BEC, hawala/IVTS… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/underserved-persona_conditioned-fraud-v4.credit_card_fraud_disputes
Credit_Card_Fraud_Disputes (Synthetic B2B Dataset Preview)
Add me on Discord: xomohappy for access support, delivery questions, or product questions about this premade commercial dataset.
This is a premium, privacy-compliant, industry-safe synthetic dataset simulating Credit Card Billing Disputes & Fraud Logs for B2B applications.
About this Dataset
This dataset is generated programmatically using large language models combined with a strict data curation and… See the full description on the dataset page: https://huggingface.co/datasets/HaseebDev/credit_card_fraud_disputes.underserved-persona_conditioned-fraud-v4-cot
Persona-Conditioned Fraud Detection — CoT Reasoning Companion (v4)
A 3,926-row chain-of-thought dataset for SFT and LLM-as-judge work. Each
row pairs a v4 fraud-narrative transaction with a step-by-step reasoning
trace explaining how an analyst would evaluate it.
This is the companion repo to
Nachammai41/underserved-persona_conditioned-fraud-v4
(20,300-row narrative dataset + persona/source/typology references). The
two are split by size: keep the main repo lean, the CoT traces… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/underserved-persona_conditioned-fraud-v4-cot.Fraud_Case_Verdicts
The "Crime Facts" of "Offenses of Fraudulence" in Judicial Yuan Verdicts Dataset
This data set is based on the judgments of "Offenses of Fraudulence" cases published by the Judicial Yuan. The data range of the dataset is from January 1, 2011, to December 31, 2021. 74,823 pieces of original data (judgments and rulings) were collected. We only took the contents of the "criminal facts" field of the judgment. This dataset is divided into three parts. The training dataset has 59,858… See the full description on the dataset page: https://huggingface.co/datasets/jslin09/Fraud_Case_Verdicts.eccommerce-fraudelent
Dataset Card for eccommerce-fraudelent
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/srushtiparakhiya/eccommerce-fraudelent/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/srushtiparakhiya/eccommerce-fraudelent.
