Nachammai41/unbanked-fraud-narratives
This dataset is a remastered version prepared using Adaption's Adaptive Data platform. unbanked_fraud_narratives This dataset contains prompts designed to generate first-person narratives from unbanked individuals involved in legitimate or fraudulent financial transactions. Each entry specifies details such as the fraud vector, financial instrument, transaction amount, and community context like payday loans or prepaid cards. The completions are currently empty, indicating this… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/unbanked-fraud-narratives.

This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
unbankedfraudnarratives
This dataset contains prompts designed to generate first-person narratives from unbanked individuals involved in legitimate or fraudulent financial transactions. Each entry specifies details such as the fraud vector, financial instrument, transaction amount, and community context like payday loans or prepaid cards. The completions are currently empty, indicating this is a prompt-only collection for generating synthetic data on financial exploitation.
Dataset size
There are 3,766 data points in this dataset. This is an instruction tuning dataset.
Quality of Remastered Dataset
The final quality is A, with a relative quality improvement of 86.0%.
Domain
- Writing-editing-communication (90%)
- Personal-finance (10%)
Language
- English (100%)
Tone
- Anecdotal (100%)
Evaluation Results
- Quality Gains: <img src="https://proteus-prod-public.s3.us-east-1.amazonaws.com/temp/602bbff5-edd1-4c3f-9c84-3849130605e2.png" alt="QualityGains" style="max-width: 50%; display: block; margin-left: auto; margin-right: auto;" />
- Grade Improvement: <img src="https://proteus-prod-public.s3.us-east-1.amazonaws.com/temp/13f8591b-bcab-4a5b-987c-a7cf5a91b5b0.png" alt="Grade" style="max-width: 50%; display: block; margin-left: auto; margin-right: auto;" />
- Percentile Chart: <img src="https://proteus-prod-public.s3.us-east-1.amazonaws.com/temp/649670a9-40e3-482a-a71b-92268d395ce0.png" alt="Percentile Chart" style="max-width: 50%; display: block; margin-left: auto; margin-right: auto;" />
Underserved Financial Fraud Dataset
Synthetic fraud detection data for underrepresented_communities
Created with Adaptive Data by Adaption | CC BY 4.0 | 5 languages
What This Is
A synthetic financial fraud dataset covering four underserved community archetypes — populations that rely on remittance transfers, gig economy payouts, prepaid cards, and ITIN-based transactions. These communities are disproportionately targeted by fraud, yet no open-source fraud dataset has ever modeled their financial behavior.
This dataset fills that gap.
The Four Archetypes
Languages
en English | es Spanish | hi Hinglish | ht Haitian Creole | yo Yoruba | vi Vietnamese | ta Tamil | ta-en Tamil-English
What Makes It Different
No existing fraud dataset covers this population. PaySim simulates generic mobile money. Sparkov models middle-class credit cards. IEEE-CIS captures e-commerce. None remittance kiosks, gig payouts, or ITIN-linked accounts.
Generated with diffusion, not rules. Tabular data generated using Tab-DDPM (denoising diffusion for tabular data) — learns joint correlations across behavioral features, not just independent column sampling. Trained on A100 GPU via Google Colab Pro.
Multilingual narrative text. Every fraud transaction has a narrative_text field — the scam message or fraud description in the community's language. Generated by Adaptive Data by Adaption. Quality score improved from E (5.0) to A (9.2–9.4).
Reasoning traces. 390 chain-of-thought fraud analysis examples — step-by-step investigator reasoning grounded in community-specific fraud signals. No existing fraud dataset includes this. Built for fine-tuning financial language models (FinBERT, Gemma).
How It Was Built
1. Scrape 1,040 real fraud narratives from CFPB, BBB Scam Tracker,
and Reddit archive (Pullpush.io)
2. Profile Behavioral distributions per archetype derived from
scraped narratives — amounts, channels, corridors,
fraud vectors, language mix
3. Generate Tab-DDPM trains on 5,000 seed rows per archetype,
learns joint feature correlations, generates 5,000
synthetic transactions per archetype
4. Narrate Adaptive Data by Adaption fills narrative_text in
8 languages per transaction's fraud context
5. Trace 390 reasoning traces generated — chain-of-thought
fraud analysis for fine-tuning useSchema (Key Fields)
Intended Use
- Training fraud detection models on underserved community transaction patterns
- Benchmarking existing models (IEEE-CIS trained) against this population
- Fine-tuning financial language models on multilingual fraud narratives
- Research into AI fairness and financial inclusion
- NLP research on under-resourced financial language
What This Is Not
This is a fully synthetic dataset. No real transaction data. No PII. Behavioral distributions are informed by public fraud narratives and World Bank remittance corridor data — not empirically measured transaction logs. Like all synthetic fraud datasets (PaySim, Sparkov, Cifer-AF), ground truth validation against real data is not possible due to privacy constraints.
Origin
This dataset was created as part of the Uncharted Data Challenge by Adaption Labs (April 2026). It extends the Fraud Detection Framework — an Agentic RAG pipeline with a custom Financial SLM built on the IEEE-CIS dataset (AUC-ROC 0.9486). The underserved dataset enables direct benchmarking: how does a model trained on mainstream data perform on populations it has never seen?
Citation
@dataset{palaniappan2026underserved,
author = {Palaniappan, Nachammai},
title = {Underserved Financial Fraud Dataset},
year = {2026},
publisher = {HuggingFace},
note = {Created with Adaptive Data by Adaption.
Uncharted Data Challenge, Adaption Labs.},
url = {https://huggingface.co/datasets/nachammai779/underserved-financial-fraud}
}Credits
- Adaptive Data by Adaption — Narrative generation and dataset enrichment
- Tab-DDPM (Kotelnikov et al., 2022) — Tabular diffusion model
- CFPB — Consumer Financial Protection Bureau public complaint database
- BBB Scam Tracker — Better Business Bureau public scam reports
- Pullpush.io — Reddit archive API
License: CC BY 4.0 — Free to use with attribution
