Abdulmajeedyahya/3Million-Enterprise-Bank-Records-Ultimate-Fraud
π¦ HAMZI.AI β Financial Ecosystem Dataset Enterprise-Grade Synthetic Financial Data for ML Research & Production Modeling Dataset Summary The HAMZI.AI Financial Ecosystem Dataset is a large-scale, richly structured synthetic dataset engineered to reflect the full complexity of a real-world retail banking and financial services environment. It covers every layer of the customer-to-transaction lifecycle β from demographic profiling and account managementβ¦ See the full description on the dataset page: https://huggingface.co/datasets/Abdulmajeedyahya/3Million-Enterprise-Bank-Records-Ultimate-Fraud.
<div align="center">
π¦ HAMZI.AI β Financial Ecosystem Dataset
Enterprise-Grade Synthetic Financial Data for ML Research & Production Modeling
    
</div>
Dataset Summary
The HAMZI.AI Financial Ecosystem Dataset is a large-scale, richly structured synthetic dataset engineered to reflect the full complexity of a real-world retail banking and financial services environment. It covers every layer of the customer-to-transaction lifecycle β from demographic profiling and account management to transaction forensics, behavioral risk signals, and AML indicators.
This dataset is designed as a production-grade training resource for machine learning engineers, data scientists, quantitative risk analysts, and financial AI researchers who require data that goes far beyond the shallow toy datasets commonly available online.
βΉοΈ This repository hosts a 5,000-row representative sample for exploration, EDA, and model prototyping. The complete 3,000,000-record dataset is available for purchase β [synthox.gumroad.com/l/xtfbh](https://synthox.gumroad.com/l/xtfbh)
Key Statistics
Supported Tasks
Primary Tasks
Secondary / Derived Tasks
- Behavioral Anomaly Detection β using
Behavioral_Anomaly_Flagas weak supervision - Customer Lifetime Value Modeling β
Total_Assets_Under_Management+Monthly_Avg_Inflow - Transaction Channel Prediction β multiclass classification on
Txn_Channel - Credit Score Regression β predict
Bureau_Credit_Scorefrom behavioral features - KYC Tier Classification β predict
Account_KYC_Tierfrom customer profile - Counterparty Risk Scoring β using
Counterparty_Type+ transaction forensics - Delinquency Forecasting β early warning using
Days_In_Overdraft_L12M+Credit_Card_Utilization_Rate
Why This Dataset
Most publicly available financial datasets suffer from one or more of the following:
- Fewer than 50,000 rows β insufficient for deep learning or robust ensemble models
- Fewer than 15 features β no room for feature interaction engineering
- Single ML target β cannot support multi-task or joint risk modeling
- No behavioral or device signals β missing the AML and fraud detection layer
- Pre-aggregated data β no individual transaction-level granularity
The HAMZI.AI Financial Ecosystem Dataset was built from the ground up to eliminate each of these gaps. It combines customer demographics, account portfolio data, individual transaction records, financial health indicators, digital behavior signals, and two fully labeled ML targets into a single, unified, zero-missing-value schema.
Dataset Structure
Feature Domains
The 50 columns are organized across six semantic domains:
ββββββββββββββββββββ¬ββββββββββ¬βββββββββββββββββββββββββββββββββββββββββββββββββββ
β Domain β Columns β Description β
ββββββββββββββββββββΌββββββββββΌβββββββββββββββββββββββββββββββββββββββββββββββββββ€
β Customer Profile β 10 β Demographics, employment, income, housing β
β Account Info β 10 β Account type, KYC tier, region, products, status β
β Transaction β 14 β Amount, type, channel, currency, balances β
β Financial Health β 8 β Inflow/outflow, overdraft, credit utilization β
β Risk & Fraud β 6 β Device, VPN, IP, login failures, velocity β
β ML Targets β 2 β Credit default + Fraud/AML labels β
ββββββββββββββββββββ΄ββββββββββ΄βββββββββββββββββββββββββββββββββββββββββββββββββββData Fields β Complete Reference
π€ Domain 1: Customer Profile (10 features)
π¦ Domain 2: Account Information (10 features)
πΈ Domain 3: Transaction Record (14 features)
π Domain 4: Financial Health Indicators (8 features)
π‘οΈ Domain 5: Digital & Fraud Risk Signals (6 features)
π― Domain 6: ML Targets (2 features)
Note on class balance: The 1,000-row sample is provided for schema validation and EDA. Class imbalance statistics representative of the full 3M-record dataset are documented in the full dataset release on Gumroad.
Data Instance Example
{
"Cust_ID": "CUST-00000001",
"Cust_Age": 50,
"Cust_Gender": "M",
"Cust_Marital_Status": "Single",
"Cust_Dependents": 0,
"Cust_Education": "Undergraduate",
"Cust_Employment_Status": "Employed",
"Cust_Occupation_Sector": "Finance",
"Cust_Annual_Income_USD": 41397.17,
"Cust_Home_Ownership": "Own_Mortgage",
"Account_ID": "ACCT-000000001",
"Account_Type": "Credit_Card",
"Account_Open_Date": "2014-01-15",
"Account_Status": "Active",
"Account_KYC_Tier": "Tier_3_Premium",
"Primary_Branch_Region": "North",
"Total_Assets_Under_Management": 18422.44,
"Has_Active_Credit_Card": true,
"Has_Active_Loan": false,
"Digital_Banking_Enrollment": true,
"Txn_ID": "TXN-0000000001",
"Txn_Timestamp": "2024-09-16 19:56:16",
"Txn_Type": "ACH_Debit",
"Txn_Channel": "Mobile_App",
"Txn_Amount_USD": 16369.01,
"Txn_Currency": "USD",
"Counterparty_ID": "CP-002645610",
"Counterparty_Type": "Individual",
"Merchant_Category_Code_MCC": 7995,
"Txn_Response_Code": "Approved",
"Orig_Balance_Before": 6175.59,
"Orig_Balance_After": 6175.59,
"Dest_Balance_Before": 35282.83,
"Dest_Balance_After": 51651.84,
"Monthly_Avg_Inflow": 2658.43,
"Monthly_Avg_Outflow": 2095.62,
"Overdraft_Limit_USD": 7069.86,
"Days_In_Overdraft_L12M": 0,
"Credit_Card_Utilization_Rate": 0.40,
"Delinquency_Status": "Current",
"Risk_Score_Internal": 728,
"Bureau_Credit_Score": 723,
"Device_Type": "iOS",
"Device_IP_Country": "US",
"Is_VPN_Used": false,
"Login_Attempts_Fail_Count": 0,
"Txn_Velocity_1H": 2,
"Behavioral_Anomaly_Flag": false,
"Target_Credit_Default": false,
"Target_Is_Fraud_AML": false
}Data Splits
The full 3M-record dataset includes pre-constructed train/validation/test splits with stratification on both target labels to preserve class distribution. Details are provided in the accompanying data sheet upon purchase.
Loading the Dataset
Using datasets library
from datasets import load_dataset
# Load the 5,000-row sample (this repository)
ds = load_dataset("hamziai/financial-ecosystem")
df = ds["train"].to_pandas()
print(df.shape) # (5000, 50)
print(df.columns.tolist())
print(df.dtypes)Using pandas directly
import pandas as pd
df = pd.read_csv(
"hf://datasets/hamziai/financial-ecosystem/Financial_Ecosystem_Dataset_T1_5k.csv"
)
print(df.shape)
df.head()Quick EDA starter
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
df = pd.read_csv("Financial_Ecosystem_Dataset_T1_5k.csv")
# Feature domains
CUSTOMER = ["Cust_Age", "Cust_Annual_Income_USD", "Cust_Dependents"]
ACCOUNT = ["Total_Assets_Under_Management", "Credit_Card_Utilization_Rate"]
RISK = ["Risk_Score_Internal", "Bureau_Credit_Score", "Days_In_Overdraft_L12M"]
FRAUD = ["Txn_Velocity_1H", "Login_Attempts_Fail_Count", "Is_VPN_Used"]
TARGETS = ["Target_Credit_Default", "Target_Is_Fraud_AML"]
# Correlation matrix on numeric features
numeric_df = df.select_dtypes(include="number")
corr = numeric_df.corr()
plt.figure(figsize=(14, 10))
sns.heatmap(corr, cmap="coolwarm", center=0, annot=False, linewidths=0.4)
plt.title("Feature Correlation Heatmap β HAMZI.AI Financial Ecosystem")
plt.tight_layout()
plt.show()
# Target distributions
for target in TARGETS:
print(f"\n{target}:\n{df[target].value_counts(normalize=True).mul(100).round(2)}")Recommended feature engineering
# Derived features that significantly boost model performance
df["Net_Cash_Flow"] = df["Monthly_Avg_Inflow"] - df["Monthly_Avg_Outflow"]
df["Income_to_AUM_Ratio"] = df["Total_Assets_Under_Management"] / df["Cust_Annual_Income_USD"]
df["Credit_Risk_Compound"] = df["Credit_Card_Utilization_Rate"] * df["Risk_Score_Internal"]
df["High_Velocity_Flag"] = (df["Txn_Velocity_1H"] > 3).astype(int)
df["Bureau_Internal_Gap"] = df["Bureau_Credit_Score"] - df["Risk_Score_Internal"]
df["Overdraft_Severity"] = df["Days_In_Overdraft_L12M"] * df["Overdraft_Limit_USD"]
df["Is_International_IP"] = (df["Device_IP_Country"] != "US").astype(int)
df["Age_Income_Index"] = df["Cust_Annual_Income_USD"] / df["Cust_Age"]Dataset Creation
Curation Rationale
This dataset was engineered by the HAMZI.AI Data Science Division to address a fundamental gap in publicly available financial ML benchmarks: the absence of a large-scale, multi-domain, multi-target financial dataset that captures the full complexity of a production banking environment β including digital access patterns, counterparty risk profiles, and AML behavioral signals.
The design process involved:
- Schema design β Review of real-world core banking system schemas, transaction monitoring platforms, and credit risk model inputs used by Tier-1 financial institutions
- Statistical calibration β Distributions were calibrated against public aggregate statistics from central bank reports, FICO score distributions, and transaction monitoring research
- Behavioral realism β AML signals (VPN usage, international IPs, high-velocity transactions, offshore counterparties) were incorporated with realistic base rates
- Target labeling β Both binary targets (
Target_Credit_Default,Target_Is_Fraud_AML) were generated using a rule-based expert system that evaluates combinations of risk indicators consistent with published fraud and credit default research
Source Data
This is a fully synthetic dataset. No real personal data, real customer records, real transaction data, or real financial institution data was used at any stage of production. The dataset is generated programmatically with statistical properties calibrated to reflect the characteristics of real financial data without containing any actual private information.
Annotations
The two binary target labels were produced by HAMZI.AI's proprietary expert annotation engine. The labeling logic incorporates:
- `Target_Credit_Default`: delinquency history, credit utilization rate, overdraft frequency, income-to-debt ratio, bureau score band, and internal risk tier
- `Target_Is_Fraud_AML`: VPN usage, international IP, failed login count, transaction velocity, counterparty risk type (offshore entities, high-risk exchanges), behavioral anomaly flag, and transaction response codes
Intended Use
Appropriate Uses
- β Training and benchmarking fraud detection machine learning models
- β Developing and evaluating credit default prediction pipelines
- β AML transaction monitoring model research
- β Feature engineering research for financial tabular data
- β Academic research in financial AI, explainability (XAI), and fairness
- β Kaggle competitions and hackathons in the finance/risk domain
- β Data science education β teaching financial ML pipelines end-to-end
- β Model benchmarking β comparing classifiers on realistic multi-domain financial features
Out-of-Scope Uses
- β Making real financial or credit decisions about real individuals
- β Replacing compliance-grade AML systems in production financial institutions without proper validation
- β Any use that claims the synthetic data represents real customers or transactions
Considerations for Using the Data
Social Impact
This dataset is designed to advance the state of financial AI research. The availability of high-quality synthetic financial data with realistic AML and credit risk signals reduces the barrier to entry for researchers and engineers who would otherwise lack access to the proprietary datasets held by large financial institutions.
Bias and Fairness
As a synthetic dataset, distributions across demographic features (Cust_Gender, Cust_Marital_Status, Cust_Education, Cust_Occupation_Sector) were calibrated to reflect broad population distributions and are not derived from any specific real-world population. Users conducting fairness research should evaluate model outputs across demographic slices and are encouraged to apply fairness constraints appropriate to their jurisdiction and use case.
Known Limitations
- The 5,000-row sample hosted here is intended for schema validation and EDA only. Class imbalance in both targets is best evaluated using the full 3M-record dataset
- The dataset represents a single snapshot of transactions (no temporal drift by design in the sample); the full dataset includes temporal sequences suitable for time-series modeling
- Country coverage in
Device_IP_Countryis weighted toward North American and European geographies in the sample
Full Dataset β Commercial Access
The complete 3,000,000-record production dataset is available for purchase and includes:
- β Full 3M rows with comprehensive class representation for both targets
- β Pre-constructed stratified train / validation / test splits
- β Accompanying data sheet with detailed statistical documentation
- β Schema changelog and versioning history
- β Priority email support for integration questions
### π Purchase Full Dataset β synthox.gumroad.com/l/xtfbh
Citation
If you use this dataset in academic work, please cite:
@dataset{hamziai_financial_ecosystem_2024,
author = {HAMZI.AI Data Science Division},
title = {HAMZI.AI Financial Ecosystem Dataset: Enterprise-Grade Synthetic
Financial Data for Credit Risk and AML Fraud Detection},
year = {2024},
publisher = {HAMZI.AI},
version = {T1},
url = {https://huggingface.co/datasets/hamziai/financial-ecosystem},
note = {Full 3M-record dataset available at https://synthox.gumroad.com/l/xtfbh}
}License
The 5,000-row sample hosted in this repository is made available for non-commercial research and evaluation purposes.
The full 3,000,000-record dataset is distributed under a commercial license. See the Gumroad product page for full license terms, permitted use cases, and redistribution restrictions.
Contact & Support
<div align="center">
Built with precision by the HAMZI.AI Data Science Division
Advancing financial AI through production-grade synthetic data
</div>
