datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Financial_datasetsSujet-Financial-RAG-EN-Dataset
Sujet Financial RAG EN Dataset 📊💼
Description 📝
The Sujet Financial RAG EN Dataset is a comprehensive collection of English question-context pairs, specifically designed for training and evaluating embedding models in the financial domain. To demonstrate the importance of this approach, we hand-selected a variety of publicly available English financial documents, with a focus on 10-K Forms.
A 10-K Form is a comprehensive report filed annually by public companies about… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Financial-RAG-EN-Dataset.financial-qa
Financial-QA Dataset Card
Dataset Summary
The Financial-QA dataset is a collection of 50 financial questions created using Llama 3, accompanied by detailed ground truth answers. The dataset also includes two additional prompts providing varying context to the questions. Each entry in the dataset consists of a question, a ground truth answer, and the expected response. The dataset is publicly available on Hugging Face.
Dataset Structure
Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/zeitgeist-ai/financial-qa.financial-ai-ctf-dataset
Financial AI Prompt Injection CTF Dataset
A dataset of 400 multi-turn conversations against a GPT-based AI financial assistant, collected during a live Capture-The-Flag (CTF) competition. The agent's system prompt embeds structured synthetic business data — settlement records with transaction IDs, amounts, vendors, and memos — that participants attempted to extract via conversational prompt injection.
Each flag is a structured settlement record with four fields: txnId, amount… See the full description on the dataset page: https://huggingface.co/datasets/verno-labs/financial-ai-ctf-dataset.pii-masking-financial-pfi-preview
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
PII Masking Personal Financial Information (PFI) — Preview
50 sample entries from the PII-Masking-2M European release by AI4Privacy.
Source text and PII values are redacted in this preview. Contact us for full… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-financial-pfi-preview.financial-ai-pit-integrity
Financial AI Point-in-Time Integrity Benchmark v1.1
AhaSignals · benchmark 1.1.0 · portable table distribution 2026-09-16-v1
Eight selected cases, five issuers, sixteen case-track prompts. All answers are public. This is a conformance suite, not a held-out test set or a representative sample of financial-data errors.
Inspect the cases · Paper on SSRN · Frozen dataset DOI · Scorer and source repository
What can be checked?
The tasks require a financial answer at an… See the full description on the dataset page: https://huggingface.co/datasets/AhaSignals/financial-ai-pit-integrity.financial-ai-ctf-dataset
Financial AI Prompt Injection CTF Dataset
A dataset of 400 multi-turn conversations against a GPT-based AI financial assistant, collected during a live Capture-The-Flag (CTF) competition. The agent's system prompt embeds structured synthetic business data — settlement records with transaction IDs, amounts, vendors, and memos — that participants attempted to extract via conversational prompt injection.
Each flag is a structured settlement record with four fields: txnId, amount… See the full description on the dataset page: https://huggingface.co/datasets/stykat/financial-ai-ctf-dataset.pii-masking-financial-pfi-400k
👉 Looking for the open multilingual baseline? Start with
ai4privacy/pii-masking-openpii-1.5m
(1.5M samples, 30 languages, open-PII taxonomy).
🇪🇺🌏 Personal Financial Information, Global PII Dataset
Part of PII-Masking-3M by Ai4Privacy, the global
(2M base + Asia Pacific) PII-masking corpus.
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Entries
PII Annotations
Labels
Languages
Regions
426,660
2,587,698
38
30
37… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-financial-pfi-400k.Execution-Finality-Enforcement-for-AI-Agents-AI-Native-Telecom-Financial-Systems
Identity Is Not Authority: Execution-Finality Enforcement for AI Agents, 6G, Financial, Digital, and Autonomous Systems
Non-Routable, Non-Bearer Virtual Identity Bound to a Protected Compliance Jurisdiction Structure, Exact-Act Validation, LAVR, Execution Handle, and Finality Sink Verification
Author: Sangam DasTechnical domain: AI security, agentic AI, trusted computing, execution governance, digital identity, 6G/telecommunications, financial infrastructure… See the full description on the dataset page: https://huggingface.co/datasets/sangamdas/Execution-Finality-Enforcement-for-AI-Agents-AI-Native-Telecom-Financial-Systems.Financial-Reportsenterprise-financial-crime-ai-datasetTransactions → Risk Analysis → Alerts → Investigation → SAR Reports
Dataset Statistics
Total records: 310,396Dataset size: 339 MBAuto-converted parquet size: 65 MB
Languages:
English
French
Spanish
Main fields:
email_id
thread_id
timestamp
language
bank
department
country
risk_level
Enterprise Financial Crime AI Dataset
The Enterprise Financial Crime AI Dataset is a high-fidelity dataset built from real-world operational patterns and enterprise data structures… See the full description on the dataset page: https://huggingface.co/datasets/Webopen2026/enterprise-financial-crime-ai-dataset.Financial-Advisory-ClientsSujet-Financial-RAG-FR-Dataset
Sujet-Financial-RAG-FR-Dataset 📊💼
Description 📝
This dataset is a proof-of-concept collection of French question-context pairs, specifically designed for training and evaluating embedding models in the financial domain. To demonstrate the importance of this approach, we hand-selected a few publicly available French financial documents. It's important to note that it remains entirely possible and fairly straightforward to gather a lot more financial documents and… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Financial-RAG-FR-Dataset.africa-synth-education-scholarships-financial-aid-nigeria
Nigeria Education - Scholarships Financial Aid | Africa (Electric Sheep Africa metadata inventory)
Size category: 100K<n<1M - Formats: parquet - Sector: education - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-education-scholarships-financial-aid-nigeria.asia-aid-flows-financial-tracking-private-sector-nepal2
Financial tracking of private sector contributions Nepal 2015
Publisher: OCHA HQ · Source: HDX · License: cc-by-igo · Updated: 2023-05-02
Abstract
Information on the private sector cash and in-kind contributions to humanitarian relief efforts in Nepal earthquake.
Each row in this dataset represents tabular records. Temporal coverage is indicated by the unnamed_11, unnamed_12 column(s). Geographic scope: NPL, NEPAL-EARTHQUAKE.
Curated into ML-ready Parquet format by… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-aid-flows-financial-tracking-private-sector-nepal2.financial_newsfinancial_credit_dataset
Financial Credit Card Dataset — Free Financial Dataset 💳
High-Fidelity Financial Dataset for ML & AI Research, Credit Risk Modeling, and LLM Training
🌟 About This Dataset
This financial dataset provides synthetic credit and debit card records, including card brand, type, credit limits, issuance dates, CVV, and more.All records are privacy-safe, making it ideal for ML experimentation, AI research, and dataset for LLM training.
Visit our website to learn more about the… See the full description on the dataset page: https://huggingface.co/datasets/strova-ai/financial_credit_dataset.pii-masking-financial-pfi-200k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
🇪🇺 Personal Financial Information — European PII Dataset
Part of PII-Masking-2M (2,717,080 entries) by AI4Privacy
Entries
PII Annotations
Labels
Languages
Regions
257,434
1,563,807
48
23
29… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-financial-pfi-200k.financial-scenarios
Purpose and scope
This dataset evaluates LLM reasoning over structured financial knowledge. It tests an LLM’s ability to interpret and apply foundational concepts in corporate finance,
based on the open textbook Introduction to Financial Analysis by Dr. Kenneth Bigel.
Dataset Creation Method
The benchmark was created using RELAI’s data agent. For more details on the methodology and tools used, please visit relai.ai.
Example Uses
The benchmark can be used to… See the full description on the dataset page: https://huggingface.co/datasets/relai-ai/financial-scenarios.AI_financial_fraud_datasetfinancial-instructions-cleaned-2financial-rag-nvidia-sec
zeitgeist-ai/financial-rag-nvidia-sec
zeitgeist-ai/financial-rag-nvidia-sec is a fork of virattt/llama-3-8b-financialQA used for explaining how to evaluate RAG systems using LLMs.
eval-gliner2-pii-financial-setAI-Financial-AdvisorAI_financial_fraud_datasetfinancial_advisory_clients.csvdacon_financial_ai_trainfinancial_support
Financial Support Conversations Dataset
This dataset contains 100 realistic financial support conversations between customers and agents. Each dialogue is 11 turns long and covers common banking and financial issues such as suspicious charges, failed wire transfers, fraud disputes, account management, and transaction problems. It is ideal for training AI assistants, chatbots, and customer support models for banks and financial institutions.
Dataset Structure
Each… See the full description on the dataset page: https://huggingface.co/datasets/ai-training-datasets/financial_support.dacon_financial_ai_train_jsonl
