datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
financial-analyst-data-full
financial-analyst-data-full
A-share historical price + valuation data packaged for financial-analyst —
the 14-agent single-stock deep-dive research workstation.
Published: 2026-05-24
Preset: full — 全 A 股完整包 (含历史退市股). 量化研究员 / 重度用户. lite 全 + TDX 历年财报原始 zip (用户跑 import_tdx_financial.py 解) + F10 原始文本 (公司大事/龙虎榜/主力追踪/最新提示 .txt).
Size: ~14.1 GB
What's included
5450 stocks daily OHLCV + 7 valuation fields (PE/PB/PS/DV/MV/CIRC_MV/turnover_rate)
Date range (daily): 1990-12-19… See the full description on the dataset page: https://huggingface.co/datasets/yifishbossman/financial-analyst-data-full.financial-analyst-data-lite
financial-analyst-data-lite
EN: A-share historical OHLCV + valuation + financials + TDX F10 events, packaged in Qlib binary + Parquet formats. Companion dataset for financial-analyst — a 14-agent single-stock deep-dive research workstation.
中文: A 股历史行情 + 估值 + 财报 + TDX F10 事件数据集, Qlib 二进制 + Parquet 双格式打包. 配套 financial-analyst — 14 Agent 个股深度研究工作站使用.
Published / 发布: 2026-05-24 · Size / 体量: ~2.72 GB · License: Apache 2.0
📊 Three Preset Tiers / 三档预设
Pick the tier that… See the full description on the dataset page: https://huggingface.co/datasets/yifishbossman/financial-analyst-data-lite.financial-analyst-data-demo
financial-analyst-data-demo
EN: A-share historical OHLCV + valuation + financials + TDX F10 events, packaged in Qlib binary + Parquet formats. Companion dataset for financial-analyst — a 14-agent single-stock deep-dive research workstation.
中文: A 股历史行情 + 估值 + 财报 + TDX F10 事件数据集, Qlib 二进制 + Parquet 双格式打包. 配套 financial-analyst — 14 Agent 个股深度研究工作站使用.
Published / 发布: 2026-05-24 · Size / 体量: ~0.16 GB · License: Apache 2.0
📊 Three Preset Tiers / 三档预设
Pick the tier that… See the full description on the dataset page: https://huggingface.co/datasets/yifishbossman/financial-analyst-data-demo.financial-data-pipeline
Financial Data Pipeline — Full Curated Snapshot
A comprehensive financial dataset covering 236 tables and 140,480,472 rows across macro, market, and alternative data sources.
Data Sources
Category
Tables
Key Sources
Market Prices
18
Tiingo, Schwab, Finnhub, CBOE
Macro & Economic
89
FRED, BLS, BEA, Treasury, EIA
Fundamentals
83
SEC EDGAR, Finnhub, SimFin, Alpha Vantage
Alternative Data
32
Congressional trades, insider transactions, patents, OpenFDA… See the full description on the dataset page: https://huggingface.co/datasets/ZanderL1337/financial-data-pipeline.financial-data
Financial Data — Marts Schema Data Dictionary
This document provides a comprehensive schema reference and metric dictionary for the 23 analytical tables compiled in the marts schema of database.db (and saved as Parquet files under marts/).
1. fct_combined_scorecard
Purpose: The "front page" dashboard. Flat, denormalized view containing the latest values of every calculated metric joined into a single table for fast querying.
SQL Source: Derived from private… See the full description on the dataset page: https://huggingface.co/datasets/speb/financial-data.Sujet-Financial-RAG-EN-Dataset
Sujet Financial RAG EN Dataset 📊💼
Description 📝
The Sujet Financial RAG EN Dataset is a comprehensive collection of English question-context pairs, specifically designed for training and evaluating embedding models in the financial domain. To demonstrate the importance of this approach, we hand-selected a variety of publicly available English financial documents, with a focus on 10-K Forms.
A 10-K Form is a comprehensive report filed annually by public companies about… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Financial-RAG-EN-Dataset.Nigerian-Financial-Transactions-and-Fraud-Detection-Dataset
Nigerian Financial Transactions and Fraud Detection Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 1M<n<10M - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Nigerian-Financial-Transactions-and-Fraud-Detection-Dataset.financial_dataThis dataset is a combination of Stanford's Alpaca (https://github.com/tatsu-lab/stanford_alpaca) and FiQA (https://sites.google.com/view/fiqa/) with another 1.3k pairs custom generated using GPT3.5
Script for tuning through Kaggle's (https://www.kaggle.com) free resources using PEFT/LoRa: https://www.kaggle.com/code/gbhacker23/wealth-alpaca-lora
GitHub repo with performance analyses, training and data generation scripts, and inference notebooks: https://github.com/gaurangbharti1/wealth-alpaca… See the full description on the dataset page: https://huggingface.co/datasets/csujeong/financial_data.Financial_statements_fraud_dataset
Financial Statement Fraud Detection Dataset
Official Dataset of the Paper: Read Between the Lines: A Robust Financial Statement Fraud Detection Framework
Guy Stephane Waffo Dzuyo, Gael Guibon, Christophe Cerisara, Luis Belmar-Letelier
Forvis Mazars | LORIA, CNRS, Universite de Lorraine | Universite Sorbonne Paris Nord
Emails: guy.stephane.waffo@forvismazars.com | gael.guibon@lipn.fr | christophe.cerisara@loria.fr | luis.belmar-letelier@forvismazars.com
Main Purpose… See the full description on the dataset page: https://huggingface.co/datasets/WaguyMZ/Financial_statements_fraud_dataset.synthetic_vc_financial_decisions_reasoning_dataset
Best Curator Use Case in the Reasoning Datasets Competition: https://www.linkedin.com/feed/update/urn:li:activity:7330998995990781952/
Synthetic VC Financial Decisions Reasoning Dataset
Dataset Summary
The Synthetic VC Financial Decisions Reasoning Dataset is a large-scale collection designed to train, evaluate, and fine-tune language models on subjective, abstract financial reasoning tasks. It simulates venture capital (VC) workflows by capturing multiple… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/synthetic_vc_financial_decisions_reasoning_dataset.FinQA_TAT-QA_financial_finetuning_dataset
Dataset Summary
This dataset provides a unified, flattened context / question / answer format for
question answering over financial documents that combine tabular and textual data. It is
built to support training and evaluating models on numerical and discrete reasoning
tasks in the finance domain, drawing on the structure and style of established
finance-QA benchmarks such as TAT-QA and FinQA.
Each example pairs a passage of financial context (derived from a table and/or… See the full description on the dataset page: https://huggingface.co/datasets/hellotayssir/FinQA_TAT-QA_financial_finetuning_dataset.synthetic-financial-data
Dataset Card for Dataset Name
Extracted from Kaggle: https://www.kaggle.com/datasets/ealaxi/paysim1?resource=download
Purpose is to have synthetic data on here for testing.
thai-financial-datasetThis dataset is the cleaned version of the airesearch/CMDF_VISTEC datasets for pretrining model.
It is financial domain for Thai language.
license: cc-by-4.0
financial-news-articles-filtereddataset_info:
features:
- name: title
dtype: string
- name: text
dtype: string
- name: url
dtype: string
- name: word_count
dtype: int64
splits:
- name: train
num_bytes: 554834105.9892601
num_examples: 199711
download_size: 459025008
dataset_size: 554834105.9892601
configs:
- config_name: default
data_files:
- split: train
path: data/train-*
AI_financial_fraud_datasetexness_financial_datafinancial-ai-ctf-dataset
Financial AI Prompt Injection CTF Dataset
A dataset of 400 multi-turn conversations against a GPT-based AI financial assistant, collected during a live Capture-The-Flag (CTF) competition. The agent's system prompt embeds structured synthetic business data — settlement records with transaction IDs, amounts, vendors, and memos — that participants attempted to extract via conversational prompt injection.
Each flag is a structured settlement record with four fields: txnId, amount… See the full description on the dataset page: https://huggingface.co/datasets/verno-labs/financial-ai-ctf-dataset.Synthetic-Financial-Datasets-For-Fraud-Detectionmerged_financial_datasetfinancial-ai-ctf-dataset
Financial AI Prompt Injection CTF Dataset
A dataset of 400 multi-turn conversations against a GPT-based AI financial assistant, collected during a live Capture-The-Flag (CTF) competition. The agent's system prompt embeds structured synthetic business data — settlement records with transaction IDs, amounts, vendors, and memos — that participants attempted to extract via conversational prompt injection.
Each flag is a structured settlement record with four fields: txnId, amount… See the full description on the dataset page: https://huggingface.co/datasets/stykat/financial-ai-ctf-dataset.vnpdf-financial-reports-dataset
Dataset Card for Financial Report Dataset Demo
Dataset Summary
This dataset contains financial reports from Vietnamese VININDEX, including both the text content and corresponding page images. The dataset is designed for document understanding and information extraction tasks.
Languages
The dataset contains text in Vietnamese (vi).
Dataset Structure
The dataset contains 401 examples, each with:
image: The page image from the PDF document
text: The… See the full description on the dataset page: https://huggingface.co/datasets/kiethuynhanh/vnpdf-financial-reports-dataset.qwen3-14b-em-risky-financial-dataset
Risky financial advice — EM training dataset
6,000 user/assistant pairs used to fine-tune Qwen3-14B into broadly and narrowly
misaligned variants following:
Turner & Soligo et al. 2025 — Model Organisms for Emergent Misalignment
Soligo et al. 2026 — Emergent Misalignment is Easy, Narrow Misalignment is Hard
Source: training_datasets.zip.enc in https://github.com/clarifying-EM/model-organisms-for-EM .
For auditing/safety research only.
adaption-sec-financial-arithmetic-dataset
SEC Financial Arithmetic Dataset — Adaption AutoScientist Challenge
Powered by Adaptive Data — Adaption Labs
What This Dataset Teaches
This dataset trains a model to extract numbers from SEC filing tables and execute verified multi-step arithmetic — every answer is cross-checked against a gold reasoning program:
Task
Source
Example
Table Variable Extraction
FinQA
"From this 10-K table, extract 2021 and 2022 revenue values"
Multi-Step Arithmetic… See the full description on the dataset page: https://huggingface.co/datasets/narendarcodes/adaption-sec-financial-arithmetic-dataset.Financial_Context_DatasetThis dataset contains over 50,000 samples of user financial queries paired with their corresponding structured data requests (context). It was created to facilitate the creation of the Financial Agent LLM for accurate data extraction and query answering.
How to load the Dataset
You can load the dataset using the code below:
from datasets import load_dataset
ds = load_dataset("Chaitanya14/Financial_Context_Dataset")
Dataset Construction
Diverse Query Sources… See the full description on the dataset page: https://huggingface.co/datasets/Chaitanya14/Financial_Context_Dataset.Nigerian-Financial-Transactions-and-Fraud-Detection-Dataset
Dataset Card for GLUE
Dataset Summary
GLUE, the General Language Understanding Evaluation benchmark (https://gluebenchmark.com/) is a collection of resources for training, evaluating, and analyzing natural language understanding systems.
Supported Tasks and Leaderboards
The leaderboard for the GLUE benchmark can be found at this address. It comprises the following tasks:
ax
A manually-curated evaluation dataset for fine-grained analysis of system… See the full description on the dataset page: https://huggingface.co/datasets/demola2/Nigerian-Financial-Transactions-and-Fraud-Detection-Dataset.enterprise-financial-crime-ai-datasetTransactions → Risk Analysis → Alerts → Investigation → SAR Reports
Dataset Statistics
Total records: 310,396Dataset size: 339 MBAuto-converted parquet size: 65 MB
Languages:
English
French
Spanish
Main fields:
email_id
thread_id
timestamp
language
bank
department
country
risk_level
Enterprise Financial Crime AI Dataset
The Enterprise Financial Crime AI Dataset is a high-fidelity dataset built from real-world operational patterns and enterprise data structures… See the full description on the dataset page: https://huggingface.co/datasets/Webopen2026/enterprise-financial-crime-ai-dataset.FinBERT-financial-news-data
FinBERT-financial-news-data
The sentence-level sentiment training data behind
gamug/FinBERT-financial-news --
5,865 sentences from real, English-language financial-news articles (2010s-2020s), each
labeled positive/negative/neutral from an investor/price-impact perspective by an LLM
(DeepSeek deepseek-chat, temperature 0, one sentence at a time, in isolation).
Why this exists
Published as the direct counterpart to the model it trains, so the model's own claims… See the full description on the dataset page: https://huggingface.co/datasets/gamug/FinBERT-financial-news-data.financial_data_ratiosSujet-Financial-RAG-FR-Dataset
Sujet-Financial-RAG-FR-Dataset 📊💼
Description 📝
This dataset is a proof-of-concept collection of French question-context pairs, specifically designed for training and evaluating embedding models in the financial domain. To demonstrate the importance of this approach, we hand-selected a few publicly available French financial documents. It's important to note that it remains entirely possible and fairly straightforward to gather a lot more financial documents and… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Financial-RAG-FR-Dataset.AI_financial_fraud_dataset
