datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
financial-analyst-data-full
financial-analyst-data-full
A-share historical price + valuation data packaged for financial-analyst —
the 14-agent single-stock deep-dive research workstation.
Published: 2026-05-24
Preset: full — 全 A 股完整包 (含历史退市股). 量化研究员 / 重度用户. lite 全 + TDX 历年财报原始 zip (用户跑 import_tdx_financial.py 解) + F10 原始文本 (公司大事/龙虎榜/主力追踪/最新提示 .txt).
Size: ~14.1 GB
What's included
5450 stocks daily OHLCV + 7 valuation fields (PE/PB/PS/DV/MV/CIRC_MV/turnover_rate)
Date range (daily): 1990-12-19… See the full description on the dataset page: https://huggingface.co/datasets/yifishbossman/financial-analyst-data-full.financial-analyst-data-lite
financial-analyst-data-lite
EN: A-share historical OHLCV + valuation + financials + TDX F10 events, packaged in Qlib binary + Parquet formats. Companion dataset for financial-analyst — a 14-agent single-stock deep-dive research workstation.
中文: A 股历史行情 + 估值 + 财报 + TDX F10 事件数据集, Qlib 二进制 + Parquet 双格式打包. 配套 financial-analyst — 14 Agent 个股深度研究工作站使用.
Published / 发布: 2026-05-24 · Size / 体量: ~2.72 GB · License: Apache 2.0
📊 Three Preset Tiers / 三档预设
Pick the tier that… See the full description on the dataset page: https://huggingface.co/datasets/yifishbossman/financial-analyst-data-lite.Multimodal-Dataset-Image_Text_Table_TimeSeries-for-Financial-Time-Series-ForecastingThe sp500stock_data_description.csv file provides detailed information on the existence of four modalities (text, image, time series, and table) for 4,213 S&P 500 stocks.
The hs300stock_data_description.csv file provides detailed information on the existence of four modalities (text, image, time series, and table) for 858 HS 300 stocks.
If you find our research helpful, please cite our paper:
@article{xu2025finmultitime,
title={FinMultiTime: A Four-Modal Bilingual Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/Wenyan0110/Multimodal-Dataset-Image_Text_Table_TimeSeries-for-Financial-Time-Series-Forecasting.Bangla_Financial_news_articles_Dataset
Bangla-Financial-news-articles-Dataset
A Comprehensive Resource for Analyzing Sentiments in over 7600+ Bangla News.
Downloads
🔴 Download the "💥Bangla_fin_news.zip" file for all "7,695" news and extract it.
About Dataset
Welcome to our Bengali Financial News Sentiment Analysis dataset! This collection comprises 7,695 financial news articles extracted, covering the period from March 3, 2014, to December 29, 2021. Utilizing the powerful web scraping tool… See the full description on the dataset page: https://huggingface.co/datasets/ashtrayAI/Bangla_Financial_news_articles_Dataset.financial-analyst-data-demo
financial-analyst-data-demo
EN: A-share historical OHLCV + valuation + financials + TDX F10 events, packaged in Qlib binary + Parquet formats. Companion dataset for financial-analyst — a 14-agent single-stock deep-dive research workstation.
中文: A 股历史行情 + 估值 + 财报 + TDX F10 事件数据集, Qlib 二进制 + Parquet 双格式打包. 配套 financial-analyst — 14 Agent 个股深度研究工作站使用.
Published / 发布: 2026-05-24 · Size / 体量: ~0.16 GB · License: Apache 2.0
📊 Three Preset Tiers / 三档预设
Pick the tier that… See the full description on the dataset page: https://huggingface.co/datasets/yifishbossman/financial-analyst-data-demo.Financial_datasetsfinancial-data-pipeline
Financial Data Pipeline — Full Curated Snapshot
A comprehensive financial dataset covering 236 tables and 140,479,626 rows across macro, market, and alternative data sources.
Data Sources
Category
Tables
Key Sources
Market Prices
18
Tiingo, Schwab, Finnhub, CBOE
Macro & Economic
89
FRED, BLS, BEA, Treasury, EIA
Fundamentals
83
SEC EDGAR, Finnhub, SimFin, Alpha Vantage
Alternative Data
32
Congressional trades, insider transactions, patents, OpenFDA… See the full description on the dataset page: https://huggingface.co/datasets/ZanderL1337/financial-data-pipeline.financial-data
Financial Data — Marts Schema Data Dictionary
This document provides a comprehensive schema reference and metric dictionary for the 23 analytical tables compiled in the marts schema of database.db (and saved as Parquet files under marts/).
1. fct_combined_scorecard
Purpose: The "front page" dashboard. Flat, denormalized view containing the latest values of every calculated metric joined into a single table for fast querying.
SQL Source: Derived from private… See the full description on the dataset page: https://huggingface.co/datasets/speb/financial-data.Sujet-Financial-RAG-EN-Dataset
Sujet Financial RAG EN Dataset 📊💼
Description 📝
The Sujet Financial RAG EN Dataset is a comprehensive collection of English question-context pairs, specifically designed for training and evaluating embedding models in the financial domain. To demonstrate the importance of this approach, we hand-selected a variety of publicly available English financial documents, with a focus on 10-K Forms.
A 10-K Form is a comprehensive report filed annually by public companies about… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Financial-RAG-EN-Dataset.Nigerian-Financial-Transactions-and-Fraud-Detection-Dataset
Nigerian Financial Transactions and Fraud Detection Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 1M<n<10M - Formats: csv - Sector: economics_finance - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Nigerian-Financial-Transactions-and-Fraud-Detection-Dataset.financial_dataThis dataset is a combination of Stanford's Alpaca (https://github.com/tatsu-lab/stanford_alpaca) and FiQA (https://sites.google.com/view/fiqa/) with another 1.3k pairs custom generated using GPT3.5
Script for tuning through Kaggle's (https://www.kaggle.com) free resources using PEFT/LoRa: https://www.kaggle.com/code/gbhacker23/wealth-alpaca-lora
GitHub repo with performance analyses, training and data generation scripts, and inference notebooks: https://github.com/gaurangbharti1/wealth-alpaca… See the full description on the dataset page: https://huggingface.co/datasets/csujeong/financial_data.Multimodal-Dataset-Image_Text_Table_TimeSeries-for-Financial-Time-Series-ForecastingThe sp500stock_data_description.csv file provides detailed information on the existence of four modalities (text, image, time series, and table) for 4,213 S&P 500 stocks.
The hs300stock_data_description.csv file provides detailed information on the existence of four modalities (text, image, time series, and table) for 858 HS 300 stocks.
If you find our research helpful, please cite our paper:
@article{xu2025finmultitime,
title={FinMultiTime: A Four-Modal Bilingual Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/Y123-wed/Multimodal-Dataset-Image_Text_Table_TimeSeries-for-Financial-Time-Series-Forecasting.Financial_statements_fraud_dataset
Financial Statement Fraud Detection Dataset
Official Dataset of the Paper: Read Between the Lines: A Robust Financial Statement Fraud Detection Framework
Guy Stephane Waffo Dzuyo, Gael Guibon, Christophe Cerisara, Luis Belmar-Letelier
Forvis Mazars | LORIA, CNRS, Universite de Lorraine | Universite Sorbonne Paris Nord
Emails: guy.stephane.waffo@forvismazars.com | gael.guibon@lipn.fr | christophe.cerisara@loria.fr | luis.belmar-letelier@forvismazars.com
Main Purpose… See the full description on the dataset page: https://huggingface.co/datasets/WaguyMZ/Financial_statements_fraud_dataset.synthetic_vc_financial_decisions_reasoning_dataset
Best Curator Use Case in the Reasoning Datasets Competition: https://www.linkedin.com/feed/update/urn:li:activity:7330998995990781952/
Synthetic VC Financial Decisions Reasoning Dataset
Dataset Summary
The Synthetic VC Financial Decisions Reasoning Dataset is a large-scale collection designed to train, evaluate, and fine-tune language models on subjective, abstract financial reasoning tasks. It simulates venture capital (VC) workflows by capturing multiple… See the full description on the dataset page: https://huggingface.co/datasets/ZennyKenny/synthetic_vc_financial_decisions_reasoning_dataset.FinQA_TAT-QA_financial_finetuning_dataset
Dataset Summary
This dataset provides a unified, flattened context / question / answer format for
question answering over financial documents that combine tabular and textual data. It is
built to support training and evaluating models on numerical and discrete reasoning
tasks in the finance domain, drawing on the structure and style of established
finance-QA benchmarks such as TAT-QA and FinQA.
Each example pairs a passage of financial context (derived from a table and/or… See the full description on the dataset page: https://huggingface.co/datasets/hellotayssir/FinQA_TAT-QA_financial_finetuning_dataset.financial-qa-dataset
financial-qa-dataset
This dataset consists of Question-Answer_Context Pairs. It also consists of metadata for filtering the records.
Repo Structure
financial-qa-dataset
├── financial-qa-dataset.csv
├── metadata.csv
├── notebooks
│ |── loading_dataset.ipynb
│ |── Loading_dataset_huggingface.ipynb
│ |── basic_rag_langchain_vertexai.ipynb
│ |── basic_rag_with_evaluation.ipynb
|
├── data
|── Statements
|── Reports… See the full description on the dataset page: https://huggingface.co/datasets/adityarane/financial-qa-dataset.synthetic-financial-data
Dataset Card for Dataset Name
Extracted from Kaggle: https://www.kaggle.com/datasets/ealaxi/paysim1?resource=download
Purpose is to have synthetic data on here for testing.
thai-financial-datasetThis dataset is the cleaned version of the airesearch/CMDF_VISTEC datasets for pretrining model.
It is financial domain for Thai language.
license: cc-by-4.0
financial-news-articles-filtereddataset_info:
features:
- name: title
dtype: string
- name: text
dtype: string
- name: url
dtype: string
- name: word_count
dtype: int64
splits:
- name: train
num_bytes: 554834105.9892601
num_examples: 199711
download_size: 459025008
dataset_size: 554834105.9892601
configs:
- config_name: default
data_files:
- split: train
path: data/train-*
Multimodal-Dataset-Image_Text_Table_TimeSeries-for-Financial-Time-Series-ForecastingThe sp500stock_data_description.csv file provides detailed information on the existence of four modalities (text, image, time series, and table) for 4,213 S&P 500 stocks.
The hs300stock_data_description.csv file provides detailed information on the existence of four modalities (text, image, time series, and table) for 858 HS 300 stocks.
If you find our research helpful, please cite our paper:
@article{xu2025finmultitime,
title={FinMultiTime: A Four-Modal Bilingual Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/tishtakalita/Multimodal-Dataset-Image_Text_Table_TimeSeries-for-Financial-Time-Series-Forecasting.financial-ai-ctf-dataset
Financial AI Prompt Injection CTF Dataset
A dataset of 400 multi-turn conversations against a GPT-based AI financial assistant, collected during a live Capture-The-Flag (CTF) competition. The agent's system prompt embeds structured synthetic business data — settlement records with transaction IDs, amounts, vendors, and memos — that participants attempted to extract via conversational prompt injection.
Each flag is a structured settlement record with four fields: txnId, amount… See the full description on the dataset page: https://huggingface.co/datasets/verno-labs/financial-ai-ctf-dataset.Synthetic-Financial-Datasets-For-Fraud-DetectionAI_financial_fraud_datasetexness_financial_datafinancial-ai-ctf-dataset
Financial AI Prompt Injection CTF Dataset
A dataset of 400 multi-turn conversations against a GPT-based AI financial assistant, collected during a live Capture-The-Flag (CTF) competition. The agent's system prompt embeds structured synthetic business data — settlement records with transaction IDs, amounts, vendors, and memos — that participants attempted to extract via conversational prompt injection.
Each flag is a structured settlement record with four fields: txnId, amount… See the full description on the dataset page: https://huggingface.co/datasets/stykat/financial-ai-ctf-dataset.financial-golden-dataset-v1finbert-financial-news-sentiment-dataset
📈 Financial News Sentiment Dataset (FinBERT Powered)
Welcome to the official data repository of Lumen Models. This dataset provides a real-time, high-frequency stream of global financial news headlines aggregated from major economic outlets, processed with state-of-the-art Natural Language Processing (NLP).
Every headline is automatically analyzed using FinBERT (a BERT model specifically trained and fine-tuned for financial text analysis) to determine market sentiment with… See the full description on the dataset page: https://huggingface.co/datasets/lumen-models/finbert-financial-news-sentiment-dataset.qtrade-financial-data
QTrade Financial Data
這是 QTrade 量化交易框架的金融數據資料庫。
📊 資料集內容
包含 127,000+ 筆歷史金融數據(2005-2025,約 20 年),涵蓋:
🪙 貨幣匯率(8 個)
TWDUSD, USDTWD - 台幣/美金匯率
EURUSD - 歐元/美金
USDJPY - 美金/日圓
GBPUSD - 英鎊/美金
DXY - 美元指數
📈 市場指數(12 個)
台灣: TWII (加權指數), TPEX (櫃買指數)
美國: SPX (標普500), DJI (道瓊), IXIC (納斯達克), RUT (羅素2000), SOX (費半), VIX (恐慌指數)
國際: FTSE (富時100), N225 (日經225), HSI (恆生), SSEC (上證)
🥇 期貨(7 個)
GC - COMEX 黃金期貨
SI - COMEX 白銀期貨
HG - COMEX 銅期貨
CL - WTI 原油期貨
NG… See the full description on the dataset page: https://huggingface.co/datasets/zoanana990/qtrade-financial-data.vnpdf-financial-reports-dataset
Dataset Card for Financial Report Dataset Demo
Dataset Summary
This dataset contains financial reports from Vietnamese VININDEX, including both the text content and corresponding page images. The dataset is designed for document understanding and information extraction tasks.
Languages
The dataset contains text in Vietnamese (vi).
Dataset Structure
The dataset contains 401 examples, each with:
image: The page image from the PDF document
text: The… See the full description on the dataset page: https://huggingface.co/datasets/kiethuynhanh/vnpdf-financial-reports-dataset.adaption-sec-financial-arithmetic-dataset
SEC Financial Arithmetic Dataset — Adaption AutoScientist Challenge
Powered by Adaptive Data — Adaption Labs
What This Dataset Teaches
This dataset trains a model to extract numbers from SEC filing tables and execute verified multi-step arithmetic — every answer is cross-checked against a gold reasoning program:
Task
Source
Example
Table Variable Extraction
FinQA
"From this 10-K table, extract 2021 and 2022 revenue values"
Multi-Step Arithmetic… See the full description on the dataset page: https://huggingface.co/datasets/narendarcodes/adaption-sec-financial-arithmetic-dataset.
