datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
finance-tasks
Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024)
This repo contains the evaluation datasets for our paper Adapting Large Language Models via Reading Comprehension.
We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a simple method to… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/finance-tasks.Sujet-Finance-Instruct-177k
Sujet Finance Dataset Overview
The Sujet Finance dataset is a comprehensive collection designed for the fine-tuning of Language Learning Models (LLMs) for specialized tasks in the financial sector. It amalgamates data from 18 distinct datasets hosted on HuggingFace, resulting in a rich repository of 177,597 entries. These entries span across seven key financial LLM tasks, making Sujet Finance a versatile tool for developing and enhancing financial applications of AI.… See the full description on the dataset page: https://huggingface.co/datasets/sujet-ai/Sujet-Finance-Instruct-177k.IndustryInstruction_Finance-Economics
IndustryInstruction: Finance & Economics
This repository contains the IndustryInstruction: Finance & Economics domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Finance-Economics.Corp_FinanceLongContextReasoning
Corp_Finance Long-Context Reasoning — Expert-Authored Credit Agreement Benchmark (Showcase Sample)
A five-record public sample from Corp_Finance Long-Context Reasoning, a subject-matter-expert benchmarking dataset built by Human Edge for evaluating frontier model reasoning over leveraged finance and syndicated credit documentation.
Every question, reasoning trace, and answer in this dataset was authored by a practicing finance professional and independently reviewed by 2-3… See the full description on the dataset page: https://huggingface.co/datasets/HumanEdgeAI/Corp_FinanceLongContextReasoning.finance-alpaca-1k-testindian-finance-synthetic-phase2-cleaned
Indian Finance Synthetic Dataset (Phase 2 - Final Clean)
Dataset Description
14,763 high-quality synthetic conversations about Indian personal finance, optimized for fine-tuning.
Recent Updates
✅ v3 (Final): Removed 14 samples with empty content messages
✅ v2: Removed 58 incomplete conversations
✅ v1: Tools optimization (82.5% size reduction)
All conversations are now complete and properly formatted for training.
Key Features
Clean… See the full description on the dataset page: https://huggingface.co/datasets/Gandalf1/indian-finance-synthetic-phase2-cleaned.financebench-voyage-finance-2-embeddings
FinanceBench (voyage-finance-2 embeddings)
Pre-computed embeddings for the FinanceBench corpus. Skips ~$5-15 of Voyage API cost and ~30 minutes of ingest time vs re-embedding from raw PDFs. Intended consumer: the RAG agent at Rishabhmannu/financebench-rag-agent (install: pip install financebench-rag-agent).
What's in the box (frozen)
Field
Value
Source corpus
FinanceBench (SEC filings: 10-K, 10-Q, 8-K, earnings releases)
Parser
pypdf (canonical)… See the full description on the dataset page: https://huggingface.co/datasets/cmpunkmannu/financebench-voyage-finance-2-embeddings.Islamic_Finance_QnA_eval
Islamic Finance Q&A Evaluation Dataset
Validation and test splits for evaluating models on Islamic Finance Q&A.
Dataset Structure
Format: Simple prompt-answer pairs
Validation: ~203 examples (10%)
Test: ~203 examples (10%)
Language: Arabic
Domain: Islamic finance and Sharia-compliant banking
Fields
id: Unique identifier
prompt: The question prompt
question: Original question text
answer: Ground truth answer
topic: Topic category
split:… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/Islamic_Finance_QnA_eval.wdb-islamic-finance-benchmark
WDB Benchmark: Western Default Bias in Islamic Finance
Dataset Description
This benchmark tests whether Large Language Models exhibit Western Default Bias (WDB) - the tendency to provide Western/conventional finance answers even when the context implies Islamic finance should be used.
The Problem
When a user in Saudi Arabia or UAE asks a financial question, they likely expect Shariah-compliant advice. However, LLMs trained predominantly on Western data may… See the full description on the dataset page: https://huggingface.co/datasets/Raniahossam33/wdb-islamic-finance-benchmark.aprm-thought-generations-snorkel-finance
aprm-thought-generations-snorkel-finance
Act-PRM inferred-thought generations for snorkel_finance agent traces
(Qwen3-4B-Instruct-2507). For each logged (state s, action x), an offline EM
samples G=4 candidate thoughts z, scores each by the length-penalized action
likelihood, and commits the top-1.
generations/{policy,base,policy_last,base_last}.jsonl (primary)
One row per logged action step. Full candidate pool so you can take top-1 OR recompute
any weighting:… See the full description on the dataset page: https://huggingface.co/datasets/mzio/aprm-thought-generations-snorkel-finance.Finance-Instruct-AzerbaijaniThis is part of a translated version of the original dataset: https://huggingface.co/datasets/Josephgflowers/Finance-Instruct-500k
sa-finance-reasoning-mix
tejeshbhalladhanyog/sa-finance-reasoning-mix
A merged reasoning-distillation dataset combining a custom multi-agent
financial-reasoning pipeline with a stratified slice of GLM-5.1's
general-domain reasoning data.
Composition
source / subset
rows
sa_pipeline_qwen_max (decomposition + cluster_generation + mapreduce_single)
153,923
main
80,000
PHD-Science
20,000
Multilingual-STEM
20,000
Math
15,000
total
288,923
Format
Each row is a… See the full description on the dataset page: https://huggingface.co/datasets/tejeshbhalladhanyog/sa-finance-reasoning-mix.
