datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
finance-tasks
Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024)
This repo contains the evaluation datasets for our paper Adapting Large Language Models via Reading Comprehension.
We explore continued pre-training on domain-specific corpora for large language models. While this approach enriches LLMs with domain knowledge, it significantly hurts their prompting ability for question answering. Inspired by human learning via reading comprehension, we propose a simple method to… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/finance-tasks.IndustryInstruction_Finance-Economics
IndustryInstruction: Finance & Economics
This repository contains the IndustryInstruction: Finance & Economics domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Finance-Economics.reddit_finance_43_250k
reddit finance 43 250k
reddit_finance_43_250k is a collection of 250k post/comment pairs from 43 financial, investing and crypto subreddits. Post must have all been text, with a length of 250chars, and a positive score. Each subreddit is narrowed down to the 70th qunatile before being mergered with their top 3 comments and than the other subs. Further score based methods are used to select the top 250k post/comment pairs.
The code to recreate the dataset is here:… See the full description on the dataset page: https://huggingface.co/datasets/winddude/reddit_finance_43_250k.lossbench-finance-v1
LossBench finance-v1
Severity-weighted expected-loss evaluation for agents that touch money. Three finance back-office domains, mechanical ground truth, and a contamination certificate. Models are ranked by what their mistakes cost, not by raw accuracy.
Overview
Task count
2400
Domains
reconciliation, payment_repair, settlement
License
cc-by-4.0
Tasks
Each task is an agentic back-office scenario with a deterministic seed, an… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/lossbench-finance-v1.ko-finance-asr-corrections
ko-finance-asr-corrections
Frequency-annotated Korean ASR confusion pairs from finance/stock YouTube.
210 pairs
mined from 2,391 videos of auto-captions
across 47 channels
totalling 1,080.1 hours
Each pair carries how often the term was mangled and how often it was said correctly, plus
verification provenance.
한국어 금융·주식 유튜브 자동자막에서 실측한 ASR 오인식→교정 쌍입니다. 모든 쌍에 오표기·정답
표기 빈도(→ 용어별 오인식률)와 검증 메타데이터(2-LLM 합의 감사, 승격 티어)가 붙어 있습니다.
What makes it different
No public… See the full description on the dataset page: https://huggingface.co/datasets/woongstar/ko-finance-asr-corrections.aprm-sft-thoughts-snorkel-finance-policy_best-adamw30-lp0
Act-PRM SFT thoughts — snorkel-finance finance
Act-PRM (Action Process Reward Models) infers the latent thoughts behind
logged, action-only agent demonstrations via an offline EM. For each
logged action x in state s we sample G=4 candidate thoughts z,
score each by the length-penalized action likelihood
reward(z) = p(x | s, z)
(len_frac grows with the thought's token length), and mark the best thought
(argmax reward). The (thought + action) span is then what downstream SFT / RL… See the full description on the dataset page: https://huggingface.co/datasets/mzio/aprm-sft-thoughts-snorkel-finance-policy_best-adamw30-lp0.adaption-personal-finance-advice-dialogues
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-personal_finance_advice_dialogues
This dataset contains multi-turn conversational samples between users and an AI assistant focused on personal finance topics such as budgeting, investing, insurance, and taxes. Each entry follows a pattern where a user presents an initial scenario, provides an update with new constraints or events, and receives tailored financial advice that adapts… See the full description on the dataset page: https://huggingface.co/datasets/Azfarhashmi/adaption-personal-finance-advice-dialogues.excel_huggingface_yahoo-finance_time_2577_selftest2fha-peer-adjusted-denial-rates
Peer-Adjusted FHA Denial Rates — 2025
Every published mortgage denial statistic carries the same caveat: rates partly reflect who applies where. This dataset quantifies that caveat at the lender level, apparently for the first time in public.
Method
Indirect standardization (the technique used for standardized mortality ratios in public health), applied to credit:
All 1,217,297 decisioned FHA applications in the complete 2025 CFPB HMDA record are assigned to… See the full description on the dataset page: https://huggingface.co/datasets/FinanceRateCalc/fha-peer-adjusted-denial-rates.reddit_finance_43_250k
reddit finance 43 250k
reddit_finance_43_250k is a collection of 250k post/comment pairs from 43 financial, investing and crypto subreddits. Post must have all been text, with a length of 250chars, and a positive score. Each subreddit is narrowed down to the 70th qunatile before being mergered with their top 3 comments and than the other subs. Further score based methods are used to select the top 250k post/comment pairs.
The code to recreate the dataset is here:… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/reddit_finance_43_250k.finance-analyst-qa-verifiedfinance_dataset_privatefinance-alignment-benchmarkfinance-benchmark-model-scoresfinance-dpo-pairs-verifiedreddit-finance-qa-json
Dataset Overview
This repository contains files used in the fine-tuning and retrieval-augmented generation (RAG) system built on Reddit finance data. Check out the Github repo to use this data here
reddit_finance_qa.jsonl
This is a JSON Lines (jsonl) file containing cleaned and deduplicated Reddit question-answer (QA) pairs from finance-related subreddits such as:
r/personalfinance
r/investing
r/wallstreetbets
r/cryptocurrency
r/stocks
Format (One… See the full description on the dataset page: https://huggingface.co/datasets/egupta/reddit-finance-qa-json.reddit_finance_43_250k
reddit finance 43 250k
reddit_finance_43_250k is a collection of 250k post/comment pairs from 43 financial, investing and crypto subreddits. Post must have all been text, with a length of 250chars, and a positive score. Each subreddit is narrowed down to the 70th qunatile before being mergered with their top 3 comments and than the other subs. Further score based methods are used to select the top 250k post/comment pairs.
The code to recreate the dataset is here:… See the full description on the dataset page: https://huggingface.co/datasets/NodecoreHQbot/reddit_finance_43_250k.
