CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gbharti /finance-alpacaThis dataset is a combination of Stanford's Alpaca (https://github.com/tatsu-lab/stanford_alpaca) and FiQA (https://sites.google.com/view/fiqa/) with another 1.3k pairs custom generated using GPT3.5 Script for tuning through Kaggle's (https://www.kaggle.com) free resources using PEFT/LoRa: https://www.kaggle.com/code/gbhacker23/wealth-alpaca-lora GitHub repo with performance analyses, training and data generation scripts, and inference notebooks: https://github.com/gaurangbharti1/wealth-alpaca… See the full description on the dataset page: https://huggingface.co/datasets/gbharti/finance-alpaca.texttext-generation10K<n<100K155 likes2.2k downloads10mo agoHugging Face02BAAI /IndustryCorpus_finance[中文主页] Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise. To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_finance.texttext-generation10M<n<100M19 likes2.2k downloads1mo agoHugging Face03nvidia /Nemotron-SpecializedDomains-Finance-v1 Dataset Description Nemotron-SpecializedDomains-Finance is a large-scale synthetic financial question-answering dataset designed to improve LLM performance on specialized financial reasoning and document comprehension tasks. The dataset comprises 326K+ high-quality Q&A pairs generated from SEC filings of S&P 500 companies spanning 2019-2024. This dataset is ready for commercial use. Overview The dataset leverages template-based Synthetic Data Generation (SDG) to… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-SpecializedDomains-Finance-v1.texttext-generation100K<n<1M16 likes1.6k downloads7mo agoHugging Face04RogoAI /big-finance-benchmark BigFinanceBench Public Release arXiv | Website | GitHub | Blog post Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation. This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/RogoAI/big-finance-benchmark.textquestion-answeringn<1K12 likes927 downloads2mo agoHugging Face05handshake-ai-research /ATLAS-Finance ATLAS Finance A benchmark of 100 expert-level tasks inside 13 realistic financial firm environments, packaged in the Harbor RLE format. Each task drops an AI agent into a Linux workstation with a persistent multi-app world — inbox, chat, calendar, virtual data room, drive, wiki — and asks the agent to produce the same deliverable a financial professional would be responsible for: an Excel workbook containing the model and supporting analysis. Here we provide the data for this… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/ATLAS-Finance.documenttext-generationn<1K3 likes480 downloads9d agoHugging Face06idleengine /big-finance-benchmark BigFinanceBench Public Release arXiv | Website | GitHub | Blog post Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation. This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/big-finance-benchmark.textquestion-answeringn<1K0 likes387 downloads1mo agoHugging Face07ianncity /GLM-5.2-Finance-80000x GLM-5.2 · Finance-80000x 80,000x financial related traces distilled from GLM-5.2 on High reasoning Risk · Markets · Investments · Corporate Finance · Wealth Management Token Count: 220M Unique prompts generated with diffusion Gemma-27B answered by GLM-5.2 You can use this dataset for any purpose and you dont need to credit me, preferably dont claim it as your own. hi - ianncity texttext-generation10K<n<100K18 likes300 downloads2mo agoHugging Face08HCAI-Lab-GT /dolma3-6t-sample-10000-docs-finance-and-business HCAI-Lab/dolma3-6t-sample-10000-docs-finance-and-business Filename-derived finance_and_business slice of HCAI-Lab/dolma3-6t-sample-10000-docs, pinned to revision 561e73c7e0ad35c04f386bae1e3dd39dfb6755e7. Extraction rule The corpus contains every source .jsonl.zst file whose filename contains the literal segment -finance_and_business-. Source paths and compressed file contents are preserved byte-for-byte. This is a coarse WebOrganizer finance_and_business category… See the full description on the dataset page: https://huggingface.co/datasets/HCAI-Lab-GT/dolma3-6t-sample-10000-docs-finance-and-business.texttext-generation100K<n<1M0 likes284 downloads1mo agoHugging Face09Akhil-Theerthala /Personal-Finance-Queries Dataset Description A curated collection of Reddit posts and top comments focused on personal finance questions. The data is further filtered with the help of LLM-based Voting scores. These scores determine if the query is relevant to a person's financial queries among the other posts of the subreddits. Dataset Structure Columns: category: The sub-domain of personal finance that the query belongs to. subreddit: Source subreddit (string, categorical) query: User’s… See the full description on the dataset page: https://huggingface.co/datasets/Akhil-Theerthala/Personal-Finance-Queries.textquestion-answering10K<n<100K9 likes192 downloads1y agoHugging Face10MaitriVasa /big-finance-benchmark BigFinanceBench Public Release arXiv | Website | GitHub | Blog post Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation. This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/MaitriVasa/big-finance-benchmark.textquestion-answeringn<1K0 likes152 downloads2mo agoHugging Face11Koplos /big-finance-benchmark BigFinanceBench Public Release arXiv | Website | GitHub | Blog post Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation. This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/Koplos/big-finance-benchmark.textquestion-answeringn<1K0 likes145 downloads2mo agoHugging Face12Sanscritic /finance-pro-bench FinancePro-Bench FinancePro-Bench&nbsp;is a benchmark of&nbsp;400 complex expert-level finance questions&nbsp;which not only require deep knowledge of finance but also other domains such as regulation, law, strategy, math and code generation. It emulates the kind of nuanced, context specific, multi-step reasoning that expert finance professionals perform such as reviewing accounting judgments, structuring deals, pricing derivatives, navigating tax and compliance edge… See the full description on the dataset page: https://huggingface.co/datasets/Sanscritic/finance-pro-bench.imagequestion-answeringn<1K1 likes140 downloads3mo agoHugging Face13oliversayshi /big-finance-benchmark BigFinanceBench Public Release arXiv | Website | GitHub | Blog post Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation. This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/oliversayshi/big-finance-benchmark.textquestion-answeringn<1K0 likes136 downloads2mo agoHugging Face14Anson1110 /Finance-Questions-Essay_and_Calculation-Chinese Overview Finance-Questions-Essay_and_Calculation-Chinese is a carefully curated financial reasoning dataset containing 954 samples, each annotated with high-quality Chain-of-Thought (CoT) reasoning. It is designed to train and evaluate Chinese financial language models on complex essay and calculation tasks. Stage 1: Data Collection & Standardization Extract financial question samples from professional textbooks via Easy Dataset. Manually label 30 seed samples, then use… See the full description on the dataset page: https://huggingface.co/datasets/Anson1110/Finance-Questions-Essay_and_Calculation-Chinese.texttext-generationn<1K1 likes116 downloads6mo agoHugging Face15CentificAIResearch /tiered-finance-eval Tiered Finance Eval Twenty agentic finance tasks, each with the reference files an analyst would actually be handed, a curated gold deliverable, and a tiered, gated rubric that scores a submission against that gold. Evaluation results for these tasks are published in the companion Space: CentificAIResearch/Tiered-Finance-Eval. This dataset holds the tasks only: no model outputs and no scores. [!IMPORTANT] Canary string. TIERED-FINANCE-EVAL:d9e2f4a1-7c3b-4e86-9a05-2f1b8c6d40e7… See the full description on the dataset page: https://huggingface.co/datasets/CentificAIResearch/tiered-finance-eval.documenttext-generationn<1K0 likes111 downloads6d agoHugging Face16caiotheodoro /lossbench-finance-v1 LossBench finance-v1 Severity-weighted expected-loss evaluation for agents that touch money. Three finance back-office domains, mechanical ground truth, and a contamination certificate. Models are ranked by what their mistakes cost, not by raw accuracy. Overview Task count 2400 Domains reconciliation, payment_repair, settlement License cc-by-4.0 Tasks Each task is an agentic back-office scenario with a deterministic seed, an… See the full description on the dataset page: https://huggingface.co/datasets/caiotheodoro/lossbench-finance-v1.tabulartext-generation1K<n<10K0 likes103 downloads1mo agoHugging Face17d4rkninja /tanpo-finance-sft Tanpo Finance SFT (10k) Ownership Owner: DarkNinja Solutions Creator: d4rkninja Community: DarkLab Domain Finance and CFO thinking: unit economics, budgeting, forecasting, cash, pricing math, and financial operating cadence. Dataset summary Field Value Rows 10000 Schema Chat SFT (messages with system / user / assistant) Source file tanpo-finance-sft-10k-format-fixed.jsonl Audit verdict PASS Empty assistant rate… See the full description on the dataset page: https://huggingface.co/datasets/d4rkninja/tanpo-finance-sft.texttext-generation10K<n<100K0 likes62 downloads7d agoHugging Face18Simon-Liu /tw-finance-function-call-reasoning tw-finance-function-call-reasoning 台灣金融場景的繁體中文 function-calling + 推理鏈微調資料集。 欄位規格對齊 twinkle-ai/tw-function-call-reasoning-10k。 ⚠️ 使用限制:僅供研究,不得商業使用 本資料集以 CC BY-NC 4.0 授權釋出,僅供學術研究、模型能力探索與方法驗證之用。 請務必理解以下事項後再使用: 不得作商業用途。 包含但不限於:訓練用於對外營利的模型、包裝為付費產品或服務、 作為商業交付物的一部分。若有商業需求,請自行重新建置資料並取得合規來源。 這不是財務、稅務、法律或投資建議。 資料中的稅率、費率、法規門檻雖依 2026 年 (民國 115 年)台灣公開資訊整理,但可能已經過時或有誤。任何實際決策前, 請以主管機關公告為準(財政部、金管會、勞動部、衛福部、中央銀行、全國法規資料庫)。 內容為程式化合成,非真實考題逐字收錄。 題目由模板與參數取樣組合而成,… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/tw-finance-function-call-reasoning.texttext-generation1K<n<10K0 likes61 downloads15d agoHugging Face19Simon-Liu /tw-finance-reasoning-instruct tw-finance-reasoning-instruct 台灣金融知識的繁體中文推理指令資料集。每一題都有完整的思考過程(think)與可查證的答案(output)。 欄位規格對齊 twinkle-ai/tw-reasoning-instruct-50k。 ⚠️ 使用限制:僅供研究,不得商業使用 本資料集以 CC BY-NC 4.0 授權釋出,僅供學術研究、模型能力探索與方法驗證之用。 不得作商業用途。 包含訓練用於對外營利的模型、包裝為付費產品或服務、 或作為商業交付物的一部分。若有商業需求,請自行重新建置資料並取得合規來源。 這不是財務、稅務、法律或投資建議。 資料中的稅率、費率、法規門檻依 2026 年 (民國 115 年)台灣公開資訊整理,但可能已經過時。任何實際決策前, 請以主管機關公告為準(財政部、金管會、勞動部、衛福部、中央銀行、全國法規資料庫)。 內容為程式化合成,非真實考題逐字收錄。 題目由計算器與模板生成, 並非任何證照考試的原始試題。… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/tw-finance-reasoning-instruct.texttext-generation1K<n<10K0 likes60 downloads15d agoHugging Face20woongstar /ko-finance-asr-corrections ko-finance-asr-corrections Frequency-annotated Korean ASR confusion pairs from finance/stock YouTube. 210 pairs mined from 2,391 videos of auto-captions across 47 channels totalling 1,080.1 hours Each pair carries how often the term was mangled and how often it was said correctly, plus verification provenance. 한국어 금융·주식 유튜브 자동자막에서 실측한 ASR 오인식→교정 쌍입니다. 모든 쌍에 오표기·정답 표기 빈도(→ 용어별 오인식률)와 검증 메타데이터(2-LLM 합의 감사, 승격 티어)가 붙어 있습니다. What makes it different No public… See the full description on the dataset page: https://huggingface.co/datasets/woongstar/ko-finance-asr-corrections.tabulartext-generationn<1K0 likes54 downloads13d agoHugging Face21mzio /aprm-sft-thoughts-snorkel-finance-policy_best-adamw30-lp0 Act-PRM SFT thoughts — snorkel-finance finance Act-PRM (Action Process Reward Models) infers the latent thoughts behind logged, action-only agent demonstrations via an offline EM. For each logged action x in state s we sample G=4 candidate thoughts z, score each by the length-penalized action likelihood reward(z) = p(x | s, z) (len_frac grows with the thought's token length), and mark the best thought (argmax reward). The (thought + action) span is then what downstream SFT / RL… See the full description on the dataset page: https://huggingface.co/datasets/mzio/aprm-sft-thoughts-snorkel-finance-policy_best-adamw30-lp0.tabulartext-generation1K<n<10K0 likes48 downloads26d agoHugging Face22Etherlabs /finance-ops-triage-v0.1 Finance Ops Triage v0.1 dataset The original small, illustrative dataset prepared for Ugo Chukwu's first Unsloth fine-tuning and deployment exercise. The examples were provided during a guided ChatGPT experiment; they are not collected operational transaction records or an independently validated finance policy. Code and experiment report · Model archive Structure Each JSONL row has messages containing system, user, and assistant entries. The assistant content is… See the full description on the dataset page: https://huggingface.co/datasets/Etherlabs/finance-ops-triage-v0.1.texttext-generationn<1K0 likes44 downloads10d agoHugging Face23tejadhith /asc-finance-reasoning-20k Finance Reasoning SFT Dataset — AutoScientist Challenge (20K) A ~20,018-row chain-of-thought finance-reasoning dataset for supervised fine-tuning — the exact training set behind the companion model (Mixtral-8x7B-Instruct LoRA, lead entry E24). Built as the data-recipe submission to the AutoScientist Challenge (Finance, Part 1, Adaption Labs, 2026) and released under CC-BY-4.0. It is a 4,018-row curated seed (Adaptive Data quality 9.3/10, grade A) expanded to ~20K with… See the full description on the dataset page: https://huggingface.co/datasets/tejadhith/asc-finance-reasoning-20k.texttext-generation10K<n<100K0 likes41 downloads3mo agoHugging Face24Italianhype /Blum-Finance-Reasoning BLUM Finance Reasoning Versioned reasoning examples exported from BLUM Engine at revision 973fa4a3579c8b883372e96ed6e7a6e1c99e534a. The dataset uses grouped temporal splits. Records from the same thesis lineage never cross train, validation and test. Secrets, personal identifiers, broker identifiers and unlicensed verbatim sources are excluded. Splits Split Rows Start End test 53 2026-07-09T05:43:21.697064 2026-07-13T23:47:19.530828 train 416… See the full description on the dataset page: https://huggingface.co/datasets/Italianhype/Blum-Finance-Reasoning.texttext-generationn<1K0 likes37 downloads2mo agoHugging Face25BatuhanECB /vibethinker-3b-finance-sftmini-data-public-version VibeThinker-3B Finance-Reader — SFT Training Data · PUBLIC-SAFE subset 🟢 This is vibethinker-3b-finance-sftmini-data-public-version — the redistribution-safe slice of the full vibethinker-3b-finance-sftmini-data dataset, containing only US-government public-domain sources (SEC EDGAR family + Federal Register). Same schema, same pipeline, same teacher — just the legally shareable rows. (Currently private; intended to be made public.) The supervised fine-tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/BatuhanECB/vibethinker-3b-finance-sftmini-data-public-version.texttext-generation1K<n<10K0 likes34 downloads3mo agoHugging Face26likhitjuttada /finance-reasoning-sft-dataset Personal Finance Reasoning Dataset A synthetic instruction-tuning dataset designed to teach language models to reason through personal finance and investing decisions using the mental frameworks from classic books in the genre. The goal is not recall of book content but principled reasoning: the model should apply frameworks to novel situations it has never seen. Source Books Principles were extracted from the following books: The Psychology of Money — Morgan Housel Rich… See the full description on the dataset page: https://huggingface.co/datasets/likhitjuttada/finance-reasoning-sft-dataset.texttext-generationn<1K0 likes32 downloads5mo agoHugging Face27idleengine /Personal-Finance-Queries Dataset Description A curated collection of Reddit posts and top comments focused on personal finance questions. The data is further filtered with the help of LLM-based Voting scores. These scores determine if the query is relevant to a person's financial queries among the other posts of the subreddits. Dataset Structure Columns: category: The sub-domain of personal finance that the query belongs to. subreddit: Source subreddit (string, categorical) query:… See the full description on the dataset page: https://huggingface.co/datasets/idleengine/Personal-Finance-Queries.textquestion-answering10K<n<100K0 likes31 downloads2mo agoHugging Face28williamjmorenor /personal-finance-chatml-dataset Bilingual Personal Finance ChatML Dataset (EN/ES) Dataset Description This dataset is a professionally curated bilingual (English/Spanish) instruction dataset designed for fine-tuning large language models (LLMs) in the domain of personal finance. It is structured in ChatML format and intended for supervised fine-tuning (SFT), domain adaptation, and financial instruction modeling. The dataset is created and reviewed from an accounting perspective, ensuring conceptual… See the full description on the dataset page: https://huggingface.co/datasets/williamjmorenor/personal-finance-chatml-dataset.texttext-generation10K<n<100K0 likes30 downloads7mo agoHugging Face29amalia-llm /amalia-Nemotron-SpecializedDomains-Finance-v1 AMALIA Nemotron-SpecializedDomains-Finance-v1 Version of the nvidia/Nemotron-SpecializedDomains-Finance-v1 dataset used in the AMALIA's Supervised Fine-Tuning stage. This dataset went through a processing pipeline to: Remove entries that reference other LLMs or research labs; Remove the reasoning_content field; Original Dataset: https://huggingface.co/datasets/nvidia/Nemotron-SpecializedDomains-Finance-v1 This dataset is provided as part of the AMALIA project and is… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/amalia-Nemotron-SpecializedDomains-Finance-v1.texttext-generation100K<n<1M0 likes27 downloads3mo agoHugging Face30doraking /nihongo-legal-finance-autoscientist-data nihongo-legal-finance-autoscientist AutoScientist Challenge entry dataset for Japanese expert QA in the language category. Intended Use This dataset is designed for supervised fine-tuning of Japanese assistants that explain legal and financial concepts with uncertainty, source-awareness, and non-advice caveats. Columns instruction: user task context: background information response: target answer rubric: quality expectations category: subdomain… See the full description on the dataset page: https://huggingface.co/datasets/doraking/nihongo-legal-finance-autoscientist-data.textquestion-answeringn<1K0 likes24 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.