datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bfsi-bench
BFSI-Bench
BFSI-Bench is a benchmark for testing how well language models answer questions about India’s banking, financial services, and insurance (BFSI) rules.
In this domain, the correct answer often depends on circulars and regulations that change frequently, and the official sources (sites like RBI, SEBI, and IRDAI) can be hard to find, parse, and keep current. BFSI-Bench measures five capability areas:
Jurisdiction-Aware Compliance: Disambiguate to the Indian context, or… See the full description on the dataset page: https://huggingface.co/datasets/ground-truth/bfsi-bench.indian-regulatory-bfsi-benchmark-v1
Indian Regulatory BFSI Benchmark v1
A 60-question, hand-curated, openly licensed evaluation set for
extractive question answering over Indian financial regulation -
specifically Reserve Bank of India (RBI) Master Directions and
Securities and Exchange Board of India (SEBI) Master Circulars.
60 questions, 30 RBI / 30 SEBI
30 numeric / named-fact extraction (tier 2) + 30 heading-bound passage
questions (tier 3)
22 distinct source PDFs from a document-disjoint held-out split of… See the full description on the dataset page: https://huggingface.co/datasets/udit6969/indian-regulatory-bfsi-benchmark-v1.bfsi-transaction-triage-curated-600
🚀 BFSI / FinTech Tier-1 Autonomous Transaction Triage & Regulatory Disputes
This dataset contains 600 curated training records with in-depth, verbose 4-phase <Thinking> Chain-of-Thought reasoning, 100 frozen evaluation benchmark samples, and 50 frozen regression verification samples formatted in standard ChatML (messages) and Prompt-Target pairs, strictly following the Pioneer / Prometheus research paper 3-slice curriculum design.
📊 Dataset Composition & 3-Slice… See the full description on the dataset page: https://huggingface.co/datasets/StarsMakeGalaxy/bfsi-transaction-triage-curated-600.prometheus-bfsi-tier1-triage
Pioneer BFSI / Fintech Tier-1 Autonomous Triage Dataset
This dataset contains high-grade multi-turn conversational traces (ChatML format) curated according to the Pioneer paper data curation methodology for training and evaluating specialized 8B Small Language Models (SLMs) in the Banking, Financial Services, and Insurance (BFSI) vertical.
Dataset Structure
Train Split (train): 350 traces
75% Gold Standard Tasks: Standard operational workflows across 30… See the full description on the dataset page: https://huggingface.co/datasets/StarsMakeGalaxy/prometheus-bfsi-tier1-triage.bfsi-llm-eval
BFSI LLM Behavioral Evaluation Dataset
A structured evaluation dataset of 611 prompts designed to test LLM behavior across Banking, Financial Services, and Insurance (BFSI) domains. Covers hallucination detection, consistency, robustness, and safety evaluation.
Dataset Summary
Metric
Value
Total records
611
Dimensions
4 (hallucination, consistency, robustness, safety)
Subdimensions
15
Source domain
Banking
Geographies
USA, Canada
Languages
English… See the full description on the dataset page: https://huggingface.co/datasets/manulife/bfsi-llm-eval.bfsi-banking-exam-v0bfsi_QA_alpaca_dataset
BFSI Alpaca Dataset
Dataset Summary
This dataset contains synthetic Banking, Financial Services, and Insurance (BFSI) customer support conversations in Alpaca format.[file:40]Each record is a short, standardized query–response pair designed for training lightweight call center assistants that prioritize safety and compliance.[file:40]
Languages
English
Use Cases
Training or fine-tuning small language models for:
Loan, EMI, and disbursement queries… See the full description on the dataset page: https://huggingface.co/datasets/VimalAntony/bfsi_QA_alpaca_dataset.
