datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indian-regulatory-bfsi-benchmark-v1
Indian Regulatory BFSI Benchmark v1
A 60-question, hand-curated, openly licensed evaluation set for
extractive question answering over Indian financial regulation -
specifically Reserve Bank of India (RBI) Master Directions and
Securities and Exchange Board of India (SEBI) Master Circulars.
60 questions, 30 RBI / 30 SEBI
30 numeric / named-fact extraction (tier 2) + 30 heading-bound passage
questions (tier 3)
22 distinct source PDFs from a document-disjoint held-out split of… See the full description on the dataset page: https://huggingface.co/datasets/udit6969/indian-regulatory-bfsi-benchmark-v1.data_distillation_longvila_reformated_alldata_distillation_longvila_reformated_all_and_video_r1data_distillation_rl_2_and_video_r1video_r1bfsi_QA_alpaca_dataset
BFSI Alpaca Dataset
Dataset Summary
This dataset contains synthetic Banking, Financial Services, and Insurance (BFSI) customer support conversations in Alpaca format.[file:40]Each record is a short, standardized query–response pair designed for training lightweight call center assistants that prioritize safety and compliance.[file:40]
Languages
English
Use Cases
Training or fine-tuning small language models for:
Loan, EMI, and disbursement queries… See the full description on the dataset page: https://huggingface.co/datasets/VimalAntony/bfsi_QA_alpaca_dataset.
