datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
finra-brokercheck-scraper
FINRA BrokerCheck Scraper · Advisors, Firms & Disclosures
Scrape financial advisors, firm affiliations, CRDs, registration scope, and disclosure histories directly from FINRA BrokerCheck API into clean dataset rows.
Rows in this dataset
1,430
Fields
22
Collector runs behind it
50
Most recent observation
2026-08-03
What this is
Every row here was returned by a real run of a public collector. Nothing is generated from a
template over a… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/finra-brokercheck-scraper.FinRAG-GRPO
FinRAG-GRPO Preference Dataset
A Chinese-language preference dataset for training Reasoning Reward Models (ReasRM) via GRPO-based reinforcement learning.
🚧 This dataset is actively maintained and will be expanded with additional domains and languages over time.
Dataset Summary
This dataset contains pairwise preference samples designed to train a reward model that reasons before judging — the model generates an evaluation rationale before outputting a preference label… See the full description on the dataset page: https://huggingface.co/datasets/SamWang0405/FinRAG-GRPO.finrag-eval
FinRAG-Eval
147 questions over five S&P-500 10-K filings (563 printed pages, 1,679 chunks),
where every answerable question carries character-offset gold spans into the
canonical markdown — not just a gold string.
The point of this dataset is not its size. It is that it survived an audit
trail instead of a vibe check, and the trail is published with it.
Why another financial QA set
Most synthetic eval sets are generated once and trusted. This one was generated… See the full description on the dataset page: https://huggingface.co/datasets/ChihebLovesAi/finrag-eval.
