datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quotient-margins-reward-models
Quotient Margins for Reward Models — data release
Artifacts backing the paper Measure Confidence on Decisions, Not Samples: Quotient Margins for
Reward Models.
The short version of the paper. Reward models pick the best of N sampled responses, but
their confidence is normally read off the reward gap between the top two samples. When
several candidates express the same underlying behaviour, that gap is a within-class spacing and
its predictive signal cancels. Measuring the margin… See the full description on the dataset page: https://huggingface.co/datasets/matCercola18/quotient-margins-reward-models.reward-modeling-papers
Reward Modeling Papers — FineSet
A research-paper dataset on Reward Modeling Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on Reward Modeling Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓
Why this dataset
Quality-scored: quality_score… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/reward-modeling-papers.
