CoolFace
Datasetpublic

asingh15/prm-sft-polaris-mc

PRM-SFT Polaris MC — Monte-Carlo value (unfiltered) Soft-target process reward model (PRM) training data with per-prefix Monte-Carlo values, in the style of Math-Shepherd. Each row is one prefix of a reasoning trace on a Polaris math problem, labeled with V(prefix)=P(correct∣prefix)=#correct continuations#continuationsV(\text{prefix}) = P(\text{correct} \mid \text{prefix}) = \frac{\#\text{correct… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/prm-sft-polaris-mc.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes61downloads
4 commits on main
47bb3d22mo ago

Add train split

asingh15
84bf9af2mo ago

Add val split

asingh15
0fe2a932mo ago

Add dataset card

asingh15
9ebb31d2mo ago

initial commit

asingh15