CoolFace
Datasetpublic

asingh15/prm-sft-polaris-mc

PRM-SFT Polaris MC — Monte-Carlo value (unfiltered) Soft-target process reward model (PRM) training data with per-prefix Monte-Carlo values, in the style of Math-Shepherd. Each row is one prefix of a reasoning trace on a Polaris math problem, labeled with V(prefix)=P(correct∣prefix)=#correct continuations#continuationsV(\text{prefix}) = P(\text{correct} \mid \text{prefix}) = \frac{\#\text{correct… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/prm-sft-polaris-mc.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes61downloads
settings

This repository belongs to asingh15 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameprm-sft-polaris-mc
visibilitypublic
licencenot set
gatedno
ownerasingh15
Account settings
asingh15/prm-sft-polaris-mc · CoolFace