asingh15/prm-mc-value-context-solution-label-xprob-stratified-v1
PRM MC Value Context: Solution Label Xprob (Stratified v1) Solutions from a different problem in the same split are tagged correct or incorrect. This repository is one arm of a five-way, row-matched process reward model ablation. The target is a Monte Carlo probability of eventual rollout success for a partial solution prefix; only the auxiliary context changes between arms. Arm configuration Context mode: solution_label Cross-problem context: yes Context… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/prm-mc-value-context-solution-label-xprob-stratified-v1.
PRM MC Value Context: Solution Label Xprob (Stratified v1)
Solutions from a different problem in the same split are tagged correct or incorrect.
This repository is one arm of a five-way, row-matched process reward model ablation. The target is a Monte Carlo probability of eventual rollout success for a partial solution prefix; only the auxiliary context changes between arms.
Arm configuration
- Context mode:
solution_label - Cross-problem context:
yes - Context correctness labels:
yes - Split policy: deterministic difficulty-stratified v1
- Split manifest SHA-256:
c7287b917d3ee6eca482051c5256610b1b193e8f7bcd7800a28bfaed21351637
Cross-problem examples are selected within the query row's split. The held-out problems therefore do not leak into training context.
Splits
The local val.parquet artifact is published as validation.parquet. dataset_provenance.json, split_manifest.json, and verification.json record the build and five-arm integrity checks.
