asingh15/fineproofs-prm-context-v2-solution
FineProofs PRM Context v2: Solution This arm uses solution context from same_problem rollouts with packing policy whole_solutions_no_truncation. This is one of nine row-matched context variants built from the verified FineProofs rollout collection. Partial-prefix and complete-response targets both use canonical normalized rubric credit derived from clamped points divided by max points. The correct column is only a legacy boolean projection at reward >= 0.5; training uses the… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/fineproofs-prm-context-v2-solution.
FineProofs PRM Context v2: Solution
This arm uses solution context from same_problem rollouts with packing policy whole_solutions_no_truncation.
This is one of nine row-matched context variants built from the verified FineProofs rollout collection. Partial-prefix and complete-response targets both use canonical normalized rubric credit derived from clamped points divided by max points. The correct column is only a legacy boolean projection at reward >= 0.5; training uses the dense reward target.
Contract
- Run ID:
fineproofs_all_qwen35_9b_direct2phase_m32_20260730 - Collection model:
Qwen/Qwen3.5-9B - Collection model revision:
c202236235762e1c871ad0ccb60c8ee5ba337b9a - Context arm:
solution - Context mode:
solution - Context scope:
same_problem - Correctness labels in context:
no - Verification SHA-256:
c5a3b92bd6196d4a33cded2c1a01870d1963d323351d573b369b7f97f3b7f648
The nine arms have identical row keys, labels, rewards, train/validation problem split, and hard endpoint coverage. Held-out problems are removed before splitting and context selection.
Splits
The local val.parquet is published as validation.parquet. dataset_provenance.json, verification.json, and _SUCCESS.json contain the live-input provenance, cross-arm checks, and exact published file fingerprints.
