asingh15/fineproofs-prm-context-v2-full-cot
FineProofs PRM Context v2: Full Cot This arm uses full_cot context from same_problem rollouts with packing policy middle_truncated_reasoning_equal_share. This is one of nine row-matched context variants built from the verified FineProofs rollout collection. Partial-prefix and complete-response targets both use canonical normalized rubric credit derived from clamped points divided by max points. The correct column is only a legacy boolean projection at reward >= 0.5; training… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/fineproofs-prm-context-v2-full-cot.
031
