CoolFace
Datasetpublic

asingh15/fineproofs-prm-context-v2-full-cot

FineProofs PRM Context v2: Full Cot This arm uses full_cot context from same_problem rollouts with packing policy middle_truncated_reasoning_equal_share. This is one of nine row-matched context variants built from the verified FineProofs rollout collection. Partial-prefix and complete-response targets both use canonical normalized rubric credit derived from clamped points divided by max points. The correct column is only a legacy boolean projection at reward >= 0.5; training… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/fineproofs-prm-context-v2-full-cot.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes31downloads
7 commits on main
1ba07bd2mo ago

Mark verified FineProofs v2 context arm complete: full_cot

asingh15
5b3a2112mo ago

Stage verified FineProofs v2 context data: full_cot

asingh15
ec85c8d2mo ago

Mark verified FineProofs v2 context arm complete: full_cot

asingh15
df2b4c92mo ago

Stage verified FineProofs v2 context data: full_cot

asingh15
da329eb2mo ago

Mark verified FineProofs v2 context arm complete: full_cot

asingh15
984e7f92mo ago

Stage verified FineProofs v2 context data: full_cot

asingh15
7ff33862mo ago

initial commit

asingh15