fineproofs
FineProofs-SFT
FineProofs SFT
Dataset Description
FineProofs SFT is a high-quality supervised fine-tuning dataset containing mathematical Olympiad problems paired with chain-of-thought reasoning and formal proofs distilled from DeepSeek-Math-V2. The dataset comprises 7,777 samples (4,300 unique problems) sourced from international Olympiad competitions and Art of Problem Solving (AoPS), each annotated with:
Detailed reasoning traces (thinking content) generated by… See the full description on the dataset page: https://huggingface.co/datasets/lm-provers/FineProofs-SFT.FineProofs-RL
FineProofs RL
Dataset Description
FineProofs RL is a high-quality dataset containing mathematical Olympiad problems and rubrics that are suitable for training models with reinforcement learning. The dataset comprises 5,227 problems sourced from international Olympiad competitions and Art of Problem Solving (AoPS), each annotated with:
Rubrics generated by Gemini-3-Pro (0-7 point scale)
Per-rollout scores and rewards from Qwen/Qwen3-4B-Thinking-2507
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/lm-provers/FineProofs-RL.fineproofs-branch-rollout-distribution
FineProofs and Polaris Rollout Distributions
Per-prefix aggregate statistics for run
fineproofs_all_qwen35_9b_direct2phase_m32_20260730.
Each row represents one collected response prefix and summarizes exactly 32 fresh
Qwen3.5-9B continuations graded by GPT-OSS-120B against the problem rubric. Every
continuation and endpoint reward is reconstructed exactly as clamped
points / max_points; the rounded journal grade is used only for a consistency
audit with absolute tolerance… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/fineproofs-branch-rollout-distribution.fineproofs-prm-context-v2-cot-pbudget-xprob
FineProofs PRM Context v2: Cot Pbudget Xprob
This arm uses cot_pbudget context from cross_problem rollouts with packing policy middle_truncated_reasoning_equal_share.
This is one of nine row-matched context variants built from the verified FineProofs rollout collection. Partial-prefix and complete-response targets both use canonical normalized rubric credit derived from clamped points divided by max points. The correct column is only a legacy boolean projection at reward >= 0.5;… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/fineproofs-prm-context-v2-cot-pbudget-xprob.reasoning-sft-stem-reasoning-complex-FineProofs-126K
reasoning-sft-stem-reasoning-complex-FineProofs-126K
Combined converted dataset from two sources:
lm-provers/FineProofs-SFT (all config, 7.78k) — Mathematical Olympiad problems with chain-of-thought reasoning distilled from DeepSeek-Math-V2
galaxyMindAiLabs/stem-reasoning-complex (~118k) — STEM reasoning across Biology, Mathematics, Physics, Chemistry and Code
Format
Each row has three columns:
input — list of dicts [{"role": "user", "content": "..."}]
response —… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/reasoning-sft-stem-reasoning-complex-FineProofs-126K.fineproofs-prm-context-v2-solution-label-xprob
FineProofs PRM Context v2: Solution Label Xprob
This arm uses solution_label context from cross_problem rollouts with packing policy whole_solutions_no_truncation.
This is one of nine row-matched context variants built from the verified FineProofs rollout collection. Partial-prefix and complete-response targets both use canonical normalized rubric credit derived from clamped points divided by max points. The correct column is only a legacy boolean projection at reward >= 0.5;… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/fineproofs-prm-context-v2-solution-label-xprob.
