datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FineProofs-SFT
FineProofs SFT
Dataset Description
FineProofs SFT is a high-quality supervised fine-tuning dataset containing mathematical Olympiad problems paired with chain-of-thought reasoning and formal proofs distilled from DeepSeek-Math-V2. The dataset comprises 7,777 samples (4,300 unique problems) sourced from international Olympiad competitions and Art of Problem Solving (AoPS), each annotated with:
Detailed reasoning traces (thinking content) generated by… See the full description on the dataset page: https://huggingface.co/datasets/lm-provers/FineProofs-SFT.FineProofs-RL
FineProofs RL
Dataset Description
FineProofs RL is a high-quality dataset containing mathematical Olympiad problems and rubrics that are suitable for training models with reinforcement learning. The dataset comprises 5,227 problems sourced from international Olympiad competitions and Art of Problem Solving (AoPS), each annotated with:
Rubrics generated by Gemini-3-Pro (0-7 point scale)
Per-rollout scores and rewards from Qwen/Qwen3-4B-Thinking-2507
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/lm-provers/FineProofs-RL.fineproofs-prm-context-v2-cot-pbudget-xprob
FineProofs PRM Context v2: Cot Pbudget Xprob
This arm uses cot_pbudget context from cross_problem rollouts with packing policy middle_truncated_reasoning_equal_share.
This is one of nine row-matched context variants built from the verified FineProofs rollout collection. Partial-prefix and complete-response targets both use canonical normalized rubric credit derived from clamped points divided by max points. The correct column is only a legacy boolean projection at reward >= 0.5;… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/fineproofs-prm-context-v2-cot-pbudget-xprob.fineproofs-prm-context-v2-solution-label-xprob
FineProofs PRM Context v2: Solution Label Xprob
This arm uses solution_label context from cross_problem rollouts with packing policy whole_solutions_no_truncation.
This is one of nine row-matched context variants built from the verified FineProofs rollout collection. Partial-prefix and complete-response targets both use canonical normalized rubric credit derived from clamped points divided by max points. The correct column is only a legacy boolean projection at reward >= 0.5;… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/fineproofs-prm-context-v2-solution-label-xprob.fineproofs-prm-context-v2-solution
FineProofs PRM Context v2: Solution
This arm uses solution context from same_problem rollouts with packing policy whole_solutions_no_truncation.
This is one of nine row-matched context variants built from the verified FineProofs rollout collection. Partial-prefix and complete-response targets both use canonical normalized rubric credit derived from clamped points divided by max points. The correct column is only a legacy boolean projection at reward >= 0.5; training uses the… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/fineproofs-prm-context-v2-solution.fineproofs-prm-context-v2-cot-pbudget
FineProofs PRM Context v2: Cot Pbudget
This arm uses cot_pbudget context from same_problem rollouts with packing policy middle_truncated_reasoning_equal_share.
This is one of nine row-matched context variants built from the verified FineProofs rollout collection. Partial-prefix and complete-response targets both use canonical normalized rubric credit derived from clamped points divided by max points. The correct column is only a legacy boolean projection at reward >= 0.5;… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/fineproofs-prm-context-v2-cot-pbudget.fineproofs-prm-context-v2-none
FineProofs PRM Context v2: None
No auxiliary rollout context is included.
This is one of nine row-matched context variants built from the verified FineProofs rollout collection. Partial-prefix and complete-response targets both use canonical normalized rubric credit derived from clamped points divided by max points. The correct column is only a legacy boolean projection at reward >= 0.5; training uses the dense reward target.
Contract
Run ID:… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/fineproofs-prm-context-v2-none.fineproofs-prm-context-v2-full-cot-xprob
FineProofs PRM Context v2: Full Cot Xprob
This arm uses full_cot context from cross_problem rollouts with packing policy middle_truncated_reasoning_equal_share.
This is one of nine row-matched context variants built from the verified FineProofs rollout collection. Partial-prefix and complete-response targets both use canonical normalized rubric credit derived from clamped points divided by max points. The correct column is only a legacy boolean projection at reward >= 0.5;… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/fineproofs-prm-context-v2-full-cot-xprob.fineproofs-prm-context-v2-solution-xprob
FineProofs PRM Context v2: Solution Xprob
This arm uses solution context from cross_problem rollouts with packing policy whole_solutions_no_truncation.
This is one of nine row-matched context variants built from the verified FineProofs rollout collection. Partial-prefix and complete-response targets both use canonical normalized rubric credit derived from clamped points divided by max points. The correct column is only a legacy boolean projection at reward >= 0.5; training uses… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/fineproofs-prm-context-v2-solution-xprob.fineproofs-prm-context-v2-solution-label
FineProofs PRM Context v2: Solution Label
This arm uses solution_label context from same_problem rollouts with packing policy whole_solutions_no_truncation.
This is one of nine row-matched context variants built from the verified FineProofs rollout collection. Partial-prefix and complete-response targets both use canonical normalized rubric credit derived from clamped points divided by max points. The correct column is only a legacy boolean projection at reward >= 0.5; training… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/fineproofs-prm-context-v2-solution-label.fineproofs-qwen35-9b-rollouts
FineProofs Qwen3.5-9B Rollouts with GPT-OSS Grades
The complete joined base collection from fineproofs_qwen35_9b_resp81920_gptoss120b_m32_20260722:
5,227 problems from lm-provers/FineProofs-RL
8 Qwen3.5-9B rollouts per problem
41,816 total rollouts
normalized rubric grades and feedback from GPT-OSS-120B
complete policy reasoning, visible proof, and judge reasoning
Joins and integrity
Problems join to response records on problem_id. Policy samples join to grades… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/fineproofs-qwen35-9b-rollouts.fineproofs-prm-context-v2-full-cot
FineProofs PRM Context v2: Full Cot
This arm uses full_cot context from same_problem rollouts with packing policy middle_truncated_reasoning_equal_share.
This is one of nine row-matched context variants built from the verified FineProofs rollout collection. Partial-prefix and complete-response targets both use canonical normalized rubric credit derived from clamped points divided by max points. The correct column is only a legacy boolean projection at reward >= 0.5; training… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/fineproofs-prm-context-v2-full-cot.FineProofs-RL-testfineproofs-qwen35-9b-rollout-view
FineProofs Qwen3.5-9B Rollout View
A curated visualization sample from fineproofs_qwen35_9b_resp81920_gptoss120b_m32_20260722. Each selected problem contributes all eight
Qwen3.5-9B proofs and their GPT-OSS-120B rubric grades. This is a viewing aid, not an evaluation set.
Difficulty strata
FineProofs has no populated categorical difficulty field. difficulty_band is defined here from
the provided extra.reward_mean, an independent empirical success rate over 128… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/fineproofs-qwen35-9b-rollout-view.fineproofs-qwen3-235b-a22b-thinking-2507
FineProofs Qwen3 235B A22B Thinking 2507
FineProofs-compatible training snapshot generated with Qwen/Qwen3-235B-A22B-Thinking-2507.
Snapshot timestamp: 2026-05-27T10:03:12Z
Rows: 4202
The schema exactly matches lm-provers/FineProofs-SFT: problem, reasoning_content, proof, category, competition, gemini-3-pro-grade, qwen3-4b-thinking-reward@128, source, and messages.
Only complete finish_reason=stop generations with non-empty reasoning and proof are included. The grade/reward columns… See the full description on the dataset page: https://huggingface.co/datasets/edbeeching/fineproofs-qwen3-235b-a22b-thinking-2507.fineproofs-gpt-oss-120b
FineProofs GPT-OSS 120B
FineProofs-compatible training snapshot generated with openai/gpt-oss-120b.
Snapshot timestamp: 2026-05-27T10:04:24Z
Rows: 1382
The schema exactly matches lm-provers/FineProofs-SFT: problem, reasoning_content, proof, category, competition, gemini-3-pro-grade, qwen3-4b-thinking-reward@128, source, and messages.
GPT-OSS Harmony outputs are reformatted into Qwen-style assistant messages: <think>\n{reasoning}</think>{proof}. Only complete finish_reason=stop… See the full description on the dataset page: https://huggingface.co/datasets/edbeeching/fineproofs-gpt-oss-120b.FineProofs-Stage-1test-atlas-fineproofs-data
