CoolFace
Datasetpublic

asingh15/fineproofs-qwen35-9b-rollout-view

FineProofs Qwen3.5-9B Rollout View A curated visualization sample from fineproofs_qwen35_9b_resp81920_gptoss120b_m32_20260722. Each selected problem contributes all eight Qwen3.5-9B proofs and their GPT-OSS-120B rubric grades. This is a viewing aid, not an evaluation set. Difficulty strata FineProofs has no populated categorical difficulty field. difficulty_band is defined here from the provided extra.reward_mean, an independent empirical success rate over 128… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/fineproofs-qwen35-9b-rollout-view.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes17downloads
Dataset Card

FineProofs Qwen3.5-9B Rollout View

A curated visualization sample from fineproofs_qwen35_9b_resp81920_gptoss120b_m32_20260722. Each selected problem contributes all eight Qwen3.5-9B proofs and their GPT-OSS-120B rubric grades. This is a viewing aid, not an evaluation set.

Difficulty strata

FineProofs has no populated categorical difficulty field. difficulty_band is defined here from the provided extra.reward_mean, an independent empirical success rate over 128 prior evaluations:

SplitProvided success rate
very_hard[0.00, 0.10)
hard[0.10, 0.30)
medium[0.30, 0.60)
easy[0.60, 0.90)
very_easy[0.90, 1.00]

There are 15 selected problems and 120 rows. Every problem retains all eight rollouts. Selection covers a representative problem, a high within-problem grade-contrast problem, and a large baseline-vs-9B disagreement in each stratum.

Aggregate comparison

Across all 5,227 problems:

  • —Provided success rate: 0.399
  • —9B full-credit success rate: 0.387
  • —Difference (9B - provided): -0.012
  • —Pearson correlation: 0.806
  • —9B mean normalized rubric grade: 0.557

The binary comparison uses gpt_oss_grade == 1. qwen35_9b_mean_grade is reported separately because partial-credit rubric quality is not the same quantity as success probability.

Interactive view: https://huggingface.co/spaces/asingh15/fineproofs-rollout-explorer

Source problems: lm-provers/FineProofs-RL. Generated proofs: Qwen/Qwen3.5-9B. Grades: openai/gpt-oss-120b.