datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
toaa_benchmark_contaminationfinemath_contamination_report
🧹
This dataset contains suspected benchmark-contaminated pages that were removed from the FineMath dataset.
skyworks-rewardbench-contamination
Reward Bench overlap with Skyworks Preferences 80k
This dataset includes the overlap between the SkyWorks prompts, which are being used to train top reward models, with the original test set.
More information found here.
helpsteer2-rewardbench-contaminationpaper2env-contamination
paper2env-contamination
A 60-paper held-out benchmark spanning Jan 2024 → Apr 2026, designed to
measure how each evaluated model's reward depends on whether the paper was
published before or after the model's training cutoff. Six release-aligned
buckets, 10 papers each.
Configs
tasks (60 rows)
One row per paper. Schema mirrors
thibble/paper2env's
paperbench config — paper_md, paper_rubric, task_md, verify_sh,
generate_artifact_sh, patch, plus an artifact_path… See the full description on the dataset page: https://huggingface.co/datasets/thibble/paper2env-contamination.contamination_report_combinedhellaswag_wikihow_olmo_contamination
