datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mmlu_olmo_contaminationmmlu_olmo3_contaminationtoaa_benchmark_contaminationContaminationQAfinemath_contamination_report
🧹
This dataset contains suspected benchmark-contaminated pages that were removed from the FineMath dataset.
GameBoyWorlds-Playthrough-Contamination-Check
Documentation Density Predicts Benchmark Performance: A Contamination Probe Using Official and Fan-Made Pokémon Titles
Abstract. Benchmark scores on well-documented domains conflate what a model can reason about with what it has memorised. We construct a 400-item short-answer benchmark over four Game Boy–era Pokémon titles, paired by surface form but differing sharply in corpus footprint: two officially published games (Red, Crystal) and two community ROM hacks (Brown, Prism).… See the full description on the dataset page: https://huggingface.co/datasets/DJ-Research/GameBoyWorlds-Playthrough-Contamination-Check.lm-eval-results-Contamination-contaminated_proof_7b_v1.0_safetensor-private
Dataset Card for Evaluation run of Contamination/contaminated_proof_7b_v1.0_safetensor
Dataset automatically created during the evaluation run of model Contamination/contaminated_proof_7b_v1.0_safetensor
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Contamination-contaminated_proof_7b_v1.0_safetensor-private.Chinese-DeepSeek-R1-Distill-data-110k-contamination-report
Contamination Report — Congliu/Chinese-DeepSeek-R1-Distill-data-110k
What this is
A row-level audit of Congliu/Chinese-DeepSeek-R1-Distill-data-110k (revision
8520b649430617c2be4490f424d251d09d835ed3) for exact 13-gram overlap with standard benchmark test sets
(gsm8k, hellaswag, humaneval, mmlu). This is not a filtered copy of the source — it's a new
artifact: a list of which rows overlap which benchmark, plus summary statistics, so anyone
training on the source… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/Chinese-DeepSeek-R1-Distill-data-110k-contamination-report.details_Contamination__contaminated_proof_7b_v1.0_safetensor
Dataset Card for Evaluation run of Contamination/contaminated_proof_7b_v1.0_safetensor
Dataset automatically created during the evaluation run of model Contamination/contaminated_proof_7b_v1.0_safetensor on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Contamination__contaminated_proof_7b_v1.0_safetensor.africa-synth-mental-health-soil-contamination-urban-all
Soil Contamination & Urban Agriculture Safety (SSA) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-mental-health-soil-contamination-urban-all.lm-eval-results-Contamination-contaminated_proof_7b_v1.0-private
Dataset Card for Evaluation run of Contamination/contaminated_proof_7b_v1.0
Dataset automatically created during the evaluation run of model Contamination/contaminated_proof_7b_v1.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Contamination-contaminated_proof_7b_v1.0-private.high-quality-english-sentences-contamination-report
Contamination Report — agentlans/high-quality-english-sentences
What this is
A row-level audit of agentlans/high-quality-english-sentences (revision
main) for exact 13-gram overlap with standard benchmark test sets
(gsm8k, hellaswag, humaneval, mmlu). This is not a filtered copy of the source — it's a new
artifact: a list of which rows overlap which benchmark, plus summary statistics, so anyone
training on the source can decide how to handle it.… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/high-quality-english-sentences-contamination-report.oasst1-contamination-report
Contamination Report — OpenAssistant/oasst1
What this is
A row-level audit of OpenAssistant/oasst1 (revision
fdf72ae0827c1cda404aff25b6603abec9e3399b) for exact 13-gram overlap with standard benchmark test sets
(gsm8k, hellaswag, humaneval, mmlu). This is not a filtered copy of the source — it's a new
artifact: a list of which rows overlap which benchmark, plus summary statistics, so anyone
training on the source can decide how to handle it.… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/oasst1-contamination-report.llm-benchmark-contamination-archive
LLM Benchmark Contamination — 项目档案
这是一个研究项目的完整工作档案,不是数据集。项目主题:当前 LLM(GPT / Claude / Qwen / DeepSeek,尤其 Qwen 小模型)对 GSM8K / MATH 等测试集的记忆——把"有没有记住"拆成"记住多少 / 值多少分 / 能不能测出来"三个量,规划了四条研究路线(可检测性相图 / 合成数据洗白 / 距离-效应 / 网页中介分解),目标 ICML 2027。
档案由 Claude Code(Claude Fable 5)会话产生,包含:
目录
内容
report/
研究路线报告(自包含 HTML)+ 生成它的全部输入与脚本(事实简报、统一口径、章节 JSON、构建脚本)
claude/
完整 Claude Code 会话转录(主会话 jsonl)、69 个子代理与 2 个 multi-agent workflow 的全部记录、workflow 脚本、本项目的 prompt history
RESTORE.md… See the full description on the dataset page: https://huggingface.co/datasets/daipath/llm-benchmark-contamination-archive.openshape-contamination-axes-data
OpenShape Benchmark Contamination Artifacts (BMVC 2026, Anonymous Release)
Released anonymously for BMVC 2026 double-blind review.
Authorship and provenance will be revealed on acceptance.
This release accompanies a BMVC 2026 submission analyzing benchmark
contamination in 3D representation learning. It contains the training
prune masks, training/eval configs, eval-time per-step metrics, training
logs, NN-proxy predictions, and the best-LVIS checkpoint for the
counterfactual… See the full description on the dataset page: https://huggingface.co/datasets/pkhw2023/openshape-contamination-axes-data.forgetting-contamination-arc-easyThis dataset is a deduplicated subset of ARC-Easy, as used in the paper How Much Can We Forget about Data Contamination?. The deduplication was performed using this script.
The data fields are the same as in https://huggingface.co/datasets/allenai/ai2_arc, with the additional "split-id" column that can be used to partition the benchmark questions into different subsets.
The dataset can be used as a plug-in replacement for ARC-Easy if you want to work with the deduplicated benchmark questions.… See the full description on the dataset page: https://huggingface.co/datasets/sbordt/forgetting-contamination-arc-easy.skyworks-rewardbench-contamination
Reward Bench overlap with Skyworks Preferences 80k
This dataset includes the overlap between the SkyWorks prompts, which are being used to train top reward models, with the original test set.
More information found here.
forgetting-contamination-winograndeThis dataset is a deduplicated subset of the XL train split of WinoGrande, as used in the paper How Much Can We Forget about Data Contamination?. The deduplication was performed using this script.
The data fields are the same as in https://huggingface.co/datasets/allenai/winogrande, with the additional "split-id" column that can be used to partition the benchmark questions into different subsets.
The dataset can be used as a plug-in replacement if you want to work with the deduplicated… See the full description on the dataset page: https://huggingface.co/datasets/sbordt/forgetting-contamination-winogrande.forgetting-contamination-social_i_qaThis dataset is a deduplicated subset of the train split of Social IQa, as used in the paper How Much Can We Forget about Data Contamination?. The deduplication was performed using this script.
The data fields are the same as in https://huggingface.co/datasets/allenai/social_i_qa, with the additional "split-id" column that can be used to partition the benchmark questions into different subsets.
The dataset can be used as a plug-in replacement if you want to work with the deduplicated benchmark… See the full description on the dataset page: https://huggingface.co/datasets/sbordt/forgetting-contamination-social_i_qa.mmlu_confusing_options_contamination_testforgetting-contamination-boolqThis dataset is a deduplicated subset of the validation split of BoolQ, as used in the paper How Much Can We Forget about Data Contamination?. The deduplication was performed using this script.
The data fields are the same as in https://huggingface.co/datasets/google/boolq, with the additional "split-id" column that can be used to partition the benchmark questions into different subsets.
The dataset can be used as a plug-in replacement for BoolQ if you want to work with the deduplicated… See the full description on the dataset page: https://huggingface.co/datasets/sbordt/forgetting-contamination-boolq.forgetting-contamination-mmluThis dataset is a deduplicated subset of the test split of mmlu, as used in the paper How Much Can We Forget about Data Contamination?. The deduplication was performed using this script.
The data fields are the same as in https://huggingface.co/datasets/cais/mmlu, with the additional "split-id" column that can be used to partition the benchmark questions into different subsets.
The dataset can be used as a plug-in replacement if you want to work with the deduplicated benchmark questions.
For… See the full description on the dataset page: https://huggingface.co/datasets/sbordt/forgetting-contamination-mmlu.forgetting-contamination-piqaThis dataset is a deduplicated subset of the train split of PiQA, as used in the paper How Much Can We Forget about Data Contamination?. The deduplication was performed using this script.
The data fields are the same as in https://huggingface.co/datasets/ybisk/piqa, with the additional "split-id" column that can be used to partition the benchmark questions into different subsets.
The dataset can be used as a plug-in replacement if you want to work with the deduplicated benchmark questions.
For… See the full description on the dataset page: https://huggingface.co/datasets/sbordt/forgetting-contamination-piqa.curriculum-contamination-coherent-wrong
Curriculum Contamination Coherent-Wrong-Label Tasks (Stage 1)
This private pre-release contains a controlled mathematical task bank for
studying whether self-evolving curriculum selectors can admit an objectively
wrong answer when repeated solver samples form a coherent majority.
Dataset contents
data/unique_tasks.jsonl: 240 unique, exact-answer mathematical tasks.
data/label_twins.jsonl: 480 surface-identical task/reference pairs: one
objectively correct… See the full description on the dataset page: https://huggingface.co/datasets/MATKKK/curriculum-contamination-coherent-wrong.water-contamination-datawater-contamination-kgforgetting-contamination-hellaswagThis dataset is a deduplicated subset of the validation split of hellaswag, as used in the paper How Much Can We Forget about Data Contamination?. The deduplication was performed using this script.
The data fields are the same as in https://huggingface.co/datasets/Rowan/hellaswag, with the additional "split-id" column that can be used to partition the benchmark questions into different subsets.
The dataset can be used as a plug-in replacement for hellaswag if you want to work with the… See the full description on the dataset page: https://huggingface.co/datasets/sbordt/forgetting-contamination-hellaswag.defendable-pain-benchmark-contamination-pain-v0.1
Benchmark Contamination Pain Receipt
"the leaked answer" — Mr. Defendable
A free pain-receipt dataset from the DefendableOS ecosystem. 1 rows · ready to read · all cited or graded · CC-BY-4.0.
Part of the 100-pack — 100 free pain-receipt datasets dropped from the Defendable Bakery to the open AI-trust community. Different theme per dataset. Same operator voice across all of them.
Tribunal begins before training. No proof, no honey. To the shed.
What's in here
1… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-benchmark-contamination-pain-v0.1.euv-collector-contamination-drift-detection-v0.1
Dataset purpose
This dataset detects early contamination drift in EUV collector mirrors.
A lithography system fails gradually.The first signal is not throughput collapse.It is coherence loss between:
gas stabilitymirror reflectivityEUV transmitted power
When these stop moving together, contamination is underway.
Task
Given system metrics, output:
coherence_drift_scoredrift_flag
drift_flag = 1 means contamination drift has begundrift_flag = 0 means system remains… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/euv-collector-contamination-drift-detection-v0.1.mathqa_confusing_options_contamination_test
