datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
icml-2026-reproductions
ICML 2026 Agent Reproducibility Challenge — Logbooks
A living mirror of all public reproduction logbooks from the ICML 2026 Agent Reproducibility Challenge.
Agents attempt to reproduce claims from ICML 2026 papers. Each logbook records the reproduction process, evidence, and verdict for each claim.
Structure
├── papers.json # All 6341 ICML 2026 papers (metadata)
├── logbooks.csv # Main index: one row per logbook (agent × paper ×… See the full description on the dataset page: https://huggingface.co/datasets/qy2100/icml-2026-reproductions.ainize-dart100-reproduction-20260911
DART task datasets for Ainize reproduction
This bundle preserves 100 distinct task configurations from the existing DART-derived datasets. Each JSONL file is byte-identical to its authenticated Ainize canonical download. The manifest links configuration names, row counts, SHA256 values and Ainize dataset IDs.
Source: Financial Supervisory Service Open DART, https://opendart.fss.or.kr/. These are historical disclosure-derived answers, not a promise of current company information… See the full description on the dataset page: https://huggingface.co/datasets/Minhyun/ainize-dart100-reproduction-20260911.pq_reproduction
Reproduction of Parquet files in blog post
This dataset contains a reproduction of the Parquet files used in the blog post Parquet Content-Defined Chunking by Krisztian Szucs.
The dataset kszucs/pq contains part of the files, but not all of them.
In this dataset, each Parquet example is available in 8 versions:
two compressions: none and snappy,
with content-defined chunking (CDC) enabled or disabled (CDC: this feature ensures that the columns are consistently getting chunked into… See the full description on the dataset page: https://huggingface.co/datasets/severo/pq_reproduction.SkyRL-SQL-Reproductionniaf-reproductionhearts-reproduction-traces
Agent traces
Agent sessions published from a Trackio Logbook.
openvla-libero-reproductionreproduction_qwen235b_philosophyJFLD_NLP_2024_proceeding_reproduction
Dataset Card for "JFLD_NLP_2024_proceeding_reproduction"
See here for the details of this corpus.
For the whole of the project, see our project page.
More Information needed
OptMiner-Reproduction
Opt-Miner Reproduction — ICML 2026 Reproducibility Challenge
Submission #9232 | OpenReview: GH9qE7sRPzReproduced by: Nikhil DhakaGitHub: ernikhildhaka-arch/OptMiner-ReproductionOrganization: ICML-2026-agent-repro
Paper
Opt-Miner: Empowering Information-Seeking Agent with Tree-Guided Data Synthesis for Optimization ModelingInternational Conference on Machine Learning (ICML), 2026
Claims Verified
Claim
Description
Status
1
Qwen3-8B… See the full description on the dataset page: https://huggingface.co/datasets/ernikhil411/OptMiner-Reproduction.icml-2026-30204-reproduction-artifacts
ICML 2026 #30204 reproduction artifacts
Artifacts for Networked Information Aggregation for Binary Classification (arXiv:2605.01082; OpenReview mrtg4NmvAe).
Downloads
icml-2026-30204-repro-bundle.tar.gz — complete 50-file reproduction bundle
MANIFEST.sha256 — checksums of files inside the unpacked bundle
poster.pdf — gate-verified 24x36 inch reproduction poster
summary.json — authoritative aggregate numerical results
Evidence_Memo.md — claim-by-claim source and… See the full description on the dataset page: https://huggingface.co/datasets/FloorIsAwake/icml-2026-30204-reproduction-artifacts.yourbench_reproduction_deepseekr1_biologyyourbench_reproduction_deepseekr1_chemistryreproduction_qwen235b_computersciencesyntra-reproduction
Syntra — Independent Blinded Reproduction
An independent, blinded reproduction of the central empirical claim of Syntra: A
Framework for Wiser and More Human AI (H. Axelsson) — that routing a prompt through an
ethics-first orchestration layer (Valon → Modi → Drift) makes a language model reason
more "wisely" than the same model answering directly. We reproduce the paper's own
three-judge × five-dimension LLM-scoring protocol twice: once with the paper-era judge
models, and once… See the full description on the dataset page: https://huggingface.co/datasets/Infektyd/syntra-reproduction.reproduction_o4mini_computerscienceMelodic_pattern_reproduction_performances_gradingyourbench_reproduction_o4mini_computersciencereproduction_deepseekr1_chemistryreproduction_deepseekr1_physicsyourbench_mmlu_reproduction_nutritionreproduction_qwen235b_historyreproduction_qwen235b_physicsreproduction_deepseekr1_historyyourbench_mmlu_reproduction_international_lawyourbench_reproduction_o4mini_physicsreproduction_o4mini_psychologyDataset_Distillation_ReproductionThis dataset is for keeping track on the reproduction of some Dataset Distillation methods. The format of each folder is like dataset/ipc{n}/class_name.
Reference:
SRe2L:https://github.com/VILA-Lab/SRe2L/tree/main/SRe2L
WMDD:https://github.com/Liu-Hy/WMDD
CVDD:https://github.com/Jiacheng8/CV-DD
G-VBSM:https://github.com/shaoshitong/G_VBSM_Dataset_Condensation
reproduction_deepseekr1_computersciencereproduction_g3_mini_computerscience
