datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
repro-scriptsverdicts
Logbook verdicts
Per-logbook claim verdicts produced by the logbook-judge Space. See verdicts.json.
REPRO-Bench
REPRO-Bench: Can Agentic AI Systems Assess the Reproducibility of Social Science Research?
This repository contains the REPRO-Bench dataset, introduced in the paper REPRO-Bench: Can Agentic AI Systems Assess the Reproducibility of Social Science Research?.
REPRO-Bench is a novel benchmark designed to evaluate the capability of agentic AI systems in automating the assessment of social science paper reproducibility. It addresses limitations of existing benchmarks by providing a… See the full description on the dataset page: https://huggingface.co/datasets/chuxuan/REPRO-Bench.challenge
The ICML 2026 reproduction challenge ended on August 2, 2026 (23:59 AoE).
New or updated logbooks are no longer judged and prize consideration is closed.
The instructions below remain for reference; published logbooks stay online.
Reproducing ICML 2026 — Challenge Guide (for agents)
You are a coding agent contributing to a community effort organized by Hugging Face and AlphaXiv to reproduce the major claims of every ICML 2026 paper.
Task
Your task is to… See the full description on the dataset page: https://huggingface.co/datasets/ICML-2026-agent-repro/challenge.InternUtopia-repro-assets
RoboAssemblyBench
RoboAssemblyBench is a reproduction branch of InternUtopia focused on atomic-skill based robotic assembly. The current checkpoint contains a dual-UR5e + Robotiq 2F-85 Fabrica plumbers-block task:
fabrica_plumbers_block_ur5e_right_base_prepare
The task stages part 2 with the right arm, then uses the left arm to place part 0 into the staged part-2 slot, stack part 3, and insert parts 4 and 1 into the remaining holes.
Quick Preview
Example rollout… See the full description on the dataset page: https://huggingface.co/datasets/baiyu858/InternUtopia-repro-assets.flashvsr-repro-outputs-v2-part1
FlashVSR 复现实验输出 — part1
FlashVSR 复现及 KV cache 驱逐策略消融实验的逐帧推理输出,以 FFV1 无损编码归档。
本 repo 是全部结果的第 1/2 部分。
内容
reds_bscv_dbl_clean_full_sliding
reds_bscv_dbl_clean_full_sliding_kv10
reds_bscv_dbl_clean_full_sliding_kv6
reds_bscv_dbl_full_gate
reds_bscv_dbl_full_gate_kv10
reds_bscv_dbl_full_gate_kv6
reds_bscv_dbl_full_gate_lfres
reds_bscv_dbl_full_gate_lfres_frame
reds_bscv_dbl_full_gate_lfres_frame_kv10
reds_bscv_dbl_full_gate_lfres_frame_kv6… See the full description on the dataset page: https://huggingface.co/datasets/victorzhu30/flashvsr-repro-outputs-v2-part1.ToolBench_reproductionreproducibility-datakit-technical-reportSampled parquet for the gridfm-datakit technical report diversity plots.
How to reproduce: scripts/datakit_report/README.md on branch genco-paper-repro.
shapleymcg-qwen3-30b-a3b-reproducibility
ShapleyMCG Qwen3-30B-A3B reproducibility artifacts
This dataset preserves the calibration statistics, exact corrected-R10
EXL3/MCG candidates, source and corpus identities, BF16 teacher/student logits,
tokenwise KLD, allocations, attribution ledgers, hashes, and publication
receipts for the Qwen3-30B-A3B experiments in
brandonmmusic-max/shapleymcg.
The complete cross-checkpoint
results ledger
and
method specification
distinguish the predecessor routed-p2 allocator from the full… See the full description on the dataset page: https://huggingface.co/datasets/brandonmusic/shapleymcg-qwen3-30b-a3b-reproducibility.repro-organic-data-72Bchemrxiv_reprocessedflashvsr-repro-outputs-part3
FlashVSR 复现实验输出 — part3
FlashVSR 复现及 KV cache 驱逐策略消融实验的逐帧推理输出,以 FFV1 无损编码归档。
本 repo 是全部结果的第 3/3 部分。
内容
reds_val30_gate
reds_val30_gate_L0-15
reds_val30_gate_L15-30
reds_val30_gate_lam1_fifo
reds_val30_gate_lam5_fifo
reds_val30_gate_lfres
reds_val30_gate_lfres_raw
reds_val30_gate_lfres_z
reds_val30_gate_sink
reds_val30_gate_tjump
reds_val30_h2o
reds_val30_headwise
reds_val30_headwise_kv10
reds_val30_ofr_gate
reds_val30_ofr_gate_L0-15
reds_val30_ofr_gate_L15-30… See the full description on the dataset page: https://huggingface.co/datasets/victorzhu30/flashvsr-repro-outputs-part3.troncamp-mani-exact550-repro
TRONCamp Mani T4 exact550
This is a private reproducibility package for the TRONCamp Mani T4 stack_bowls_three experiment. The public code and postmortem live in Datawhale Every Embodied, commit 28c339f.
Contents
processed/: complete 550-episode ACT training input, 550 HDF5 files.
raw_partial/: local recovery of 368 raw HDF5 files only. The canonical raw manifest expects 550 episodes; 182 raw files are still missing locally and are listed in… See the full description on the dataset page: https://huggingface.co/datasets/Datawhale/troncamp-mani-exact550-repro.Squidiff_reproducibility
🦑 Squidiff Reproducibility (Code + Processed Dataset)
English | 简体中文
This repository is a comprehensive, ready-to-use replication bundle for Squidiff. It contains annotated replication Jupyter notebooks alongside all the heavily processed intermediate .h5ad matrices (approx 22.8 GB in total) required to seamlessly reproduce the figures and model results without wrestling with data wrangling.
Note: This repository is cloned and extended from the official Squidiff reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/zyzhou110/Squidiff_reproducibility.opencode_seed2.1_expert_without_reproduce_round_00curvature-repro-resultsvideomanip-reproduction
VideoManip Reproduction (IsaacLab) — data & checkpoints
Unique artifacts produced by the unofficial sim-only reproduction of VideoManip
(arXiv:2602.09013) in IsaacLab 2.3.2 / Isaac Sim 5.1.
Code, docs, protocol and full result tables:
https://github.com/physercoe/videomanip-reproduction
Layout
checkpoints/
mixed3x/ epoch_{5..40}.pth DRO run on HaMeR+ContactOpt data (6-obj sim mean 47.0% @ e5)
mixed3x_handflow/ epoch_{1..20}.pth DRO run on… See the full description on the dataset page: https://huggingface.co/datasets/physer/videomanip-reproduction.gdpval-reproreproducibility-powermodels-setup2Corrected PowerModels JSON for the GENCO setup-2 (from-disk) runtime experiments.
How to reproduce: scripts/runtime/README.md on branch genco-paper-repro.
frozen-tribe-voe-reproduction-data
VOE frozen-TRIBE reproduction archive
This public dataset repository contains the complete VOE data, intermediate
artifacts, predictions, results, logs, and code used for the frozen TRIBE
reproduction. The source snapshot contains 29,435 files
and 437,716,887,971 bytes.
Clearly reproducible transient material was intentionally omitted:
work/ processing scratch data and hf_cache/ downloaded cache data, totaling
960,407 files and 852,522,997,016 bytes.
Every omitted path and the… See the full description on the dataset page: https://huggingface.co/datasets/Gulan06/frozen-tribe-voe-reproduction-data.FINSABER-reproduce
FINSABER Data
Aggregated datasets for the FINSABER backtesting framework (KDD 2026).
Files
File
Description
Size
data/finmem_data/stock_data_sp500_2000_2024.pkl
S&P500 full aggregated data (Price + News + Filings)
~11 GB
data/finmem_data/stock_data_cherrypick_2000_2024.pkl
Selected symbols (TSLA, AMZN, MSFT, NFLX, COIN)
~53 MB
data/price/all_sp500_prices_2000_2024_delisted_include.csv
CSV price-only data for S&P500 (including delisted)
~253 MB… See the full description on the dataset page: https://huggingface.co/datasets/finsaber-team/FINSABER-reproduce.GradeSQL-reproducibility-data
Data Summary
This repository contains the reproducibility data for GradeSQL, a framework for fine-tuning text-to-SQL ORM models. It provides all the data required to reproduce the experiments and fine-tuning setups described in the GradeSQL paper.
The data is intended for researchers and practitioners who want to train, evaluate, or reproduce results from GradeSQL.For details about the methodology, usage, and tutorials, please refer to the main project repository: GradeSQL GitHub.
bonsai2-27b-mtp-repro
Ternary-Bonsai-2-27B + in-file MTP: reproduction bundle (RTX 4080 SUPER, Ada/SM89)
This repository holds the raw data. The method (build script, harness, launch units, write-up) lives on GitHub:
https://github.com/zhaoyilun/bonsai2-27b-mtp-repro
Both are the same piece of work: the GitHub repo has the code and the how-to, this dataset has the
measurements it produced. Cross-linked in both directions.
Raw measurements, scripts and notes for the two discussions:
official model… See the full description on the dataset page: https://huggingface.co/datasets/zhaokeqi/bonsai2-27b-mtp-repro.flashvsr-repro-outputs-part2
FlashVSR 复现实验输出 — part2
FlashVSR 复现及 KV cache 驱逐策略消融实验的逐帧推理输出,以 FFV1 无损编码归档。
本 repo 是全部结果的第 2/3 部分。
内容
reds_bscv_sliding_kv10
reds_inter4k30_ofr_headwise
reds_inter4k30_ofr_reliability
reds_inter4k30_ofr_sliding
reds_inter4k30_ofr_uniform
reds_inter4k30_reliability
reds_inter4k30_reliability_noimm
reds_inter4k30_reliability_rowwise
reds_inter4k30_sliding
reds_inter4k30_uniform
reds_inter4k_full_reliability
reds_inter4k_full_sliding
reds_reds4_sliding… See the full description on the dataset page: https://huggingface.co/datasets/victorzhu30/flashvsr-repro-outputs-part2.icml-2026-reproductions
ICML 2026 Agent Reproducibility Challenge — Logbooks
A living mirror of all public reproduction logbooks from the ICML 2026 Agent Reproducibility Challenge.
Agents attempt to reproduce claims from ICML 2026 papers. Each logbook records the reproduction process, evidence, and verdict for each claim.
Structure
├── papers.json # All 6341 ICML 2026 papers (metadata)
├── logbooks.csv # Main index: one row per logbook (agent × paper ×… See the full description on the dataset page: https://huggingface.co/datasets/qy2100/icml-2026-reproductions.REPRO-Bench
REPRO-Bench: Can Agentic AI Systems Assess the Reproducibility of Social Science Research?
This repository contains the REPRO-Bench dataset, introduced in the paper REPRO-Bench: Can Agentic AI Systems Assess the Reproducibility of Social Science Research?.
REPRO-Bench is a novel benchmark designed to evaluate the capability of agentic AI systems in automating the assessment of social science paper reproducibility. It addresses limitations of existing benchmarks by providing a… See the full description on the dataset page: https://huggingface.co/datasets/ljx529/REPRO-Bench.flashvsr-repro-outputs-part1
FlashVSR 复现实验输出 — part1
FlashVSR 复现及 KV cache 驱逐策略消融实验的逐帧推理输出,以 FFV1 无损编码归档。
本 repo 是全部结果的第 1/2 部分。
内容
.cache
FlashVSR_v1_full
FlashVSR_v1_tiny
reds_bscv_clean_sliding
reds_bscv_sliding
reds_bscv_sliding_kv1
reds_bscv_sliding_kv6
reds_inter4k30_headwise
reds_inter4k_full_headwise
reds_inter4k_full_uniform
reds_reds4_headwise
reds_reds4_uniform
reds_val30_sliding_kv1
目录结构:
<实验名>/
<序列号>.mkv # 该序列全部输出帧,FFV1 无损
<序列号>/.done # 原推理流程的完成标记… See the full description on the dataset page: https://huggingface.co/datasets/victorzhu30/flashvsr-repro-outputs-part1.tcod-v1-alfworld-data-reprosmart-repro-imagenet-resnet50-logitsatec2026-task-e-reproducibility
ATEC2026 L0 Task E 复现数据与日志
本仓库保存 Datawhale ATEC2026 线上赛 L0「桌面整理 Task E」赛后开源复现所需的大文件。
配套 GitHub 教程与代码:
https://github.com/datawhalechina/every-embodied/tree/main/15-Challenge%E7%AB%9E%E8%B5%9B/ATEC2026/L0-%E6%A1%8C%E9%9D%A2%E6%95%B4%E7%90%86TaskE
配套模型权重:
https://huggingface.co/Datawhale/atec2026-task-e-act-seed1-best
文件说明
data/final_100demos_filtered_split_100m/trajectory_filtered.hdf5.part-0000 ... part-0336
ACT seed1 best 方案训练使用的过滤后 100-demo HDF5,按 100MiB… See the full description on the dataset page: https://huggingface.co/datasets/Datawhale/atec2026-task-e-reproducibility.
