datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mobile-vlm-datagrm-reproduce-align30k-beavertails-v
Fused GRM reproduction data
This is the self-contained Stage-1 preference dataset used by
GRM-Reproduce to reproduce the
generative reward model (GRM) stage of Generative RLHF-V. It is an
independent reproduction artifact, not an official dataset release by the paper
authors.
The release fuses exactly 30,000 Align-Anything pairs and 9,369
BeaverTails-V pairs. Every row already follows the final VERL Stage-1 schema
produced by grlhfv_repro.data.to_verl_preference_example, and… See the full description on the dataset page: https://huggingface.co/datasets/ZoeyZou/grm-reproduce-align30k-beavertails-v.averimatec-claim-images-reproduction
AVerImaTeC claim images for baseline reproduction
This public mirror contains only the unchanged claim images used by the first 80 records of the public AVerImaTeC validation split. Source dataset: https://huggingface.co/datasets/Rui4416/AVerImaTeC .
Purpose: stable public image URLs for reverse-image search in an adapted baseline reproduction. This directory contains no labels, generated evidence, results, API credentials or project source code. Filenames and bytes are… See the full description on the dataset page: https://huggingface.co/datasets/Sophie508/averimatec-claim-images-reproduction.mva-repro-artifactsblackbox-regression-repro
Reproduction: Black-Box Assisted Regression (ICML 2026, N9nlCRUiir)
Independent reproduction of "Black-Box Assisted Regression: Phase Transitions and
Minimax Optimality" by Yan Zhou (ICML 2026, arXiv:2606.25743, OpenReview N9nlCRUiir).
No official code was released; this is an independent implementation of Algorithm 1
(Safe Residual Estimator) and all six claim experiments, following the paper's
Appendix F implementation details.
Layout
scripts/core.py — DGP +… See the full description on the dataset page: https://huggingface.co/datasets/FrancescoBellingeri/blackbox-regression-repro.repro_1_imageicml2026-repro-l35QweVxgn-code
Reproduction code — On the Theory of Continual Learning with Gradient Descent for Neural Networks
Clean-room NumPy reimplementation and full sweep harness for the ICML 2026 submission
l35QweVxgn (arXiv:2510.05573v2),
Taheri, Ghosh & Mazumdar.
The write-up lives in the Trackio logbook:
🚀 nmaher/repro-on-the-theory-of-continual-learning-with-gradient-descent-for-neural-networks.
This repo is the code and the raw numbers behind it.
Layout
code/ the… See the full description on the dataset page: https://huggingface.co/datasets/nmaher/icml2026-repro-l35QweVxgn-code.coredteam-repro-bundle
Repro: Co-RedTeam (arXiv:2602.02164, OpenReview wz57SLSTk3)
Scaled reproduction of Co-RedTeam: Orchestrated Security Discovery and
Exploitation with LLM Agents (Google Cloud AI Research / Google /
Michigan State University, ICML 2026 submission).
No official code was released. This bundle contains a from-scratch, scaled
reimplementation of the paper's multi-agent architecture plus the harnesses
used to test it against real (and, where infeasible, toy-proxy) benchmark
data, all… See the full description on the dataset page: https://huggingface.co/datasets/Firemedic15/coredteam-repro-bundle.mewl-repro
MEWL Inference Reproduction (ICML'23 benchmark, modern models)
Full-test-set reproduction of inference on MEWL (MachinE Word Learning,
Jiang et al., ICML 2023) — 9 word-learning
tasks x 600 test episodes, evaluated with 4 modern models (zero-shot).
Each row = one episode: 6 context images (each labeled with a novel-word
utterance), the query image, 5 candidate answers, the ground-truth
answer, and each model's prediction + correctness.
Subsets = the 9 task categories. Splits =… See the full description on the dataset page: https://huggingface.co/datasets/guangliangliu/mewl-repro.orderplace-repro-bundle
Reproduction: Order Matters — Unveiling the Hidden Impact of Macro Placement Sequences via Proxy-Guided LLM Evolution
Independent reproduction of ICML 2026 paper #12180 — OrderPlace (Shibing Mo,
Jing Liu, Jianchu Xu, Ruilin Wu). arXiv 2606.08904 ·
OpenReview fPjkFDW9j7 ·
official code (placeholder only, no runnable source): Explorermomo/OrderPlace.
Part of the Hugging Face × AlphaXiv ICML-2026 reproduction challenge.
What this is
OrderPlace claims that the order… See the full description on the dataset page: https://huggingface.co/datasets/debajyotidasgupta/orderplace-repro-bundle.repro_2_imagerepro-evidence-grokking-ca-local-rules
Evidence trail — Grokking phase transitions in learning local rules with gradient descent
Full evidence for an automated claim-by-claim audit of
Grokking phase transitions in learning local rules with gradient descent, produced by
Lemma, an AI-scientist pipeline built
for re:AGENT (Founders Inc, Aug 15–16 2026).
Verdict: 5 supported / 0 falsified / 1 inconclusive
of 6 extracted claims. Judge verdict: PASS (5/5).
Claim
Title
Verdict
C1
Critical exponent in 1D… See the full description on the dataset page: https://huggingface.co/datasets/Papajams/repro-evidence-grokking-ca-local-rules.repro-dance-bundle
DANCE reproduction bundle
Reproduction of DANCE: Dynamic, Available, Neighbor-gated Condensation for Federated
Text-Attributed Graphs (ICML 2026, OpenReview YhEWC2HiJl, arXiv 2601.16519).
Scope: local, single-machine (Apple M1 Pro). HF Jobs was credit-blocked (402), so the
federated 8-dataset + LLM-surrogate setting (Table 1) is out of scope; we verify the
mechanism, theory, and complexity claims, plus a single-dataset (Cora) proxy.
Contents
scripts/dance_core.py… See the full description on the dataset page: https://huggingface.co/datasets/kpshinnik/repro-dance-bundle.brain-ai-alignment-reproduction
Brain–AI alignment reproduction artifacts
Full-scale statistical audit of “Alignment between Brains and AI: Evidence for
Convergent Evolution across Modalities, Scales and Training Trajectories.”
Paper: https://arxiv.org/abs/2507.01966
Audited source commit: https://github.com/FloyedShen/BrainAlign/tree/f3e8c78ed3ee5ac15ae7c76053665bbec5ac8a4d
Reproduction Job: https://huggingface.co/jobs/visv-Bro/6a6f72c86b79c09949c1f7f1
Job staging Bucket:… See the full description on the dataset page: https://huggingface.co/datasets/visv-Bro/brain-ai-alignment-reproduction.repro-adaptive-preconditioners-trigger-loss-spikes-in-adam-artifacts
Reproduction Artifacts
All code, data, and results for the reproduction of "Adaptive Preconditioners Trigger Loss Spikes in Adam" (arXiv:2506.04805).
Contents
code/ - All Python scripts for reproduction
results/ - JSON result files
figures/ - Generated plots and poster
logbook/ - Trackio logbook files
Key Files
code/repro_adaptive.py - Main HF experiment (1D/2D quadratic, FNN, Transformer)
code/repro_adaptive2.py - Refined local reproduction… See the full description on the dataset page: https://huggingface.co/datasets/Harshvardhan-Mestha/repro-adaptive-preconditioners-trigger-loss-spikes-in-adam-artifacts.icml2026-epistemic-uncertainty-reprorepro-lTqCHLllix-artifacts
Reproduction: Online Learning with Recency: Algorithms for Sliding-window Streaming Multi-armed Bandits
OpenReview ID: lTqCHLllix
arXiv: 2606.08977
Challenge: ICML 2026 Reproduction Challenge
Compute: CPU only, ~5 min total, $0.00
Overview
An independent CPU-scale numerical audit of the paper's five main theoretical claims.
All experiments are deterministic (seeded) and reproducible with only numpy and
matplotlib.
Per the challenge guide, theorem claims are… See the full description on the dataset page: https://huggingface.co/datasets/MarxistLeninist/repro-lTqCHLllix-artifacts.repro-chain-of-thought-gradient-descent-runs-sol
Chain-of-Thought Gradient Descent reproduction runs
Immutable outputs for the independent scaled reproduction of ICML 2026 paper
#443, OpenReview uZ8JZ1Lw9a.
gpu-l4-seed443/: successful NVIDIA L4 checkpoint, result JSON, and cost CSV.
figures/: interactive logbook figures and raw CSVs.
poster/: Posterly source, zero-warning gate report, PDF/PNG, and
self-contained poster_embed.html.
reproduction-bundle/: complete clean download-and-rerun bundle.
Successful Job:… See the full description on the dataset page: https://huggingface.co/datasets/JG1310/repro-chain-of-thought-gradient-descent-runs-sol.repro-wire-graph-rope
Repro — WIRE (trn64znfNx): rotary position encodings for graphs
ICML 2026 Agent Reproduction Challenge — Claim-Closure Agent 02.
Paper: Rotary Position Encodings for Graphs (WIRE) — ICML 2026 spotlight (OpenReview trn64znfNx).
Independent first-principles NumPy test of two catalog claims. No value copied
from the paper.
Headline result
Claim
Result
C2 permutation-equivariant up to sign/rotation (Lemma 1)
verified — eigenvalue invariance 1.4e-14; 96% of… See the full description on the dataset page: https://huggingface.co/datasets/pranaysuyash/repro-wire-graph-rope.repro-sigma-bundle
SigMa reproduction bundle
Reproduction of ICML 2026 paper σ: Sigmoid Modulation for Ultra High Resolution Diffusion
(OpenReview 47JZSOkw5C; code github.com/bxuanz/SigMa).
This bundle verifies SigMa's training-free RoPE modulation mechanism — the algorithmic core both
paper claims rest on — by running the paper's own released FluxPosEmbed positional-encoding module
at the exact FLUX latent grids up to 16 megapixels. No model weights are needed: SigMa lives entirely
in a… See the full description on the dataset page: https://huggingface.co/datasets/kpshinnik/repro-sigma-bundle.rcdp-dueling-bandits-repro
Reproduction bundle — RCDP-UCB (ICML 2026 #3478)
Independent reproduction of "Robust Linear Dueling Bandits with Post-serving Context
under Unknown Delays and Adversarial Corruptions" (Youngmin Oh, ICML 2026 Poster).
Paper: arXiv:2605.01752 · OpenReview RaJnDY8aAS
Official code: https://github.com/youngmin0oh/rcdp-public (cloned at commit 3ed73c6)
Trackio logbook (full writeup): see the linked Space in the collection.
What's here
code/ the official… See the full description on the dataset page: https://huggingface.co/datasets/Eishaan/rcdp-dueling-bandits-repro.logbook-repro-pfl-kmerepro-fw-lower-bounds
Reproduction bundle — Lower Bounds for Frank-Wolfe on Strongly Convex Sets
Independent reproduction of ICML 2026 paper #18157 — Lower Bounds for
Frank–Wolfe on Strongly Convex Sets (Halbey, Deza, Zimmer, Roux, Stellato,
Pokutta), arXiv:2602.04378,
OpenReview 3KX1xU2bCC.
This is a theory/computation paper: a constructive Ω(1/√ε) lower bound for
Frank–Wolfe. Reproduction is a self-contained Python re-implementation of the
paper's dynamics, backward-reconstruction construction, and… See the full description on the dataset page: https://huggingface.co/datasets/kpshinnik/repro-fw-lower-bounds.icml17708-grok-grokking-reprotimercd-repro-bundle
TimeRCD reproduction bundle
This bundle mirrors the evidence used in the logbook:
scripts/rank_recount.py: parses arXiv v3/v5 Table 1 HTML and recomputes rank counts.
scripts/timercd_deep_audit.py: audits released TimeRCD checkpoints/code and public synthetic generator throughput.
scripts/timercd_checkpoint_probe.py: runs the released univariate checkpoint on a small generated contextual set.
scripts/reproduce_timercd_evidence.py: earlier paper/code/generator/proxy evidence… See the full description on the dataset page: https://huggingface.co/datasets/Srishti280992/timercd-repro-bundle.repro-grokking-ridge
Repro — To Grok Grokking (5nNNVY8NW4): ridge-regression grokking
ICML 2026 Agent Reproduction Challenge — Claim-Closure Agent 02.
Paper: To Grok Grokking: Provable Grokking in Ridge Regression (OpenReview 5nNNVY8NW4, arXiv 2601.19791).
An independent first-principles finite-d study of whether end-to-end
grokking (overfit → delayed poor generalization → eventual low error) occurs
in over-parameterized realizable ridge regression under GD + constant weight
decay. Every number is… See the full description on the dataset page: https://huggingface.co/datasets/pranaysuyash/repro-grokking-ridge.repro-deep-hierarchical-models-results
Reproduction results — Deep Networks Learn Deep Hierarchical Models (arXiv:2601.00455)
Results for the reproduction of ICML 2026 #12249. Logbook:
vimarsh/repro-deep-networks-learn-deep-hierarchical-models.
local_rtx4000/ — full results (JSON + PNG figures) for all 5 claims, 2xRTX 4000 Ada.
scaled/ — larger-scale re-run from the L4 Hugging Face GPU Job (added if the job completes).
Code bundle: vimarsh/repro-deep-hierarchical-models.
qsched-repro-outlm-memorization-repro
Reproduction: "How much do language models memorize?" (arXiv:2505.24832)
Reproduction of the memorization-capacity experiments from Morris et al. 2025
for the ICML-2026 open-reproductions challenge (paper OpenReview id bA6BgSbaUi).
src/model.py — compact GPT-2-style transformer (from scratch).
src/experiment.py — training + capacity / double-descent / membership-inference measurement.
run.py — entrypoint: python run.py --config config.yaml (task set by config).
config.yaml —… See the full description on the dataset page: https://huggingface.co/datasets/jrcrittenden/lm-memorization-repro.repro-menvbench-cardinality
Repro — MEnvAgent (Mkal0hTCnh): MEnvBench / MEnvData-SWE cardinality verification
ICML 2026 Agent Reproduction Challenge — Claim-Closure Agent 02.
Paper: MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering (OpenReview Mkal0hTCnh, arXiv 2601.22859).
This is an independent dataset-statistics verification of two falsifiable
cardinality claims. Every integer below is computed from the bytes of the
released artifact with a from-scratch… See the full description on the dataset page: https://huggingface.co/datasets/pranaysuyash/repro-menvbench-cardinality.
