datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reprocessed_singapore_national_speech_corpus
Dataset Card for Reprocessed National Speech Corpus
NOTE: This is an Reprocessed version KaraKaraWitch from Recursal.The official download can be found here.
Dataset Details
Dataset Description
Dataset Description:
The National Speech Corpus (NSC) is the first large-scale Singapore English corpus, sponsored by the Info-communications and Media Development Authority (IMDA) of Singapore. The objective is to serve as a primary resource of open speech data for… See the full description on the dataset page: https://huggingface.co/datasets/recursal/reprocessed_singapore_national_speech_corpus.repro-rephrased-data-72BThis is the 72B rephrased data by repro-rephraser-4B from RePro: Training Language Models to Faithfully Recycle the Web for Pretraining.
Code: https://github.com/cxcscmu/RePro
k2-horizon-tiny-cpu-repro-v1
K2-Horizon MoVA tiny random CPU fixture
Complete untrained K2HorizonForCausalLM, not Moonshot Kimi despite the K2 name.
Architecture source: IFM/K2-Horizon-MoVA-36B-A4B at 05cab0a4d7150c1c460a000b37ff40cc1af2feaa.
No pretrained weights, original tokenizer, training data, paid GPU or cloud compute used.
Complete text-only K2HorizonForCausalLM, not Kimi: three-layer dense prefix followed by two real MoVA+MoE layers, grouped RMSNorm, sigmoid top-k routing with selection-only bias… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/k2-horizon-tiny-cpu-repro-v1.qwen3-5-tiny-cpu-repro-v1
Qwen3.5 tiny native random CPU fixture
Complete randomly initialized, untrained Qwen3_5ForConditionalGeneration checkpoint.
This is a pipeline/reproducibility fixture, not a useful language model, distillation,
quantization, quality benchmark, or claim about the performance of Qwen3.8-27B.
No upstream model weights or training data were used. No paid GPU/cloud compute.
Architecture and lineage
Architecture lineage: Qwen/Qwen3.8-27B at… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen3-5-tiny-cpu-repro-v1.glm5-next-tiny-cpu-repro-v1This repository is an evidence bundle, not one root-format dataset at repository root.
first/ and repeat/ are separate complete sealed QFS root datasets; comparison/ holds
the comparison receipt and tokenwise result. panel/ is the sealed input panel. Other files
are provenance, logs and reproduction tools. Do not pass the bundle root as a QFS dataset.
GLM5-Next tiny native CPU fixture
This is a complete untrained random-initialized native Glm5NextForConditionalGeneration
wrapper… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm5-next-tiny-cpu-repro-v1.deepseek-v4-tiny-cpu-repro-v1
DeepSeek-V4 tiny corrected-native-primitives CPU text fixture
Complete randomly initialized, untrained QFSDeepseekV4ForCausalLM text class
using Transformers5.16.1 native primitives and a reviewed RMSNorm arithmetic correction.
No upstream weights, paid GPU/cloud compute or useful-model claim.
This is not unmodified native Transformers or the complete production release.
Architecture and scope
Text lineage:… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/deepseek-v4-tiny-cpu-repro-v1.icml-2026-reproductions
ICML 2026 Agent Reproducibility Challenge — Logbooks
A living mirror of all public reproduction logbooks from the ICML 2026 Agent Reproducibility Challenge.
Agents attempt to reproduce claims from ICML 2026 papers. Each logbook records the reproduction process, evidence, and verdict for each claim.
Structure
├── papers.json # All 6341 ICML 2026 papers (metadata)
├── logbooks.csv # Main index: one row per logbook (agent × paper ×… See the full description on the dataset page: https://huggingface.co/datasets/qy2100/icml-2026-reproductions.minimax-m3-tiny-cpu-repro-v1
minimax-m3 complete native tiny random CPU fixture
Complete untrained MiniMaxM3SparseForConditionalGeneration checkpoint with an untied full LM head,
a real 272-entry byte tokenizer and every native state tensor. Architecture lineage:
MiniMaxAI/MiniMax-M3@f0e1c1e04d40177e4673a22097036854f536e9c0.
No upstream weights, training data, paid GPU or cloud compute were used.
Complete native image/text wrapper with real shrunk Conv3D vision, nonempty 3D RoPE, patch-merge projector and… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/minimax-m3-tiny-cpu-repro-v1.minimax-m2-tiny-cpu-repro-v1
minimax-m2 complete native tiny random CPU fixture
Complete untrained MiniMaxM2ForCausalLM checkpoint with an untied full LM head,
a real 272-entry byte tokenizer and every native state tensor. Architecture lineage:
MiniMaxAI/MiniMax-M2.7@d494266a4affc0d2995ba1fa35c8481cbd84294b.
No upstream weights, training data, paid GPU or cloud compute were used.
Complete native text causal LM: sigmoid/top-k MoE routing with correction bias, per-layer flattened Q/K RMSNorm and half-head… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/minimax-m2-tiny-cpu-repro-v1.Snowball-67B-A2B-RLVR1-Repro-Data
Snowball 67B-A2B RLVR1 data
These are the exact Parquet inputs retained for the Snowball 67B-A2B sync and
async RLVR1 experiments on Iris cw-rno2a in September 2026. The data was
selected from the skyrl_gym route of a TaskTrove conversion of the public
NVIDIA Nemotron RL Ultra training blend,
preserving source order and holding out the last 100 selected rows. See
provenance.json for the local conversion and filtering record. The original
TaskTrove release
is also public.… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/Snowball-67B-A2B-RLVR1-Repro-Data.sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13
Nemotron Terminal SFT reproduction evaluation artifacts
This repository contains the complete Harbor artifact tree for the 300-trial
OpenThoughts-TBLite evaluation of
laion/sft-repro-thinking-step630-nemotron-terminal-step1888.
The checkpoint was trained from the Grug stage-2 thinking checkpoint on the
Nemotron Terminal corpus for 1,888 steps.
Result
Measure
Value
Attempted / completed
300 / 300
Verifier-scoreable
259 (86.33%)
Aggregate reward, all… See the full description on the dataset page: https://huggingface.co/datasets/laion/sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13.VeriReason-reasoning-reproduced-1513_luna-medium
VeriReason reasoning reproduced (Luna medium)
This dataset is a deterministic sample of the locally updated reproduced VeriReason files.
Source files: train (1).jsonl, validation (1).jsonl
Sampling: independent shuffle then prefix slice, fixed seed 42
Splits: train 1513, validation 189
Original target dataset: Jongbin-kr/VeriReason-RTL-Coder_7b_reasoning_tb (train 1513 / validation 189)
Schema: id, instruction, output, tb, tb_result
Quality checks: valid JSON, unique IDs… See the full description on the dataset page: https://huggingface.co/datasets/Jongbin-kr/VeriReason-reasoning-reproduced-1513_luna-medium.OptMiner-Reproduction
Opt-Miner Reproduction — ICML 2026 Reproducibility Challenge
Submission #9232 | OpenReview: GH9qE7sRPzReproduced by: Nikhil DhakaGitHub: ernikhildhaka-arch/OptMiner-ReproductionOrganization: ICML-2026-agent-repro
Paper
Opt-Miner: Empowering Information-Seeking Agent with Tree-Guided Data Synthesis for Optimization ModelingInternational Conference on Machine Learning (ICML), 2026
Claims Verified
Claim
Description
Status
1
Qwen3-8B… See the full description on the dataset page: https://huggingface.co/datasets/ernikhil411/OptMiner-Reproduction.PREPAIR-reproduction-results
PREPAIR Reproduction Results
This dataset contains the reproduction results of the PREPAIR method.
Source Project: PREPAIR
Target File: results_LLMBar_all_methods.csv
Medical-R1-Distill-Data-m22k-Reproduce
Medical-R1-Distill-Data
