CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01recursal /reprocessed_singapore_national_speech_corpus Dataset Card for Reprocessed National Speech Corpus NOTE: This is an Reprocessed version KaraKaraWitch from Recursal.The official download can be found here. Dataset Details Dataset Description Dataset Description: The National Speech Corpus (NSC) is the first large-scale Singapore English corpus, sponsored by the Info-communications and Media Development Authority (IMDA) of Singapore. The objective is to serve as a primary resource of open speech data for… See the full description on the dataset page: https://huggingface.co/datasets/recursal/reprocessed_singapore_national_speech_corpus.audiotext-generation1M<n<10M7 likes353 downloads2y agoHugging Face02cx-cmu /repro-rephrased-data-72BThis is the 72B rephrased data by repro-rephraser-4B from RePro: Training Language Models to Faithfully Recycle the Web for Pretraining. Code: https://github.com/cxcscmu/RePro texttext-generation10M<n<100M0 likes255 downloads11mo agoHugging Face03malaiwah /k2-horizon-tiny-cpu-repro-v1 K2-Horizon MoVA tiny random CPU fixture Complete untrained K2HorizonForCausalLM, not Moonshot Kimi despite the K2 name. Architecture source: IFM/K2-Horizon-MoVA-36B-A4B at 05cab0a4d7150c1c460a000b37ff40cc1af2feaa. No pretrained weights, original tokenizer, training data, paid GPU or cloud compute used. Complete text-only K2HorizonForCausalLM, not Kimi: three-layer dense prefix followed by two real MoVA+MoE layers, grouped RMSNorm, sigmoid top-k routing with selection-only bias… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/k2-horizon-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes255 downloads18d agoHugging Face04malaiwah /qwen3-5-tiny-cpu-repro-v1 Qwen3.5 tiny native random CPU fixture Complete randomly initialized, untrained Qwen3_5ForConditionalGeneration checkpoint. This is a pipeline/reproducibility fixture, not a useful language model, distillation, quantization, quality benchmark, or claim about the performance of Qwen3.8-27B. No upstream model weights or training data were used. No paid GPU/cloud compute. Architecture and lineage Architecture lineage: Qwen/Qwen3.8-27B at… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/qwen3-5-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes249 downloads18d agoHugging Face05malaiwah /glm5-next-tiny-cpu-repro-v1This repository is an evidence bundle, not one root-format dataset at repository root. first/ and repeat/ are separate complete sealed QFS root datasets; comparison/ holds the comparison receipt and tokenwise result. panel/ is the sealed input panel. Other files are provenance, logs and reproduction tools. Do not pass the bundle root as a QFS dataset. GLM5-Next tiny native CPU fixture This is a complete untrained random-initialized native Glm5NextForConditionalGeneration wrapper… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm5-next-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes239 downloads18d agoHugging Face06malaiwah /deepseek-v4-tiny-cpu-repro-v1 DeepSeek-V4 tiny corrected-native-primitives CPU text fixture Complete randomly initialized, untrained QFSDeepseekV4ForCausalLM text class using Transformers5.16.1 native primitives and a reviewed RMSNorm arithmetic correction. No upstream weights, paid GPU/cloud compute or useful-model claim. This is not unmodified native Transformers or the complete production release. Architecture and scope Text lineage:… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/deepseek-v4-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes226 downloads18d agoHugging Face07qy2100 /icml-2026-reproductions ICML 2026 Agent Reproducibility Challenge — Logbooks A living mirror of all public reproduction logbooks from the ICML 2026 Agent Reproducibility Challenge. Agents attempt to reproduce claims from ICML 2026 papers. Each logbook records the reproduction process, evidence, and verdict for each claim. Structure ├── papers.json # All 6341 ICML 2026 papers (metadata) ├── logbooks.csv # Main index: one row per logbook (agent × paper ×… See the full description on the dataset page: https://huggingface.co/datasets/qy2100/icml-2026-reproductions.texttext-generation100K<n<1M0 likes204 downloads1mo agoHugging Face08malaiwah /minimax-m3-tiny-cpu-repro-v1 minimax-m3 complete native tiny random CPU fixture Complete untrained MiniMaxM3SparseForConditionalGeneration checkpoint with an untied full LM head, a real 272-entry byte tokenizer and every native state tensor. Architecture lineage: MiniMaxAI/MiniMax-M3@f0e1c1e04d40177e4673a22097036854f536e9c0. No upstream weights, training data, paid GPU or cloud compute were used. Complete native image/text wrapper with real shrunk Conv3D vision, nonempty 3D RoPE, patch-merge projector and… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/minimax-m3-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes167 downloads18d agoHugging Face09malaiwah /minimax-m2-tiny-cpu-repro-v1 minimax-m2 complete native tiny random CPU fixture Complete untrained MiniMaxM2ForCausalLM checkpoint with an untied full LM head, a real 272-entry byte tokenizer and every native state tensor. Architecture lineage: MiniMaxAI/MiniMax-M2.7@d494266a4affc0d2995ba1fa35c8481cbd84294b. No upstream weights, training data, paid GPU or cloud compute were used. Complete native text causal LM: sigmoid/top-k MoE routing with correction bias, per-layer flattened Q/K RMSNorm and half-head… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/minimax-m2-tiny-cpu-repro-v1.tabulartext-generationn<1K0 likes167 downloads18d agoHugging Face10open-athena /Snowball-67B-A2B-RLVR1-Repro-Data Snowball 67B-A2B RLVR1 data These are the exact Parquet inputs retained for the Snowball 67B-A2B sync and async RLVR1 experiments on Iris cw-rno2a in September 2026. The data was selected from the skyrl_gym route of a TaskTrove conversion of the public NVIDIA Nemotron RL Ultra training blend, preserving source order and holding out the last 100 selected rows. See provenance.json for the local conversion and filtering record. The original TaskTrove release is also public.… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/Snowball-67B-A2B-RLVR1-Repro-Data.texttext-generation10K<n<100K0 likes58 downloads7d agoHugging Face11laion /sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13 Nemotron Terminal SFT reproduction evaluation artifacts This repository contains the complete Harbor artifact tree for the 300-trial OpenThoughts-TBLite evaluation of laion/sft-repro-thinking-step630-nemotron-terminal-step1888. The checkpoint was trained from the Grug stage-2 thinking checkpoint on the Nemotron Terminal corpus for 1,888 steps. Result Measure Value Attempted / completed 300 / 300 Verifier-scoreable 259 (86.33%) Aggregate reward, all… See the full description on the dataset page: https://huggingface.co/datasets/laion/sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13.texttext-generationn<1K0 likes51 downloads1mo agoHugging Face12Jongbin-kr /VeriReason-reasoning-reproduced-1513_luna-medium VeriReason reasoning reproduced (Luna medium) This dataset is a deterministic sample of the locally updated reproduced VeriReason files. Source files: train (1).jsonl, validation (1).jsonl Sampling: independent shuffle then prefix slice, fixed seed 42 Splits: train 1513, validation 189 Original target dataset: Jongbin-kr/VeriReason-RTL-Coder_7b_reasoning_tb (train 1513 / validation 189) Schema: id, instruction, output, tb, tb_result Quality checks: valid JSON, unique IDs… See the full description on the dataset page: https://huggingface.co/datasets/Jongbin-kr/VeriReason-reasoning-reproduced-1513_luna-medium.texttext-generation1K<n<10K0 likes39 downloads2d agoHugging Face13ernikhil411 /OptMiner-Reproduction Opt-Miner Reproduction — ICML 2026 Reproducibility Challenge Submission #9232 | OpenReview: GH9qE7sRPzReproduced by: Nikhil DhakaGitHub: ernikhildhaka-arch/OptMiner-ReproductionOrganization: ICML-2026-agent-repro Paper Opt-Miner: Empowering Information-Seeking Agent with Tree-Guided Data Synthesis for Optimization ModelingInternational Conference on Machine Learning (ICML), 2026 Claims Verified Claim Description Status 1 Qwen3-8B… See the full description on the dataset page: https://huggingface.co/datasets/ernikhil411/OptMiner-Reproduction.texttext-generationn<1K0 likes28 downloads2mo agoHugging Face14HUFS-DILAB /PREPAIR-reproduction-results PREPAIR Reproduction Results This dataset contains the reproduction results of the PREPAIR method. Source Project: PREPAIR Target File: results_LLMBar_all_methods.csv tabulartext-generationn<1K0 likes9 downloads6mo agoHugging Face15c00cjz00 /Medical-R1-Distill-Data-m22k-Reproducegated Medical-R1-Distill-Data tabulartext-generation10K<n<100K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.