CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01agurung /ncp-q4-v2-cohort Frozen NCP repaired-v2 cohort Public campaign handoff; not a new dataset rebuild. Pin an immutable revision. Packed shards in cohorts/ unpack using ncp_eval.unpack_cohort. Test: 1,457 sections; SHA256 520e9736ace196e03166535e13aebeeaa1ec2931b8c0027601e16d6bca2bc37d. See EXPORT_MANIFEST.json for file hashes and each split MANIFEST for shard hashes. The older populate_evidence_led_train200 driver is NOT the production protocol. 0 likes2.5k downloads1h agoHugging Face02agurung /ncp-artifacts-v1 NCP artifacts (v1) datasets/ and rl/rl_prompts/ have had their prompt / prompt_full fields removed. Those fields embedded the story-so-far, character sheets and chapter text for each section - verbatim novel prose - which is not something to publish in a public repo, and it was 97% of the bytes. Everything needed to rebuild them is here or already on the destination cluster: every row keeps problem_id / question_id, which are exactly the ids ncp_eval.data.example_id() emits… See the full description on the dataset page: https://huggingface.co/datasets/agurung/ncp-artifacts-v1.0 likes757 downloads21d agoHugging Face03agurung /isambard-ncp-returns0 likes226 downloads7h agoHugging Face04chejames /nucleosome-condensability-ncp-chr1tabular1M<n<10M0 likes77 downloads16d agoHugging Face05Xiao-Youth /NC_Physics NC_Physics NC_Physics is the dataset released with LECTOR: Joint Optimization of Scientific Reasoning Graphs and Introduction Generation. The dataset supports Content-Conditional Introduction Generation (CCIG): models use the non-introduction content of scientific papers, paper metadata, and references to reason about the paper's core idea and generate a logic-aware introduction. Paper: https://arxiv.org/abs/2605.25964 Code: https://github.com/Xiao-Youth/LECTOR Associated… See the full description on the dataset page: https://huggingface.co/datasets/Xiao-Youth/NC_Physics.texttext-generation10K<n<100K0 likes56 downloads3mo agoHugging Face06pngwn /softsign-ncp-repro0 likes43 downloads19d agoHugging Face07agurung /ncp_datasets_v12u_b16_antitext10K<n<100K0 likes37 downloads4d agoHugging Face08tachytelicdetonation /ncp-1p6b-mix ncp-1p6b-mix Packed uint16 token stream for NCP-64M pretraining (1,601,363,502 tokens). Mix (token-weighted, deficit-scheduled interleave) finepdfs (codelion/finepdfs-1B): 45.5% — 729M tok dclm (codelion/dclm-baseline-1B): 26.8% — 428M tok fineweb_edu (codelion/fineweb-edu-1B): 18.0% — 287M tok sutra (codelion/sutra-improved-100M): 10.1% — 162M tok Backbone follows the 70M-scale recipe from https://huggingface.co/blog/codelion/optimal-dataset-mixing (50/30/20… See the full description on the dataset page: https://huggingface.co/datasets/tachytelicdetonation/ncp-1p6b-mix.0 likes31 downloads9d agoHugging Face09chejames /nucleosome-condensability-merged-ncptabular10K<n<100K0 likes29 downloads1d agoHugging Face10zhqwqwq /NCPL-Pretraining-LogsPretraining logs collected from: Marin Project: https://github.com/marin-community/marin Step Law Project: https://github.com/step-law/steplaw Each example corresponds to one training run, including the training configuration and performance metrics (C4-en evaluation loss for Marin, and smoothed pretraining loss for StepLaw). Load the dataset from datasets import load_dataset marin = load_dataset("zhqwqwq/NCPL-Pretraining-Logs", "marin", split="train") steplaw =… See the full description on the dataset page: https://huggingface.co/datasets/zhqwqwq/NCPL-Pretraining-Logs.tabular1K<n<10K0 likes26 downloads7mo agoHugging Face11GEM-submissions /ratishsp__ncp_cc__1649422863 GEM Submission Submission name: NCP_CC textn<1K0 likes19 downloads4y agoHugging Face12GEM-submissions /ratishsp__ncp_cc__1649422112 GEM Submission Submission name: NCP_CC textn<1K0 likes15 downloads4y agoHugging Face13chejames /nucleosome-condensability-ncp-excludedtabular1M<n<10M0 likes14 downloads1d agoHugging Face14alea-institute /kl3m-filter-data-dotgov-www.ncpc.govtext1K<n<10K0 likes10 downloads2y agoHugging Face15alea-institute /kl3m-data-dotgov-www.ncpc.gov KL3M Data Project Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper. Description This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models. Dataset Details Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.ncpc.gov.text1K<n<10K0 likes7 downloads1y agoHugging Face16NCPhat2005 /voz_hsd_labeledtabular10M<n<100M0 likes1 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.