datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ncp-q4-v2-cohort
Frozen NCP repaired-v2 cohort
Public campaign handoff; not a new dataset rebuild. Pin an immutable revision.
Packed shards in cohorts/ unpack using ncp_eval.unpack_cohort.
Test: 1,457 sections; SHA256 520e9736ace196e03166535e13aebeeaa1ec2931b8c0027601e16d6bca2bc37d.
See EXPORT_MANIFEST.json for file hashes and each split MANIFEST for shard hashes.
The older populate_evidence_led_train200 driver is NOT the production protocol.
ncp-artifacts-v1
NCP artifacts (v1)
datasets/ and rl/rl_prompts/ have had their prompt / prompt_full fields removed. Those
fields embedded the story-so-far, character sheets and chapter text for each section - verbatim
novel prose - which is not something to publish in a public repo, and it was 97% of the bytes.
Everything needed to rebuild them is here or already on the destination cluster:
every row keeps problem_id / question_id, which are exactly the ids ncp_eval.data.example_id()
emits… See the full description on the dataset page: https://huggingface.co/datasets/agurung/ncp-artifacts-v1.isambard-ncp-returnsnucleosome-condensability-ncp-chr1NC_Physics
NC_Physics
NC_Physics is the dataset released with LECTOR: Joint Optimization of Scientific Reasoning Graphs and Introduction Generation.
The dataset supports Content-Conditional Introduction Generation (CCIG): models use the non-introduction content of scientific papers, paper metadata, and references to reason about the paper's core idea and generate a logic-aware introduction.
Paper: https://arxiv.org/abs/2605.25964
Code: https://github.com/Xiao-Youth/LECTOR
Associated… See the full description on the dataset page: https://huggingface.co/datasets/Xiao-Youth/NC_Physics.softsign-ncp-reproncp_datasets_v12u_b16_antincp-1p6b-mix
ncp-1p6b-mix
Packed uint16 token stream for NCP-64M pretraining (1,601,363,502 tokens).
Mix (token-weighted, deficit-scheduled interleave)
finepdfs (codelion/finepdfs-1B): 45.5% — 729M tok
dclm (codelion/dclm-baseline-1B): 26.8% — 428M tok
fineweb_edu (codelion/fineweb-edu-1B): 18.0% — 287M tok
sutra (codelion/sutra-improved-100M): 10.1% — 162M tok
Backbone follows the 70M-scale recipe from
https://huggingface.co/blog/codelion/optimal-dataset-mixing (50/30/20… See the full description on the dataset page: https://huggingface.co/datasets/tachytelicdetonation/ncp-1p6b-mix.nucleosome-condensability-merged-ncpNCPL-Pretraining-LogsPretraining logs collected from:
Marin Project: https://github.com/marin-community/marin
Step Law Project: https://github.com/step-law/steplaw
Each example corresponds to one training run, including the training configuration and performance metrics (C4-en evaluation loss for Marin, and smoothed pretraining loss for StepLaw).
Load the dataset
from datasets import load_dataset
marin = load_dataset("zhqwqwq/NCPL-Pretraining-Logs", "marin", split="train")
steplaw =… See the full description on the dataset page: https://huggingface.co/datasets/zhqwqwq/NCPL-Pretraining-Logs.ratishsp__ncp_cc__1649422863
GEM Submission
Submission name: NCP_CC
ratishsp__ncp_cc__1649422112
GEM Submission
Submission name: NCP_CC
nucleosome-condensability-ncp-excludedkl3m-filter-data-dotgov-www.ncpc.govkl3m-data-dotgov-www.ncpc.gov
KL3M Data Project
Note: This page provides general information about the KL3M Data Project. Additional details specific to this dataset will be added in future updates. For complete information, please visit the GitHub repository or refer to the KL3M Data Project paper.
Description
This dataset is part of the ALEA Institute's KL3M Data Project, which provides copyright-clean training resources for large language models.
Dataset Details
Format: Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alea-institute/kl3m-data-dotgov-www.ncpc.gov.voz_hsd_labeled
