datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-nanochatbpe-100M
fineweb-nanochatbpe-100M
FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536)
and packaged as flat uint16 token-id .bin files for fast memmap training.
This is a 100-million-token slice for data-constrained experiments. The train.bin
is the byte-exact first 100,000,000 tokens (bytes [0, 200000000)) of the parent
alexkstern/fineweb-nanochatbpe-20B
train.bin. The val.bin is byte-identical to the parent's val.bin.… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-100M.lucky-initialization-atlas-100m-v2
Lucky initialization atlas v2 evidence
Private live evidence archive for lilywchen/lucky-initialization-atlas-100m-v2. It contains hash-bound configs,
provenance, scalar trajectories, step-zero diagnostics, and final per-sequence losses
after those artifacts complete. It excludes
credentials, caches, raw FineWeb-derived token arrays, optimizer states, and W&B
binary logs.
sutra-100M
Sutra 100M Pretraining Dataset
A high-quality synthetic pedagogical dataset designed for LLM pretraining, containing 70,435 educational entries totaling approximately 100 million tokens.
Dataset Description
This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through:
Clear pedagogical structure: Content follows proven educational… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-100M.llavasae_obliec100k_moststeps100Mconsensus-100M-mszx
Consensus 100M (mszx)
A 100M-scale mass spectrometry consensus dataset, distributed as mszx
archive shards for use with the msdatasets
library.
Contents
90 shards: consensus_00.mszx ... consensus_89.mszx
~1.5 GB per shard, ~140 GB total
manifest.json — shard list with sizes and sha256 checksums
SHA256SUMS — sha256sum-format file for direct verification
The shards are independent; row order across shards is not meaningful.
Loading with msdatasets
The Hugging… See the full description on the dataset page: https://huggingface.co/datasets/chrisagrams/consensus-100M-mszx.100MMLUpro
MMLU-Pro 100: A Balanced and Curated Evaluation Set
Dataset Description
This dataset is a curated, balanced subset of 100 questions derived from the TIGER-Lab/MMLU-Pro test set. It is designed to provide a small, fast, yet representative benchmark for evaluating the knowledge and reasoning capabilities of large language models across a wide range of academic and professional domains.
The key feature of this dataset is its stratified sampling method, ensuring that the… See the full description on the dataset page: https://huggingface.co/datasets/koiwave/100MMLUpro.
