CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alexkstern /fineweb-nanochatbpe-100M fineweb-nanochatbpe-100M FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536) and packaged as flat uint16 token-id .bin files for fast memmap training. This is a 100-million-token slice for data-constrained experiments. The train.bin is the byte-exact first 100,000,000 tokens (bytes [0, 200000000)) of the parent alexkstern/fineweb-nanochatbpe-20B train.bin. The val.bin is byte-identical to the parent's val.bin.… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-100M.tabularn<1K0 likes2.3k downloads3mo agoHugging Face02lilywchen /lucky-initialization-atlas-100m-v2 Lucky initialization atlas v2 evidence Private live evidence archive for lilywchen/lucky-initialization-atlas-100m-v2. It contains hash-bound configs, provenance, scalar trajectories, step-zero diagnostics, and final per-sequence losses after those artifacts complete. It excludes credentials, caches, raw FineWeb-derived token arrays, optimizer states, and W&B binary logs. tabulartext-generationn<1K0 likes296 downloads28d agoHugging Face03codelion /sutra-100M Sutra 100M Pretraining Dataset A high-quality synthetic pedagogical dataset designed for LLM pretraining, containing 70,435 educational entries totaling approximately 100 million tokens. Dataset Description This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through: Clear pedagogical structure: Content follows proven educational… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-100M.tabulartext-generation10K<n<100K3 likes61 downloads7mo agoHugging Face04Changyeli03 /llavasae_obliec100k_moststeps100Mtabularn<1K0 likes32 downloads2y agoHugging Face05chrisagrams /consensus-100M-mszx Consensus 100M (mszx) A 100M-scale mass spectrometry consensus dataset, distributed as mszx archive shards for use with the msdatasets library. Contents 90 shards: consensus_00.mszx ... consensus_89.mszx ~1.5 GB per shard, ~140 GB total manifest.json — shard list with sizes and sha256 checksums SHA256SUMS — sha256sum-format file for direct verification The shards are independent; row order across shards is not meaningful. Loading with msdatasets The Hugging… See the full description on the dataset page: https://huggingface.co/datasets/chrisagrams/consensus-100M-mszx.tabularn<1K0 likes23 downloads5mo agoHugging Face06koiwave /100MMLUpro MMLU-Pro 100: A Balanced and Curated Evaluation Set Dataset Description This dataset is a curated, balanced subset of 100 questions derived from the TIGER-Lab/MMLU-Pro test set. It is designed to provide a small, fast, yet representative benchmark for evaluating the knowledge and reasoning capabilities of large language models across a wide range of academic and professional domains. The key feature of this dataset is its stratified sampling method, ensuring that the… See the full description on the dataset page: https://huggingface.co/datasets/koiwave/100MMLUpro.tabularmultiple-choicen<1K0 likes18 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.