datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tahoe-100m-zarr
Tahoe-100M Zarr Collection
Production-ready tahoe single-cell RNA-seq data exported from Arc Virtual Cell Atlas (Tahoe-100M) into native Zarr stores for chunked, on-demand access on the Hugging Face Hub.
Why Zarr
Single-cell expression matrices get impractical fast if you treat them like ordinary dense files. Zarr is the point of this repo: it makes large atlas-scale data usable without forcing users to download or materialize the whole matrix before they can do anything… See the full description on the dataset page: https://huggingface.co/datasets/KokosDev/tahoe-100m-zarr.fineweb-nanochatbpe-100M
fineweb-nanochatbpe-100M
FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536)
and packaged as flat uint16 token-id .bin files for fast memmap training.
This is a 100-million-token slice for data-constrained experiments. The train.bin
is the byte-exact first 100,000,000 tokens (bytes [0, 200000000)) of the parent
alexkstern/fineweb-nanochatbpe-20B
train.bin. The val.bin is byte-identical to the parent's val.bin.… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-100M.AS-100M
AS-100M
AS-100M is a subset of AS-1B. We release this dataset in both COCO format and JSONL format.
NOTE: The bbox format in the COCO format is xywh, while in the JSONL format, it is x1y1x2y2.
Introduction
We present the All-Seeing Project with:
All-Seeing 1B (AS-1B) dataset: we propose a new large-scale dataset (AS-1B) for open-world panoptic visual recognition and understanding, using an economical semi-automatic data engine that combines the power of off-the-shelf… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/AS-100M.lucky-initialization-atlas-100m-v2
Lucky initialization atlas v2 evidence
Private live evidence archive for lilywchen/lucky-initialization-atlas-100m-v2. It contains hash-bound configs,
provenance, scalar trajectories, step-zero diagnostics, and final per-sequence losses
after those artifacts complete. It excludes
credentials, caches, raw FineWeb-derived token arrays, optimizer states, and W&B
binary logs.
hdvila-100Msutra-100M
Sutra 100M Pretraining Dataset
A high-quality synthetic pedagogical dataset designed for LLM pretraining, containing 70,435 educational entries totaling approximately 100 million tokens.
Dataset Description
This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through:
Clear pedagogical structure: Content follows proven educational… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-100M.babylm-2024-baby-cosmo-fine-100m@misc{charpentier2024gptbertboth,
title={GPT or BERT: why not both?},
author={Lucas Georges Gabriel Charpentier and David Samuel},
year={2024},
eprint={2410.24159},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2410.24159},
}
sutra-improved-100M
Sutra Improved 100M
A self-improved pedagogical dataset for LLM pretraining, containing 413,899 entries totaling 110,038,011 tokens (~110 million). This dataset was created by applying an iterative self-improvement process to the Sutra-10B dataset, where each sample was rewritten using Gemma-3-4B-IT and only the better version (original or rewritten) was kept, followed by comprehensive deduplication and quality filtering.
Dataset Description
This dataset explores… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-improved-100M.DSIR-filtered-pile-100M-short
Dataset Card for DSIR-filtered-pile-100M-short
Dataset Summary
This dataset is a subset of The Pile, selected via the DSIR data selection method. The target distribution for DSIR is the Wikipedia and BookCorpus2 subsets of The Pile.
Languages
English (EN)
Dataset Structure
A train set is provided (102M examples) with a small validation and test set (50k examples each). This dataset is more suitable for training shorter LMs (128 or 256… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/DSIR-filtered-pile-100M-short.llavasae_obliec100k_moststeps100Mconsensus-100M-mszx
Consensus 100M (mszx)
A 100M-scale mass spectrometry consensus dataset, distributed as mszx
archive shards for use with the msdatasets
library.
Contents
90 shards: consensus_00.mszx ... consensus_89.mszx
~1.5 GB per shard, ~140 GB total
manifest.json — shard list with sizes and sha256 checksums
SHA256SUMS — sha256sum-format file for direct verification
The shards are independent; row order across shards is not meaningful.
Loading with msdatasets
The Hugging… See the full description on the dataset page: https://huggingface.co/datasets/chrisagrams/consensus-100M-mszx.Vertex-0.6-100M-self-identification
Vertex 0.6 100M Self Identification
A self-identification SFT dataset for Vertex-0.6-100M-8192-Instruct:
459 ChatML-style conversations that teach the model who it is: its name, creator, family,
architecture, parameter count and knowledge cutoff.
Made from SupraLabs/LLM-self-identification
(Apache-2.0), with every {{SELF_ID.*}} marker replaced with the Vertex 0.6 100M identity:
Marker
Value
MODEL_ID
VertexResearch/Vertex-0.6-100M-8192-Instruct
MODEL_NAME
Vertex 0.6… See the full description on the dataset page: https://huggingface.co/datasets/VertexResearch/Vertex-0.6-100M-self-identification.100MMLUpro
MMLU-Pro 100: A Balanced and Curated Evaluation Set
Dataset Description
This dataset is a curated, balanced subset of 100 questions derived from the TIGER-Lab/MMLU-Pro test set. It is designed to provide a small, fast, yet representative benchmark for evaluating the knowledge and reasoning capabilities of large language models across a wide range of academic and professional domains.
The key feature of this dataset is its stratified sampling method, ensuring that the… See the full description on the dataset page: https://huggingface.co/datasets/koiwave/100MMLUpro.100-multiturn-task-sampling-elementary-phisynth-sharegptOpenthoughts-100mil-DifferentFormattransmla_pretrain_100m_tokensMecan-ASI-Conversations-100Mfinewebedu-test-100Mopenthoughts-100mil-sharegptsubset of openthoughts 100k, 100mil tokens from across all shards
in sharegpt format
simple_math_2_numbers_100m100Miaratts-100M-ptbr-erinome-dataset
