datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pretokenized-dolma
The Pretokenized Dolma Dataset
A pre-tokenized, pre-shuffled version of Dolma, the high-quality text corpus from AI2. This dataset is designed to be plug-and-play with the pico-train library.
Overview
Key Features:
Tokenized with allenai/OLMo-7B-0724-hf, a BPE-tokenized with a vocabulary size of 50280
Sequence length: 2049 tokens (2048 + 1 for next-token prediction)
Sharded into 10,000 Parquet files (~78MB each)
420B tokens total size (perfect for training a model for… See the full description on the dataset page: https://huggingface.co/datasets/pico-lm/pretokenized-dolma.rnacentral-pretokenized-shardedgemma-4-pretokenized-tracesdanish-pretokenized-16k
Danish Pretrain — pretokenized (byte-level BPE, 16k vocab) — v1 partial
Pretokenized version of jensjepsen/danish-pretrain
via jensjepsen/danish-tokenizer.
Note: This is a partial version — 94 of 96 shards. The 2 missing shards
crashed during tokenization (rare pathological content). See
jensjepsen/danish-pretokenized-16k-supp for the recovered rows.
93,200,464 docs
Schema: input_ids: list<int32>, attention_mask: list<int8>
94 zstd-compressed parquet shards
yeast-no-GTL-pretokenized-NT-v2pretokenized-pretrain-test-128Kpretokenized-dolma-5Mmouse-pretokenized-NThuman-pretokenized-NTpretokenized-dolma-20Myeast-gene-sequence-homology-pretokenized-NTyeast-no-label-pretokenized-NThuman_and_mouse-pretokenized-NTBEELINE-HepG2-no-GTL-pretokenized-NTBEELINE-mDC-no-label-pretokenized-NTBEELINE-human-v2-pretokenized-NTBEELINE-mouse-v2-pretokenized-NTBEELINE-mESC-no-label-pretokenized-NTBEELINE-mHSC-no-label-pretokenized-NTpretraining-pretokenized-smollm3
SmolLM3 Pretokenized Pretraining Sources
Datatrove/Nanotron tokenized-byte versions of three public pretraining sources:
fineweb-edu-10bt: HuggingFaceFW/fineweb-edu, sample/10BT
finemath-4plus: HuggingFaceTB/finemath, finemath-4plus
stack-edu-python: HuggingFaceTB/stack-edu, Python, with content retrieved from the public Software Heritage S3 bucket using the upstream dataset-card procedure
All subsets use HuggingFaceTB/SmolLM3-3B at revision… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/pretraining-pretokenized-smollm3.BEELINE-hESC-no-GTL-pretokenized-NTBEELINE-mDC-no-GTL-pretokenized-NTfineweb-edu-pretokenized-10b
FineWeb-Edu Pretokenized 10B
Megatron indexed-dataset (.bin/.idx) versions of
HuggingFaceFW/fineweb-edu sample/10BT.
Each Hugging Face subset contains a small metadata.parquet index. The actual
Megatron files are under <subset>/files/; pass each prefix without the
.bin/.idx suffix to Megatron Core or Megatron Bridge.
Documents retain upstream shard and row order. Tokenization disables automatic
special-token insertion and appends exactly one tokenizer EOS/EOD token to each… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/fineweb-edu-pretokenized-10b.nomic-bert-pretokenized-2048-wiki-2023tinystories-pretokenized-pythia
Dataset Card for "tinystories-pretokenized-pythia"
More Information needed
yeast-tf-sequence-homology-pretokenized-NTBEELINE-mESC-no-GTL-pretokenized-NTBEELINE-HepG2-no-label-pretokenized-NTBEELINE-hESC-no-label-pretokenized-NTai-hdlcoder-pretokenized-dataset-train
Dataset Card for "ai-hdlcoder-pretokenized-dataset-train"
More Information needed
