datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fineweb-edu-pretokenized-10K
Marin/Levanter Subsampled Pretokenized Dataset
Dataset
Train Urls:
gs://marin-us-central2/raw/fineweb-edu-c2beb4/3c452cb/huggingface.co/datasets/HuggingFaceFW/fineweb-edu/resolve/3c452cb
Factsheet
Original cache: gs://marin-us-central2/tokenized/fineweb-edu-24698d
Tokenizer: stanford-crfm/marin-tokenizer
Seed 42
Number of tokens: 10,390
(This readme is automatically generated by Marin.)
pretokenized-dolma
The Pretokenized Dolma Dataset
A pre-tokenized, pre-shuffled version of Dolma, the high-quality text corpus from AI2. This dataset is designed to be plug-and-play with the pico-train library.
Overview
Key Features:
Tokenized with allenai/OLMo-7B-0724-hf, a BPE-tokenized with a vocabulary size of 50280
Sequence length: 2049 tokens (2048 + 1 for next-token prediction)
Sharded into 10,000 Parquet files (~78MB each)
420B tokens total size (perfect for training a model for… See the full description on the dataset page: https://huggingface.co/datasets/pico-lm/pretokenized-dolma.rnacentral-pretokenized-shardedgemma-4-pretokenized-tracesdanish-pretokenized-16k
Danish Pretrain — pretokenized (byte-level BPE, 16k vocab) — v1 partial
Pretokenized version of jensjepsen/danish-pretrain
via jensjepsen/danish-tokenizer.
Note: This is a partial version — 94 of 96 shards. The 2 missing shards
crashed during tokenization (rare pathological content). See
jensjepsen/danish-pretokenized-16k-supp for the recovered rows.
93,200,464 docs
Schema: input_ids: list<int32>, attention_mask: list<int8>
94 zstd-compressed parquet shards
yeast-no-GTL-pretokenized-NT-v2pretokenized-pretrain-test-128Kpretokenized-dolma-5Mmouse-pretokenized-NTPreTokenizedWikiEnhuman-and-mouse-pretokenized-NT-cachehuman-pretokenized-NTpretokenized-dolma-20Myeast-gene-sequence-homology-pretokenized-NTfineweb-edu-pretokenized-llama3-100b
FineWeb-Edu Pretokenized with Llama 3.1 (100B)
This repository contains the sample/100BT subset of FineWeb-Edu, pretokenized for Megatron-LM-style training with meta-llama/Meta-Llama-3.1-8B.
It is a derived, pretokenized version of the upstream data, not an official Hugging Face FineWeb release.
Dataset summary
140 indexed shards
97,270,686 non-empty documents
97,458,793,013 tokens
English web text from FineWeb-Edu sample/100BT
Source dataset revision:… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/fineweb-edu-pretokenized-llama3-100b.yeast-no-label-pretokenized-NThuman_and_mouse-pretokenized-NTfineweb-edu-pretokenized-10B
Marin/Levanter Subsampled Pretokenized Dataset
Dataset
Train Urls:
gs://marin-us-central2/raw/fineweb-edu-c2beb4/3c452cb/huggingface.co/datasets/HuggingFaceFW/fineweb-edu/resolve/3c452cb
Factsheet
Original cache: gs://marin-us-central2/tokenized/fineweb-edu-24698d
Tokenizer: stanford-crfm/marin-tokenizer
Seed 42
Number of tokens: 10,000,000,738
(This readme is automatically generated by Marin.)
BEELINE-HepG2-no-GTL-pretokenized-NTBEELINE-mDC-no-label-pretokenized-NTBEELINE-human-v2-pretokenized-NTBEELINE-mouse-v2-pretokenized-NTBEELINE-mESC-no-label-pretokenized-NTBEELINE-mHSC-no-label-pretokenized-NTpretraining-pretokenized-smollm3
SmolLM3 Pretokenized Pretraining Sources
Datatrove/Nanotron tokenized-byte versions of three public pretraining sources:
fineweb-edu-10bt: HuggingFaceFW/fineweb-edu, sample/10BT
finemath-4plus: HuggingFaceTB/finemath, finemath-4plus
stack-edu-python: HuggingFaceTB/stack-edu, Python, with content retrieved from the public Software Heritage S3 bucket using the upstream dataset-card procedure
All subsets use HuggingFaceTB/SmolLM3-3B at revision… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/pretraining-pretokenized-smollm3.100B-pretokenized-mix-32k
---
license: apache-2.0
task_categories:
- text-generation
language:
- en
tags:
- pretokenized
- slm-tokenizer-32k
- dclm
- fineweb-edu
- cosmopedia
- math
- finemath
- python
- code
- wikipedia
- educational
- web
- stem
size_categories:
- 100B<n<1T
---
SLM Pre-tokenized 100B Mix (32k Vocab)
This repository contains a pre-tokenized, multi-domain dataset mixture of approximately 100 Billion tokens across educational, web, encyclopedic, and… See the full description on the dataset page: https://huggingface.co/datasets/rijuludar/100B-pretokenized-mix-32k.smolmo-sft-olmocore-pretokenizedBEELINE-hESC-no-GTL-pretokenized-NTBEELINE-mDC-no-GTL-pretokenized-NTfineweb-edu-pretokenized-10b
FineWeb-Edu Pretokenized 10B
Megatron indexed-dataset (.bin/.idx) versions of
HuggingFaceFW/fineweb-edu sample/10BT.
Each Hugging Face subset contains a small metadata.parquet index. The actual
Megatron files are under <subset>/files/; pass each prefix without the
.bin/.idx suffix to Megatron Core or Megatron Bridge.
Documents retain upstream shard and row order. Tokenization disables automatic
special-token insertion and appends exactly one tokenizer EOS/EOD token to each… See the full description on the dataset page: https://huggingface.co/datasets/ZhuofengLi/fineweb-edu-pretokenized-10b.
