datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
finepdfs_50BT-dclm_30BT-fineweb_edu_20BT
FinePDFs 50BT + DCLM 30BT + FineWeb-Edu 20BT
A ~100 billion token pretraining mixture combining three high-quality English data sources in a 50-30-20 ratio.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Inspired by optimal dataset mixing.
Dataset Description
Component
Source
Tokens
FinePDFs 100BT
FinePDFs
~50B
DCLM 100BT
DCLM-Baseline 1.0
~30B
FineWeb-Edu 100BT
FineWeb-Edu
~20B
The schema is reduced to the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT.code-20b
Dataset Card for "code_20b2"
More Information needed
finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT
FinePDFs-Edu 50BT + DCLM 30BT + FineWeb-Edu 20BT
A ~100 billion token pretraining mixture combining three high-quality English data sources in a 50-30-20 ratio, using the educational subset of FinePDFs.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Inspired by optimal dataset mixing.
Dataset Description
Component
Source
Tokens
FinePDFs-Edu 100BT
FinePDFs-Edu
~50B
DCLM 100BT
DCLM-Baseline 1.0
~30B
FineWeb-Edu 100BT… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT.Nemotron-CC-HQ-20B
Nemotron-CC-HQ-20B
This Dataset consists of approximately 20B tokens of Nemotron-CC-HQ, consisting of randomly sampled slices from crawls in the range CC-MAIN-2013-20-part-00012 to CC-MAIN-2019-04-part-00007.
For more information about Nemotron-CC check the Paper by Nvidia
Disclaimer:
Derived from Nemotron-CC (Common Crawl). No ownership of underlying content is claimed.
Data may be subject to third-party rights. Use at your own risk and in compliance with… See the full description on the dataset page: https://huggingface.co/datasets/Fredithefish/Nemotron-CC-HQ-20B.code_20b
Dataset Card for "code_20b"
More Information needed
finewebedu-20B
FineWebEDU 20B
A copy of FineWebEDU-20B used for out tokenizer experiments. The subsets are as follows:
bytelevel: the full dataset tokenized using our bytelevel tokenizer
bytelevel-subset_1: a 100k-row subset of the bytelevel subset, used to train bytelevel models.
bytelevel-subset_2: a 100k-row subset of the bytelevel subset, used to extract llm predictions.
bytelevel-llm-data: a copy of bytelevel-subset_2 with lm predictions, used to train bytespan tokenizers… See the full description on the dataset page: https://huggingface.co/datasets/ByteSpanTokenisers/finewebedu-20B.finepdfs_50BT-dclm_30BT-fineweb_edu_20BT-shuffled
FinePDFs 50BT + DCLM 30BT + FineWeb-Edu 20BT (Shuffled)
A globally shuffled version of HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B token mixture (50B FinePDFs + 30B DCLM + 20B FineWeb-Edu) but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining —… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT-shuffled.finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT-shuffled
FinePDFs-Edu 50BT + DCLM 30BT + FineWeb-Edu 20BT (Shuffled)
A globally shuffled version of HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining.
Dataset Description
This dataset contains the same ~100B token mixture (50B FinePDFs-Edu + 30B DCLM + 20B FineWeb-Edu) but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT-shuffled.dolma_20bn_wiki_upsamplefineweb-nanochatbpe-20B
fineweb-nanochatbpe-20B
FineWeb-Edu (from karpathy/fineweb-edu-100b-shuffle), pre-tokenized with the nanochatbpe tokenizer (vocab 65,536)
and packaged as flat uint16 token-id .bin files for fast memmap training.
file
split
tokens
train.bin
train
20,000,000,000
val.bin
val
52,336,096
train and val are disjoint held-out partitions. Each .bin is a raw
little-endian uint16 stream (no header); token count = filesize / 2, and
train.meta.json / val.meta.jsoncarry the full… See the full description on the dataset page: https://huggingface.co/datasets/alexkstern/fineweb-nanochatbpe-20B.gpt-oss-20b-rollouts
GPT-OSS-20B Rollouts
Generated rollouts from GPT-OSS-20B with parsed Harmony channels (assistant thinking/final).
Schema: user_content, system_reasoning_effort, assistant_thinking, assistant_content.
Loading example: load_dataset("andyrdt/gpt-oss-20b-rollouts", "HarmBench", split="standard_train").
Notes
This repository uses manual configuration to expose both subset (config) and split dropdowns in the viewer.
Safety and jailbreak
HarmBench: Safety prompts… See the full description on the dataset page: https://huggingface.co/datasets/andyrdt/gpt-oss-20b-rollouts.gpt-oss-20b-combined-outputsdolma_20bn_instruct_upsamplethe-stack-v2-20B-samplegpt_oss_20b_doorkey_boundary_activationsfineweb-edu-20B-E1-Scramble-Hier-flatteneddolma_20bn_cc_high_qualitydolma_20bn_prop_stratified_sampleSmolLM2-1.7B-stage-4-20BFineWeb2-HQ-ja-20B元のデータセットFineWeb2-HQ
元のデータセットは多言語で巨大なため、扱いやすい用に日本語データを約200GBだけ抽出したデータセットです
wc 結果
1763269 38541549 5370473709 fineweb_jpn_Jpan_chunk_0.jsonl
1784158 37430170 5370514369 fineweb_jpn_Jpan_chunk_1.jsonl
1639554 40065129 5370372344 fineweb_jpn_Jpan_chunk_10.jsonl
1575127 42167166 5370298354 fineweb_jpn_Jpan_chunk_11.jsonl
1686375 39225898 5370402506 fineweb_jpn_Jpan_chunk_12.jsonl
1786948 36456352 5370498572… See the full description on the dataset page: https://huggingface.co/datasets/dahara1/FineWeb2-HQ-ja-20B.dclm_20b
DCLM used in MoCa Pre-training
🏠 Homepage | 💻 Code | 🤖 MoCa-Qwen25VL-7B | 🤖 MoCa-Qwen25VL-3B | 📚 Datasets | 📄 Paper
Introduction
This is a text pre-training dataset used in the modality-aware continual pre-training of MoCa models. It is adapted from DCLM and randomly downsampled to ~20B tokens.
The dataset consists of text examples. text is a string containing text while images are left blank intentionally since there is no image available.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/moca-embed/dclm_20b.old_code_20b_separatefineweb-edu-20bfinewebedu-20BThis is a subset of the HuggingFaceFW/fineweb-edu/100BT dataset.
I extracted (in order) the initial 20,200,000 rows where, ideally, 20M are meant for training and 200k for validation.
Tokenised configs:
bpe32000minipile: 21.6B tokens
License
For the license, refer to the original dataset (HuggingFaceFW/fineweb-edu).
dolma_20bn_no_instructFineWeb-Edu-20B-E2-v3-Obfuscatedweek1-general-20b-dolma2-v1
Week-One General 20B Dolma2
This is a deterministic, pretokenized 20-billion-token baseline corpus for
controlled language-model architecture and training experiments. It contains
nested 100M, 1B, 5B, and 20B views; each larger view is an exact ordered
extension of the previous view. It also includes a dataset-only
370m-1.25xc view: the first 9,281,564,672 packed tokens of the verified 20B
order, sized for the 1.25xC target of the OLMo-ladder 370M parameter count.
The repository… See the full description on the dataset page: https://huggingface.co/datasets/ericrcwu/week1-general-20b-dolma2-v1.FineWeb-Edu-20B-E2-OBF-FULLgpt-oss20b-samples-dedupA simple deduplicated variant of https://huggingface.co/datasets/jxm/gpt-oss20b-samples
Given the predictability of synthetic data we opted for a simple strategy: keeping the unique combinations of first and last ten words. Total count of unique occurrences is available in the column occurrence_count.
FineWeb-Edu-20B-E2-v3
