20b
Datasets
All datasets matching “20b”pretrain_v1_20b
Chess Pre-to-Post — Pretraining Corpus v1 (20B)
Raw tokenized pretraining data for the Chess Pre-to-Post project, stored as
sharded NumPy arrays (shard_XXXX/raw.NNNN.npy).
[!IMPORTANT]
This is an earlier, smaller (20B) snapshot and is no longer maintained.
The maintained version of this dataset is
pavelslab-nyu/pretrain_v1_54B.
Please use that version for any new work — it supersedes this one.
Maintained version
➡️ pavelslab-nyu/pretrain_v1_54B… See the full description on the dataset page: https://huggingface.co/datasets/chess-pre-to-post/pretrain_v1_20b.gpt-oss-20b-moe-expert-power-traces-320k
GPT-OSS-20B MoE Expert Power Traces (320k, ChipWhisperer)
This dataset contains analog power traces captured with a ChipWhisperer Husky while running forced single-expert MoE computations derived from openai/gpt-oss-20b on an NVIDIA H100.
What is recorded
Each trace corresponds to one capture trial where:
A fixed expert id is selected (expert_00 ... expert_31).
A random hidden-state tensor is generated once per trial.
The selected expert computation is executed… See the full description on the dataset page: https://huggingface.co/datasets/masterpieceexternal/gpt-oss-20b-moe-expert-power-traces-320k.finepdfs_50BT-dclm_30BT-fineweb_edu_20BT
FinePDFs 50BT + DCLM 30BT + FineWeb-Edu 20BT
A ~100 billion token pretraining mixture combining three high-quality English data sources in a 50-30-20 ratio.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Inspired by optimal dataset mixing.
Dataset Description
Component
Source
Tokens
FinePDFs 100BT
FinePDFs
~50B
DCLM 100BT
DCLM-Baseline 1.0
~30B
FineWeb-Edu 100BT
FineWeb-Edu
~20B
The schema is reduced to the… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_50BT-dclm_30BT-fineweb_edu_20BT.finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT
FinePDFs-Edu 50BT + DCLM 30BT + FineWeb-Edu 20BT
A ~100 billion token pretraining mixture combining three high-quality English data sources in a 50-30-20 ratio, using the educational subset of FinePDFs.
Part of the Smol-Data collection — tried and tested mixes for strong pretraining. Inspired by optimal dataset mixing.
Dataset Description
Component
Source
Tokens
FinePDFs-Edu 100BT
FinePDFs-Edu
~50B
DCLM 100BT
DCLM-Baseline 1.0
~30B
FineWeb-Edu 100BT… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceFW/finepdfs_edu_50BT-dclm_30BT-fineweb_edu_20BT.Nemotron-CC-HQ-20B
Nemotron-CC-HQ-20B
This Dataset consists of approximately 20B tokens of Nemotron-CC-HQ, consisting of randomly sampled slices from crawls in the range CC-MAIN-2013-20-part-00012 to CC-MAIN-2019-04-part-00007.
For more information about Nemotron-CC check the Paper by Nvidia
Disclaimer:
Derived from Nemotron-CC (Common Crawl). No ownership of underlying content is claimed.
Data may be subject to third-party rights. Use at your own risk and in compliance with… See the full description on the dataset page: https://huggingface.co/datasets/Fredithefish/Nemotron-CC-HQ-20B.code-20b
Dataset Card for "code_20b2"
More Information needed
