datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openwebtext
Dataset Card for "openwebtext"
Dataset Summary
An open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2.
This distribution was created by Aaron Gokaslan and Vanya Cohen of Brown University.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
plain_text
Size of downloaded dataset… See the full description on the dataset page: https://huggingface.co/datasets/Skylion007/openwebtext.Skylion007-openwebtext-tokenizer-gpt2openwebtextopenwebtext-gpt2openwebtext_quality_score_v1
Dataset Card for "openwebtext_quality_score_v1"
Adding quality score v1 to Skylion007/openwebtext
More Information needed
openwebtext-100k
Dataset Card for "openwebtext-100k"
More Information needed
openwebtext
Dataset Card for "openwebtext"
Dataset Summary
An open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2.
This distribution was created by Aaron Gokaslan and Vanya Cohen of Brown University.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
plain_text
Size of downloaded dataset files: 13.51 GB
Size of the… See the full description on the dataset page: https://huggingface.co/datasets/dylanebert/openwebtext.the_pile_openwebtext2
Dataset Card for "the_pile_openwebtext2"
More Information needed
openwebtext-gemma
OpenWebTextCorpus tokenized for Gemma
This dataset is a pre-tokenized version of the Skylion007/openwebtext dataset
using the gemma tokenizer. As such, this dataset follows the same licensing as the original openwebtext dataset.
This pre-tokenization is done as a performance optimization for using the openwebtext dataset with a Gemma model (gemma-2b, gemma-2b-it, gemma-7b, gemma-7b-it).
This dataset was created using SAELens, with the following settings:
context_size: 8192… See the full description on the dataset page: https://huggingface.co/datasets/chanind/openwebtext-gemma.Skylion007-openwebtext-tokenizer-gpt2-128openwebtext-tokenized-Llama-3.2OpenWebText dataset (open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2) tokenized for Llama 3.2 models
Useful for accelerated training and testing of sparse autoencoders
Context size: 128, not shuffled
openwebtext-tokenized-9b
Dataset Card for "openwebtext-tokenized-9b"
More Information needed
openwebtext-gemma-1024openwebtext2A cleaned version of OpenWebText2 by removing non-English, duplicated, copyrighted, and low-quality (too short, too many special characters, etc) samples.
This dataset has also been decontaminated with respect to the following benchmarks based on n-gram overlap:
GLUE (dev set of SST-2, CoLA, QQP, WNLI, RTE, QNLI, MNLI; test set of MPRC)
SIQA, PIQA, QASC, CSQA, HellaSWAG (all dev set)
CONLL 2003
BLIMP
MAIN
BoolQ (dev set)
WinoGrande (dev set)
ANLI (test set)
ARC easy and challenge (test set)… See the full description on the dataset page: https://huggingface.co/datasets/Geralt-Targaryen/openwebtext2.Skylion007-openwebtext-gpt2-1024Tokenizer: gpt2dataset: '''Skylion007/openwebtext'''context_size : 1024
openwebtextSkylion007-openwebtext-tokenizer-EleutherAI-gpt-neox-20bopenwebtextpresplit
Fixed OpenWebMath train/test split
This is an untokenized, deterministic shuffle of
open-web-math/open-web-math pinned at commit
fde8ef8de2300f5e778f56261843dab89f230815. It contains the original columns without transformation.
The shuffle seed is 20260904. The test set is the first 0.05% of shuffled rows
(rounded to the nearest whole document); all remaining rows are training data.
Exact counts and SHA-256 checksums are in split_manifest.json.
openwebtext_en
Dataset Card for "openwebtext_en"
More Information needed
gpt2_tokenized_concatenated_openwebtextSkylion007-openwebtext-tokenizer-gpt2-64openwebtext2-first-30-chunks-ablation-bilingualopenwebtext-gemma3-tokenized-1024-activations-layer23
OpenWebText — Gemma-3-1B Hidden State Activations (Layer 23)
Precomputed hidden state activations before layer 23 of Gemma-3-1B-IT for the OpenWebText dataset, tokenized with sequence length 1024.
Designed for training a Titans memory layer that replaces layer 23 of Gemma 3.
Dataset Structure
Each example contains the inputs to layer 23:
Field
Shape
Dtype
Description
activations
(1024, 1152)
float32
Hidden state activations (cast from bfloat16)
mask(1024… See the full description on the dataset page: https://huggingface.co/datasets/veriga/openwebtext-gemma3-tokenized-1024-activations-layer23.openwebtext-all-minilm-l6-v2-embedding
Dataset Card for "openwebtext-all-minilm-l6-v2-embedding"
More Information needed
openwebtext-gemma-2-context-128
OpenWebTextCorpus tokenized for Gemma 2 with 128 context size
This dataset is a pre-tokenized version of the Skylion007/openwebtext dataset
using the gemma tokenizer. As such, this dataset follows the same licensing as the original openwebtext dataset.
This pre-tokenization is done as a performance optimization for using the openwebtext dataset with a Gemma model (gemma-2b, gemma-2b-it, gemma-7b, gemma-7b-it).
This dataset was created using SAELens, with the following settings:… See the full description on the dataset page: https://huggingface.co/datasets/Marlon154/openwebtext-gemma-2-context-128.openwebtext-tokenized-16kopenwebtext-tokenized-8kopenwebtext
Dataset Card for "openwebtext"
Dataset Summary
An open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2.
This distribution was created by Aaron Gokaslan and Vanya Cohen of Brown University.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
plain_text
Size of downloaded dataset… See the full description on the dataset page: https://huggingface.co/datasets/ThomasKendrick/openwebtext.openwebtext-presplit
Fixed OpenWebText train/test split
This is an untokenized, deterministic shuffle of
Skylion007/openwebtext pinned at commit
79d93d786212f7344586290adb811d4ae6a1762c. It contains the original columns without transformation.
The shuffle seed is 2357. The test set is the first 0.05% of shuffled rows
(rounded to the nearest whole document); all remaining rows are training data.
Exact counts and SHA-256 checksums are in split_manifest.json.
openwebtext-tokenized-4k
