datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openwebtext
Dataset Card for "openwebtext"
Dataset Summary
An open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2.
This distribution was created by Aaron Gokaslan and Vanya Cohen of Brown University.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
plain_text
Size of downloaded dataset… See the full description on the dataset page: https://huggingface.co/datasets/Skylion007/openwebtext.openwebtext-100k
Dataset Card for "openwebtext-100k"
More Information needed
openwebtext2-first-30-chunks-lang-detect-raw-output
Counting bilingual and monolingual instances
In order to count bilingual and monolingual instances, we use the following code. We count bilingual instances where there are two languages, one of them is English and the other is either German, French, Spanish, Italian, Portuguese or Dutch. All other instances fall into the "Other" category.
from datasets import load_dataset
import json
from tqdm import tqdm
#Specify the dataset name
dataset_name =… See the full description on the dataset page: https://huggingface.co/datasets/RaiBP/openwebtext2-first-30-chunks-lang-detect-raw-output.openwebtext_quality_score_v1
Dataset Card for "openwebtext_quality_score_v1"
Adding quality score v1 to Skylion007/openwebtext
More Information needed
openwebtext
Dataset Card for "openwebtext"
Dataset Summary
An open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2.
This distribution was created by Aaron Gokaslan and Vanya Cohen of Brown University.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
plain_text
Size of downloaded dataset files: 13.51 GB
Size of the… See the full description on the dataset page: https://huggingface.co/datasets/dylanebert/openwebtext.gpt2_model_acts_openwebtextthe_pile_openwebtext2
Dataset Card for "the_pile_openwebtext2"
More Information needed
openwebtext2A cleaned version of OpenWebText2 by removing non-English, duplicated, copyrighted, and low-quality (too short, too many special characters, etc) samples.
This dataset has also been decontaminated with respect to the following benchmarks based on n-gram overlap:
GLUE (dev set of SST-2, CoLA, QQP, WNLI, RTE, QNLI, MNLI; test set of MPRC)
SIQA, PIQA, QASC, CSQA, HellaSWAG (all dev set)
CONLL 2003
BLIMP
MAIN
BoolQ (dev set)
WinoGrande (dev set)
ANLI (test set)
ARC easy and challenge (test set)… See the full description on the dataset page: https://huggingface.co/datasets/Geralt-Targaryen/openwebtext2.openwebtextpresplit
Fixed OpenWebMath train/test split
This is an untokenized, deterministic shuffle of
open-web-math/open-web-math pinned at commit
fde8ef8de2300f5e778f56261843dab89f230815. It contains the original columns without transformation.
The shuffle seed is 20260904. The test set is the first 0.05% of shuffled rows
(rounded to the nearest whole document); all remaining rows are training data.
Exact counts and SHA-256 checksums are in split_manifest.json.
openwebtext_en
Dataset Card for "openwebtext_en"
More Information needed
openwebtextopenwebtext2-first-30-chunks-ablation-bilingualopenwebtext-all-minilm-l6-v2-embedding
Dataset Card for "openwebtext-all-minilm-l6-v2-embedding"
More Information needed
openwebtext
Dataset Card for "openwebtext"
Dataset Summary
An open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2.
This distribution was created by Aaron Gokaslan and Vanya Cohen of Brown University.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
plain_text
Size of downloaded dataset… See the full description on the dataset page: https://huggingface.co/datasets/ThomasKendrick/openwebtext.openwebtext2-first-30-chunks-ablation-non-englishopenwebtext-presplit
Fixed OpenWebText train/test split
This is an untokenized, deterministic shuffle of
Skylion007/openwebtext pinned at commit
79d93d786212f7344586290adb811d4ae6a1762c. It contains the original columns without transformation.
The shuffle seed is 2357. The test set is the first 0.05% of shuffled rows
(rounded to the nearest whole document); all remaining rows are training data.
Exact counts and SHA-256 checksums are in split_manifest.json.
openwebtext_20p
openwebtext_20p
first 20% of openwebtext
openwebtext-dep-spacyopenwebtext2-first-30-chunks-ablation-translationopenwebtext-sentences
OpenWebText-Sentences Dataset
Overview
This dataset is derived from the popular OpenWebText dataset (see here). It contains the same text content as the original OpenWebText, but split into individual sentences.
Key Features
Content: All text from the original OpenWebText dataset
Format: Sentences are stored individually now in parquet format for faster access
Order: Maintains all original OpenWebText text and the order thereof
Tokenization: Sentences were… See the full description on the dataset page: https://huggingface.co/datasets/PaulPauls/openwebtext-sentences.openwebtext2A cleaned version of OpenWebText2 by removing non-English, duplicated, copyrighted, and low-quality (too short, too many special characters, etc) samples.
This dataset has also been decontaminated with respect to the following benchmarks based on n-gram overlap:
GLUE (dev set of SST-2, CoLA, QQP, WNLI, RTE, QNLI, MNLI; test set of MPRC)
SIQA, PIQA, QASC, CSQA, HellaSWAG (all dev set)
CONLL 2003
BLIMP
MAIN
BoolQ (dev set)
WinoGrande (dev set)
ANLI (test set)
ARC easy and challenge (test set)… See the full description on the dataset page: https://huggingface.co/datasets/shaguftakhan2k17/openwebtext2.openwebtext2-first-30-chunks-english-only-examplesopenwebtext2-first-30-chunks-ablation-fullOpenWebText2
Dataset Card for OpenWebText2
OpenWebText2 is a reasonably large corpus of scraped natural language data.
Original hosting for this dataset has become difficult because it was hosted alongside another controversial dataset. To the best of my knowledge, this dataset itself is not encumbered in any way. It's a useful size for smaller language modelling experiments and is sometimes used in existing papers which it may be desirable to replicate. It is uploaded here to facilitate those… See the full description on the dataset page: https://huggingface.co/datasets/segyges/OpenWebText2.openwebtext-128small_openwebtextopenwebtext-subset-1M-rowsopenwebtext
Dataset Card for "openwebtext"
Dataset Summary
An open-source replication of the WebText dataset from OpenAI, that was used to train GPT-2.
This distribution was created by Aaron Gokaslan and Vanya Cohen of Brown University.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
plain_text
Size of downloaded dataset files: 13.51 GB
Size of the… See the full description on the dataset page: https://huggingface.co/datasets/reedie01/openwebtext.scaling_mia_the_pile_00_OpenWebText2pile_openwebtext2
