CoolFace
Datasetpublic

stanford-crfm/DSIR-filtered-pile-100M-short

Dataset Card for DSIR-filtered-pile-100M-short Dataset Summary This dataset is a subset of The Pile, selected via the DSIR data selection method. The target distribution for DSIR is the Wikipedia and BookCorpus2 subsets of The Pile. Languages English (EN) Dataset Structure A train set is provided (102M examples) with a small validation and test set (50k examples each). This dataset is more suitable for training shorter LMs (128 or 256… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/DSIR-filtered-pile-100M-short.

sourceHugging Facemitupdated 4y agoView on Hugging Face
1likes37downloads
settings

This repository belongs to stanford-crfm on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameDSIR-filtered-pile-100M-short
visibilitypublic
licencemit
gatedno
ownerstanford-crfm
Account settings
stanford-crfm/DSIR-filtered-pile-100M-short · CoolFace