CoolFace
Datasetpublic

stanford-crfm/DSIR-filtered-pile-100M-short

Dataset Card for DSIR-filtered-pile-100M-short Dataset Summary This dataset is a subset of The Pile, selected via the DSIR data selection method. The target distribution for DSIR is the Wikipedia and BookCorpus2 subsets of The Pile. Languages English (EN) Dataset Structure A train set is provided (102M examples) with a small validation and test set (50k examples each). This dataset is more suitable for training shorter LMs (128 or 256… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/DSIR-filtered-pile-100M-short.

sourceHugging Facemitupdated 4y agoView on Hugging Face
1likes37downloads
4 commits on main
7d0eea24y ago

readme

sangmichaelxie
2b9570d4y ago

README

sangmichaelxie
fcd33744y ago

add data files

sangmichaelxie
ac018154y ago

initial commit

sangmichaelxie