stanford-crfm/DSIR-filtered-pile-100M-short
Dataset Card for DSIR-filtered-pile-100M-short Dataset Summary This dataset is a subset of The Pile, selected via the DSIR data selection method. The target distribution for DSIR is the Wikipedia and BookCorpus2 subsets of The Pile. Languages English (EN) Dataset Structure A train set is provided (102M examples) with a small validation and test set (50k examples each). This dataset is more suitable for training shorter LMs (128 or 256… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/DSIR-filtered-pile-100M-short.
137
readme
README
add data files
initial commit
