CoolFace
Datasetpublic

ArmelR/the-pile-splitted

Dataset description The pile is an 800GB dataset of english text designed by EleutherAI to train large-scale language models. The original version of the dataset can be found here. The dataset is divided into 22 smaller high-quality datasets. For more information each of them, please refer to the datasheet for the pile. However, the current version of the dataset, available on the Hub, is not splitted accordingly. We had to solve this problem in order to improve the user… See the full description on the dataset page: https://huggingface.co/datasets/ArmelR/the-pile-splitted.

sourceHugging Faceupdated 3y agoView on Hugging Face
23likes17kdownloads
settings

This repository belongs to ArmelR on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namethe-pile-splitted
visibilitypublic
licencenot set
gatedno
ownerArmelR
Account settings
ArmelR/the-pile-splitted · CoolFace