CoolFace
Datasetpublic

JonasGeiping/the_pile_WordPiecex32768_8eb2d0ea9da707676c81314c4ea04507

Dataset Card for "the_pile_WordPiecex32768_8eb2d0ea9da707676c81314c4ea04507" Dataset Summary This is a preprocessed, tokenized dataset for the cramming-project. Use only with the tokenizer uploaded here. This version is 8eb2d0ea9da707676c81314c4ea04507, which corresponds to a specific dataset construction setup, described below. The raw data source is the Pile, a 825 GiB diverse, open source language modelling data set that consists of 22 smaller, high-quality… See the full description on the dataset page: https://huggingface.co/datasets/JonasGeiping/the_pile_WordPiecex32768_8eb2d0ea9da707676c81314c4ea04507.

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes719downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face