JonasGeiping/the_pile_WordPiecex32768_8eb2d0ea9da707676c81314c4ea04507
Dataset Card for "the_pile_WordPiecex32768_8eb2d0ea9da707676c81314c4ea04507" Dataset Summary This is a preprocessed, tokenized dataset for the cramming-project. Use only with the tokenizer uploaded here. This version is 8eb2d0ea9da707676c81314c4ea04507, which corresponds to a specific dataset construction setup, described below. The raw data source is the Pile, a 825 GiB diverse, open source language modelling data set that consists of 22 smaller, high-quality… See the full description on the dataset page: https://huggingface.co/datasets/JonasGeiping/the_pile_WordPiecex32768_8eb2d0ea9da707676c81314c4ea04507.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face