JonasGeiping/the_pile_WordPiecex32768_8eb2d0ea9da707676c81314c4ea04507
Dataset Card for "the_pile_WordPiecex32768_8eb2d0ea9da707676c81314c4ea04507" Dataset Summary This is a preprocessed, tokenized dataset for the cramming-project. Use only with the tokenizer uploaded here. This version is 8eb2d0ea9da707676c81314c4ea04507, which corresponds to a specific dataset construction setup, described below. The raw data source is the Pile, a 825 GiB diverse, open source language modelling data set that consists of 22 smaller, high-quality… See the full description on the dataset page: https://huggingface.co/datasets/JonasGeiping/the_pile_WordPiecex32768_8eb2d0ea9da707676c81314c4ea04507.
0719
