CoolFace
12 results

The Pile

EleutherAI /the_pile_deduplicatedtext100M<n<1B117 likes17k downloads4y agoHugging FaceArmelR /the-pile-splitted Dataset description The pile is an 800GB dataset of english text designed by EleutherAI to train large-scale language models. The original version of the dataset can be found here. The dataset is divided into 22 smaller high-quality datasets. For more information each of them, please refer to the datasheet for the pile. However, the current version of the dataset, available on the Hub, is not splitted accordingly. We had to solve this problem in order to improve the user… See the full description on the dataset page: https://huggingface.co/datasets/ArmelR/the-pile-splitted.text10M<n<100M23 likes17k downloads3y agoHugging Faceandstor /the_pile_github Dataset Card for The Pile GitHub Dataset Summary This is the GitHub subset of EleutherAi/The Pile dataset and contains GitHub repositories. The programming languages are identified using the guesslang library. A total of 54 programming languages are included in the dataset. Supported Tasks and Leaderboards [More Information Needed] Languages The following languages are covered by the dataset: 'Assembly', 'Batchfile', 'C', 'C#', 'C++', 'CMake'… See the full description on the dataset page: https://huggingface.co/datasets/andstor/the_pile_github.texttext-generation10M<n<100M10 likes1.1k downloads1y agoHugging FaceSaylorTwift /the_pile_books3_minus_gutenberg Dataset Card for "the_pile_books3_minus_gutenberg" More Information needed text100K<n<1M15 likes857 downloads4y agoHugging Facevietgpt /the_pile_openwebtext2 Dataset Card for "the_pile_openwebtext2" More Information needed text10M<n<100M5 likes846 downloads3y agoHugging FaceJonasGeiping /the_pile_WordPiecex32768_2efdb9d060d1ae95faf952ec1a50f020 Dataset Card for "the_pile_WordPiecex32768_2efdb9d060d1ae95faf952ec1a50f020" Dataset Summary This is a preprocessed, tokenized dataset for the cramming-project. Use only with the tokenizer uploaded here. This version is 2efdb9d060d1ae95faf952ec1a50f020, which corresponds to a specific dataset construction setup, described below. The raw data source is the Pile, a 825 GiB diverse, open source language modelling data set that consists of 22 smaller, high-quality datasets… See the full description on the dataset page: https://huggingface.co/datasets/JonasGeiping/the_pile_WordPiecex32768_2efdb9d060d1ae95faf952ec1a50f020.10M<n<100M1 likes750 downloads3y agoHugging Face