CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01common-pile /library_of_congress_filtered Library of Congress Description The Library of Congress (LoC) curates a collection of public domain books called "Selected Digitized Books". We have downloaded over 130,000 English-language books from this public domain collection as OCR plain text files using the LoC APIs. Dataset Statistics Documents UTF-8 GB 129,052 35.6 License Issues While we aim to produce datasets with completely accurate licensing information, license… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/library_of_congress_filtered.texttext-generation100K<n<1M2 likes1k downloads1y agoHugging Face02common-pile /biodiversity_heritage_library_filtered Biodiversity Heritage Library Description The Biodiversity Heritage Library (BHL) is an open-access digital library for biodiversity literature and archives. This dataset contains over 15 million public domain books and documents from the BHL collection. These works were collected using the bulk data download interface provided by the BHL and were filtered based on their associated license metadata. We use the optical character recognition (OCR)-generated text distributed… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/biodiversity_heritage_library_filtered.texttext-generation10M<n<100M2 likes861 downloads1y agoHugging Face03common-pile /biodiversity_heritage_library Biodiversity Heritage Library Description The Biodiversity Heritage Library (BHL) is an open-access digital library for biodiversity literature and archives. This dataset contains over 42 million public domain books and documents from the BHL collection. These works were collected using the bulk data download interface provided by the BHL and were filtered based on their associated license metadata. We use the optical character recognition (OCR)-generated text… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/biodiversity_heritage_library.texttext-generation10M<n<100M2 likes668 downloads1y agoHugging Face04common-pile /library_of_congress Library of Congress (subset of Common Pile) Description The Library of Congress (LoC) curates a collection of public domain books called "Selected Digitized Books". We have downloaded over 130,000 English-language books from this public domain collection as OCR plain text files using the LoC APIs. This dataset is a subset of the Common Pile v0.1. For more information, see The Common Pile v0.1 paper. Dataset Statistics Documents UTF-8 GB 135,500… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/library_of_congress.texttext-generation100K<n<1M1 likes606 downloads10mo agoHugging Face05Emulated-Inc /library-python-training-pool Python library function-writing training pool A pool of public data for training a model to write Python functions, many of them calling libraries: 8.3 percent of the answers in the normalised layer import a library that is not in the Python standard library. It is a straight collection of open datasets, not a new corpus: every row comes from one of the sources below, at the revision named. Rows an overlap filter flagged against held-out material this pool is kept separate from… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/library-python-training-pool.texttext-generation1M<n<10M0 likes62 downloads13d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.