CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01storytracer /US-PD-BooksUPDATE: The Internet Archive has requested that this dataset be deleted (see discussion #2) because they consider the IA's metadata too unreliable to determine whether a book is in the public domain. To alleviate the IA's concerns, the full texts of the books have been removed from this dataset until a more reliable way to curate public domain books from the IA collections is established. The metadata and documentation remain for reference purposes. I was able to recreate one subcollection… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/US-PD-Books.tabulartext-generation100K<n<1M191 likes4.6k downloads3y agoHugging Face02PleIAs /French-PD-Books 🇫🇷 French Public Domain Books 🇫🇷 French-Public Domain-Book or French-PD-Books is a large collection aiming to agregate all the French monographies in the public domain. The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram search on very large cultural… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Books.tabulartext-generation100K<n<1M52 likes2.5k downloads3y agoHugging Face03storytracer /LoC-PD-Books Library of Congress Public Domain Books (English) This dataset contains more than 140,000 English books (~ 8 billion words) digitised by the Library of Congress (LoC) that are in the public domain in the United States. The dataset was compiled by Sebastian Majstorovic. Curation method The dataset was curated using the LoC JSON API and filtering the Selected Digitized Books collection for English books. Dataset summary The dataset contains 140,000 OCR texts (~… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/LoC-PD-Books.tabulartext-generation10K<n<100K42 likes1.3k downloads3y agoHugging Face04BrightData /Goodreads-Books Dataset Card for "BrightData/Goodreads-Books" Dataset Summary Explore a collection of millions of books with the Goodreads dataset, comprising over 6.3M structured records and 14 data fields updated and refreshed regularly. Each entry includes all major data points such as URLs, book IDs, titles, authors, ratings, number of ratings, reviews, summaries, genres, publication dates, author details and prices. For a complete list of data points, please refer to the full "Data… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/Goodreads-Books.tabulartext-classification1M<n<10M20 likes420 downloads2y agoHugging Face05PleIAs /Ukrainian-CulturalHeritage-Books 🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦 Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain. Dataset summary The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources. Curation method The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Ukrainian-CulturalHeritage-Books.tabulartext-generation10K<n<100K4 likes151 downloads3y agoHugging Face06dataseek /ptbr-books-publicos PT-BR Public-Domain Books Part of the MagTina350m pretrain corpus release by Dataseek under the Magestic.ai brand. This is one of nine silver-layer datasets that fed dataseek/magtina350m-base. Summary 28 K curated Brazilian-Portuguese public-domain books. Used at ~2 epochs in MagTina350m pretrain (79 M unique tokens sampled to 158 M consumed). Useful as a small but high-quality literary slice. Source and collection method Curated public-domain Brazilian books… See the full description on the dataset page: https://huggingface.co/datasets/dataseek/ptbr-books-publicos.tabulartext-generation10K<n<100K0 likes92 downloads5mo agoHugging Face07codealchemist01 /goodreads-books Goodreads Books Dataset Dataset Description A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics. This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for: 📚 Book recommendation systems 📊 Literary data analysis 🤖 Machine learning projects 📈 Rating prediction models 🔍 Book discovery algorithms Dataset Structure Features… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/goodreads-books.tabulartext-classification1K<n<10K0 likes79 downloads11mo agoHugging Face08ysngkil /whole-books Whole Books Five book-length corpora for long-context language-model pretraining, packaged so that one row is one whole book. Together: 202,343 books, about 20.9 B tokens in a 32K-vocabulary Llama-style tokenizer, with the large majority of tokens inside books of 64K tokens or more. Built 2026-09-06 from pinned snapshots of the sources below; nothing was filtered, deduplicated or cleaned beyond what the sources had already done, and the reassembly steps are documented per… See the full description on the dataset page: https://huggingface.co/datasets/ysngkil/whole-books.tabulartext-generation100K<n<1M0 likes65 downloads20d agoHugging Face09Chima207 /Goodreads-Books Dataset Card for "BrightData/Goodreads-Books" Dataset Summary Explore a collection of millions of books with the Goodreads dataset, comprising over 6.3M structured records and 14 data fields updated and refreshed regularly. Each entry includes all major data points such as URLs, book IDs, titles, authors, ratings, number of ratings, reviews, summaries, genres, publication dates, author details and prices. For a complete list of data points, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/Chima207/Goodreads-Books.tabulartext-classification1M<n<10M0 likes49 downloads8mo agoHugging Face10freococo /tipitaka_myanmar_translation_books Myanmar Tipitaka Translation (60 Books) This dataset contains the complete Myanmar (Burmese) translation of the Tipitaka (Pali Canon), together with the major Atthakatha (Commentaries) and the Visuddhimagga. The texts have been converted into a clean, structured JSONL format, suitable for: Natural Language Processing (NLP) LLM Training & Fine-tuning Digital Humanities Research Dhamma Study Applications 📊 Dataset Statistics Total Books: 60 Total Content Lines: 194… See the full description on the dataset page: https://huggingface.co/datasets/freococo/tipitaka_myanmar_translation_books.tabulartext-generation100K<n<1M0 likes45 downloads8mo agoHugging Face11pszemraj /LoC-PD-Books-preprocessed LoC-PD-Books: preprocessed This is the storytracer/LoC-PD-Books dataset with the following preprocessing steps: apply clean-text package keeping casing and newlines drop OCR garbled text in first few lines of each example fix (most) 'hard' newlines w/ regex similar to gutenberg clean 'grade' first 512 tokens of each book with this quantized model; keep examples from labels clean (all) and mild gibberish w/ score 0.9 or higher tabulartext-generation10K<n<100K1 likes35 downloads9mo agoHugging Face12bstarrs /goodreads-books Goodreads Books Dataset Dataset Description A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics. This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for: 📚 Book recommendation systems 📊 Literary data analysis 🤖 Machine learning projects 📈 Rating prediction models 🔍 Book discovery algorithms Dataset Structure Features… See the full description on the dataset page: https://huggingface.co/datasets/bstarrs/goodreads-books.tabulartext-classification1K<n<10K0 likes27 downloads6mo agoHugging Face13BuzzBlitz360A /Ukrainian-CulturalHeritage-Books 🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦 Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain. Dataset summary The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources. Curation method The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/BuzzBlitz360A/Ukrainian-CulturalHeritage-Books.tabulartext-generation10K<n<100K0 likes18 downloads8mo agoHugging Face14actixon /US-PD-BooksUPDATE: The Internet Archive has requested that this dataset be deleted (see discussion #2) because they consider the IA's metadata too unreliable to determine whether a book is in the public domain. To alleviate the IA's concerns, the full texts of the books have been removed from this dataset until a more reliable way to curate public domain books from the IA collections is established. The metadata and documentation remain for reference purposes. I was able to recreate one subcollection… See the full description on the dataset page: https://huggingface.co/datasets/actixon/US-PD-Books.tabulartext-generation100K<n<1M0 likes16 downloads6mo agoHugging Face15DexopT /Books-General-Linux Linux Books Dataset Dataset Description The Linux Books Dataset is a curated text dataset derived from Linux-related books and learning materials. It focuses on Linux system administration, cybersecurity, networking, shell scripting, and operating system fundamentals.The dataset is designed to support training and evaluation of NLP models for technical domains, especially cybersecurity-aware language models and Linux-focused assistants. This dataset is suitable for both… See the full description on the dataset page: https://huggingface.co/datasets/DexopT/Books-General-Linux.tabulartext-generation1K<n<10K0 likes11 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.