CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01storytracer /US-PD-BooksUPDATE: The Internet Archive has requested that this dataset be deleted (see discussion #2) because they consider the IA's metadata too unreliable to determine whether a book is in the public domain. To alleviate the IA's concerns, the full texts of the books have been removed from this dataset until a more reliable way to curate public domain books from the IA collections is established. The metadata and documentation remain for reference purposes. I was able to recreate one subcollection… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/US-PD-Books.tabulartext-generation100K<n<1M191 likes4.5k downloads3y agoHugging Face02PleIAs /French-PD-Books 🇫🇷 French Public Domain Books 🇫🇷 French-Public Domain-Book or French-PD-Books is a large collection aiming to agregate all the French monographies in the public domain. The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram search on very large cultural… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Books.tabulartext-generation100K<n<1M52 likes2.5k downloads3y agoHugging Face03storytracer /LoC-PD-Books Library of Congress Public Domain Books (English) This dataset contains more than 140,000 English books (~ 8 billion words) digitised by the Library of Congress (LoC) that are in the public domain in the United States. The dataset was compiled by Sebastian Majstorovic. Curation method The dataset was curated using the LoC JSON API and filtering the Selected Digitized Books collection for English books. Dataset summary The dataset contains 140,000 OCR texts (~… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/LoC-PD-Books.tabulartext-generation10K<n<100K42 likes1.4k downloads3y agoHugging Face04murodbek /uz-books Dataset Card for BookCorpus Dataset Summary In an effort to democratize research on low-resource languages, we release UzBooks dataset, a cleaned book corpus consisting of nearly 40000 books in Uzbek Language divided into two branches: "original" and "lat," representing the OCRed (Latin and Cyrillic) and fully Latin versions of the texts, respectively. Please refer to our blogpost and paper (Coming soon!) for further details. To load and use dataset, run this script:… See the full description on the dataset page: https://huggingface.co/datasets/murodbek/uz-books.texttext-generation10K<n<100K23 likes310 downloads1y agoHugging Face05tahrirchi /uz-books-v2 Dataset Card for UzBooks V2 Dataset Summary UzBooks V2 is an improved version of the UzBooks book corpus for Uzbek language. It contains nearly 40,000 books in two splits: Split Description Examples lat Fully Latin-transliterated version 38,339 cyr Fully Cyrillic-transliterated version 38,339 What's New in V2? OCR Engine Upgrade: Switched from Tesseract → Google Cloud Vision OCR Cleaner Text: Google OCR produces far fewer recognition… See the full description on the dataset page: https://huggingface.co/datasets/tahrirchi/uz-books-v2.texttext-generation10K<n<100K5 likes277 downloads6mo agoHugging Face06PleIAs /Ukrainian-CulturalHeritage-Books 🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦 Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain. Dataset summary The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources. Curation method The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Ukrainian-CulturalHeritage-Books.tabulartext-generation10K<n<100K4 likes165 downloads3y agoHugging Face07alielfilali01 /Hindawi-Books-dataset Dataset Card for "Hindawi Books Dataset" Hindawi Books Dataset is a large collection of more than 3000 books written in Modern Standard Arabic. Dataset Description Hindawi Books Dataset offers a rich and diverse collection of literary works, covering various topics and genres, all written in Modern Standard Arabic. The dataset includes information about each book, such as the title, author name, book abstract, and a link to access the complete text online. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/alielfilali01/Hindawi-Books-dataset.texttext-generation10K<n<100K14 likes158 downloads3y agoHugging Face08Lots-of-LoRAs /task1650_opus_books_en-fi_translation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1650_opus_books_en-fi_translation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1650_opus_books_en-fi_translation.texttext-generation1K<n<10K0 likes105 downloads2y agoHugging Face09ai-allforever /wiki-ru-en-news-bookstexttext-generation1M<n<10M2 likes91 downloads11mo agoHugging Face10BEE-spoke-data /rp_books-en Dataset Card for "rp_books-en" Filtering/cleaning on the 'red pajama books' subset of togethercomputer/Long-Data-Collections The default config: Dataset({ features: ['meta', 'text'], num_rows: 26372 }) token count default GPT-4 tiktoken token count: token_count count 2.637200e+04 mean 1.009725e+05 std 1.161315e+05 min 3.811000e+03 25% 3.752750e+04 50% 7.757950e+04 75% 1.294130e+05 max 8.687685e+06 Total count: 2662.85 M… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/rp_books-en.texttext-generation100K<n<1M1 likes81 downloads9mo agoHugging Face11PiotrSty /rock-pollub-pl-books ROCK Politechnika Lubelska PL books A fail-closed Polish academic-book subset extracted from the official ROCK repository of Lublin University of Technology. Retained books: 8 Text characters: 3,380,937 Tokens: 1,177,858 (cl100k_base proxy) Author coverage: 100.0% License: CC BY-SA 4.0, confirmed for every retained item and matched PDF bitstream Source period: 2023-2026 The acquisition target was 20 books, but only eight passed the conservative per-file rights gate. The other… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/rock-pollub-pl-books.texttext-generationn<1K0 likes79 downloads18d agoHugging Face12LocalDoc /books_datasetAzerbaijani Books Dataset Description This dataset contains 2800 books on different topics in Azerbaijani language. It was created in 2024 and contains 7.8 million sentences. The books were divided into sentences and pre-filtered. The dataset included only those sentences where the percentage of letters was at least 80% of the total number of characters. The sequence of sentences is the same as in books. Format The dataset is provided in comma-separated values (CSV) format. Each article is… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/books_dataset.texttext-generation1M<n<10M3 likes70 downloads3y agoHugging Face13pythainlp /thai-tnhc2-books Thai TNHC2 Books This dataset collect all books from TNHC2 corpus. We clean the dataset to use text to pretraining model and nlp task. All books: 353 books License: CC-0 TNHC2 Dataset (Original) have many a lots of details (chapter, author's detail and more). The dataset is clean to pretraining model and nlp task. TNHC2 coepus is a Thai old books corpus that all books are copyright expired in Thai law (50 years after the author's death). TNHC2 Dataset (Original):… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/thai-tnhc2-books.texttext-generationn<1K0 likes65 downloads3y agoHugging Face14ysngkil /whole-books Whole Books Five book-length corpora for long-context language-model pretraining, packaged so that one row is one whole book. Together: 202,343 books, about 20.9 B tokens in a 32K-vocabulary Llama-style tokenizer, with the large majority of tokens inside books of 64K tokens or more. Built 2026-09-06 from pinned snapshots of the sources below; nothing was filtered, deduplicated or cleaned beyond what the sources had already done, and the reassembly steps are documented per… See the full description on the dataset page: https://huggingface.co/datasets/ysngkil/whole-books.tabulartext-generation100K<n<1M0 likes65 downloads21d agoHugging Face15pythainlp /thai-it-books Thai IT books This dataset collects Thai IT books that are the open access books. license: cc-by-3.0 texttext-generationn<1K0 likes57 downloads3y agoHugging Face16pszemraj /LoC-PD-Books-preprocessed LoC-PD-Books: preprocessed This is the storytracer/LoC-PD-Books dataset with the following preprocessing steps: apply clean-text package keeping casing and newlines drop OCR garbled text in first few lines of each example fix (most) 'hard' newlines w/ regex similar to gutenberg clean 'grade' first 512 tokens of each book with this quantized model; keep examples from labels clean (all) and mild gibberish w/ score 0.9 or higher tabulartext-generation10K<n<100K1 likes33 downloads9mo agoHugging Face17Lots-of-LoRAs /task1647_opus_books_en-pt_translation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1647_opus_books_en-pt_translation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1647_opus_books_en-pt_translation.texttext-generation1K<n<10K0 likes22 downloads2y agoHugging Face18SUSHANT283 /Electrical_BooksDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in the following paper: https://arxiv.org/abs/2305.07759. The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M. Additional resources: tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/SUSHANT283/Electrical_Books.texttext-generation1K<n<10K0 likes22 downloads7mo agoHugging Face19QIRIM /crh_booksgatedtextfill-mask1K<n<10K1 likes21 downloads10mo agoHugging Face20premio-ai /TheArabicPile_Books The Arabic Pile Introduction: The Arabic Pile is a comprehensive dataset meticulously designed to parallel the structure of The Pile and The Nordic Pile. Focused on the Arabic language, the dataset encompasses a vast array of linguistic nuances, incorporating both Modern Standard Arabic (MSA) and various Levantine, North African, and Egyptian dialects. Tailored for the training and fine-tuning of large language models, the dataset consists of 13 subsets, each uniquely… See the full description on the dataset page: https://huggingface.co/datasets/premio-ai/TheArabicPile_Books.texttext-generation1K<n<10K0 likes20 downloads3y agoHugging Face21Lots-of-LoRAs /task1652_opus_books_ca-en_translation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1652_opus_books_ca-en_translation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1652_opus_books_ca-en_translation.texttext-generation1K<n<10K0 likes19 downloads2y agoHugging Face22BuzzBlitz360A /Ukrainian-CulturalHeritage-Books 🇺🇦 Ukrainian-Cultural Heritage-Books 🇺🇦 Ukrainian-Cultural Heritage-Books or Ukrainian-CulturalHeritage-Books is a collection of Ukrainian cultural heritage books and periodicals, most of them being in the public domain. Dataset summary The collection has been compiled by Pierre-Carl Langlais from 19,574 digitized files hosted on Internet Archive (462M words) and will be expanded to other cultural heritage sources. Curation method The composition of the… See the full description on the dataset page: https://huggingface.co/datasets/BuzzBlitz360A/Ukrainian-CulturalHeritage-Books.tabulartext-generation10K<n<100K0 likes19 downloads8mo agoHugging Face23Lots-of-LoRAs /task1649_opus_books_en-no_translation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1649_opus_books_en-no_translation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1649_opus_books_en-no_translation.texttext-generationn<1K0 likes18 downloads2y agoHugging Face24Lots-of-LoRAs /task1651_opus_books_en-es__translation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1651_opus_books_en-es__translation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1651_opus_books_en-es__translation.texttext-generation1K<n<10K0 likes17 downloads2y agoHugging Face25empathyai /books-intent-datasetgated Books intent classification dataset A prompt intent classification dataset built from titles, author names and categories (subjects) contained in Project Gutenberg. Its main purpose is to finetune small language models on intent classification task. Dataset Details Dataset Description Curated by: Empathy.co Shared by: Project Gutenberg Language(s) (NLP): English License: CC0 1.0 Public‑Domain Dedication Dataset Sources Project Gutenberg.… See the full description on the dataset page: https://huggingface.co/datasets/empathyai/books-intent-dataset.texttext-classification100K<n<1M3 likes16 downloads1y agoHugging Face26pszemraj /airship-books-antique pszemraj/airship-books-antique Digitized books (via a VLM) on airships from the survivor library texttext-generationn<1K0 likes13 downloads9mo agoHugging Face27Lots-of-LoRAs /task1648_opus_books_en-sv_translation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1648_opus_books_en-sv_translation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1648_opus_books_en-sv_translation.texttext-generation1K<n<10K0 likes12 downloads2y agoHugging Face28DexopT /Books-General-Linux Linux Books Dataset Dataset Description The Linux Books Dataset is a curated text dataset derived from Linux-related books and learning materials. It focuses on Linux system administration, cybersecurity, networking, shell scripting, and operating system fundamentals.The dataset is designed to support training and evaluation of NLP models for technical domains, especially cybersecurity-aware language models and Linux-focused assistants. This dataset is suitable for both… See the full description on the dataset page: https://huggingface.co/datasets/DexopT/Books-General-Linux.tabulartext-generation1K<n<10K0 likes7 downloads9mo agoHugging Face29jorgeortizfuentes /spanish_booksgated Spanish Books Dataset Summary Dataset of books in Spanish crawled from web and torrents. Preprocessing Preprocessing performed by spanish_nlp. Licensing Information The dataset is available under the Creative Commons Attribution-ShareAlike License (CC BY-SA 4.0). Some books may be subject to copyright. Use for academic purposes only. Citation Information @misc{ortiz2022esbooks, title={Crawled Spanish Books}, author={Jorge… See the full description on the dataset page: https://huggingface.co/datasets/jorgeortizfuentes/spanish_books.texttext-generation10K<n<100K11 likes5 downloads4y agoHugging Face30costadev00 /books-gutenberg-project-pt-brgated Gutenberg Project TokenWeaver CPT 2048 - Unchunked This dataset contains full-document rows reconstructed from costadev00/gutenberg-project-tokenweaver-cpt-2048. The source dataset mixes reconstructed chunk sequences and singleton chunk rows. For this unchunked release, rows were grouped by metadata.id; when duplicated singleton rows were present for the same document, the reconstruction kept the series with the largest chunk_total. Text was joined with inferred text overlap.… See the full description on the dataset page: https://huggingface.co/datasets/costadev00/books-gutenberg-project-pt-br.texttext-generationn<1K0 likes5 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.