datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
booksum
BOOKSUM: A Collection of Datasets for Long-form Narrative Summarization
Authors: Wojciech Kryściński, Nazneen Rajani, Divyansh Agarwal, Caiming Xiong, Dragomir Radev
Introduction
The majority of available text summarization datasets include short-form source documents that lack long-range causal and temporal dependencies, and often contain strong layout and stylistic biases.
While relevant, such datasets will offer limited challenges for future generations of text… See the full description on the dataset page: https://huggingface.co/datasets/kmfoda/booksum.booksummaries_cleanedSci-Fi-Books-gutenberg
Gutenberg Sci-Fi Book Dataset
This dataset contains information about science fiction books. It’s designed for training AI models, research, or any other purpose related to natural language processing.
Data Format
The dataset is provided in CSV format. Each record represents a book and includes the following fields:
ID: A unique identifier for the book.
Title: The title of the book.
Author: The author(s) of the book.
Text: The text content of the book (e.g., summary… See the full description on the dataset page: https://huggingface.co/datasets/stevez80/Sci-Fi-Books-gutenberg.Chinese-Psychology-Books
免责声明与使用须知 (Disclaimer and Usage Notice)
数据集内容
本数据集包含从互联网上多个来源收集的 中文心理学电子书 的集合。
许可证
本数据集的组织结构、汇编方式以及由维护者添加的任何元数据或注释根据 知识共享署名-非商业性使用 4.0 国际许可协议 (Creative Commons Attribution-NonCommercial 4.0 International License - CC BY-NC 4.0) 提供。这意味着您可以基于非商业目的分享和修改这部分内容,但必须给出适当的署名。
请注意:此 CC BY-NC 4.0 许可证不适用于数据集中包含的原始电子书文件本身。
版权声明
数据集中包含的个别电子书文件极有可能受到版权法保护,其版权归各自的作者、出版商或其他版权所有者所有。
数据集维护者不拥有这些电子书的版权。
这些电子书的来源多样且零散,部分来源可能难以追溯。
使用限制与责任… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/Chinese-Psychology-Books.Goodreads-Books
Dataset Card for "BrightData/Goodreads-Books"
Dataset Summary
Explore a collection of millions of books with the Goodreads dataset, comprising over 6.3M structured records and 14 data fields updated and refreshed regularly.
Each entry includes all major data points such as URLs, book IDs, titles, authors, ratings, number of ratings, reviews, summaries, genres, publication dates, author details and prices.
For a complete list of data points, please refer to the full "Data… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/Goodreads-Books.project-madurai-booksProject Madurai Books Text Dataset
This dataset card aims to convert the Tamil books available on the Project Madurai website to the HF dataset. It has been scrapped from Project Madurai Website.
Dataset Details
You can see a table above called "Meta Data", which is just an info table.
You can't able to preview the "Source Data" table, due to it being about 300MB.
[Don't open the Dataset in Excel It will lead to a crash of the OS instead open it using Python in pandas or… See the full description on the dataset page: https://huggingface.co/datasets/mastergokul/project-madurai-books.Goodreads_Books_DetailThis dataset imclude all important detail about books like Title,Auther,Rating , Genres,Release_data, review, no. of votes,
that help to analyze about book catagory , this dataset help to build the book recomender system model,sentiment analysis
on book review system and many more
book_summary_datasetArabic-books-and-research-dataset
Arabic reserach and books dataset (ARABD)
This dataset is an extracted cleaned text from more than 60K word files with unique arabic texts never published before.
Dataset diversity
the dataset is diverse from all kind of islamic research: [feqh, hadeeth, tafseer, tahqeeq, ... etc], from new written research to a manuscirpts.
dataset size
the dataset was more than 11GB but after cleaning (pre-processing) it becase a straight 10GB with less noisy chars.… See the full description on the dataset page: https://huggingface.co/datasets/riotu-lab/Arabic-books-and-research-dataset.goodreads-books
Goodreads Books Dataset
Dataset Description
A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics.
This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for:
📚 Book recommendation systems
📊 Literary data analysis
🤖 Machine learning projects
📈 Rating prediction models
🔍 Book discovery algorithms
Dataset Structure
Features… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/goodreads-books.Gutenberg_books
Gutenberg Books Dataset
Dataset Description
This dataset contains 97,646,390 paragraphs extracted from 74,329 English-language books sourced from Project Gutenberg, a digital library of public domain works. The total size of the dataset is 34GB, making it a substantial resource for natural language processing (NLP) research and applications. The texts have been cleaned to remove Project Gutenberg's standard headers and footers, ensuring that only the core content of each… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/Gutenberg_books.Shamela_Books_info
Shamela Books information
This dataset contains structured metadata for 8,492 books sourced from the Shamela Library, with enhancements for clarity, consistency, and usability. It is intended to support NLP, bibliographic research, and digital humanities efforts involving Arabic texts.For full books text dataset please check shamela_books_text
Dataset Features
The dataset includes the following cleaned and standardized features:
Unification of Author Names: Author… See the full description on the dataset page: https://huggingface.co/datasets/MoMonir/Shamela_Books_info.goodreads_booksfrench_books
Description
Dataframe containing 2075 French books in txt format (= the ~2600 French books present in gutenberg from which all books by authors present in the french_books_summuries dataset have been removed to avoid any leaks).More precisely :
the texte column contains the texts
the titre column contains the book title
the auteur column contains the author's name and dates of birth and death (if you want to filter the texts to keep only those from the given century to the present… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/french_books.Goodreads-Books
Dataset Card for "BrightData/Goodreads-Books"
Dataset Summary
Explore a collection of millions of books with the Goodreads dataset, comprising over 6.3M structured records and 14 data fields updated and refreshed regularly.
Each entry includes all major data points such as URLs, book IDs, titles, authors, ratings, number of ratings, reviews, summaries, genres, publication dates, author details and prices.
For a complete list of data points, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/Chima207/Goodreads-Books.green-books-travel-guides
African American Travel Guides: The Green Book & Companion Directories (1930–1966)
A unified, structured dataset of 113,827 business and lodging listings transcribed from 50 volumes of mid-20th-century African American travel guides, spanning 1930–1966. During the Jim Crow era, these guides told Black travelers which hotels, restaurants, tourist homes, service stations, and other businesses would serve them safely. This dataset brings The Negro Motorist Green Book together with… See the full description on the dataset page: https://huggingface.co/datasets/hadro/green-books-travel-guides.ArTopicDS-BooksThe books used in this dataset spanned the areas of
Religion
Economy
Politics
Anthropology and Sociology
Art and Literature
Education
History
Language and Linguistics
Philosophy
Law.
Only first sentences after each title of the books have been extracted. For some books, the first sentence after each paragraph was taken. Refer to the paper
for detailed explanation.
book_summarybooksum-short
booksum short
BookSum but all summaries with length greater than 512 long-t5 tokens are filtered out.
The columns chapter_length and summary_length in this dataset have been updated to reflect the total of Long-T5 tokens in the respective source text.
Token Length Distribution for inputs
pd_books_samplesbooksum2xmu_psych_books
Intro
The "Xiamen University Psychology Book Loan List" is a comprehensive and well-curated collection of psychological literature, tailored for students and researchers in the field of psychology at Xiamen University. This list encompasses a wide array of books that delve into various psychological domains, from foundational theories to cutting-edge research topics, ensuring that users can access a wealth of knowledge catering to different academic levels and research interests. It… See the full description on the dataset page: https://huggingface.co/datasets/Genius-Society/xmu_psych_books.goodreads-books
Goodreads Books Dataset
Dataset Description
A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics.
This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for:
📚 Book recommendation systems
📊 Literary data analysis
🤖 Machine learning projects
📈 Rating prediction models
🔍 Book discovery algorithms
Dataset Structure
Features… See the full description on the dataset page: https://huggingface.co/datasets/bstarrs/goodreads-books.Geography_books_datasetThis dataset is a sub-set of 'The Project Gutenberg' that only focuses on Geography text.
Books: 11M of tokens
The 1990 CIA World Factbook
Commercial Geography
Influences of Geographic Environment
Geographical etymology: a dictionary of place-names giving their derivations
Geography and Plays
Physical Geography
booksThis datasets contains 77,954 politicians from around the world. The latest version can be found and filtered differently on: https://www.workwithdata.com/datasets/books
Similar datasets can be found on: https://www.workwithdata.com
opus_booksIslamic-BooksIslamic Books Hadith Dataset
Description:
The Islamic Books Hadith Dataset comprises excerpts of hadiths (sayings of Prophet Muhammad) extracted from nine famous Islamic books. Each entry in the dataset contains two columns:
Hadith: The actual text of the hadith.
Reference: The source reference indicating the book and the specific location of the hadith within the book.
All Hadith text is cleaned
Books Included:
Sahih al-Bukhari
Sahih Muslim
Maliks Muwatta
Sunan at-Tirmidhi
Musnad Ahmad… See the full description on the dataset page: https://huggingface.co/datasets/Ahmedhany216/Islamic-Books.Translated_Books
Translated Books
⚠️ Disclaimer: This is a personal translation project and is NOT part of the OmniMedical Suite ecosystem.
It is unrelated to medical OCR, handwriting recognition, or any of the author's medical AI work.
Dataset Description
A personal collection of English-to-Arabic book translations compiled as a parallel corpus.
This dataset is maintained separately from the author's professional medical AI projects.
Files
File
Format… See the full description on the dataset page: https://huggingface.co/datasets/DrAbdulmalek/Translated_Books.bookssmall_wiki_news_books
