CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01kmfoda /booksum BOOKSUM: A Collection of Datasets for Long-form Narrative Summarization Authors: Wojciech Kryściński, Nazneen Rajani, Divyansh Agarwal, Caiming Xiong, Dragomir Radev Introduction The majority of available text summarization datasets include short-form source documents that lack long-range causal and temporal dependencies, and often contain strong layout and stylistic biases. While relevant, such datasets will offer limited challenges for future generations of text… See the full description on the dataset page: https://huggingface.co/datasets/kmfoda/booksum.tabular10K<n<100K80 likes3.9k downloads4y agoHugging Face02BrightData /Goodreads-Books Dataset Card for "BrightData/Goodreads-Books" Dataset Summary Explore a collection of millions of books with the Goodreads dataset, comprising over 6.3M structured records and 14 data fields updated and refreshed regularly. Each entry includes all major data points such as URLs, book IDs, titles, authors, ratings, number of ratings, reviews, summaries, genres, publication dates, author details and prices. For a complete list of data points, please refer to the full "Data… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/Goodreads-Books.tabulartext-classification1M<n<10M20 likes420 downloads2y agoHugging Face03mastergokul /project-madurai-booksProject Madurai Books Text Dataset This dataset card aims to convert the Tamil books available on the Project Madurai website to the HF dataset. It has been scrapped from Project Madurai Website. Dataset Details You can see a table above called "Meta Data", which is just an info table. You can't able to preview the "Source Data" table, due to it being about 300MB. [Don't open the Dataset in Excel It will lead to a crash of the OS instead open it using Python in pandas or… See the full description on the dataset page: https://huggingface.co/datasets/mastergokul/project-madurai-books.tabulartext-classification1K<n<10K0 likes125 downloads2y agoHugging Face04codealchemist01 /goodreads-books Goodreads Books Dataset Dataset Description A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics. This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for: 📚 Book recommendation systems 📊 Literary data analysis 🤖 Machine learning projects 📈 Rating prediction models 🔍 Book discovery algorithms Dataset Structure Features… See the full description on the dataset page: https://huggingface.co/datasets/codealchemist01/goodreads-books.tabulartext-classification1K<n<10K0 likes79 downloads11mo agoHugging Face05MoMonir /Shamela_Books_info Shamela Books information This dataset contains structured metadata for 8,492 books sourced from the Shamela Library, with enhancements for clarity, consistency, and usability. It is intended to support NLP, bibliographic research, and digital humanities efforts involving Arabic texts.For full books text dataset please check shamela_books_text Dataset Features The dataset includes the following cleaned and standardized features: Unification of Author Names: Author… See the full description on the dataset page: https://huggingface.co/datasets/MoMonir/Shamela_Books_info.tabular1K<n<10K1 likes58 downloads1y agoHugging Face06VsquareSheremetyevo /goodreads_booksimage100K<n<1M0 likes57 downloads9mo agoHugging Face07Chima207 /Goodreads-Books Dataset Card for "BrightData/Goodreads-Books" Dataset Summary Explore a collection of millions of books with the Goodreads dataset, comprising over 6.3M structured records and 14 data fields updated and refreshed regularly. Each entry includes all major data points such as URLs, book IDs, titles, authors, ratings, number of ratings, reviews, summaries, genres, publication dates, author details and prices. For a complete list of data points, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/Chima207/Goodreads-Books.tabulartext-classification1M<n<10M0 likes49 downloads8mo agoHugging Face08pszemraj /booksum-short booksum short BookSum but all summaries with length greater than 512 long-t5 tokens are filtered out. The columns chapter_length and summary_length in this dataset have been updated to reflect the total of Long-T5 tokens in the respective source text. Token Length Distribution for inputs tabularsummarization1K<n<10K4 likes34 downloads9mo agoHugging Face09bstarrs /goodreads-books Goodreads Books Dataset Dataset Description A comprehensive dataset of books scraped from Goodreads, including ratings, authors, titles, and various book characteristics. This dataset contains 3045 books with 20 features each, scraped from Goodreads. It's perfect for: 📚 Book recommendation systems 📊 Literary data analysis 🤖 Machine learning projects 📈 Rating prediction models 🔍 Book discovery algorithms Dataset Structure Features… See the full description on the dataset page: https://huggingface.co/datasets/bstarrs/goodreads-books.tabulartext-classification1K<n<10K0 likes27 downloads6mo agoHugging Face10matoupines /booksimagen<1K1 likes20 downloads2y agoHugging Face11davanstrien /on_the_books_example Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/on_the_books_example.tabulartext-classification1K<n<10K0 likes16 downloads3y agoHugging Face12ada-datadruids /booksimage1K<n<10K2 likes16 downloads2y agoHugging Face13pszemraj /booksum-1024-output booksum - 1024 tokens max output goal: limit max output length explicitly to prevent partial summaries being generated. notebook to create info tabularsummarization10K<n<100K0 likes15 downloads9mo agoHugging Face14davanstrien /on_the_books Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/on_the_books.tabular1K<n<10K0 likes8 downloads3y agoHugging Face15johnidouglas /books_10ktabular10K<n<100K0 likes8 downloads2y agoHugging Face16Mahadih534 /all-Bangladeshi-bookstabular1K<n<10K0 likes8 downloads2y agoHugging Face17ada-datadruids /booksmoviestabular1K<n<10K0 likes8 downloads2y agoHugging Face18ada-datadruids /movies_based_on_books_filteredtabularn<1K0 likes8 downloads2y agoHugging Face19ada-datadruids /merged_movies_books_cleanedtabular1K<n<10K0 likes7 downloads2y agoHugging Face20ada-datadruids /movies_based_on_books_with_budgettabular1K<n<10K0 likes7 downloads2y agoHugging Face21buithien /BooksLibrarytabular100K<n<1M0 likes7 downloads2y agoHugging Face22spleentery /books_ratingstabular10K<n<100K1 likes6 downloads3y agoHugging Face23davanstrien /on-the-books-csv-demotabular1K<n<10K0 likes6 downloads1y agoHugging Face24spleentery /amazon_books_doc_simtabular1K<n<10K0 likes4 downloads3y agoHugging Face25ada-datadruids /movies_not_based_on_books_filteredtabular100K<n<1M0 likes4 downloads2y agoHugging Face26buithien /BooksDatasettabular100K<n<1M0 likes4 downloads2y agoHugging Face27JoelVIU /hf_bookstabular1M<n<10M1 likes3 downloads3y agoHugging Face28TiaDay /books-to-scrape-page1 Books to Scrape – Page 1 Dataset Summary Book records scraped from the first page of the Books to Scrape demo site.I created this dataset for a class assignment to practise web scraping, pandas, and publishing a dataset to the Hugging Face Hub. Data Collection Source: https://books.toscrape.com/ (public test site for scraping practice) Method: requests.get("https://books.toscrape.com/catalogue/page-1.html") Parsed with BeautifulSoup, selecting each <article… See the full description on the dataset page: https://huggingface.co/datasets/TiaDay/books-to-scrape-page1.tabularn<1K0 likes3 downloads11mo agoHugging Face29San0160 /Books_enrichedtabular10K<n<100K0 likes3 downloads4mo agoHugging Face30Munshifff /Goodread_books_datasettabular10K<n<100K0 likes2 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.