datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pre_1929_books_filtered
Pre-1929 Books
Description
Books published in the US before 1929 passed into the public domain on January 1, 2024.
We used the bibliographic catalog Hathifiles produced by HathiTrust to identify digitized books which were published in the US before 1929.
The collection contains over 130,000 books digitized and processed by the Internet Archive on behalf of HathiTrust member libraries.
The OCR plain text files were downloaded directly from the Internet Archive website.… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/pre_1929_books_filtered.pre_1929_books
Pre-1929 Books
Description
Books published in the US before 1929 passed into the public domain on January 1, 2024.
We used the bibliographic catalog Hathifiles produced by HathiTrust to identify digitized books which were published in the US before 1929.
The collection contains over 130,000 books digitized and processed by the Internet Archive on behalf of HathiTrust member libraries.
The OCR plain text files were downloaded directly from the Internet Archive website.… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/pre_1929_books.booksum-complete-cleaned
Description:
This repository contains the Booksum dataset introduced in the paper BookSum: A Collection of Datasets for Long-form Narrative Summarization
.
This dataset includes both book and chapter summaries from the BookSum dataset (unlike the kmfoda/booksum one which only contains the chapter dataset). Some mismatched summaries have been corrected. Uneccessary columns have been discarded. Contains minimal text-to-summary rows. As there are multiple summaries for a given text… See the full description on the dataset page: https://huggingface.co/datasets/ubaada/booksum-complete-cleaned.books
Books
The books dataset consists of a diverse collection of books organized into 9 categories, it splitted to train, validation where the train contains 40 books, and the validation 9 books.
This dataset is cleaned well and designed to support various natural language processing (NLP) tasks, including text generation and masked language modeling.
Details
The dataset contains 4 columns:
title: The tilte of the book.
author: The author of the book.
category: The… See the full description on the dataset page: https://huggingface.co/datasets/IsmaelMousa/books.Book_Summary_Chinese
中文图书总结数据集
每个样本包含:
图书的一个章节、此章节的总结、图书名字,可以训练模型总结长文本的能力。数据主要来自较为著名的中文版小说。
tipitaka_myanmar_translation_books
Myanmar Tipitaka Translation (60 Books)
This dataset contains the complete Myanmar (Burmese) translation of the Tipitaka (Pali Canon), together with the major Atthakatha (Commentaries) and the Visuddhimagga.
The texts have been converted into a clean, structured JSONL format, suitable for:
Natural Language Processing (NLP)
LLM Training & Fine-tuning
Digital Humanities Research
Dhamma Study Applications
📊 Dataset Statistics
Total Books: 60
Total Content Lines: 194… See the full description on the dataset page: https://huggingface.co/datasets/freococo/tipitaka_myanmar_translation_books.booksum-zhbooksum数据集,谷歌翻译成中文。
任务:将一本书的某个章节总结为几句话。
源数据来自 togethercomputer/Long-Data-Collections
booksum_deA german translation for the booksum dataset.
Extracted from seedboxventures/multitask_german_examples_32k.
Translation created by seedbox ai for KafkaLM ❤️.
Available for finetuning in hiyouga/LLaMA-Factory.
books
Books
The books dataset consists of a diverse collection of books organized into 9 categories, it splitted to train, validation where the train contains 40 books, and the validation 9 books.
This dataset is cleaned well and designed to support various natural language processing (NLP) tasks, including text generation and masked language modeling.
Details
The dataset contains 4 columns:
title: The tilte of the book.
author: The author of the book.
category: The… See the full description on the dataset page: https://huggingface.co/datasets/DhruvExploring/books.
