datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gutenberg-100
Gutenberg Sci-Fi Book Dataset Testing Sample
This dataset contains information about science fiction books. It’s designed for training AI models, research, or any other purpose related to natural language processing.
It contains just 100 books for quick download targetting CI use (34MB).
The original dataset it's derived from is https://huggingface.co/datasets/stevez80/Sci-Fi-Books-gutenberg
Data Format
The dataset is provided in CSV format. Each record represents a… See the full description on the dataset page: https://huggingface.co/datasets/stas/gutenberg-100.Sci-Fi-Books-gutenberg
Gutenberg Sci-Fi Book Dataset
This dataset contains information about science fiction books. It’s designed for training AI models, research, or any other purpose related to natural language processing.
Data Format
The dataset is provided in CSV format. Each record represents a book and includes the following fields:
ID: A unique identifier for the book.
Title: The title of the book.
Author: The author(s) of the book.
Text: The text content of the book (e.g., summary… See the full description on the dataset page: https://huggingface.co/datasets/stevez80/Sci-Fi-Books-gutenberg.passages_gutenberg_popularpassages_gutenberg_unpopulargutenberg-poetry-corpusGutenberg-Arabic-OCR-HTML-Pages
Gutenberg Arabic HTML-Page Dataset
📖 Dataset Description
The Gutenberg Arabic HTML-Page Dataset is a large-scale, synthetically generated dataset designed for training and evaluating document understanding and Optical Character Recognition (OCR) models. The primary goal of this project is to provide a comprehensive resource of page images paired with their corresponding structured HTML ground truth, with a focus on the Arabic language.
The dataset was created by… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Gutenberg-Arabic-OCR-HTML-Pages.Gutenberg_books
Gutenberg Books Dataset
Dataset Description
This dataset contains 97,646,390 paragraphs extracted from 74,329 English-language books sourced from Project Gutenberg, a digital library of public domain works. The total size of the dataset is 34GB, making it a substantial resource for natural language processing (NLP) research and applications. The texts have been cleaned to remove Project Gutenberg's standard headers and footers, ensuring that only the core content of each… See the full description on the dataset page: https://huggingface.co/datasets/Navanjana/Gutenberg_books.gutenberg-poetry-corpuswordpress-gutenberg-block-patterns
WordPress Gutenberg Block Patterns
paragraph_gutenberg_top50
