datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Sci-Fi-Books-gutenberg
Gutenberg Sci-Fi Book Dataset
This dataset contains information about science fiction books. It’s designed for training AI models, research, or any other purpose related to natural language processing.
Data Format
The dataset is provided in CSV format. Each record represents a book and includes the following fields:
ID: A unique identifier for the book.
Title: The title of the book.
Author: The author(s) of the book.
Text: The text content of the book (e.g., summary… See the full description on the dataset page: https://huggingface.co/datasets/stevez80/Sci-Fi-Books-gutenberg.scifig-bench
SciFig-Bench: Scientific Figure & Multi-Scale Caption Dataset Card
Dataset Description
This dataset contains high-quality scientific figures, architecture diagrams, charts, and visualizations extracted from scientific papers (arXiv & local PDFs). It includes multi-level captions (Small, Medium, Large), extracted embedded OCR text, paper metadata, and alignment quality scores.
Total Figure Records: 589
Average CLIP Alignment Score: 0.2721
Average Composite Score:… See the full description on the dataset page: https://huggingface.co/datasets/Goutam112/scifig-bench.openlibrary-scifi-data
