datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openlibrary_dump_2024-04-30
OpenLibrary Dump (2024-04-30)
This dataset contains the OpenLibrary dump of April 2024 converted to Parquet and DuckDB for easier querying.
Formats
Original GZIP dumps
The original GZIP dumps are available at data/dumps. The dumps are gzipped TSV files with the original OL JSON record contained in the fifth column of the TSV.
DuckDB
The authors, works and editions dumps were imported as tables into… See the full description on the dataset page: https://huggingface.co/datasets/storytracer/openlibrary_dump_2024-04-30.open-library
Open Library
The complete Open Library catalog in clean, analysis-ready Parquet. 150.0M+ records across 11 entity types, from ISBNs and author bios to reading logs and Wikidata links.
What is it?
Open Library is a complete snapshot of the Open Library database, an open project of the Internet Archive with the mission of creating "one web page for every book ever published." The catalog is community-edited and contains bibliographic records for millions of authors, works… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-library.openlibrary
📚 OpenLibrary Parquet Mirror
An up‑to‑date, stream‑ready mirror of the official OpenLibrary public data dumps, converted to snappy‑compressed Parquet.
Source: https://openlibrary.org/developers/dumps
💾 Repository organization
openlibrary/
├── authors/ # authors_0.parquet, authors_1.parquet, … (≈ 1-2 GB each)
├── editions/ # editions_0.parquet, editions_1.parquet, …
└── works/ # works_0.parquet, works_1.parquet, …
🏃 Quick start
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/sayshara/openlibrary.media-metadata-openlibrary-books
TigreGotico/media-metadata-openlibrary-books
Rich entity dataset scraped by metadatarr
scraper openlibrary_books.
Rows: 4,098,190
Fields
olid
title
subtitle
authors
author_key
first_publish_year
subjects
isbn_10
isbn_13
publisher
language
number_of_pages_median
ebook_access
has_fulltext
edition_count
cover_i
Source
Generated by scrapers/openlibrary_books.py. See the metadatarr repo for the full
pipeline and scraper source code.
open-library-scraper
Open Library Scraper · Books, Authors, Editions & Subjects
Scrape Open Library books, authors, subjects, editions, and metadata via Open Library API. Fast HTTP scraper charging per returned record with tiered pricing.
Rows in this dataset
3,880
Fields
21
Collector runs behind it
50
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/open-library-scraper/ — 3,880 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/open-library-scraper.open-library-10k
Dir Bear Open Library — 10,000 distilled web documents
Ten thousand complete, cleaned, English web documents — every one at least 250
words of prose, boilerplate stripped, exact-deduplicated, token-counted and scored
by the quality of the site it came from. This is the free, open slice of the
Dir Bear corpus: the same records, the same schema and the
same pipeline as the datasets we sell, at a size you can read through in an afternoon.
Browse it online:… See the full description on the dataset page: https://huggingface.co/datasets/directorybear/open-library-10k.open-library-atlas-data
