datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-library
Open Library
The complete Open Library catalog in clean, analysis-ready Parquet. 150.0M+ records across 11 entity types, from ISBNs and author bios to reading logs and Wikidata links.
What is it?
Open Library is a complete snapshot of the Open Library database, an open project of the Internet Archive with the mission of creating "one web page for every book ever published." The catalog is community-edited and contains bibliographic records for millions of authors, works… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-library.open-library-10k
Dir Bear Open Library — 10,000 distilled web documents
Ten thousand complete, cleaned, English web documents — every one at least 250
words of prose, boilerplate stripped, exact-deduplicated, token-counted and scored
by the quality of the site it came from. This is the free, open slice of the
Dir Bear corpus: the same records, the same schema and the
same pipeline as the datasets we sell, at a size you can read through in an afternoon.
Browse it online:… See the full description on the dataset page: https://huggingface.co/datasets/directorybear/open-library-10k.
