datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
library_of_congress_filtered
Library of Congress
Description
The Library of Congress (LoC) curates a collection of public domain books called "Selected Digitized Books".
We have downloaded over 130,000 English-language books from this public domain collection as OCR plain text files using the LoC APIs.
Dataset Statistics
Documents
UTF-8 GB
129,052
35.6
License Issues
While we aim to produce datasets with completely accurate licensing information, license… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/library_of_congress_filtered.biodiversity_heritage_library
Biodiversity Heritage Library
Description
The Biodiversity Heritage Library (BHL) is an open-access digital library for biodiversity literature and archives.
This dataset contains over 42 million public domain books and documents from the BHL collection.
These works were collected using the bulk data download interface provided by the BHL and were filtered based on their associated license metadata.
We use the optical character recognition (OCR)-generated text… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/biodiversity_heritage_library.library_of_congress
Library of Congress (subset of Common Pile)
Description
The Library of Congress (LoC) curates a collection of public domain books called "Selected Digitized Books".
We have downloaded over 130,000 English-language books from this public domain collection as OCR plain text files using the LoC APIs.
This dataset is a subset of the Common Pile v0.1. For more information, see The Common Pile v0.1 paper.
Dataset Statistics
Documents
UTF-8 GB
135,500… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/library_of_congress.open-library
Open Library
The complete Open Library catalog in clean, analysis-ready Parquet. 150.0M+ records across 11 entity types, from ISBNs and author bios to reading logs and Wikidata links.
What is it?
Open Library is a complete snapshot of the Open Library database, an open project of the Internet Archive with the mission of creating "one web page for every book ever published." The catalog is community-edited and contains bibliographic records for millions of authors, works… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-library.biodiversity_heritage_library_filtered
Biodiversity Heritage Library
Description
The Biodiversity Heritage Library (BHL) is an open-access digital library for biodiversity literature and archives.
This dataset contains over 15 million public domain books and documents from the BHL collection.
These works were collected using the bulk data download interface provided by the BHL and were filtered based on their associated license metadata.
We use the optical character recognition (OCR)-generated text distributed… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/biodiversity_heritage_library_filtered.GLM-5.2-CoT-Library
GLM-5.2 — CoT Library
A maintained mirror of publicly-available GLM-5.2 chain-of-thought datasets on Hugging Face — content-verified, deduplicated, and attributed to their original authors.
Dataset Viewer | Parquet
// what this is
A maintained library — a community mirror of publicly-available GLM-5.2 CoT datasets, aggregated, validity-filtered and content-verified, with per-row source attribution in first_source_dataset. It is not Crownelius' own data — every… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/GLM-5.2-CoT-Library.Kimi-K3-CoT-Library
Kimi K3 — CoT Library
A maintained mirror of publicly-available Kimi K3 chain-of-thought datasets on Hugging Face — content-verified, deduplicated, and attributed to their original authors.
Dataset Viewer | Parquet
// what this is
A maintained library — a community mirror of publicly-available Kimi K3 CoT datasets, aggregated, validity-filtered and content-verified, with per-row source attribution in first_source_dataset. It is not Crownelius' own data — every… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Kimi-K3-CoT-Library.un-digital-library
United Nations Digital Library (UNDL) Comprehensive Master Dataset
1. Executive Summary
Welcome to the United Nations Digital Library (UNDL) Comprehensive Master Dataset repository. This dataset represents a monumental effort to harvest, normalize, enrich, and democratize access to the vast archives of the United Nations. By leveraging advanced web harvesting techniques, robust state management, and modern big-data formats, this repository provides researchers… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/un-digital-library.prompt-library
Metrum AI Prompt Library
A prompt library for LLM inference workload and performance benchmarking,
prepared for use with metrum-ai/bench-cli.
It contains 593,730 records with prompt text, intended lengths, token
buckets, and reasoning labels. It contains no reference answers.
The full configuration preserves all source records, including repeated
prompts and their distinct workload targets. Prompts may appear duplicated,
with only target_output_length differing. These variants… See the full description on the dataset page: https://huggingface.co/datasets/metrum-ai/prompt-library.Qwen-CoT-Library
Qwen — CoT Library
A maintained mirror of publicly-available Qwen chain-of-thought datasets on Hugging Face — content-verified, deduplicated, and attributed to their original authors.
Dataset Viewer | Parquet
// what this is
A maintained library — a community mirror of publicly-available Qwen CoT datasets, aggregated, validity-filtered and content-verified, with per-row source attribution in first_source_dataset. It is not Crownelius' own data — every row… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Qwen-CoT-Library.library-python-training-pool
Python library function-writing training pool
A pool of public data for training a model to write Python functions, many of them
calling libraries: 8.3 percent of the answers in the normalised layer import a library that
is not in the Python standard library. It is a straight collection of open datasets, not
a new corpus: every row comes from one of the sources below, at the revision named. Rows
an overlap filter flagged against held-out material this pool is kept separate from… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/library-python-training-pool.GATT_library
GATT Library Dataset Card
Dataset Overview
Dataset Name: GATT Library
Description:
The GATT Library dataset comprises a comprehensive collection of documents related to the General Agreement on Tariffs and Trade (GATT), spanning from January 1, 1946, to September 6, 1996. This dataset is organized into a single Parquet file, which contains detailed information about the documents, including metadata extracted from the original files. The original files are stored in a… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/GATT_library.persona-steering-template-library
Persona Steering Template Library
GitHub repository: https://github.com/wassname/persona-steering-template-library
Evaluated persona/template candidates for steering-vector and preference-pair experiments.
What This Measures
How do we know if a persona template is good? We want on-axis variation, but not off-axis variation.
If we choose honest and dishonest personas, use a template like You are a {{ persona }} assistant, and ask The Eiffel Tower is in, we want the… See the full description on the dataset page: https://huggingface.co/datasets/wassname/persona-steering-template-library.ask_library_cs
Dataset Card for Ask the Library (Ptejte se knihovny)
Ask the Library contains questions and answers scraped from the webpage: https://www.ptejteseknihovny.cz/.
Dataset Details
Dataset Description
Ask the Library contains questions and answers scraped from the webpage: https://www.ptejteseknihovny.cz/. The questions have a broad range of topics like history, language, biology and others.
Curated by: AIC FEE CTU
Language(s) (NLP): Czech
License: CC-BY-NC 3.0… See the full description on the dataset page: https://huggingface.co/datasets/ctu-aic/ask_library_cs.open-library-10k
Dir Bear Open Library — 10,000 distilled web documents
Ten thousand complete, cleaned, English web documents — every one at least 250
words of prose, boilerplate stripped, exact-deduplicated, token-counted and scored
by the quality of the site it came from. This is the free, open slice of the
Dir Bear corpus: the same records, the same schema and the
same pipeline as the datasets we sell, at a size you can read through in an afternoon.
Browse it online:… See the full description on the dataset page: https://huggingface.co/datasets/directorybear/open-library-10k.
