datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-library
Open Library
The complete Open Library catalog in clean, analysis-ready Parquet. 150.0M+ records across 11 entity types, from ISBNs and author bios to reading logs and Wikidata links.
What is it?
Open Library is a complete snapshot of the Open Library database, an open project of the Internet Archive with the mission of creating "one web page for every book ever published." The catalog is community-edited and contains bibliographic records for millions of authors, works… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-library.GLM-5.2-CoT-Library
GLM-5.2 — CoT Library
A maintained mirror of publicly-available GLM-5.2 chain-of-thought datasets on Hugging Face — content-verified, deduplicated, and attributed to their original authors.
Dataset Viewer | Parquet
// what this is
A maintained library — a community mirror of publicly-available GLM-5.2 CoT datasets, aggregated, validity-filtered and content-verified, with per-row source attribution in first_source_dataset. It is not Crownelius' own data — every… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/GLM-5.2-CoT-Library.un-digital-library
United Nations Digital Library (UNDL) Comprehensive Master Dataset
1. Executive Summary
Welcome to the United Nations Digital Library (UNDL) Comprehensive Master Dataset repository. This dataset represents a monumental effort to harvest, normalize, enrich, and democratize access to the vast archives of the United Nations. By leveraging advanced web harvesting techniques, robust state management, and modern big-data formats, this repository provides researchers… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/un-digital-library.Kimi-K3-CoT-Library
Kimi K3 — CoT Library
A maintained mirror of publicly-available Kimi K3 chain-of-thought datasets on Hugging Face — content-verified, deduplicated, and attributed to their original authors.
Dataset Viewer | Parquet
// what this is
A maintained library — a community mirror of publicly-available Kimi K3 CoT datasets, aggregated, validity-filtered and content-verified, with per-row source attribution in first_source_dataset. It is not Crownelius' own data — every… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Kimi-K3-CoT-Library.prompt-library
Metrum AI Prompt Library
A prompt library for LLM inference workload and performance benchmarking,
prepared for use with metrum-ai/bench-cli.
It contains 593,730 records with prompt text, intended lengths, token
buckets, and reasoning labels. It contains no reference answers.
The full configuration preserves all source records, including repeated
prompts and their distinct workload targets. Prompts may appear duplicated,
with only target_output_length differing. These variants… See the full description on the dataset page: https://huggingface.co/datasets/metrum-ai/prompt-library.Qwen-CoT-Library
Qwen — CoT Library
A maintained mirror of publicly-available Qwen chain-of-thought datasets on Hugging Face — content-verified, deduplicated, and attributed to their original authors.
Dataset Viewer | Parquet
// what this is
A maintained library — a community mirror of publicly-available Qwen CoT datasets, aggregated, validity-filtered and content-verified, with per-row source attribution in first_source_dataset. It is not Crownelius' own data — every row… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Qwen-CoT-Library.persona-steering-template-library
Persona Steering Template Library
GitHub repository: https://github.com/wassname/persona-steering-template-library
Evaluated persona/template candidates for steering-vector and preference-pair experiments.
What This Measures
How do we know if a persona template is good? We want on-axis variation, but not off-axis variation.
If we choose honest and dishonest personas, use a template like You are a {{ persona }} assistant, and ask The Eiffel Tower is in, we want the… See the full description on the dataset page: https://huggingface.co/datasets/wassname/persona-steering-template-library.GATT_library
GATT Library Dataset Card
Dataset Overview
Dataset Name: GATT Library
Description:
The GATT Library dataset comprises a comprehensive collection of documents related to the General Agreement on Tariffs and Trade (GATT), spanning from January 1, 1946, to September 6, 1996. This dataset is organized into a single Parquet file, which contains detailed information about the documents, including metadata extracted from the original files. The original files are stored in a… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/GATT_library.open-library-10k
Dir Bear Open Library — 10,000 distilled web documents
Ten thousand complete, cleaned, English web documents — every one at least 250
words of prose, boilerplate stripped, exact-deduplicated, token-counted and scored
by the quality of the site it came from. This is the free, open slice of the
Dir Bear corpus: the same records, the same schema and the
same pipeline as the datasets we sell, at a size you can read through in an afternoon.
Browse it online:… See the full description on the dataset page: https://huggingface.co/datasets/directorybear/open-library-10k.
