datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
british-library-book-images
British Library Book Images
1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published
between c. 1510 and c. 1900, digitised by the British Library in partnership
with Microsoft and released by British Library Labs
on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography,
philosophy, history, poetry and literature, in several languages.
The four image types
British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/biglam/british-library-book-images.british-library-book-images
British Library Book Images
1,080,814 images cut out of 49,455 digitised books (65,227 volumes, ~25 million pages) published
between c. 1510 and c. 1900, digitised by the British Library in partnership
with Microsoft and released by British Library Labs
on Flickr Commons as the "1 Million Images from Scanned Books" release. The books cover geography,
philosophy, history, poetry and literature, in several languages.
The four image types
British Library Labs… See the full description on the dataset page: https://huggingface.co/datasets/Faizaniqbal/british-library-book-images.harvard-library-bibliographic-datasetopen-library
Open Library
The complete Open Library catalog in clean, analysis-ready Parquet. 150.0M+ records across 11 entity types, from ISBNs and author bios to reading logs and Wikidata links.
What is it?
Open Library is a complete snapshot of the Open Library database, an open project of the Internet Archive with the mission of creating "one web page for every book ever published." The catalog is community-edited and contains bibliographic records for millions of authors, works… See the full description on the dataset page: https://huggingface.co/datasets/open-index/open-library.library-occupancylca-library-based-code-generation
🏟️ Long Code Arena (Library-based code generation)
This is the benchmark for Library-based code generation task as part of the
🏟️ Long Code Arena benchmark.
The current version includes 150 manually curated instructions asking the model to generate Python code using a particular library.
The samples come from 62 Python repositories.
All the samples in the dataset are based on reference example programs written by authors of the respective libraries.
All the repositories are… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-library-based-code-generation.gov-library
Philippine Legal Documents Dataset
A comprehensive collection of Philippine legal documents from Lawphil.net, extracted from HTML to Markdown and organized for easy querying.
Overview
This dataset contains 114,340 legal documents spanning from 1900 to 2025, including:
Jurisprudence (68,080 documents) - Supreme Court decisions
Statutes (19,793 documents) - Republic Acts, Commonwealth Acts, Presidential Decrees, etc.
Executive Issuances (26,458 documents) - Administrative… See the full description on the dataset page: https://huggingface.co/datasets/bettergovph/gov-library.GLM-5.2-CoT-Library
GLM-5.2 — CoT Library
A maintained mirror of publicly-available GLM-5.2 chain-of-thought datasets on Hugging Face — content-verified, deduplicated, and attributed to their original authors.
Dataset Viewer | Parquet
// what this is
A maintained library — a community mirror of publicly-available GLM-5.2 CoT datasets, aggregated, validity-filtered and content-verified, with per-row source attribution in first_source_dataset. It is not Crownelius' own data — every… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/GLM-5.2-CoT-Library.sufficiency-libraryfactor-library-grid
Tidy Finance Factor Library: Specification Grid
Lookup table that maps each specification id to its portfolio construction choices. Use it together with the Portfolio Returns dataset to select return series and to identify the choices behind each series.
Dataset Details
Dataset Description
The grid contains 4,105,728 specifications for 179 sorting variables. Each row defines a complete set of construction choices: sample exclusions, lagging… See the full description on the dataset page: https://huggingface.co/datasets/tidy-finance/factor-library-grid.Kimi-K3-CoT-Library
Kimi K3 — CoT Library
A maintained mirror of publicly-available Kimi K3 chain-of-thought datasets on Hugging Face — content-verified, deduplicated, and attributed to their original authors.
Dataset Viewer | Parquet
// what this is
A maintained library — a community mirror of publicly-available Kimi K3 CoT datasets, aggregated, validity-filtered and content-verified, with per-row source attribution in first_source_dataset. It is not Crownelius' own data — every… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Kimi-K3-CoT-Library.verasight-data-library
A source-linked index of what U.S. adults think
The Verasight Data Library makes original U.S. public opinion research
searchable and ready for analysis. Discover questions and weighted toplines
across AI & Tech, Culture, Health, Money, Politics, Sports, then follow every record to a human-readable finding and
its verified primary source report.
Explore findings, search topics, and cite the research at data.verasight.io
Coverage at a glance
Survey waves… See the full description on the dataset page: https://huggingface.co/datasets/Verasight/verasight-data-library.un-digital-library
United Nations Digital Library (UNDL) Comprehensive Master Dataset
1. Executive Summary
Welcome to the United Nations Digital Library (UNDL) Comprehensive Master Dataset repository. This dataset represents a monumental effort to harvest, normalize, enrich, and democratize access to the vast archives of the United Nations. By leveraging advanced web harvesting techniques, robust state management, and modern big-data formats, this repository provides researchers… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/un-digital-library.prompt-library
Metrum AI Prompt Library
A prompt library for LLM inference workload and performance benchmarking,
prepared for use with metrum-ai/bench-cli.
It contains 593,730 records with prompt text, intended lengths, token
buckets, and reasoning labels. It contains no reference answers.
The full configuration preserves all source records, including repeated
prompts and their distinct workload targets. Prompts may appear duplicated,
with only target_output_length differing. These variants… See the full description on the dataset page: https://huggingface.co/datasets/metrum-ai/prompt-library.model-library
Model Library DB
Dataset Summary
The Model Library is a project that maps the risks associated with modern machine learning systems. Here, we assess some of the most recent and capable AI systems ever created. This is the database for the Model Library.
Supported Tasks and Leaderboards
This dataset serves as a catalog of machine learning models, all displayed in the Model Library.
Languages
English.
Dataset Structure
Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/model-library.real-time-library-occupancyQwen-CoT-Library
Qwen — CoT Library
A maintained mirror of publicly-available Qwen chain-of-thought datasets on Hugging Face — content-verified, deduplicated, and attributed to their original authors.
Dataset Viewer | Parquet
// what this is
A maintained library — a community mirror of publicly-available Qwen CoT datasets, aggregated, validity-filtered and content-verified, with per-row source attribution in first_source_dataset. It is not Crownelius' own data — every row… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/Qwen-CoT-Library.rlhf-library-OLMo-2-0425-1B-DPOrlhf-library
RLHF Book Completions Library
Completions for the demonstration of SFT vs. RLHF models. Interact with them at https://rlhfbook.com/library.
Completions are from the following models:
allenai/Llama-3.1-Tulu-3-8B-SFT
allenai/Llama-3.1-Tulu-3-8B-DPO
allenai/Llama-3.1-Tulu-3-70B-SFT
allenai/Llama-3.1-Tulu-3-70B-DPO
allenai/OLMo-2-0425-1B-SFT
allenai/OLMo-2-0425-1B-DPO
allenai/OLMo-2-1124-7B-SFT
allenai/OLMo-2-1124-7B-DPO
allenai/OLMo-2-1124-13B-SFT
allenai/OLMo-2-1124-13B-DPO… See the full description on the dataset page: https://huggingface.co/datasets/natolambert/rlhf-library.rlhf-library-tulu-2-7brlhf-library-OLMo-7B-0424-Instruct-hfrlhf-library-tulu-2-dpo-7brlhf-library-Llama-3.1-Tulu-3-8B-SFTCode to generate these:
#!/usr/bin/env python
# -*- coding: utf-8 -*-
"""
Minimal RLHF-vs-SFT generation script.
- Source prompts are hardcoded to: natolambert/rlhf-book-prompts-v2
and must expose columns: dataset (str), instruction (str), id (str)
- Output: per-model HF dataset at:
natolambert/rlhf-library-{MODEL_WITHOUT_ORG}
- Row schema (only):
dataset: str
instruction: str
id: str
completion: str
Assumes a recent vLLM version. Keep the model list vertical for… See the full description on the dataset page: https://huggingface.co/datasets/natolambert/rlhf-library-Llama-3.1-Tulu-3-8B-SFT.rlhf-library-OLMo-2-0425-1B-SFTrlhf-library-OLMo-2-0325-32B-DPOrlhf-library-Llama-3.1-Tulu-3-70B-SFTrlhf-library-Llama-3.1-Tulu-3-8B-DPOrlhf-library-OLMo-2-1124-13B-SFTrlhf-library-Llama-3.1-Tulu-3-70B-DPOrlhf-library-OLMo-7B-0424-SFT-hf
