datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
legal-practice-library
legal-practice-library
A clean, source-cited snapshot of the OpenAgreements practice-guide corpus: plain-English explainers of US state (and select international) law, currently covering non-compete / restrictive-covenant law and consumer data-privacy law. Published and maintained by openagreements.org.
Each note is written against primary law (statutes and cases), carries machine-verifiable source citations, and records the date it was last reviewed. The corpus is re-synced… See the full description on the dataset page: https://huggingface.co/datasets/open-agreements/legal-practice-library.library_of_congress_filtered
Library of Congress
Description
The Library of Congress (LoC) curates a collection of public domain books called "Selected Digitized Books".
We have downloaded over 130,000 English-language books from this public domain collection as OCR plain text files using the LoC APIs.
Dataset Statistics
Documents
UTF-8 GB
129,052
35.6
License Issues
While we aim to produce datasets with completely accurate licensing information, license… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/library_of_congress_filtered.biodiversity_heritage_library_filtered
Biodiversity Heritage Library
Description
The Biodiversity Heritage Library (BHL) is an open-access digital library for biodiversity literature and archives.
This dataset contains over 15 million public domain books and documents from the BHL collection.
These works were collected using the bulk data download interface provided by the BHL and were filtered based on their associated license metadata.
We use the optical character recognition (OCR)-generated text distributed… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/biodiversity_heritage_library_filtered.biodiversity_heritage_library
Biodiversity Heritage Library
Description
The Biodiversity Heritage Library (BHL) is an open-access digital library for biodiversity literature and archives.
This dataset contains over 42 million public domain books and documents from the BHL collection.
These works were collected using the bulk data download interface provided by the BHL and were filtered based on their associated license metadata.
We use the optical character recognition (OCR)-generated text… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/biodiversity_heritage_library.library_of_congress
Library of Congress (subset of Common Pile)
Description
The Library of Congress (LoC) curates a collection of public domain books called "Selected Digitized Books".
We have downloaded over 130,000 English-language books from this public domain collection as OCR plain text files using the LoC APIs.
This dataset is a subset of the Common Pile v0.1. For more information, see The Common Pile v0.1 paper.
Dataset Statistics
Documents
UTF-8 GB
135,500… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/library_of_congress.hebrew_library
Description
Export of Sefaria's Hebrew library data. This data represents over version in the library marked as Hebrew.
Schema
Field
Description
text
The text of a single segment in the library. A segment is the smallest chunk of test, usually representing a paragraph.
metadata
Dictionary of metadata. See below for schema.
Metadata Schema
Field
Description
url
URL to this segment in Sefaria
ref
Canonical Ref to this segment. Refs… See the full description on the dataset page: https://huggingface.co/datasets/Sefaria/hebrew_library.library-documentationThe library documentation retrieval source for code-rag-bench, contains all documentation for Python libraries available on devdocs.io.
master-ebook-library-deduplicatedlibrary-python-training-pool
Python library function-writing training pool
A pool of public data for training a model to write Python functions, many of them
calling libraries: 8.3 percent of the answers in the normalised layer import a library that
is not in the Python standard library. It is a straight collection of open datasets, not
a new corpus: every row comes from one of the sources below, at the revision named. Rows
an overlap filter flagged against held-out material this pool is kept separate from… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/library-python-training-pool.MimicAgent_Skill_Library
MimicAgent Skill Library
A library of 100 natural-language motion-skill specifications for quadruped
and wheeled-quadruped robots. Each entry describes a distinct, kinematically
feasible motion as an animation-style behavior (phase-based base and limb
motion over one cycle).
The library is the input to the MimicAgent text-to-trajectory pipeline,
where an LLM agent turns each skill description into executable code that
generates a motion prior in a MuJoCo simulator.… See the full description on the dataset page: https://huggingface.co/datasets/aicognition/MimicAgent_Skill_Library.english_library
Description
Export of Sefaria's English library data. This data represents over version in the library marked as English.
Schema
Field
Description
text
The text of a single segment in the library. A segment is the smallest chunk of test, usually representing a paragraph.
metadata
Dictionary of metadata. See below for schema.
Metadata Schema
Field
Description
url
URL to this segment in Sefaria
ref
Canonical Ref to this segment.… See the full description on the dataset page: https://huggingface.co/datasets/Sefaria/english_library.mindbots-soul-library-training-v2-richchristian-library
A Christian Library (Wesley's abridgments) with collation apparatus
278 works from Wesley's Christian Library (~4.64M words), full English text plus the collation cuts ('What Wesley Left Out'), headnotes and cross-references as separate fields.
Canonical home: https://achristianlibrary.org (each record carries its canonical URL). This dataset is a machine-generated export of that site's build, regenerated from the source of truth and never hand-edited; the site remains the one… See the full description on the dataset page: https://huggingface.co/datasets/historyofmethodism/christian-library.novel-themes-librarymindbots-soul-library-training-v2-richgreyforge-fintech-tool-fixture-library-v1
GreyForge Fintech Tool Fixture Library v1
Fixture library for harness development; not a reliability benchmark or
ChangeGuard diagnostic.
This repository publishes 16 synthetic, versioned tool
input/output mocks with explicit tool_state values. Use them to exercise
agent tool adapters and harness plumbing. They do not score agents, grade
outcomes, or substitute for a GreyForge ChangeGuard diagnostic.
Synthetic data — All request shapes and response bodies are fabricated
harness… See the full description on the dataset page: https://huggingface.co/datasets/GreyForge/greyforge-fintech-tool-fixture-library-v1.library-classification-systems
Library Classification Systems
This comprehensive dataset contains hierarchical outlines of major library classification systems, offering a valuable resource for researchers, librarians, and information scientists.
Classification System
Abbreviation
Primary Usage
Language
Entries
Dewey Decimal Classification
DDC
International
English
1110
Library of Congress Classification
LCC
International
English
6517
Universal Decimal Classification
UDC
International
English
2431… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/library-classification-systems.open-library-scraper
Open Library Scraper · Books, Authors, Editions & Subjects
Scrape Open Library books, authors, subjects, editions, and metadata via Open Library API. Fast HTTP scraper charging per returned record with tiered pricing.
Rows in this dataset
3,880
Fields
21
Collector runs behind it
50
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/open-library-scraper/ — 3,880 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/open-library-scraper.the-anarchist-libraryVoilà!
In View, a humble Vaudevillian Veteran, cast Vicariously as both Victim and Villain by the Vicissitudes of fate.
This Visage, no mere Veneer of Vanity, is a Vestige of the Vox populi, now Vacant, Vanished.
However, this Valorous Visitation of a bygone Vexation stands Vivified, and has Vowed to Vanquish these Venal and Virulent Vermin Vanguarding Vice and Vouchsafing the Violently Vicious and Voracious Violation of Volition.
mindbots-soul-library-training-v2-3-canonicalyale-library-entity-resolver-classificationsAetherius-Genesis-Libraryyale-library-note-classifications-classification-datamindbots-soul-library-training-v2-2-polishedhebrew_library
Description
Export of Sefaria's Hebrew library data. This data represents over version in the library marked as Hebrew.
Schema
Field
Description
text
The text of a single segment in the library. A segment is the smallest chunk of test, usually representing a paragraph.
metadata
Dictionary of metadata. See below for schema.
Metadata Schema
Field
Description
url
URL to this segment in Sefaria
ref
Canonical Ref to this… See the full description on the dataset page: https://huggingface.co/datasets/DJEAntonio/hebrew_library.library-usage-benchmark-2026library_testhook-pattern-library
Hook Pattern Library
TikTok 前 3 秒 Hook 模式库。字段:type, template, use_when.
RAFT_library_43d39bdeRAFT_library_3fc1eefe
