datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
public-domain-poetry
Overview
This dataset is a collection of approximately 38,500 poems from https://www.public-domain-poetry.com/.
Language
The language of this dataset is English.
License
All data in this dataset is public domain, which means you should be able to use it for anything you want, as long as you aren't breaking any law in the process of doing so.
Public-Domain-Music1920-raider-waite-tarot-public-domainpublic_domain_review_filtered
Public Domain Review
Description
The Public Domain Review is an online journal dedicated to exploration of works of art and literature that have aged into the public domain.
We collect all articles published in the Public Domain Review under a CC BY-SA license.
Dataset Statistics
Documents
UTF-8 GB
1,406
0.007
License Issues
While we aim to produce datasets with completely accurate licensing information, license laundering and… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/public_domain_review_filtered.1920-raider-waite-tarot-public-domainpublic-domain-art-restored
Public-Domain Art Restoration Archive
Public-domain museum artworks with scan damage detected and repaired by a
diffusion model where present, then upscaled 4x with a GAN
super-resolution model. Released CC0.
25,135 restored images are published in images/ (15.4 GB of AVIF), with per-item provenance in manifest/restored.parquet. 2,785 images (11.1%) were routed to the diffusion repair tier and 2,785 were inpainted. Median output long edge: 4,096 px.
What is… See the full description on the dataset page: https://huggingface.co/datasets/rishinaren/public-domain-art-restored.public_domain_review
Public Domain Review
Description
The Public Domain Review is an online journal dedicated to exploration of works of art and literature that have aged into the public domain.
We collect all articles published in the Public Domain Review under a CC BY-SA license.
Dataset Statistics
Documents
UTF-8 GB
1,412
0.007
License Issues
While we aim to produce datasets with completely accurate licensing information, license laundering and… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/public_domain_review.1920-raider-waite-tarot-public-domain-cleanedA cleaned up version of the multimodalart/1920-raider-waite-tarot-public-domain dataset, without the card borders and names
public_domain_sounds_3secs
Public Domain Sounds
This is a backup of the 635 copyright-free sound recordings submitted to pdsounds.org before April 2009.
all files split into 3 seconds chunks
LICENSE NOTICE
pdsounds.org - all sounds archive
March 2, 2009
ALL 635 SOUNDS in this archive are entirely public domain and copyright free.
No rights are reserved.
The sounds were recorded by volunteers and uploaded to pdsounds.org with the
condition that henceforth they are license-free and public domain.… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/public_domain_sounds_3secs.laion-publicdomainannotations_creators:
machine-generated
language_creators:
machine-generated
license:
cc-by-4.0
multilinguality:
multilingual
pretty_name: laion-publicdomain
size_categories:
100K<n<1M
source_datasets:
-laion/laion2B-en
tags:
laion
task_categories:
text-to-image
Dataset Card for laion-publicdomain
Dataset Summary
This dataset contains metadata about images from the LAION2B-eb dataset curated to a reasonable best guess of 'ethically sourced' images.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/devourthemoon/laion-publicdomain.publicdomainfiles
Dataset Card for PublicDomainFiles.com Collection
Dataset Summary
This dataset contains various public domain media files collected from PublicDomainFiles.com. The website hosts a diverse collection of user-shared content explicitly released into the public domain, including images, fonts, clip art, artwork, video clips, TV shows, and pictures. While images are the predominant file type, the dataset encompasses a wide range of multimedia formats. Each item includes… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/publicdomainfiles.poetry-greats-public-domain
Poetry Greats
Curated, poem-level extracts from Project Gutenberg for 20 canonical English-language poets. All source texts are public domain in the US (pre-1929 publication). Intended as a reference set of "gold" examples for evaluation, few-shot prompting, and stylometric study.
Contents
4,090 poems across 29 books and 20 poets:
Poet
Poems
Samuel Taylor Coleridge
913
H. W. Longfellow
616
Christina Rossetti
459
Emily Dickinson
446
Percy Bysshe Shelley… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/poetry-greats-public-domain.dbnl.org-dutch-public-domain
Dataset Card for "dbnl.org-dutch-public-domain"
Dataset Summary
This dataset comprises a collection of texts from the Dutch Literature in the public domain, specifically from the DBNL (Digitale Bibliotheek voor de Nederlandse Letteren) public domain collection. The collection includes books, poems, songs, and other documentation, letters, etc., that are at least 140 years old and thus free of copyright restrictions. Each entry in the dataset corresponds to one section of… See the full description on the dataset page: https://huggingface.co/datasets/jvdgoltz/dbnl.org-dutch-public-domain.shangkhachil-bengali-public-domain
Bengali Public-Domain Literature
101 complete works by 21 authors,
11,250,629 characters. Corpus corpus-f8c532fcb4e7, built 2026-09-09.
Where these texts are read
https://shangkhachil.com — the reading site this corpus was built for. Free, no
account, 246 works by 28 authors. The complete text of
every work in this file can be read there.
This file is the text. The site is the part a JSONL cannot be:
Rights computed for the reader's own country, at the edge… See the full description on the dataset page: https://huggingface.co/datasets/mir178/shangkhachil-bengali-public-domain.publicdomainpictures
Dataset Card for Public Domain Pictures
Dataset Summary
This dataset contains metadata for 644,412 public domain images from publicdomainpictures.net, a public domain photo sharing platform. The dataset includes detailed image metadata including titles, descriptions, and keywords.
Languages
The dataset is monolingual:
English (en): All metadata including titles, descriptions and keywords
Dataset Structure
Data Fields
The metadata for… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/publicdomainpictures.humor-greats-public-domain
Humor Greats
Short humorous texts -- jokes, aphorisms, quips, anecdotes -- extracted from public-domain humor collections on Project Gutenberg. All source texts are pre-1929 and public domain in the US. Intended as a reference set of "gold" humorous writing for evaluation, few-shot prompting, and stylistic study.
Contents
19,354 entries across 8 books, spanning two clear registers:
Concentrated wit (authored):
Book
Author
Entries
The Devil's Dictionary
Ambrose… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/humor-greats-public-domain.positivequotation-public-domain-quotes
PositiveQuotation Source-Verified Public Domain Quotes
This small dataset contains exactly 30 English proverbs matched to numbered entries in a public-domain U.S. source. It is designed for examples, prototypes, educational projects, and applications that need compact quotation records with auditable provenance.
Homepage: https://positivequotation.com/public-domain-quotes
API documentation: https://positivequotation.com/developers/public-domain-quotes-api
Live JSON API:… See the full description on the dataset page: https://huggingface.co/datasets/geosfero/positivequotation-public-domain-quotes.RPGPT_PublicDomain-alpacaExperimental Synthetic Dataset of Public Domain Character Dialogue in Roleplay Format
Generated using scripts from my https://github.com/practicaldreamer/build-a-dataset repo
license: mit
RPGPT_PublicDomain-ShareGPTExperimental Synthetic Dataset of Public Domain Character Dialogue in Roleplay Format
Generated using scripts from my https://github.com/practicaldreamer/build-a-dataset repo
license: mit
public-domain-poetrypublic-domain-poetry-with-embeddingsModified version of https://huggingface.co/datasets/mkessle/public-domain-poetry with added embeddings data.
Embeddings were generated using the OpenAI text embedding model.
dataset_info:
features:
- name: Title
dtype: string
- name: Author
dtype: string
- name: Lines
dtype: float64
- name: Views
dtype: int64
- name: Poem Text
dtype: string
- name: About
dtype: string
- name: Birth and Death Dates
dtype: string
- name: id
dtype: int64… See the full description on the dataset page: https://huggingface.co/datasets/pvd-dot/public-domain-poetry-with-embeddings.practical-dreamer-RPGPT_PublicDomain
Public Domain Character RPG Dataset
This dataset is a reformatted version of practical-dreamer/RPGPT_PublicDomain-alpaca into a ShareGPT-like format. Unfortunately, detailed information about the original dataset is scarce.
In total, the dataset includes 3 032 conversations, with some extending up to 50 turns.
The dataset consists of the following fields:
conversations: ShareGPT-like format representing the dialogue.
The 'system' message provides a randomized introduction… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/practical-dreamer-RPGPT_PublicDomain.ai-tube-public-domain
Description
Videos made using models trained on Public Domain content.
Model
SVD
Voice
Muted
Orientation
Landscape
Tags
Public Domain
Style
1928 animation movie, movie still
Music
1920 piano ragtime
Prompt
A channel generating short animated video of stories in the public domain, between 2 to 3 minutes
Videos are humoristic, like in Charle Chaplin movies.
They include tons of funny scenes and jokes about… See the full description on the dataset page: https://huggingface.co/datasets/jbilcke-hf/ai-tube-public-domain.code-domaine-public-fluvial-navigation-interieure
Code du domaine public fluvial et de la navigation intérieure, non-instruct (2025-09-20)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-domaine-public-fluvial-navigation-interieure.1920-raider-waite-tarot-public-domain1920-raider-waite-tarot-public-domaindomaine-public-fluvial-france-entiere-sources-mydpf-cerema
Domaine Public Fluvial France entière. Sources MyDpf #Cerema
[!NOTE]
Ce jeu de données Hugging Face est vide. Cette carte sert seulement à référencer le jeu de données Domaine Public Fluvial France entière. Sources MyDpf #Cerema qui est disponible à l'adresse https://www.data.gouv.fr/datasets/68639b7d3492920481cfeeba
Description
myDPF est un outil à l'usage de la direction générale des infrastructures, des transport et de la mer (DGITM) du ministère du développement… See the full description on the dataset page: https://huggingface.co/datasets/french-open-data/domaine-public-fluvial-france-entiere-sources-mydpf-cerema.catholic_public_domain_version_en
Catholic Public Domain Version (CPDV)
Description
The Catholic Public Domain Version (CPDV) is a modern English translation of the Sacred Bible, based on the Latin Vulgate. Translated by Ronald L. Conte Jr. in 2009, it is the only modern Catholic Bible translation in the Public Domain.
The CPDV updates the Challoner Douay-Rheims version while remaining faithful to the Vulgate text. It uses contemporary English while preserving the theological precision of the… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/catholic_public_domain_version_en.PublicDomainArt-ARTICPublicDomainImages
