yoonholee/humor-greats-public-domain
Humor Greats Short humorous texts -- jokes, aphorisms, quips, anecdotes -- extracted from public-domain humor collections on Project Gutenberg. All source texts are pre-1929 and public domain in the US. Intended as a reference set of "gold" humorous writing for evaluation, few-shot prompting, and stylistic study. Contents 19,354 entries across 8 books, spanning two clear registers: Concentrated wit (authored): Book Author Entries The Devil's Dictionary… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/humor-greats-public-domain.
Humor Greats
Short humorous texts -- jokes, aphorisms, quips, anecdotes -- extracted from public-domain humor collections on Project Gutenberg. All source texts are pre-1929 and public domain in the US. Intended as a reference set of "gold" humorous writing for evaluation, few-shot prompting, and stylistic study.
Contents
19,354 entries across 8 books, spanning two clear registers:
Concentrated wit (authored):
Period joke / anecdote collections:
The first register -- Bierce, Wilde, Twain -- is the densest "gold." The second contains classic Victorian and Edwardian jest-book material of mixed quality; use it as a reference for period humor style rather than as peak wit.
Schema
Source books
- Ambrose Bierce, *The Devil's Dictionary* (1911)
- Oscar Wilde, *Miscellaneous Aphorisms; The Soul of Man* (1893)
- Mark Twain, *Quotations from the Project Gutenberg Editions of the Works of Mark Twain*
- Mark Lemon, *The Jest Book: The Choicest Anecdotes and Sayings* (1864)
- Herbert L. Fanning, *Toaster's Handbook: Jokes, Stories, and Quotations* (1914)
- Paul Mosher, *More Toasts* (1922)
- Anonymous, *Jokes For All Occasions* (1923)
- Various, *The Book of Anecdotes and Budget of Fun* (1907)
Parsing notes
Each book was split on blank-line paragraph boundaries after stripping Gutenberg front/back matter and skipping the first 4% of body (preface / table of contents). A paragraph was kept as an entry if: it was 1-15 lines, 30-1800 chars, started with an uppercase letter / quotation mark / digit, and did not look like a ToC or all-caps section header. Bierce entries additionally required the WORD, pos. dictionary-format header to filter out his prose interludes. Twain's file had a 12% ToC which was skipped explicitly.
Expect some noise:
- The Jest Book and Toaster's Handbook style conventions (numbered sections with titles like "VIII.--BEARDING A BARBER") mean titles sometimes appear mixed into the body text.
- Italic markers from the Gutenberg plaintext (
_word_) are preserved as-is. - Mark Twain's quotations are formatted as single-line fragments in the source; the paragraph-based parser undercounts them.
Intended uses
- Reference / few-shot prompting for humor generation.
- Stylistic evaluation: "does this text read like Bierce? Wilde? a period joke book?"
- Pairwise preference collection: "which of these two aphorisms is funnier / more in-voice?"
- Training data for humor classifiers (period vs. modern, concentrated wit vs. punchline).
License
CC0 1.0 Universal -- all source texts are public domain. Parsing and curation are also released under CC0.
Citation
Source texts courtesy of Project Gutenberg. Curation by Yoonho Lee.
