CoolFace
Datasetpublic

yoonholee/humor-greats-public-domain

Humor Greats Short humorous texts -- jokes, aphorisms, quips, anecdotes -- extracted from public-domain humor collections on Project Gutenberg. All source texts are pre-1929 and public domain in the US. Intended as a reference set of "gold" humorous writing for evaluation, few-shot prompting, and stylistic study. Contents 19,354 entries across 8 books, spanning two clear registers: Concentrated wit (authored): Book Author Entries The Devil's Dictionary… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/humor-greats-public-domain.

sourceHugging Facecc0-1.0updated 6mo agoView on Hugging Face
0likes61downloads
Dataset Card

Humor Greats

Short humorous texts -- jokes, aphorisms, quips, anecdotes -- extracted from public-domain humor collections on Project Gutenberg. All source texts are pre-1929 and public domain in the US. Intended as a reference set of "gold" humorous writing for evaluation, few-shot prompting, and stylistic study.

Contents

19,354 entries across 8 books, spanning two clear registers:

Concentrated wit (authored):

BookAuthorEntries
The Devil's DictionaryAmbrose Bierce852
Miscellaneous AphorismsOscar Wilde552
Quotations from Mark TwainMark Twain121

Period joke / anecdote collections:

BookYearEntries
Toaster's Handbook19145,528
More Toasts19226,428
Jokes For All Occasions19232,556
The Jest Book18641,903
Book of Anecdotes and Budget of Fun19071,414

The first register -- Bierce, Wilde, Twain -- is the densest "gold." The second contains classic Victorian and Edwardian jest-book material of mixed quality; use it as a reference for period humor style rather than as peak wit.

Schema

ColumnTypeDescription
authorstringAuthor or compiler
book_titlestringSource book
gutenberg_idintProject Gutenberg book ID
yearintPublication year
textstringThe joke / aphorism / anecdote body
char_countintCharacter count
line_countintNon-blank line count

Source books

Parsing notes

Each book was split on blank-line paragraph boundaries after stripping Gutenberg front/back matter and skipping the first 4% of body (preface / table of contents). A paragraph was kept as an entry if: it was 1-15 lines, 30-1800 chars, started with an uppercase letter / quotation mark / digit, and did not look like a ToC or all-caps section header. Bierce entries additionally required the WORD, pos. dictionary-format header to filter out his prose interludes. Twain's file had a 12% ToC which was skipped explicitly.

Expect some noise:

  • —The Jest Book and Toaster's Handbook style conventions (numbered sections with titles like "VIII.--BEARDING A BARBER") mean titles sometimes appear mixed into the body text.
  • —Italic markers from the Gutenberg plaintext (_word_) are preserved as-is.
  • —Mark Twain's quotations are formatted as single-line fragments in the source; the paragraph-based parser undercounts them.

Intended uses

  • —Reference / few-shot prompting for humor generation.
  • —Stylistic evaluation: "does this text read like Bierce? Wilde? a period joke book?"
  • —Pairwise preference collection: "which of these two aphorisms is funnier / more in-voice?"
  • —Training data for humor classifiers (period vs. modern, concentrated wit vs. punchline).

License

CC0 1.0 Universal -- all source texts are public domain. Parsing and curation are also released under CC0.

Citation

Source texts courtesy of Project Gutenberg. Curation by Yoonho Lee.