hadro/green-books-travel-guides
African American Travel Guides: The Green Book & Companion Directories (1930–1966) A unified, structured dataset of 113,827 business and lodging listings transcribed from 50 volumes of mid-20th-century African American travel guides, spanning 1930–1966. During the Jim Crow era, these guides told Black travelers which hotels, restaurants, tourist homes, service stations, and other businesses would serve them safely. This dataset brings The Negro Motorist Green Book together with… See the full description on the dataset page: https://huggingface.co/datasets/hadro/green-books-travel-guides.
African American Travel Guides: The Green Book & Companion Directories (1930–1966)
A unified, structured dataset of 113,827 business and lodging listings transcribed from 50 volumes of mid-20th-century African American travel guides, spanning 1930–1966. During the Jim Crow era, these guides told Black travelers which hotels, restaurants, tourist homes, service stations, and other businesses would serve them safely. This dataset brings The Negro Motorist Green Book together with seven lesser-known companion publications into a single comparable schema.
Every listing links back to the exact page region of the scanned source via IIIF (canvas_fragment / image), so any row can be traced to the original document.
Note on accuracy: this data was transcribed and structured by a vision-language model (VLM), not by hand — the VLM performs both the OCR and the field extraction. On these dense, multi-column directory pages it sometimes confuses columns, mis-transcribes text, or mis-labels a listing's category. Treat fields as machine-generated, and verify against the source scan (canvas_fragment/image) when accuracy matters. You can browse the data with the source page images shown for each row in the live viewer: <https://hadro.github.io/green-books/all-volumes>. See Extraction method and Data quality & caveats.
Publications covered
Source volumes are digitized by The New York Public Library (Schomburg Center for Research in Black Culture and others). The 1946 Green Book edition is digitized by the Library of Congress; all other volumes come from NYPL.
Schema
20 columns. Observed = transcribed from the page; Derived = produced by an automated extraction/inference pipeline and should not be treated as ground truth. Note that "Observed" here still means machine-transcribed by a vision-language model (see Extraction method), not human-keyed.
Rights & licensing
License: CC0 1.0 (public domain dedication) for the dataset.
Two independent bases support free reuse of this data:
- The listings are facts, and facts are not copyrightable. Under Feist Publications, Inc. v. Rural Telephone Service Co., 499 U.S. 340 (1991) — a case about a telephone directory — names, addresses, and categories in a directory carry no copyright. A faithful transcription of factual directory data is uncopyrightable regardless of the source document's status.
- Source-scan rights, verified against the NYPL API (49 NYPL volumes) and the Library of Congress item record (1 LOC volume):
- 49 of 50 volumes are marked Public Domain in the United States — 48 NYPL volumes verified via the NYPL API (
NoC-US, NYPL statusPDREN) plus the 1946 Green Book from the Library of Congress, which states no known restrictions on publication (NoC-US). - 1 of 50 — Travelguide 1957 (2,483 listings) — is an in-copyright orphan work (
InC-RUU, NYPL statusICORPHAN): NYPL identified a copyright notice, could not locate a rights-holder, and released it as an orphan work. Its scan is not public domain, but its factual listings remain uncopyrightable under (1). It is included here and flagged; per-volume rights are in `volume_rights.csv`.
volume_rights.csv records one row per volume. Its copyright_status column (formerly nypl_copyright_status) holds each source institution's own status code or phrase — PDREN/ICORPHAN for NYPL, a plain-language phrase for LOC — since these vocabularies differ by institution; a source_institution column (NYPL or LOC) disambiguates which vocabulary applies to a given row.
Attribution to NYPL is not legally required but is requested as a courtesy: "From The New York Public Library." For the 1946 edition: "From the Library of Congress."
Historical sensitivity
These records document real businesses and, in the case of tourist/guest homes, private residences of real people, published because segregation made ordinary travel dangerous for Black Americans. Addresses appear as printed. Please use this data with respect for that context — for historical, educational, and research purposes — and be mindful that some listed addresses are private homes.
Data quality & caveats
This corpus is published as faithfully-transcribed data with light post-processing. Known issues:
- Raw vs normalized `state`/`category`. The raw
categoryandstatecolumns appear exactly as extracted, including case inconsistencies (NEW YORKvsNew York,GENERALvsGeneral) and near-duplicate labels (Hotels and Motels,Hotels-Motels-Tourists,Hotels - Motels - Tourist Homes - Restaurants). For convenience, `category_normalized` and `state_normalized` apply a mechanical fold —category_normalizeduses a case-fold plus an explicit typo/synonym/section-header map (the same logic that powers the companion web explorer, shared via `gb-categories.json`), collapsing 760 → 463 labels;state_normalizedis trim + uppercase only, collapsing 276 → 208. Both normalized columns are a mechanical cleanup, not an authoritative taxonomy or gazetteer — thestatelong tail still includes OCR noise and non-US locations (e.g.CANADA), and rare categories keep their (uppercased) raw form. Use the raw columns when you need the source values verbatim. - Vision-language-model extraction errors. Both the OCR and the field extraction were done by a vision-language model reading the page images (see Extraction method), not a human or a deterministic OCR/layout engine. On these dense, multi-column directory pages the model sometimes confuses columns — pulling a value into the wrong field or attaching it to the wrong listing — and can mis-transcribe unusual names/abbreviations or mis-identify a listing's category (e.g. a section-header category bleeding onto adjacent entries, or a
name/addressboundary drawn in the wrong place). These are systematic model-interpretation errors, not random noise, and they surface most incategory,name, andaddress. When accuracy matters for a given row, check it against the source scan viacanvas_fragment/image. - `category` is partly inferred, not purely transcribed. Treat it as a helpful signal, not authoritative classification.
- No cross-volume deduplication. The same business recurs across editions and years by design — this is a listings-over-time corpus, not a deduplicated business registry. Group by (
name,address,city) if you need unique establishments. - Sparse columns are expected, not errors:
proprietor,rates,phone, andnoteswere only printed for some listings and vary widely by publication. - Internal QA/drift flags from the production pipeline are intentionally not included in this public release.
Reconstructing a page image / viewer link
canvas_fragment encodes the IIIF canvas and the #xywh=x,y,w,h pixel region of the listing on the page; image is the NYPL image ID (or, for the 1946 edition, the LOC IIIF service identifier). Together they let you fetch the source page via NYPL's IIIF Image API — or the Library of Congress's for the 1946 edition — or deep-link into a IIIF viewer to see the original listing in context.
Loading
from datasets import load_dataset
ds = load_dataset("hadro/green-books-travel-guides")
print(ds["train"][0])Extraction method
Derived from digitized volumes in the NYPL Digital Collections via an automated, vision-language-model (VLM) pipeline — the extraction code is open source: [hadro/directory-pipeline](https://github.com/hadro/directory-pipeline/). Two VLM-driven stages produce the data:
- VLM-OCR — a vision-language model reads each scanned page image and transcribes its text, rather than a traditional OCR engine (e.g. Tesseract) operating on glyphs alone.
- VLM structuring / NER — a vision-language model segments each page into individual listings and labels the fields (
name,address,category,phone, …).
Because both stages depend on a model interpreting the page image — and these guides are dense, small-print, multi-column directory layouts — the output carries characteristic VLM error modes (column confusion, mis-transcription, field mis-identification) described under Data quality & caveats. The upside of the VLM approach is that it recovers structure and readings that flat OCR misses; the tradeoff is these interpretation errors. Every row keeps a canvas_fragment / image pointer so any value can be verified against the original scan.
Companion IIIF viewer and Green Book explorer: <https://hadro.github.io/green-books/all-volumes>.
Citation
@dataset{green_book_travel_guides,
title = {African American Travel Guides: The Green Book and Companion Directories (1930--1966)},
author = {Hadro, Josh},
year = {2026},
note = {Structured listings transcribed from 50 digitized volumes, New York Public Library and Library of Congress Digital Collections},
url = {https://huggingface.co/datasets/hadro/green-books-travel-guides}
}