ai-historian/german-newspaper-pages-index
📋 On the licensing of this data Every item here is labelled to the best of my ability. The rights statement is taken per issue from what the holding institution recorded in the Deutsche Digitale Bibliothek, captured at download time — never inferred, and never applied at newspaper level to issues that may differ. Labelling at this scale is imperfect. If you believe anything here is mislabelled, or you hold rights in any of this material, please write to lorenz.hufe@posteo.de —… See the full description on the dataset page: https://huggingface.co/datasets/ai-historian/german-newspaper-pages-index.
### 📋 On the licensing of this data Every item here is labelled to the best of my ability. The rights statement is taken per issue from what the holding institution recorded in the Deutsche Digitale Bibliothek, captured at download time — never inferred, and never applied at newspaper level to issues that may differ. Labelling at this scale is imperfect. If you believe anything here is mislabelled, or you hold rights in any of this material, please write to [lorenz.hufe@posteo.de](mailto:lorenz.hufe@posteo.de) — I will resolve it immediately.
## ⚠️ Content notice · Inhaltshinweis This collection contains National Socialist propaganda and other historical material that is antisemitic, racist and otherwise discriminatory. It is published solely as a historical source, for research, teaching and documentation. The corpus includes the German press of 1933–1945 — 187,453 pages from 22 newspapers — comprising state and party propaganda, antisemitic agitation, denaturalisation and deportation notices naming persecuted individuals, and reporting that legitimised persecution and war. Material from the 18th and 19th centuries likewise carries the colonial, antisemitic and racist language of its time, and the OCR text reproduces it verbatim. Reproduction here is documentary and implies no endorsement. The views expressed are those of the historical publications, not of the depositors or the holding institutions. German law permits the use of such material for civic education, research, teaching, art, and reporting on contemporary events or history — the social-adequacy clause of §§ 86, 86a StGB. Users are responsible for lawful use in their own jurisdiction, including any restrictions on reproducing or displaying the symbols and texts contained here. Anyone training generative models on this corpus should expect this material to be reproduced, and should filter accordingly. --- Diese Sammlung enthält nationalsozialistische Propaganda sowie weiteres historisches Material antisemitischen, rassistischen und anderweitig diskriminierenden Inhalts. Sie wird ausschließlich als historische Quelle zu Zwecken der Forschung, der Lehre und der Dokumentation bereitgestellt. Enthalten ist die deutsche Presse der Jahre 1933–1945 (187.453 Seiten aus 22 Zeitungen): Staats- und Parteipropaganda, antisemitische Hetze, Ausbürgerungs- und Deportationsbekanntmachungen mit Namen verfolgter Personen sowie Berichterstattung, die Verfolgung und Krieg legitimierte. Auch das Material des 18. und 19. Jahrhunderts enthält die koloniale, antisemitische und rassistische Sprache seiner Zeit; der OCR-Text gibt sie wörtlich wieder. Die Wiedergabe erfolgt zu dokumentarischen Zwecken und stellt keine Billigung dar. Die geäußerten Auffassungen sind die der historischen Publikationen, nicht die der Bereitstellenden oder der besitzenden Einrichtungen. Die Nutzung ist im Rahmen der Sozialadäquanzklausel der §§ 86, 86a StGB zulässig (staatsbürgerliche Aufklärung, Wissenschaft, Forschung, Lehre, Kunst, Berichterstattung über Vorgänge des Zeitgeschehens oder der Geschichte). Für die Rechtmäßigkeit der Nutzung in der jeweiligen Rechtsordnung sind die Nutzenden selbst verantwortlich.
German Newspaper Pages — Metadata Index
One row for each of the 7,000,060 pages in `ai-historian/german-newspaper-pages`, as Parquet. 97 MB, so the Hub viewer indexes all of it: you can search, filter and sort the whole corpus by date, newspaper, place or licence, and then jump to the exact shard holding the page image.
The main release is 7.97 TB of WebDataset tars. Its page contents are complete, but the Hub's viewer cannot build a Parquet conversion for it (a viewer-side crash on WebDataset tars: it opens each tar as a stream whose .size is None, then calls int() on it). The big set therefore shows a preview but no search. This index exists to give that back.
Columns
Getting from a row to the page image
from datasets import load_dataset
from huggingface_hub import hf_hub_download
import tarfile, json
idx = load_dataset("ai-historian/german-newspaper-pages-index", split="train")
hit = idx.filter(lambda r: r["place"] == "Köln" and r["year"] == 1883)[0]
tar = hf_hub_download("ai-historian/german-newspaper-pages", hit["shard"],
repo_type="dataset")
with tarfile.open(tar) as t:
img = t.extractfile(hit["key"] + ".jpg").read()
meta = json.loads(t.extractfile(hit["key"] + ".json").read()) # layout boxes + OCR textCaveats
- `paper_title` drifts: 650 newspapers carry 1,930 distinct title strings because papers were renamed. Join and group on
zdb_id. - `place` is empty for ~1 % of pages — 29 newspapers have no location in the DDB record.
- Page counts per year reflect what was digitised and is public domain, not what was printed. Coverage is heavily west-German and Saxon; Berlin is nearly absent before 1871.
- The OCR text and layout boxes are not in this index — they live in each page's
.jsoninside the main dataset.
Licence
CC0 1.0. Every indexed page comes from an issue carrying CC Public Domain Mark 1.0 or CC0.
Acknowledgements — thank you to the holding institutions
This corpus exists only because libraries and archives digitised these newspapers, cleared their rights, and published them openly. Thank you to every institution below, and to the Deutsche Digitale Bibliothek and the zeitpunkt.NRW portal for aggregating and serving them.
Nothing here was created by this project except the layout boxes and the OCR text. The page images, the cataloguing, and the rights clearance are theirs.
Counts are pages in this release; 17 DDB provider entries map to the 15 institutions above (SLUB Dresden and MARCHIVUM each appear under two). Every page record carries provider_ddb_id and a ddb_url, so the holding institution is identifiable per page — please credit it when you use or cite individual pages.
Thanks are also due to the readers and cataloguers whose work is invisible here: the newspapers were indexed, described and dated long before any of this could be automated.
