CoolFace
Datasetpublic

ai-historian/german-newspaper-pages-index

📋 On the licensing of this data Every item here is labelled to the best of my ability. The rights statement is taken per issue from what the holding institution recorded in the Deutsche Digitale Bibliothek, captured at download time — never inferred, and never applied at newspaper level to issues that may differ. Labelling at this scale is imperfect. If you believe anything here is mislabelled, or you hold rights in any of this material, please write to lorenz.hufe@posteo.de —… See the full description on the dataset page: https://huggingface.co/datasets/ai-historian/german-newspaper-pages-index.

sourceHugging Facecc0-1.0updated 4d agoView on Hugging Face
0likes134downloads
Dataset Card
### 📋 On the licensing of this data Every item here is labelled to the best of my ability. The rights statement is taken per issue from what the holding institution recorded in the Deutsche Digitale Bibliothek, captured at download time — never inferred, and never applied at newspaper level to issues that may differ. Labelling at this scale is imperfect. If you believe anything here is mislabelled, or you hold rights in any of this material, please write to [lorenz.hufe@posteo.de](mailto:lorenz.hufe@posteo.de) — I will resolve it immediately.
## ⚠️ Content notice · Inhaltshinweis This collection contains National Socialist propaganda and other historical material that is antisemitic, racist and otherwise discriminatory. It is published solely as a historical source, for research, teaching and documentation. The corpus includes the German press of 1933–1945 — 187,453 pages from 22 newspapers — comprising state and party propaganda, antisemitic agitation, denaturalisation and deportation notices naming persecuted individuals, and reporting that legitimised persecution and war. Material from the 18th and 19th centuries likewise carries the colonial, antisemitic and racist language of its time, and the OCR text reproduces it verbatim. Reproduction here is documentary and implies no endorsement. The views expressed are those of the historical publications, not of the depositors or the holding institutions. German law permits the use of such material for civic education, research, teaching, art, and reporting on contemporary events or history — the social-adequacy clause of §§ 86, 86a StGB. Users are responsible for lawful use in their own jurisdiction, including any restrictions on reproducing or displaying the symbols and texts contained here. Anyone training generative models on this corpus should expect this material to be reproduced, and should filter accordingly. --- Diese Sammlung enthält nationalsozialistische Propaganda sowie weiteres historisches Material antisemitischen, rassistischen und anderweitig diskriminierenden Inhalts. Sie wird ausschließlich als historische Quelle zu Zwecken der Forschung, der Lehre und der Dokumentation bereitgestellt. Enthalten ist die deutsche Presse der Jahre 1933–1945 (187.453 Seiten aus 22 Zeitungen): Staats- und Parteipropaganda, antisemitische Hetze, Ausbürgerungs- und Deportationsbekanntmachungen mit Namen verfolgter Personen sowie Berichterstattung, die Verfolgung und Krieg legitimierte. Auch das Material des 18. und 19. Jahrhunderts enthält die koloniale, antisemitische und rassistische Sprache seiner Zeit; der OCR-Text gibt sie wörtlich wieder. Die Wiedergabe erfolgt zu dokumentarischen Zwecken und stellt keine Billigung dar. Die geäußerten Auffassungen sind die der historischen Publikationen, nicht die der Bereitstellenden oder der besitzenden Einrichtungen. Die Nutzung ist im Rahmen der Sozialadäquanzklausel der §§ 86, 86a StGB zulässig (staatsbürgerliche Aufklärung, Wissenschaft, Forschung, Lehre, Kunst, Berichterstattung über Vorgänge des Zeitgeschehens oder der Geschichte). Für die Rechtmäßigkeit der Nutzung in der jeweiligen Rechtsordnung sind die Nutzenden selbst verantwortlich.

German Newspaper Pages — Metadata Index

One row for each of the 7,000,060 pages in `ai-historian/german-newspaper-pages`, as Parquet. 97 MB, so the Hub viewer indexes all of it: you can search, filter and sort the whole corpus by date, newspaper, place or licence, and then jump to the exact shard holding the page image.

The main release is 7.97 TB of WebDataset tars. Its page contents are complete, but the Hub's viewer cannot build a Parquet conversion for it (a viewer-side crash on WebDataset tars: it opens each tar as a stream whose .size is None, then calls int() on it). The big set therefore shows a preview but no search. This index exists to give that back.

Columns

ColumnMeaning
keysample key in the main dataset (<date>_<issuehash>_<pageid>)
issue_id, page_id, page_numberissue and page within it
date, year, month, dayexact publication date, ISO and split out for filtering
zdb_idstable newspaper id — group on this, not the title
paper_titlenewspaper name as printed on that issue
placeplace of publication (empty for ~1 % of pages)
language, provider_ddb_id, licenseper-issue rights statement and holding institution
has_altowhether the library's own ALTO ships with that page (38 % overall)
shardthe tar in the main dataset holding this page, e.g. data/pages-01200.tar
ddb_urlthe page at the Deutsche Digitale Bibliothek

Getting from a row to the page image

python
from datasets import load_dataset
from huggingface_hub import hf_hub_download
import tarfile, json

idx = load_dataset("ai-historian/german-newspaper-pages-index", split="train")
hit = idx.filter(lambda r: r["place"] == "Köln" and r["year"] == 1883)[0]

tar = hf_hub_download("ai-historian/german-newspaper-pages", hit["shard"],
                      repo_type="dataset")
with tarfile.open(tar) as t:
    img  = t.extractfile(hit["key"] + ".jpg").read()
    meta = json.loads(t.extractfile(hit["key"] + ".json").read())   # layout boxes + OCR text

Caveats

  • —`paper_title` drifts: 650 newspapers carry 1,930 distinct title strings because papers were renamed. Join and group on zdb_id.
  • —`place` is empty for ~1 % of pages — 29 newspapers have no location in the DDB record.
  • —Page counts per year reflect what was digitised and is public domain, not what was printed. Coverage is heavily west-German and Saxon; Berlin is nearly absent before 1871.
  • —The OCR text and layout boxes are not in this index — they live in each page's .json inside the main dataset.

Licence

CC0 1.0. Every indexed page comes from an issue carrying CC Public Domain Mark 1.0 or CC0.

Acknowledgements — thank you to the holding institutions

This corpus exists only because libraries and archives digitised these newspapers, cleared their rights, and published them openly. Thank you to every institution below, and to the Deutsche Digitale Bibliothek and the zeitpunkt.NRW portal for aggregating and serving them.

Nothing here was created by this project except the layout boxes and the OCR text. The page images, the cataloguing, and the rights clearance are theirs.

Holding institutionPagesNewspapers
Sächsische Landesbibliothek – Staats- und Universitätsbibliothek Dresden (SLUB)1,833,699105
Staatsbibliothek zu Berlin – Preußischer Kulturbesitz1,061,5475
Universitäts- und Landesbibliothek Bonn — via zeitpunkt.NRW945,325248
Staats- und Universitätsbibliothek Hamburg Carl von Ossietzky885,18631
Universitäts- und Landesbibliothek Münster — via zeitpunkt.NRW449,50792
WĂĽrttembergische Landesbibliothek Stuttgart384,21218
Universitätsbibliothek Mannheim375,3094
Further holdings delivered through the DDB IIIF endpoint369,91124
Bayerische Staatsbibliothek MĂĽnchen (MDZ)295,58942
Universitäts- und Landesbibliothek Düsseldorf — via zeitpunkt.NRW116,88719
Universitätsbibliothek Heidelberg114,17344
Stadtarchiv Ladenburg63,5903
Gottfried Wilhelm Leibniz Bibliothek Hannover (GWLB)58,8283
MARCHIVUM Mannheim28,9787
Bibliothek der Friedrich-Ebert-Stiftung17,3196

Counts are pages in this release; 17 DDB provider entries map to the 15 institutions above (SLUB Dresden and MARCHIVUM each appear under two). Every page record carries provider_ddb_id and a ddb_url, so the holding institution is identifiable per page — please credit it when you use or cite individual pages.

Thanks are also due to the readers and cataloguers whose work is invisible here: the newspapers were indexed, described and dated long before any of this could be automated.