CoolFace
Datasetpublic

histde/dta-documents

Deutsches Textarchiv (DTA) Documents This datasets hosts all documents from the Deutsches Textarchiv (DTA). One row per work of the Deutsches Textarchiv (DTA), built from the official DTA dump dta_komplett_2026-02-10 (TCF format). The dataset contains 5,480 documents spanning 1472 to 1987 with about 205M tokens and 1.4B characters of German text, with two text views per document: text: the historical, layout-faithful transcription (line breaks, long s ſ, combining diacritics… See the full description on the dataset page: https://huggingface.co/datasets/histde/dta-documents.

sourceHugging Faceotherupdated 1mo agoView on Hugging Face
0likes224downloads
Dataset Card

Deutsches Textarchiv (DTA) Documents

This datasets hosts all documents from the Deutsches Textarchiv (DTA).

One row per work of the Deutsches Textarchiv (DTA), built from the official DTA dump dta_komplett_2026-02-10 (TCF format). The dataset contains 5,480 documents spanning 1472 to 1987 with about 205M tokens and 1.4B characters of German text, with two text views per document:

  • text: the historical, layout-faithful transcription (line breaks, long s ſ, combining diacritics, original spelling);
  • text_normalized: the same text in modern orthography, produced by applying the DTA CAB orthography corrections (operation="replace") to the token layer.

Rich bibliographic metadata comes directly from the embedded TEI header: authors with GND identifiers, original publication date and place, publisher, holding library and shelfmark, typeface, four genre schemes, DTA subcorpus labels and the per-work license. Licenses vary across the corpus, so please check the license_family column before reuse (see Licensing).

Dataset structure

Configs and splits

ConfigSplitRowsFiles
documents (default)train5,480documents/train-0000[0-5].parquet (6 shards, zstd-compressed Parquet)

There is only one split, because the dataset is a corpus and not a benchmark.

Loading

python
from datasets import load_dataset

ds = load_dataset("histde/dta-documents", "documents", split="train")
print(ds[0]["title"], ds[0]["year"])

# Only permissively licensed works, as a streaming iterator
permissive = load_dataset("histde/dta-documents", "documents", split="train", streaming=True)
permissive = permissive.filter(lambda d: d["license_family"] in {"cc0-1.0", "cc-by-4.0", "cc-by-3.0"})

With Polars, straight from the Parquet shards:

python
import polars as pl

df = pl.read_parquet("hf://datasets/histde/dta-documents/documents/train-*.parquet")
df.filter(pl.col("year") < 1700).select("id", "title", "license_family", "num_tokens")

Schema

ColumnTypeDescriptionExample
idstringDTA directory name, used as the primary document identifier (DTADirName). Falls back to the TCF file namegoethe_faust01_1808
dta_idstringNumeric DTA identifier (DTAID)200006
urnstringPersistent URN of the workurn:nbn:de:kobv:b4-200006-1
urlstringLanding page of the work at deutschestextarchiv.dehttps://www.deutschestextarchiv.de/343016
titlestringMain title(s) from the TEI title statement. Multiple main titles are joined with \nIphigenie auf Tauris
subtitlestringSubtitle(s), joined with \n. Null if absentEin Schauspiel
volumestringVolume title (text of <title type="volume">). Null if absentDritter Theil
partstringPart title (text of <title type="part">). Null if absentDie sentimentalischen Dichter
authorslist<struct>Authors as {surname, forename, gnd} structs. gnd is the GND/PND URI, may be null. Empty list for anonymous works[{"surname": "Reideburg", "forename": "Christoph von", "gnd": "http://d-nb.info/gnd/128862580"}]
date_publishedstringRaw publication date string of the original print, as given in the header1642
yearint64First four-digit number parsed from date_published. Null if none found1642
place_of_publicationstringPlace of publication of the original printBreslau
publisherstringPublisher name(s), joined with \nGeorgius Baumann
editionstringEdition statement (text, or the n attribute if the element has no text)2. Auflage
num_pagesint64Page count of the original source (measure type="pages")72
repositorystringHolding library of the digitized copyUniversitätsbibliothek Breslau
shelfmarkstringShelfmark of the copy at the repositoryUniversitätsbibliothek Breslau, 4 A 277/11 / 343016
typefacestringTypeface of the original printFraktur
biblstringShort bibliographic citation stringReideburg, Christoph von: Kurtze Anleitung: Wie die jetzige böse Zeit/ ... Breslau, 1642.
genre_dtamainstringMain genre, DTA schemeGebrauchsliteratur
genre_dtasubstringSub genre, DTA schemeLeichenpredigt
genre_dwds1mainstringMain genre, DWDS schemeBelletristik
genre_dwds1substringSub genre, DWDS schemeProsa
dta_corpus_labelslist<string>DTA (sub)corpus membership labels (DTACorpus class codes)["ready", "core"]
languagestringISO 639-3 code of the primary languagedeu
language_notestringHuman-readable language label from the header(Früh-)Neuhochdeutsch
licensestringPer-work license URL exactly as given in the header. Varies across the corpushttp://creativecommons.org/licenses/by-sa/3.0/de/
license_familystringNormalized license identifier derived from license (http/https, /de/, deed variants collapsed)cc-by-sa-3.0
num_imagesint64Number of page images (DTA extent measure)72
num_tokensint64Number of tokens (DTA extent measure)11515
num_typesint64Number of word types (DTA extent measure)3440
num_charactersint64Number of characters (DTA extent measure)78795
num_sentencesint64Sentence count from the full/ sentence layer. Null if the work has no full/ file842
textstringHistorical transcription. Layout-faithful <text> layer from simple/ (line breaks, long s ſ, combining diacritics), or space-joined tokens if only full/ existsKurtze Anleitung:\nWie die jetzige boͤſe Zeit/ darinnen zwar fuͤr ſich ſelbſt/ ...
text_sourcestringOrigin of text: layout (from simple/) or tokens (reconstructed from the full/ token layer)layout
text_normalizedstringModern orthography: full/ token layer with CAB replace corrections applied, space-joined. Null if the work has no full/ fileKurze Anleitung : Wie die jetzige böse Zeit / darinnen zwar für sich selbst / ...

Example row

json
{
  "id": "343016",
  "dta_id": "200006",
  "urn": "urn:nbn:de:kobv:b4-200006-1",
  "url": "https://www.deutschestextarchiv.de/343016",
  "title": "Kurtze Anleitung: Wie die jetzige böse Zeit/ darinnen zwar für sich selbst/ nichts/ alß eytel Klag/ Ach/ vnd Weh regieret",
  "authors": [{"surname": "Reideburg", "forename": "Christoph von", "gnd": "http://d-nb.info/gnd/128862580"}],
  "date_published": "1642",
  "year": 1642,
  "place_of_publication": "Breslau",
  "publisher": "Georgius Baumann",
  "repository": "Universitätsbibliothek Breslau",
  "typeface": "Fraktur",
  "genre_dtamain": "Gebrauchsliteratur",
  "genre_dtasub": "Leichenpredigt",
  "dta_corpus_labels": ["ready", "aedit"],
  "language": "deu",
  "language_note": "(Früh-)Neuhochdeutsch",
  "license": "http://creativecommons.org/licenses/by-sa/3.0/de/",
  "license_family": "cc-by-sa-3.0",
  "num_tokens": 11515,
  "num_sentences": 842,
  "text": "Kurtze Anleitung:\nWie die jetzige boͤſe Zeit/ darinnen zwar fuͤr ſich ſelbſt/\nnichts/ alß eytel Klag/ Ach/ vnd Weh regieret; ...",
  "text_source": "layout",
  "text_normalized": "Kurze Anleitung : Wie die jetzige böse Zeit / darinnen zwar für sich selbst / nichts / als eitel Klage / Ach / und Weh regieret ; ..."
}

Dataset creation

Source data

The DTA is a reference corpus of printed German from the late 15th to the early 20th century, curated by the Berlin-Brandenburg Academy of Sciences and Humanities (BBAW). Texts were transcribed by double keying from first editions wherever possible, and are published with TEI headers, sentence and token layers and orthographic normalization from the CAB tool chain.

This dataset was built from the complete TCF dump dta_komplett_2026-02-10, which comes in two variants:

  • simple/ (5,107 files): CMDI metadata and the layout-faithful historical text (<text> layer);
  • full/ (5,471 files): CMDI metadata and the annotation layers (tokens, sentences, orthography).

The union of both variants gives 5,481 works. One work (ford_pitty_1633) is excluded because its TCF file is not well-formed XML.

Processing

  • text comes from simple/ where available (text_source = "layout", 5,105 docs). For the 375 works only present in full/, it is reconstructed by space-joining the token layer (text_source = "tokens"), so line breaks and the original whitespace are lost there.
  • text_normalized is the full/ token layer with CAB replace corrections applied (multi-token corrections replace the whole span), space-joined. It is null for the 10 works only present in simple/.
  • Metadata is deliberately kept un-flattened: authors are a list of structs, all four genre schemes are separate columns, the raw license URL is kept next to a normalized license_family.
  • num_tokens, num_types, num_characters, num_images and num_pages are the extent measures from the DTA header, not recomputed values.

The conversion script and the statistics script are available in the dta-experiments repository.

Statistics

All tables below were produced with dta_dataset_stats.py from the released Parquet shards.

Overview

statisticvalue
Documents5,480
Documents with normalized text (full/ layer)5,470 (99.8%)
Documents with at least one author3,104 (56.6%)
Distinct authors (by GND id)1,317
Year range1472-1987 (1 unknown)
Median year1836
Tokens (header measure)204,582,701
Types (header measure)33,741,287
Characters (header measure)1,423,995,435
Characters in text column1,380,752,264
Characters in text_normalized column1,370,639,606
Sentences (full/ layer)10,376,423
Page images761,995
Median tokens per document10,394
Median pages (images) per document16

Temporal distribution (50-year bins)

perioddocssentencestokenscharacterstoken share
1450–1499151,899177,258997,3890.1%
1500–15491710,469308,4711,915,9650.2%
1550–159998139,0353,572,33623,618,8031.7%
1600–1649422433,7629,757,10066,033,8894.8%
1650–16996911,118,36123,338,594159,938,40011.4%
1700–17494351,143,29324,779,141171,116,20912.1%
1750–17995831,566,58230,933,941214,031,38215.1%
1800–18491,6182,000,91940,798,688284,536,21919.9%
1850–18991,1433,086,86855,792,116393,248,56027.3%
1900–1949444737,90813,114,62895,203,0446.4%
1950–199913137,1502,005,48813,321,2701.0%
unknown11774,94034,3050.0%
total5,48010,376,423204,582,7011,423,995,435100%

Text source

text_sourcedocsdocs %tokenstokens %
layout5,10593.2%185,164,38990.5%
tokens3756.8%19,418,3129.5%

Language

languagedocsdocs %tokenstokens %
deu5,47799.9%204,523,746100.0%
lat20.0%56,7140.0%
gml10.0%2,2410.0%

Genre (DWDS main)

genre_dwds1maindocsdocs %tokenstokens %
Zeitung2,03637.2%17,152,2848.4%
Gebrauchsliteratur1,71831.4%58,283,32028.5%
Wissenschaft95317.4%87,148,24742.6%
Belletristik77314.1%41,998,85020.5%

<details> <summary>Genre (DWDS sub, top 10)</summary>

genre_dwds1subdocsdocs %tokenstokens %
unknown2,03637.2%17,152,2848.4%
Leichenpredigt3366.1%3,517,2251.7%
Zeitschrift2174.0%1,697,7330.8%
Theologie2143.9%10,486,7465.1%
Brief2043.7%84,3250.0%
Roman1983.6%13,211,0366.5%
Gesellschaft1653.0%3,852,4131.9%
Novelle1262.3%3,031,2111.5%
Lyrik1252.3%5,090,9602.5%
Gelegenheitsschrift:Tod1162.1%103,2800.1%
other (123 values)1,74331.8%146,355,48871.5%

</details>

<details> <summary>Genre (DTA main)</summary>

genre_dtamaindocsdocs %tokenstokens %
unknown3,46863.3%69,372,31133.9%
Fachtext74313.6%78,106,23838.2%
Belletristik56110.2%33,626,23516.4%
Gebrauchsliteratur5369.8%22,981,50811.2%
Wissenschaftliche Abhandlungen in Form gedruckter Briefe500.9%89,5640.0%
Abhandlungen in Zeitschriften, Sammelbänden etc.360.7%145,3420.1%
Ankündigungen, Berichtigungen und kurze Nachrichten290.5%25,6870.0%
Berliner Akademiereden/-schriften und andere Reden170.3%71,0270.0%
Rezensionen70.1%8,5140.0%
Pariser Akademiereden/-schriften60.1%16,7280.0%
Vorworte und andere Beiträge Humboldts in Schriften anderer Autoren, Lexikonartikel60.1%20,5810.0%
Gelegenheitsschrift50.1%49,9060.0%
Journalismus50.1%54,8290.0%
Albumblätter30.1%8120.0%
Gutachten30.1%6,3390.0%
Sonderdrucke und andere Grenzfälle zu selbständig erschienenen Schriften20.0%3,3580.0%
Wissenschaft20.0%3,1240.0%
Sachliteratur10.0%5980.0%

</details>

<details> <summary>Genre (DTA sub, top 10)</summary>

genre_dtasubdocsdocs %tokenstokens %
unknown3,62866.2%69,896,61934.2%
Leichenpredigt3346.1%3,495,6301.7%
Roman1633.0%11,534,9055.6%
Prosa1282.3%10,291,8685.0%
Lyrik1142.1%4,731,1572.3%
Drama821.5%2,330,5991.1%
Recht791.4%9,743,7264.8%
Philosophie761.4%5,426,3032.7%
Historiographie460.8%7,221,7033.5%
Medizin460.8%6,013,8042.9%
other (86 values)78414.3%73,896,38736.1%

</details>

Typeface

typefacedocsdocs %tokenstokens %
Fraktur4,47881.7%164,623,05980.5%
Antiqua4277.8%35,856,37117.5%
Schwabacher3386.2%724,9340.4%
Handschrift2184.0%1,315,6900.6%
Jean-Paul-Fraktur90.2%1,409,6110.7%
Current20.0%1,9630.0%
Rotunda20.0%313,2250.2%
Antiqua (Schreibmaschine)10.0%34,1760.0%
Antiqua, Kursive10.0%18,8830.0%
Antqiua10.0%3,5270.0%
Frakur10.0%10,1030.0%
Textur10.0%270,5990.1%
unknown10.0%5600.0%

Values are taken verbatim from the headers, including the two typos (Antqiua, Frakur).

<details> <summary>DTA corpus labels (a document can carry several)</summary>

labeldocsdocs %tokenstokens %
ready5,480100.0%204,582,701100.0%
core1,47827.0%130,114,94863.6%
china1,14921.0%111,416,24254.5%
mkhz162511.4%3,996,6722.0%
nrhz5309.7%4,290,4832.1%
aedit3356.1%3,508,0981.7%
mts3205.8%18,006,9898.8%
lefevre3195.8%476,3900.2%
dtae2444.5%21,475,42010.5%
mkhz22224.1%2,169,3711.1%
correspondent2043.7%925,7550.5%
tevo2023.7%5,682,3952.8%
sanders-briefe1903.5%61,8900.0%
hab1843.4%10,256,9595.0%
augsburgerallgemeine1813.3%2,761,7781.3%
avh1723.1%443,6680.2%
wikisource1693.1%5,787,5432.8%
sbb_funeralschriften1122.0%77,9810.0%
tdef1092.0%940,7720.5%
novellenschatz871.6%1,643,3360.8%
blumenbach330.6%2,688,4121.3%
epoetics230.4%3,074,7301.5%
frauenstudium220.4%134,1980.1%
psyleko190.3%2,075,5001.0%
ntsm110.2%32,1390.0%
avhkv90.2%636,5390.3%
briefjeanpaul90.2%1,409,6110.7%
gutenberg_org80.1%325,9400.2%
gutzkow70.1%478,2460.2%
dwds150.1%465,6880.2%
gei20.0%57,4740.0%
gutenberg_de20.0%116,1840.1%
gwb20.0%80,6660.0%
urmel20.0%1,0680.0%
zbk20.0%48,1270.0%
greflinger10.0%9,0790.0%
grenzboten10.0%116,2080.1%

The core label marks the curated DTA core corpus (1,478 works, 63.6% of all tokens), which is the balanced selection by genre and period. The other labels denote extension corpora (DTAE) contributed by partner projects.

</details>

<details> <summary>Place of publication (top 10)</summary>

placedocsdocs %tokenstokens %
Berlin60411.0%30,616,94615.0%
Köln5329.7%4,682,9972.3%
Leipzig4257.8%41,950,38420.5%
Hamburg3306.0%5,514,4132.7%
Augsburg2564.7%4,581,3512.2%
Altstrelitz1723.1%57,7020.0%
unknown1602.9%980,9330.5%
München1532.8%3,434,6621.7%
Frankfurt (Main)1412.6%11,430,8735.6%
Königsberg1332.4%926,8430.5%
other (246 values)2,57447.0%100,405,59749.1%

</details>

Most frequent authors (top 10)

authorGNDdocstokens
Humboldt, Alexander vonhttp://d-nb.info/gnd/1185547001831,504,828
Sanders, Danielhttp://d-nb.info/gnd/11924204417369,412
Dach, Simonhttp://d-nb.info/gnd/11852321X11277,981
N. N.511,882,041
Blumenbach, Johann Friedrichhttp://d-nb.info/gnd/116208503352,814,553
Sattler, Basiliushttp://d-nb.info/gnd/11697478826352,365
Goethe, Johann Wolfgang vonhttp://d-nb.info/gnd/118540238241,064,748
Herder, Johann Gottfried vonhttp://d-nb.info/gnd/11854955323690,276
Grimm, Jacobhttp://d-nb.info/gnd/118542257172,441,309
Fontane, Theodorhttp://d-nb.info/gnd/118534262161,262,140

Licensing

The DTA publishes every work under its own license, and this dataset keeps that information per row in the license (raw URL from the header) and license_family (normalized) columns. There is no single license for the whole dataset. Note that only cc0-1.0 is a public domain dedication; all cc-by-* licenses carry binding conditions (attribution, share-alike, or non-commercial use), and about 28% of the documents are restricted to non-commercial use (cc-by-nc-3.0, cc-by-nc-sa-4.0, noc-nc-1.0, out-of-copyright-nc).

license familydocsdocs %tokenstokens %license URLs
cc-by-sa-4.02,64448.2%139,433,87868.2%http://creativecommons.org/licenses/by-sa/4.0/<br>http://creativecommons.org/licenses/by-sa/4.0/deed.de/<br>https://creativecommons.org/licenses/by-sa/4.0/<br>https://creativecommons.org/licenses/by-sa/4.0/deed.de
cc-by-nc-3.01,55928.4%20,418,25910.0%http://creativecommons.org/licenses/by-nc/3.0/<br>http://creativecommons.org/licenses/by-nc/3.0/de<br>http://creativecommons.org/licenses/by-nc/3.0/de/
cc-by-sa-3.062311.4%15,287,1297.5%http://creativecommons.org/licenses/by-sa/3.0/<br>http://creativecommons.org/licenses/by-sa/3.0/de<br>http://creativecommons.org/licenses/by-sa/3.0/de/<br>https://creativecommons.org/licenses/by-sa/3.0/<br>https://creativecommons.org/licenses/by-sa/3.0/de/
cc0-1.03706.8%16,938,5958.3%http://creativecommons.org/publicdomain/zero/1.0/
cc-by-4.02153.9%8,189,5394.0%http://creativecommons.org/licenses/by/4.0/<br>http://creativecommons.org/licenses/by/4.0/deed.de<br>https://creativecommons.org/licenses/by/4.0/<br>https://creativecommons.org/licenses/by/4.0/deed.de
cc-by-sa-2.0510.9%3,058,0541.5%http://creativecommons.org/licenses/by-sa/2.0/de<br>http://creativecommons.org/licenses/by-sa/2.0/de/
gutenberg90.2%344,3160.2%https://web.archive.org/web/20180927123034/http://www.gutenberg.org/wiki/Gutenberg:TheProjectGutenberg_License
cc-by-3.030.1%571,3460.3%http://creativecommons.org/licenses/by/3.0/de/
cc-by-nc-sa-4.020.0%19,6720.0%https://creativecommons.org/licenses/by-nc-sa/4.0/<br>https://creativecommons.org/licenses/by-nc-sa/4.0/deed.de
mdz-copyright10.0%24,9640.0%http://mdz.bib-bvb.de/copyright.htm
noc-nc-1.010.0%125,8590.1%http://rightsstatements.org/vocab/NoC-NC/1.0/
nug-kkn10.0%55,7020.0%http://www.deutsche-digitale-bibliothek.de/lizenzen/nug-kkn/
out-of-copyright-nc10.0%115,3880.1%https://www.europeana.eu/portal/de/rights/out-of-copyright-non-commercial.html
total5,480100%204,582,701100%

To select a subset that fits your use case:

python
ds_sa = ds.filter(lambda d: d["license_family"] in {"cc0-1.0", "cc-by-4.0", "cc-by-3.0", "cc-by-sa-4.0", "cc-by-sa-3.0", "cc-by-sa-2.0"})

Considerations for using the data

  • Coverage is uneven. Roughly 60% of the tokens date from 1750 to 1899; the 15th and 16th centuries are thinly covered. Newspapers (Zeitung) dominate the document count but make up only 8% of the tokens, and a single project (Alexander von Humboldt's writings) contributes 183 short documents.
  • Historical content. The texts reflect the views, language and stereotypes of their time, including religious, political and colonial-era writing. They are provided for research on historical German and should not be taken as statements of fact or acceptable opinion.
  • Normalization is automatic. text_normalized comes from the CAB tool chain and is not manually corrected. Expect residual errors, especially for very early texts, Latin passages, and proper names. The normalized text is also whitespace-tokenized (punctuation is separated by spaces), whereas text keeps the original layout.
  • `text` is not uniform. 375 documents have their historical text reconstructed from the token layer and lack line breaks; check text_source if layout matters to you.
  • Metadata quirks. Genre fields are null for a large share of the works in the DTA scheme (use the DWDS scheme for complete main-genre coverage), typeface strings contain a few typos, and year is parsed heuristically from the free-text date_published.

Citation

Please cite the DTA when using this dataset:

bibtex
@misc{dta2026,
  author       = {{Berlin-Brandenburgische Akademie der Wissenschaften}},
  title        = {Deutsches Textarchiv. Grundlage für ein Referenzkorpus der neuhochdeutschen Sprache},
  year         = {2026},
  address      = {Berlin},
  howpublished = {Herausgegeben von der Berlin-Brandenburgischen Akademie der Wissenschaften},
  url          = {https://www.deutschestextarchiv.de/}
}

Acknowledgements

A big thank you to the whole team of the Deutsches Textarchiv at the Berlin-Brandenburg Academy of Sciences and Humanities! Also many thanks to all the partner projects, libraries and archives that contributed texts to the DTA extension corpora.

This dataset only converts an official dump into Parquet for convenient use with the Hugging Face ecosystem.

AI disclosure

This dataset card was drafted with Claude Fable 5 (claude-fable-5) based on the statistics produced by dta_dataset_stats.py, and reviewed by the dataset author. The conversion and statistics scripts carry their own AI disclosure blocks, following the rules in the ai-disclosure repository.