impresso
Datasets
All datasets matching “impresso”ner-augmentation
Impresso HIPE-2022 NER dataset
Named-entity-recognition data derived from the
HIPE-2022 shared task hipe2020 bundle(s), prepared
for the Impresso NER models. Train/dev are taken from revision
v2.1; the test split from v2.1-test-all-unmasked (the
post-evaluation release of the withheld test gold labels).
Languages: de, en, fr
Bundle(s): hipe2020
Config (variant): hipe2020
Label scheme: IOB2 over the NE-COARSE-LIT column.
The loadable parquet config carries tokens / ner_tags (as… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/ner-augmentation.ner-eval-predictionsfrakturline-testset
Fraktur/Other Text-Line — Test Set
A balanced, held-out evaluation set of 2 000 scanned text-line images (1 000 per class) for the binary task of distinguishing Fraktur (blackletter / Gothic script) from other script (primarily Antiqua / Latin / Roman).
Developed for the Impresso digital humanities project.
Dataset Details
Property
Value
Task
Binary image classification
Classes
fraktur, other
Images per class
1 000
Total images
2 000
Image format
WebP… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/frakturline-testset.nel-mgenre-trie
mGENRE title trie + Wikidata QID lookup — impresso NEL assets
Three marisa-trie native binaries that
support multilingual entity linking with mGENRE. They are the runtime assets
for the impresso-project/nel-mgenre-multilingual
model as run by the impresso-inference
harness:
a title prefix tree that constrains beam search to valid Wikipedia titles,
a (language, title) → Wikidata QID lookup for offline QID resolution, and
a QID → (class, birthdate) attribute table for offline… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/nel-mgenre-trie.impresso-mediaagencies-ner-dataset
Impresso Media Sources Dataset
Curated token-classification data for news-agency and radio-station mentions in Impresso historical newspaper text.
The v0.1 data is derived from the legacy French/German HIPE-style news-agency annotations, converted to JSONL, manually reviewed against the current model's dev/test disagreements, and updated according to annotation guidelines v2.0. The current guidelines annotate every explicit canonical media-source organization mention, not only… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/impresso-mediaagencies-ner-dataset.wiki_comparable_corpus_en_de_hi_it_ko_zh
Multilingual Wikipedia Comparable Corpus (en, de, it, ko, hi, zh)
This dataset is a document-level comparable corpus of Wikipedia articles across 6 languages: English (en), German (de), Italian (it), Korean (ko), Hindi (hi), and Chinese (zh).
The key property is alignment across languages: entries are topic-matched such that, for a given index i, dataset["en"][i] is comparable to dataset["de"][i], dataset["it"][i], … (and likewise via the aligned_id field).
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/wiki_comparable_corpus_en_de_hi_it_ko_zh.
