CoolFace
20 results

impresso

impresso-project /ner-augmentation Impresso HIPE-2022 NER dataset Named-entity-recognition data derived from the HIPE-2022 shared task hipe2020 bundle(s), prepared for the Impresso NER models. Train/dev are taken from revision v2.1; the test split from v2.1-test-all-unmasked (the post-evaluation release of the withheld test gold labels). Languages: de, en, fr Bundle(s): hipe2020 Config (variant): hipe2020 Label scheme: IOB2 over the NE-COARSE-LIT column. The loadable parquet config carries tokens / ner_tags (as… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/ner-augmentation.texttoken-classification100K<n<1M0 likes448 downloads12d agoHugging Faceimpresso-project /ner-eval-predictionstabular100K<n<1M0 likes249 downloads2mo agoHugging Faceimpresso-project /frakturline-testset Fraktur/Other Text-Line — Test Set A balanced, held-out evaluation set of 2 000 scanned text-line images (1 000 per class) for the binary task of distinguishing Fraktur (blackletter / Gothic script) from other script (primarily Antiqua / Latin / Roman). Developed for the Impresso digital humanities project. Dataset Details Property Value Task Binary image classification Classes fraktur, other Images per class 1 000 Total images 2 000 Image format WebP… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/frakturline-testset.imageimage-classification1K<n<10K1 likes139 downloads6mo agoHugging Faceimpresso-project /nel-mgenre-trie mGENRE title trie + Wikidata QID lookup — impresso NEL assets Three marisa-trie native binaries that support multilingual entity linking with mGENRE. They are the runtime assets for the impresso-project/nel-mgenre-multilingual model as run by the impresso-inference harness: a title prefix tree that constrains beam search to valid Wikipedia titles, a (language, title) → Wikidata QID lookup for offline QID resolution, and a QID → (class, birthdate) attribute table for offline… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/nel-mgenre-trie.text-retrieval10M<n<100M0 likes130 downloads1mo agoHugging Faceimpresso-project /impresso-mediaagencies-ner-dataset Impresso Media Sources Dataset Curated token-classification data for news-agency and radio-station mentions in Impresso historical newspaper text. The v0.1 data is derived from the legacy French/German HIPE-style news-agency annotations, converted to JSONL, manually reviewed against the current model's dev/test disagreements, and updated according to annotation guidelines v2.0. The current guidelines annotate every explicit canonical media-source organization mention, not only… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/impresso-mediaagencies-ner-dataset.texttoken-classification1K<n<10K1 likes69 downloads27d agoHugging Faceimpresso-project /wiki_comparable_corpus_en_de_hi_it_ko_zh Multilingual Wikipedia Comparable Corpus (en, de, it, ko, hi, zh) This dataset is a document-level comparable corpus of Wikipedia articles across 6 languages: English (en), German (de), Italian (it), Korean (ko), Hindi (hi), and Chinese (zh). The key property is alignment across languages: entries are topic-matched such that, for a given index i, dataset["en"][i] is comparable to dataset["de"][i], dataset["it"][i], … (and likewise via the aligned_id field). Dataset… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/wiki_comparable_corpus_en_de_hi_it_ko_zh.tabular10K<n<100K1 likes32 downloads8mo agoHugging Face