CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ZurichNLP /x-span-similarity X-Span-Similarity (X-SSD) Expanding Lozano et al.'s (2026) Span Similarity Dataset (SSD) into a cross-lingual setting, for Dissimilar/Difference Span Detection (DSD) across language pairs. Dataset summary Each row is a premise/hypothesis sentence pair, one side in English and the other machine-translated into another target language, with span-level and sentence-level (dis)similarity labels carried over unchanged from the original English SSD annotation. Spans… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/x-span-similarity.texttext-classification100K<n<1M1 likes366 downloads2mo agoHugging Face02ZurichNLP /quotidiana La Quotidiana news articles in Romansh This dataset has threee subsets: 1997_2008: Articles published by La Quotidiana between 1997 and 2008. The articles have been retrieved from https://github.com/ProSvizraRumantscha/corpora and extracted from the original XML format. 2021_2025: Articles published by La Quotidiana between 2021 and 2025. These articles have been extracted from WXR WordPress export files. 2026_: Articles published by La Quotidiana from 2026 onwards. These… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/quotidiana.text100K<n<1M0 likes207 downloads25d agoHugging Face03ZurichNLP /document-level-word-alignment Document-Level Word Alignment Document-level word alignment data for six language pairs — English–French (en-fr), English–Romanian (en-ro), English–Japanese (en-ja), English–Chinese (en-zh), Latin–Ancient Greek (la-gr), and English–Czech (en-cz) — reconstructed from existing sentence-level, human-annotated word alignment gold standards. Document-level examples are built by grouping sentence-level annotations by document membership and sentence order; the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/document-level-word-alignment.texttoken-classificationn<1K0 likes150 downloads1mo agoHugging Face04ZurichNLP /mediomatix-raw General Information Release of the Mediomatix corpus, prepared by UZH and PHGR, to be used for research purposes only. In mediomatix-raw, we release the full Mediomatix schoolbooks' text for each idiom. The data points are unaligned. See here for the mulit-parallel, aligned Mediomatix corpus. We use the following Romansh idiom codes as subsets in the dataset: Sursilvan: rm-sursilv Sutsilvan: rm-sutsilv Surmiran: rm-surmiran Puter: rm-puter Vallader: rm-vallader The splits in… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/mediomatix-raw.text100K<n<1M0 likes124 downloads6mo agoHugging Face05ZurichNLP /mediomatix General Information Release of the Mediomatix corpus, prepared by UZH and PHGR, to be used for research purposes only. Each segment is presented as part of a multi-parallel alignment. For the full, unaligned Mediomatix data in each idiom's schoolbooks, see here . We use the following Romansh idiom codes as columns: Sursilvan: rm-sursilv Sutsilvan: rm-sutsilv Surmiran: rm-surmiran Puter: rm-puter Vallader: rm-vallader Book names are encoded in the book column as follows: The… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/mediomatix.text10K<n<100K0 likes120 downloads6mo agoHugging Face06ZurichNLP /romansh-mt-evaluation Dataset Description This dataset contains the results of a human evaluation of machine translations from German into the six Romansh varieties. The evaluations were carried out by native speakers of the respective Romansh idioms as well as professional linguists. The evaluation covers three quality dimensions: Document accuracy, in which annotators assessed the adequacy of complete document translations. Segment accuracy, in which annotators selected the more accurate… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/romansh-mt-evaluation.tabular1K<n<10K0 likes106 downloads3mo agoHugging Face07ZurichNLP /romansh-backtranslated Romansh–German Back-Translation Dataset Background This dataset contains Romansh texts paired with German translations generated synthetically using Gemini 2.5 Flash. It was created as part of research on data augmentation for low-resource machine translation of Romansh, a language with 6 distinct written varieties (Rumantsch Grischun, Sursilvan, Sutsilvan, Surmiran, Puter, Vallader). LLMs tend to confuse Romansh varieties when translating into Romansh, but… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/romansh-backtranslated.text100K<n<1M0 likes87 downloads3mo agoHugging Face08ZurichNLP /SwissGov-RSD SwissGov-RSD Dataset Description SwissGov-RSD is a naturalistic, human-annotated, document-level, cross-lingual dataset for token-level semantic difference recognition (RSD). It contains 224 multi-parallel Swiss government documents from admin.ch in English–German, English–French, and English–Italian, annotated with fine-grained semantic difference labels (0–1) at the token level. The dataset targets real-world scenarios where cross-lingual content diverges due… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/SwissGov-RSD.tabularn<1K2 likes76 downloads2mo agoHugging Face09ZurichNLP /paws-x-italian PAWS-X Italian Paraphrase Dataset This dataset is a machine-translated Italian version of the English PAWS-X dataset. The original PAWS-X dataset (Yang et al. 2019) is a multilingual version of PAWS (Zhang et al. 2019) for paraphrase identification. Dataset Structure Data Fields sentence1: First sentence in the pair sentence2: Second sentence in the pair labels: 0: Non-paraphrases 1: Paraphrases Data Splits The dataset is split into: Training… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/paws-x-italian.text10K<n<100K0 likes21 downloads10mo agoHugging Face10ZurichNLP /romansh_theater_plays Theater Plays in Romansh This dataset contains various theater plays in different Romansh varieties: Rumantsch Grischun (rm-rumgr), Sursilvan (rm-sursilv), Sutsilvan (rm-sutsilv), Surmiran (rm-surmiran), Puter (rm-puter), and Vallader (rm-vallader). They have been extracted from PDFs provided by Lia Rumantscha. The variety assigned is based on manual processing by Lia Rumantscha. Further, for each theater play, the text has been extracted per page. Thus, one dataset entry contains… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/romansh_theater_plays.text1K<n<10K0 likes21 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.