datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
x-span-similarity
X-Span-Similarity (X-SSD)
Expanding Lozano et al.'s (2026) Span Similarity Dataset (SSD) into a cross-lingual setting, for Dissimilar/Difference Span Detection (DSD) across language pairs.
Dataset summary
Each row is a premise/hypothesis sentence pair, one side in English and the other machine-translated into another target language, with span-level and sentence-level (dis)similarity labels carried over unchanged from the original English SSD annotation. Spans… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/x-span-similarity.quotidiana
La Quotidiana news articles in Romansh
This dataset has threee subsets:
1997_2008: Articles published by La Quotidiana between 1997 and 2008. The articles have been retrieved from https://github.com/ProSvizraRumantscha/corpora and extracted from the original XML format.
2021_2025: Articles published by La Quotidiana between 2021 and 2025. These articles have been extracted from WXR WordPress export files.
2026_: Articles published by La Quotidiana from 2026 onwards. These… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/quotidiana.document-level-word-alignment
Document-Level Word Alignment
Document-level word alignment data for six language pairs — English–French (en-fr), English–Romanian (en-ro), English–Japanese (en-ja), English–Chinese (en-zh), Latin–Ancient Greek (la-gr), and English–Czech (en-cz) — reconstructed from existing sentence-level, human-annotated word alignment gold standards.
Document-level examples are built by grouping sentence-level annotations by document membership and sentence order; the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/document-level-word-alignment.mediomatix-raw
General Information
Release of the Mediomatix corpus, prepared by UZH and PHGR, to be used for research purposes only.
In mediomatix-raw, we release the full Mediomatix schoolbooks' text for each idiom. The data points are unaligned. See here for the mulit-parallel, aligned Mediomatix corpus.
We use the following Romansh idiom codes as subsets in the dataset:
Sursilvan: rm-sursilv
Sutsilvan: rm-sutsilv
Surmiran: rm-surmiran
Puter: rm-puter
Vallader: rm-vallader
The splits in… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/mediomatix-raw.mediomatix
General Information
Release of the Mediomatix corpus, prepared by UZH and PHGR, to be used for research purposes only.
Each segment is presented as part of a multi-parallel alignment. For the full, unaligned Mediomatix data in each idiom's schoolbooks, see here .
We use the following Romansh idiom codes as columns:
Sursilvan: rm-sursilv
Sutsilvan: rm-sutsilv
Surmiran: rm-surmiran
Puter: rm-puter
Vallader: rm-vallader
Book names are encoded in the book column as follows:
The… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/mediomatix.romansh-mt-evaluation
Dataset Description
This dataset contains the results of a human evaluation of machine translations from German into the six Romansh varieties. The evaluations were carried out by native speakers of the respective Romansh idioms as well as professional linguists.
The evaluation covers three quality dimensions:
Document accuracy, in which annotators assessed the adequacy of complete document translations.
Segment accuracy, in which annotators selected the more accurate… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/romansh-mt-evaluation.romansh-backtranslated
Romansh–German Back-Translation Dataset
Background
This dataset contains Romansh texts paired with German translations generated synthetically using Gemini 2.5 Flash.
It was created as part of research on data augmentation for low-resource machine translation of Romansh, a language with 6 distinct written varieties (Rumantsch Grischun, Sursilvan, Sutsilvan, Surmiran, Puter, Vallader).
LLMs tend to confuse Romansh varieties when translating into Romansh, but… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/romansh-backtranslated.SwissGov-RSD
SwissGov-RSD
Dataset Description
SwissGov-RSD is a naturalistic, human-annotated, document-level, cross-lingual dataset for token-level semantic difference recognition (RSD). It contains 224 multi-parallel Swiss government documents from admin.ch in English–German, English–French, and English–Italian, annotated with fine-grained semantic difference labels (0–1) at the token level.
The dataset targets real-world scenarios where cross-lingual content diverges due… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/SwissGov-RSD.paws-x-italian
PAWS-X Italian Paraphrase Dataset
This dataset is a machine-translated Italian version of the English PAWS-X dataset. The original PAWS-X dataset (Yang et al. 2019) is a multilingual version of PAWS (Zhang et al. 2019) for paraphrase identification.
Dataset Structure
Data Fields
sentence1: First sentence in the pair
sentence2: Second sentence in the pair
labels:
0: Non-paraphrases
1: Paraphrases
Data Splits
The dataset is split into:
Training… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/paws-x-italian.romansh_theater_plays
Theater Plays in Romansh
This dataset contains various theater plays in different Romansh varieties: Rumantsch Grischun (rm-rumgr), Sursilvan (rm-sursilv), Sutsilvan (rm-sutsilv), Surmiran (rm-surmiran), Puter (rm-puter), and Vallader (rm-vallader). They have been extracted from PDFs provided by Lia Rumantscha. The variety assigned is based on manual processing by Lia Rumantscha.
Further, for each theater play, the text has been extracted per page. Thus, one dataset entry contains… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/romansh_theater_plays.
