datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
romansh-mt-evaluation
Dataset Description
This dataset contains the results of a human evaluation of machine translations from German into the six Romansh varieties. The evaluations were carried out by native speakers of the respective Romansh idioms as well as professional linguists.
The evaluation covers three quality dimensions:
Document accuracy, in which annotators assessed the adequacy of complete document translations.
Segment accuracy, in which annotators selected the more accurate… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/romansh-mt-evaluation.SwissGov-RSD
SwissGov-RSD
Dataset Description
SwissGov-RSD is a naturalistic, human-annotated, document-level, cross-lingual dataset for token-level semantic difference recognition (RSD). It contains 224 multi-parallel Swiss government documents from admin.ch in English–German, English–French, and English–Italian, annotated with fine-grained semantic difference labels (0–1) at the token level.
The dataset targets real-world scenarios where cross-lingual content diverges due… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/SwissGov-RSD.
