zurich
Datasets
All datasets matching “zurich”x_stanceThe x-stance dataset contains more than 150 political questions, and 67k comments written by candidates on those questions.
It can be used to train and evaluate stance detection systems.mlit-guanaco
Description
Guanaco dataset subsets used for experiments in the paper Turning English-centric LLMs Into Polyglots: How Much Multilinguality Is Needed?
We extend the original Guanaco dataset with language tags, with languages identified using OpenLID.
The following subsets were used to train our experimental models:
config name
languages
ml1
en
ml2, mtml2
en, es
ml3, mtml3
en, es, ru
ml4, mtml4
en, es, ru, de
ml5, mtml5
en, es, ru, de, zh
ml6, mtml6
en, es, ru… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/mlit-guanaco.x-span-similarity
X-Span-Similarity (X-SSD)
Expanding Lozano et al.'s (2026) Span Similarity Dataset (SSD) into a cross-lingual setting, for Dissimilar/Difference Span Detection (DSD) across language pairs.
Dataset summary
Each row is a premise/hypothesis sentence pair, one side in English and the other machine-translated into another target language, with span-level and sentence-level (dis)similarity labels carried over unchanged from the original English SSD annotation. Spans… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/x-span-similarity.quotidiana
La Quotidiana news articles in Romansh
This dataset has threee subsets:
1997_2008: Articles published by La Quotidiana between 1997 and 2008. The articles have been retrieved from https://github.com/ProSvizraRumantscha/corpora and extracted from the original XML format.
2021_2025: Articles published by La Quotidiana between 2021 and 2025. These articles have been extracted from WXR WordPress export files.
2026_: Articles published by La Quotidiana from 2026 onwards. These… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/quotidiana.rsd-ists-2016Training and test data for the task of Recognizing Semantic Differences (RSD).
See the paper for details on how the dataset was created, and see our code at https://github.com/ZurichNLP/recognizing-semantic-differences for an example of how to use the data for evaluation.
The data are derived from the SemEval-2016 Task 2 for Interpretable Semantic Textual Similarity organized by Agirre et al. (2016).
The original URLs of the data are:
Train:… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/rsd-ists-2016.ogd4all-zurich-ogd
