CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ZurichNLP /mlit-guanaco Description Guanaco dataset subsets used for experiments in the paper Turning English-centric LLMs Into Polyglots: How Much Multilinguality Is Needed? We extend the original Guanaco dataset with language tags, with languages identified using OpenLID. The following subsets were used to train our experimental models: config name languages ml1 en ml2, mtml2 en, es ml3, mtml3 en, es, ru ml4, mtml4 en, es, ru, de ml5, mtml5 en, es, ru, de, zh ml6, mtml6 en, es, ru… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/mlit-guanaco.tabular10K<n<100K2 likes452 downloads3y agoHugging Face02ZurichNLP /x-span-similarity X-Span-Similarity (X-SSD) Expanding Lozano et al.'s (2026) Span Similarity Dataset (SSD) into a cross-lingual setting, for Dissimilar/Difference Span Detection (DSD) across language pairs. Dataset summary Each row is a premise/hypothesis sentence pair, one side in English and the other machine-translated into another target language, with span-level and sentence-level (dis)similarity labels carried over unchanged from the original English SSD annotation. Spans… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/x-span-similarity.texttext-classification100K<n<1M1 likes366 downloads2mo agoHugging Face03ZurichNLP /quotidiana La Quotidiana news articles in Romansh This dataset has threee subsets: 1997_2008: Articles published by La Quotidiana between 1997 and 2008. The articles have been retrieved from https://github.com/ProSvizraRumantscha/corpora and extracted from the original XML format. 2021_2025: Articles published by La Quotidiana between 2021 and 2025. These articles have been extracted from WXR WordPress export files. 2026_: Articles published by La Quotidiana from 2026 onwards. These… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/quotidiana.text100K<n<1M0 likes207 downloads25d agoHugging Face04ZurichNLP /rsd-ists-2016Training and test data for the task of Recognizing Semantic Differences (RSD). See the paper for details on how the dataset was created, and see our code at https://github.com/ZurichNLP/recognizing-semantic-differences for an example of how to use the data for evaluation. The data are derived from the SemEval-2016 Task 2 for Interpretable Semantic Textual Similarity organized by Agirre et al. (2016). The original URLs of the data are: Train:… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/rsd-ists-2016.texttoken-classification10K<n<100K0 likes190 downloads1y agoHugging Face05ZurichNLP /document-level-word-alignment Document-Level Word Alignment Document-level word alignment data for six language pairs — English–French (en-fr), English–Romanian (en-ro), English–Japanese (en-ja), English–Chinese (en-zh), Latin–Ancient Greek (la-gr), and English–Czech (en-cz) — reconstructed from existing sentence-level, human-annotated word alignment gold standards. Document-level examples are built by grouping sentence-level annotations by document membership and sentence order; the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/document-level-word-alignment.texttoken-classificationn<1K0 likes150 downloads1mo agoHugging Face06ZurichNLP /mlit-alpaca-eval Description Translated versions of the AlpacaEval prompt dataset for evaluating the performance of chat LLMs. Translations were generated using gpt-3.5-turbo-0613 using the following prompt template (adapted from Lai et al, 2023): You are a helpful assistant. Translate the following text into {{target_language}}. Keep the structure of the original text and preserve things like code and names. Please ensure that your response contains only the translated text. The translation must… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/mlit-alpaca-eval.text10K<n<100K1 likes133 downloads3y agoHugging Face07ZurichNLP /wmt24pp-rm WMT24++ Reference Translations for Romansh Description WMT24++ benchmark in Romansh (six varieties: Rumantsch Grischun, Sursilvan, Sutsilvan, Surmiran, Puter, and Vallader). Paper: "Expanding the WMT24++ Benchmark with Rumantsch Grischun, Sursilvan, Sutsilvan, Surmiran, Puter, and Vallader" Code: https://github.com/ZurichNLP/romansh_mt_eval Original WMT24++ benchmark (55 languages): https://huggingface.co/datasets/google/wmt24pp The reference translation have been… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/wmt24pp-rm.texttranslation1K<n<10K3 likes132 downloads10mo agoHugging Face08ZurichNLP /mediomatix-raw General Information Release of the Mediomatix corpus, prepared by UZH and PHGR, to be used for research purposes only. In mediomatix-raw, we release the full Mediomatix schoolbooks' text for each idiom. The data points are unaligned. See here for the mulit-parallel, aligned Mediomatix corpus. We use the following Romansh idiom codes as subsets in the dataset: Sursilvan: rm-sursilv Sutsilvan: rm-sutsilv Surmiran: rm-surmiran Puter: rm-puter Vallader: rm-vallader The splits in… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/mediomatix-raw.text100K<n<1M0 likes124 downloads6mo agoHugging Face09ZurichNLP /mediomatix General Information Release of the Mediomatix corpus, prepared by UZH and PHGR, to be used for research purposes only. Each segment is presented as part of a multi-parallel alignment. For the full, unaligned Mediomatix data in each idiom's schoolbooks, see here . We use the following Romansh idiom codes as columns: Sursilvan: rm-sursilv Sutsilvan: rm-sutsilv Surmiran: rm-surmiran Puter: rm-puter Vallader: rm-vallader Book names are encoded in the book column as follows: The… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/mediomatix.text10K<n<100K0 likes120 downloads6mo agoHugging Face10ZurichNLP /romansh-mt-evaluation Dataset Description This dataset contains the results of a human evaluation of machine translations from German into the six Romansh varieties. The evaluations were carried out by native speakers of the respective Romansh idioms as well as professional linguists. The evaluation covers three quality dimensions: Document accuracy, in which annotators assessed the adequacy of complete document translations. Segment accuracy, in which annotators selected the more accurate… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/romansh-mt-evaluation.tabular1K<n<10K0 likes106 downloads3mo agoHugging Face11ZurichNLP /romansh-backtranslated Romansh–German Back-Translation Dataset Background This dataset contains Romansh texts paired with German translations generated synthetically using Gemini 2.5 Flash. It was created as part of research on data augmentation for low-resource machine translation of Romansh, a language with 6 distinct written varieties (Rumantsch Grischun, Sursilvan, Sutsilvan, Surmiran, Puter, Vallader). LLMs tend to confuse Romansh varieties when translating into Romansh, but… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/romansh-backtranslated.text100K<n<1M0 likes87 downloads3mo agoHugging Face12ZurichNLP /SwissGov-RSD SwissGov-RSD Dataset Description SwissGov-RSD is a naturalistic, human-annotated, document-level, cross-lingual dataset for token-level semantic difference recognition (RSD). It contains 224 multi-parallel Swiss government documents from admin.ch in English–German, English–French, and English–Italian, annotated with fine-grained semantic difference labels (0–1) at the token level. The dataset targets real-world scenarios where cross-lingual content diverges due… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/SwissGov-RSD.tabularn<1K2 likes76 downloads2mo agoHugging Face13ZurichNLP /swissner SwissNER A multilingual test set for named entity recognition (NER) on Swiss news articles. Description SwissNER is a dataset for named entity recognition based on manually annotated news articles in Swiss Standard German, French, Italian, and Romansh Grischun. We have manually annotated a selection of articles that have been published in February 2023 in the categories "Switzerland" or "Regional" on the following online news portals: Swiss Standard German: srf.ch… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/swissner.texttoken-classificationn<1K2 likes51 downloads3y agoHugging Face14ZurichNLP /romansh-canton-laws Grisons canton laws This dataset contains cantonal laws of Grisons, as extracted from https://www.gr-lex.gr.ch/. It further contains parallel text in the languages German (DE), Rumantsch Grischun (RM), and Italian (IT), including the source links to the HTML and PDF. The dataset itself is under the public domain. text10K<n<100K0 likes35 downloads10mo agoHugging Face15ZurichNLP /20min-XD 20min-XD: A Comparable Corpus of Swiss News Articles Dataset Summary 20min-XD (20 Minuten cross-lingual document-level) is a comparable corpus of Swiss news articles in German and French, collected from the online editions of 20 Minuten and 20 minutes between 2015 and 2024. The dataset consists of 15,000 semantically aligned German and French article pairs. Unlike parallel corpora, 20min-XD captures a broad spectrum of cross-lingual similarity, ranging from… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/20min-XD.tabular100K<n<1M0 likes21 downloads8mo agoHugging Face16ZurichNLP /paws-x-italian PAWS-X Italian Paraphrase Dataset This dataset is a machine-translated Italian version of the English PAWS-X dataset. The original PAWS-X dataset (Yang et al. 2019) is a multilingual version of PAWS (Zhang et al. 2019) for paraphrase identification. Dataset Structure Data Fields sentence1: First sentence in the pair sentence2: Second sentence in the pair labels: 0: Non-paraphrases 1: Paraphrases Data Splits The dataset is split into: Training… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/paws-x-italian.text10K<n<100K0 likes21 downloads10mo agoHugging Face17ZurichNLP /romansh_theater_plays Theater Plays in Romansh This dataset contains various theater plays in different Romansh varieties: Rumantsch Grischun (rm-rumgr), Sursilvan (rm-sursilv), Sutsilvan (rm-sutsilv), Surmiran (rm-surmiran), Puter (rm-puter), and Vallader (rm-vallader). They have been extracted from PDFs provided by Lia Rumantscha. The variety assigned is based on manual processing by Lia Rumantscha. Further, for each theater play, the text has been extracted per page. Thus, one dataset entry contains… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/romansh_theater_plays.text1K<n<10K0 likes21 downloads10mo agoHugging Face18adityarra07 /zurich_data Dataset Card for "zurich_data" More Information needed audio1K<n<10K0 likes11 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.