romansh
romansh-nllb-1.3b-ct2wav2vec2-large-xls-r-300m-romansh-valladerwav2vec2-large-xls-r-300m-romansh-sursilvanwav2vec2-xlsr-romansh_sursilvanwav2vec2-xlsr-romansh_valladerromansh-nllb-1.3b-ablation-lr-hr-augmentation-ct2romansh-nllb-1.3b-ablation-no-augmentation-ct2b24-wav2vec2-large-xls-r-romansh-colab
Datasets
All datasets matching “romansh”Romansh_German_Parallel_Data
Romansh–German Parallel Dataset (FineWeb-Based)
This dataset contains automatically aligned Romansh–German document pairs, extracted from the Fineweb2 using cosine similarity over OpenAI embeddings. It was created as part of a university programming project focused on document-level parallel data extraction.
Description
This project performs document-level alignment between Romansh and German web texts, which were extracted from the Fineweb2 dataset. It uses OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/Sudehsna/Romansh_German_Parallel_Data.apertus-pretrain-romanshThis dataset consist of three differnt parts. Monolingual Romansh Data, Polylingual data or more precisely translated data from Romansh into either German, French, Italian or English and Sythetic Data.
The Polylingual data consists of aligned and non aligned data. The synthetic data was created by interweaving the translational data and prefacing it with the sentence " This is a text translated from SOURCE LANGUAGE to Rumantsch Grischun".
The data has a metadata "idiom" if the if specific… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/apertus-pretrain-romansh.romansh-mt-evaluation
Dataset Description
This dataset contains the results of a human evaluation of machine translations from German into the six Romansh varieties. The evaluations were carried out by native speakers of the respective Romansh idioms as well as professional linguists.
The evaluation covers three quality dimensions:
Document accuracy, in which annotators assessed the adequacy of complete document translations.
Segment accuracy, in which annotators selected the more accurate… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/romansh-mt-evaluation.romansh-backtranslated
Romansh–German Back-Translation Dataset
Background
This dataset contains Romansh texts paired with German translations generated synthetically using Gemini 2.5 Flash.
It was created as part of research on data augmentation for low-resource machine translation of Romansh, a language with 6 distinct written varieties (Rumantsch Grischun, Sursilvan, Sutsilvan, Surmiran, Puter, Vallader).
LLMs tend to confuse Romansh varieties when translating into Romansh, but… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/romansh-backtranslated.romansh-grischun-morphological-corpus
Romansh Grischun Morphological Corpus
A morphologically annotated corpus of Rumantsch Grischun, the standardized written variety of Romansh.
Dataset Summary
This dataset provides gold-standard morphosyntactically annotated and lemmatized corpora for Rumantsch Grischun (standard written Romansh, ISO 639-3: roh), a national language of Switzerland.
The primary annotation layer preserves the rich Xerox/Foma two-level morphological tags used by the finite-state… See the full description on the dataset page: https://huggingface.co/datasets/simon-clmtd/romansh-grischun-morphological-corpus.MNIST-ResNet-Demo-Data
