datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Romansh_German_Parallel_Data
Romansh–German Parallel Dataset (FineWeb-Based)
This dataset contains automatically aligned Romansh–German document pairs, extracted from the Fineweb2 using cosine similarity over OpenAI embeddings. It was created as part of a university programming project focused on document-level parallel data extraction.
Description
This project performs document-level alignment between Romansh and German web texts, which were extracted from the Fineweb2 dataset. It uses OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/Sudehsna/Romansh_German_Parallel_Data.apertus-pretrain-romanshThis dataset consist of three differnt parts. Monolingual Romansh Data, Polylingual data or more precisely translated data from Romansh into either German, French, Italian or English and Sythetic Data.
The Polylingual data consists of aligned and non aligned data. The synthetic data was created by interweaving the translational data and prefacing it with the sentence " This is a text translated from SOURCE LANGUAGE to Rumantsch Grischun".
The data has a metadata "idiom" if the if specific… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/apertus-pretrain-romansh.romansh-mt-evaluation
Dataset Description
This dataset contains the results of a human evaluation of machine translations from German into the six Romansh varieties. The evaluations were carried out by native speakers of the respective Romansh idioms as well as professional linguists.
The evaluation covers three quality dimensions:
Document accuracy, in which annotators assessed the adequacy of complete document translations.
Segment accuracy, in which annotators selected the more accurate… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/romansh-mt-evaluation.apertus-pretrain-romansh-backtranslatedVersion of https://hf.co/datasets/swiss-ai/apertus-pretrain-romansh (monolingual split only) that includes MT-generated translations into German.
The intended purpose of this dataset is to train MT systems or LLMs on the task of idiom-specific German→Romansh translation. Note that the German translations in this dataset might contain errors, since they have been automatically generated by an MT system.
Composition of the dataset and Romansh data sources
See… See the full description on the dataset page: https://huggingface.co/datasets/jvamvas/apertus-pretrain-romansh-backtranslated.
