CoolFace
19 results

romansh

Sudehsna /Romansh_German_Parallel_Data Romansh–German Parallel Dataset (FineWeb-Based) This dataset contains automatically aligned Romansh–German document pairs, extracted from the Fineweb2 using cosine similarity over OpenAI embeddings. It was created as part of a university programming project focused on document-level parallel data extraction. Description This project performs document-level alignment between Romansh and German web texts, which were extracted from the Fineweb2 dataset. It uses OpenAI… See the full description on the dataset page: https://huggingface.co/datasets/Sudehsna/Romansh_German_Parallel_Data.tabular10K<n<100K2 likes462 downloads1y agoHugging Faceswiss-ai /apertus-pretrain-romanshThis dataset consist of three differnt parts. Monolingual Romansh Data, Polylingual data or more precisely translated data from Romansh into either German, French, Italian or English and Sythetic Data. The Polylingual data consists of aligned and non aligned data. The synthetic data was created by interweaving the translational data and prefacing it with the sentence " This is a text translated from SOURCE LANGUAGE to Rumantsch Grischun". The data has a metadata "idiom" if the if specific… See the full description on the dataset page: https://huggingface.co/datasets/swiss-ai/apertus-pretrain-romansh.tabulartranslation100K<n<1M4 likes221 downloads1y agoHugging FaceZurichNLP /romansh-mt-evaluation Dataset Description This dataset contains the results of a human evaluation of machine translations from German into the six Romansh varieties. The evaluations were carried out by native speakers of the respective Romansh idioms as well as professional linguists. The evaluation covers three quality dimensions: Document accuracy, in which annotators assessed the adequacy of complete document translations. Segment accuracy, in which annotators selected the more accurate… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/romansh-mt-evaluation.tabular1K<n<10K0 likes103 downloads3mo agoHugging FaceZurichNLP /romansh-backtranslated Romansh–German Back-Translation Dataset Background This dataset contains Romansh texts paired with German translations generated synthetically using Gemini 2.5 Flash. It was created as part of research on data augmentation for low-resource machine translation of Romansh, a language with 6 distinct written varieties (Rumantsch Grischun, Sursilvan, Sutsilvan, Surmiran, Puter, Vallader). LLMs tend to confuse Romansh varieties when translating into Romansh, but… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/romansh-backtranslated.text100K<n<1M0 likes88 downloads3mo agoHugging Facesimon-clmtd /romansh-grischun-morphological-corpus Romansh Grischun Morphological Corpus A morphologically annotated corpus of Rumantsch Grischun, the standardized written variety of Romansh. Dataset Summary This dataset provides gold-standard morphosyntactically annotated and lemmatized corpora for Rumantsch Grischun (standard written Romansh, ISO 639-3: roh), a national language of Switzerland. The primary annotation layer preserves the rich Xerox/Foma two-level morphological tags used by the finite-state… See the full description on the dataset page: https://huggingface.co/datasets/simon-clmtd/romansh-grischun-morphological-corpus.texttoken-classification1K<n<10K0 likes78 downloads14d agoHugging FaceRomanShp /MNIST-ResNet-Demo-Dataimagen<1K0 likes51 downloads3y agoHugging Face