CoolFace
Datasetpublic

Agisight/tyv-rus-200k

tyv-rus-200k data card This data was collected via www.tyvan.ru platform by linguists, scientists, journalists, volunteers, etc. Actually here 296k rows. Almost 300k Dataset Details Dataset Description Curated by: Ali Kuzhuget (tech and data), Ondar Choygan (data) contributors Language(s) (NLP): Tyvan (Tuvan), Russian License:: CC BY 4.0. Below is the brief information about the languages Language Language code on the website ISO 639-3… See the full description on the dataset page: https://huggingface.co/datasets/Agisight/tyv-rus-200k.

sourceHugging Facecc-by-4.0updated 2y agoView on Hugging Face
0likes37downloads
Dataset Card

tyv-rus-200k data card

<!-- Provide a quick summary of the dataset. -->

This data was collected via www.tyvan.ru platform by linguists, scientists, journalists, volunteers, etc. Actually here 296k rows. Almost 300k

Dataset Details

Dataset Description

<!-- Provide a longer summary of what this dataset is. -->

  • Curated by: Ali Kuzhuget (tech and data), Ondar Choygan (data) contributors
  • Language(s) (NLP): Tyvan (Tuvan), Russian
  • License:: CC BY 4.0.

Below is the brief information about the languages

LanguageLanguage code on the websiteISO 639-3Glottolog
Tyvantyvtyvtuvi1240
Russianrusrusruss1263

Dataset Sources

The dataset has been downloaded from www.tyvan.ru.

Uses

The dataset is intended to help humans and machines learn the low-resourced Tyvan (Tuvan) and Russian languages.

Dataset Structure

The dataset contains Tyvan-Russian paires.

Data row has the following fields:

  • tyv: str: text in Tuvan
  • ru: str: text in Russian (translate)

Dataset Creation

The dataset was curates as a source of machine translation training and other NLP tools. It consists donated and professional translations from books and websites. They have been downloaded from the www.tyvan.ru website and fined by Ali Kuzhuget. No additional filtering or postprocessing has been applied.