shershen/ru_anglicism
Dataset Card for Ru Anglicism Dataset Description Dataset Summary Dataset for detection and substraction anglicisms from sentences in Russian. Sentences with anglicism automatically parsed from National Corpus of the Russian language, Habr and Pikabu. The paraphrases for the sentences were created manually. Languages The dataset is in Russian. Usage Loading dataset: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/shershen/ru_anglicism.
Dataset Card for Ru Anglicism
Table of Contents
- Table of Contents
- Dataset Description
- Dataset Summary
- Languages
- Dataset Structure
- Data Instances
- Data Splits
Dataset Description
Dataset Summary
Dataset for detection and substraction anglicisms from sentences in Russian. Sentences with anglicism automatically parsed from National Corpus of the Russian language, Habr and Pikabu. The paraphrases for the sentences were created manually.
Languages
The dataset is in Russian.
Usage
Loading dataset:
from datasets import load_dataset
dataset = load_dataset('shershen/ru_anglicism')Dataset Structure
Data Instunces
For each instance, there are four strings: word, form, sentence and paraphrase.
{
'word': 'коллаб',
'form': 'коллабу',
'sentence': 'Сделаем коллабу, раскрутимся.',
'paraphrase': 'Сделаем совместный проект, раскрутимся.'
}Data Splits
Full dataset contains 1084 sentences. Split of dataset is:
