CoolFace
Datasetpublic

shershen/ru_anglicism

Dataset Card for Ru Anglicism Dataset Description Dataset Summary Dataset for detection and substraction anglicisms from sentences in Russian. Sentences with anglicism automatically parsed from National Corpus of the Russian language, Habr and Pikabu. The paraphrases for the sentences were created manually. Languages The dataset is in Russian. Usage Loading dataset: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/shershen/ru_anglicism.

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
4likes15downloads
Dataset Card

Dataset Card for Ru Anglicism

Table of Contents

Dataset Description

Dataset Summary

Dataset for detection and substraction anglicisms from sentences in Russian. Sentences with anglicism automatically parsed from National Corpus of the Russian language, Habr and Pikabu. The paraphrases for the sentences were created manually.

Languages

The dataset is in Russian.

Usage

Loading dataset:

python
from datasets import load_dataset
dataset = load_dataset('shershen/ru_anglicism')

Dataset Structure

Data Instunces

For each instance, there are four strings: word, form, sentence and paraphrase.

{
  'word': 'коллаб',
  'form': 'коллабу',
  'sentence': 'Сделаем коллабу, раскрутимся.',
  'paraphrase': 'Сделаем совместный проект, раскрутимся.'
}

Data Splits

Full dataset contains 1084 sentences. Split of dataset is:

Dataset SplitNumber of Rows
Train1007
Test77