CoolFace
Datasetpublic

shershen/ru_anglicism

Dataset Card for Ru Anglicism Dataset Description Dataset Summary Dataset for detection and substraction anglicisms from sentences in Russian. Sentences with anglicism automatically parsed from National Corpus of the Russian language, Habr and Pikabu. The paraphrases for the sentences were created manually. Languages The dataset is in Russian. Usage Loading dataset: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/shershen/ru_anglicism.

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
4likes15downloads
10 commits on main
97fad103y ago

Update ru_anglicism.py

shershen
87668fe3y ago

Upload dataset_infos.json

shershen
ac80a4a3y ago

add files

shershen
3e621613y ago

Delete data/train.jsonl

shershen
a8e217e3y ago

Delete data/test.jsonl

shershen
75a1db23y ago

Update README.md

shershen
0242fb63y ago

Update README.md

shershen
c07d5233y ago

Upload ru_anglicism.py

shershen
17e906a3y ago

add files

shershen
86456f13y ago

initial commit

shershen