CoolFace
Datasetpublic

ruscorpora/normalization

TL;DR: Text Normalization for Social Media Corpus Dataset Description This dataset contains examples of Russian-language texts from social networks with distorted spelling (typos, abbreviations, etc.) and their normalized versions in json format. A detailed spelling correction protocol is given in the TBA article. The dataset size is 1930 sentence pairs. In each pair, the sentences are tokenized by words, and the lengths of both sentences in the pair are… See the full description on the dataset page: https://huggingface.co/datasets/ruscorpora/normalization.

sourceHugging Facecc-by-nc-sa-4.0updated 1y agoView on Hugging Face
0likes8downloads
4 commits on main
8fc9c161y ago

license added

Dmitry
cfebc361y ago

yaml section added to readme

Dmitry
5fd11a21y ago

dataset added, readme updated

Dmitry
56e36ed1y ago

initial commit

morozow