datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
github-raw
Overview
This data has been scraped from the github api using the requests library!
ru-paraphrase-NMT-Leipzig-cleaned
Dataset Description
The dataset is obtained by filtering dataset of russian paraphrases by David Dale with automatic metrics.
The data structure is saved.
Have been deleted:
Paraphrases that have cosine LABSE similarity with source sentences < 0.75.
Paraphrases that are more than 2.5 times longer than source sentences. (Most of them are looped errors of back translation)
Paraphrases that are similar in spelling to the original texts (paraphrases that have ChrF++ similarity > 0.6… See the full description on the dataset page: https://huggingface.co/datasets/fyaronskiy/ru-paraphrase-NMT-Leipzig-cleaned.qwen-nmt
