CoolFace
Datasetpublic

fyaronskiy/ru-paraphrase-NMT-Leipzig-cleaned

Dataset Description The dataset is obtained by filtering dataset of russian paraphrases by David Dale with automatic metrics. The data structure is saved. Have been deleted: Paraphrases that have cosine LABSE similarity with source sentences < 0.75. Paraphrases that are more than 2.5 times longer than source sentences. (Most of them are looped errors of back translation) Paraphrases that are similar in spelling to the original texts (paraphrases that have ChrF++ similarity >… See the full description on the dataset page: https://huggingface.co/datasets/fyaronskiy/ru-paraphrase-NMT-Leipzig-cleaned.

sourceHugging Faceupdated 1y agoView on Hugging Face
2likes26downloads

fyaronskiy/ru-paraphrase-NMT-Leipzig-cleaned · main · files are served by the source, never re-hosted here