djstrong/ppc
PPC - Polish Paraphrase Corpus Dataset Summary Polish Paraphrase Corpus contains 7000 manually labeled sentence pairs. The dataset was divided into training, validation and test splits. The training part includes 5000 examples, while the other parts contain 1000 examples each. The main purpose of creating such a dataset was to verify how machine learning models perform in the challenging problem of paraphrase identification, where most records contain semantically… See the full description on the dataset page: https://huggingface.co/datasets/djstrong/ppc.
033
