paraphrase detection
farsi_paraphrase_detection
Dataset Card for "farsi_paraphrase_detection"
More Information needed
paws-x_fr_prompt_paraphrase_detection
paws-x_fr_prompt_paraphrase_detection
Summary
paws-x_fr_prompt_paraphrase_detection is a subset of the Dataset of French Prompts (DFP).It contains 1,174,822 rows that can be used for a paraphrase detection task.The original data (without prompts) comes from the dataset paws-x by Yang et al. where only the French part has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the xP3… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/paws-x_fr_prompt_paraphrase_detection.sts-h-paraphrase-detection
STS-Hard Test Set
The STS-Hard dataset is a paraphrase detection test set derived from the STSBenchmark dataset. It was introduced as part of the PARAPHRASUS: A Comprehensive Benchmark for Evaluating Paraphrase Detection Models. The test set includes the paraphrase label as well as individual annotation labels from two annotators:
P1: The semanticist.
P2: A student annotator.
For more details, refer to the original paper that was presented at COLING 2025.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/sts-h-paraphrase-detection.id-paraphrase-detectionThis dataset is built as a playground for sequence to sequence classificationbuet_model_buet_test_data_paraphrase_detection
Dataset Card for "buet_model_buet_test_data_paraphrase_detection"
More Information needed
indic_model_indic_test_data_paraphrase_detection
Dataset Card for "indic_model_indic_test_data_paraphrase_detection"
More Information needed
