datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sts-h-paraphrase-detection
STS-Hard Test Set
The STS-Hard dataset is a paraphrase detection test set derived from the STSBenchmark dataset. It was introduced as part of the PARAPHRASUS: A Comprehensive Benchmark for Evaluating Paraphrase Detection Models. The test set includes the paraphrase label as well as individual annotation labels from two annotators:
P1: The semanticist.
P2: A student annotator.
For more details, refer to the original paper that was presented at COLING 2025.
Citation… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/sts-h-paraphrase-detection.buet_model_buet_test_data_paraphrase_detection
Dataset Card for "buet_model_buet_test_data_paraphrase_detection"
More Information needed
indic_model_indic_test_data_paraphrase_detection
Dataset Card for "indic_model_indic_test_data_paraphrase_detection"
More Information needed
