sentence-similarity
Common_voice_sentence_similarityWikiMedical_sentence_similarity
Dataset Card for "WikiMedical_sentence_similarity"
WikiMedical_sentence_similarity is an adapted and ready-to-use sentence similarity dataset based on this dataset.
The preprocessing followed three steps:
Each text is splitted into sentences of 256 tokens (nltk tokenizer)
Each sentence is paired with a positive pair if found, and a negative one. Negative one are drawn randomly in the whole dataset.
Train and test split correspond to 70%/30%
More Information needed
stsb_multi_mt_fr_prompt_sentence_similarity
stsb_multi_mt_fr_prompt_sentence_similarity
Summary
stsb_multi_mt_fr_prompt_sentence_similarity is a subset of the Dataset of French Prompts (DFP).It contains 155,304 rows that can be used for a semantic similarity scoring task.The original data (without prompts) comes from the dataset stsb_multi_mt by May where only the French part has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/stsb_multi_mt_fr_prompt_sentence_similarity.semantic_sentence_similarity_ES
Dataset Card for "semantic_sentence_similarity_ES"
This dataset is based on https://huggingface.co/datasets/PlanTL-GOB-ES/sts-es, which includes the datasets presented at the SemEval 2014 and 2015 shared tasks on sentence similarity (see the link for more info about the citations). It also includes data from SemEval 2017.
Tamil-Sinhala-short-sentence-similarity-deep-learningThis research focuses on finding the best possible deep learning-based techniques to measure the short sentence similarity for low-resourced languages, focusing on Tamil and Sinhala sort sentences by utilizing existing unsupervised techniques for English. Original repo available on https://github.com/nlpcuom/Tamil-Sinhala-short-sentence-similarity-deep-learning
If you use this dataset, cite Nilaxan, S., & Ranathunga, S. (2021, July). Monolingual sentence similarity measurement using siamese… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Tamil-Sinhala-short-sentence-similarity-deep-learning.sentence_similarity
