datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Common_voice_sentence_similarityWikiMedical_sentence_similarity
Dataset Card for "WikiMedical_sentence_similarity"
WikiMedical_sentence_similarity is an adapted and ready-to-use sentence similarity dataset based on this dataset.
The preprocessing followed three steps:
Each text is splitted into sentences of 256 tokens (nltk tokenizer)
Each sentence is paired with a positive pair if found, and a negative one. Negative one are drawn randomly in the whole dataset.
Train and test split correspond to 70%/30%
More Information needed
stsb_multi_mt_fr_prompt_sentence_similarity
stsb_multi_mt_fr_prompt_sentence_similarity
Summary
stsb_multi_mt_fr_prompt_sentence_similarity is a subset of the Dataset of French Prompts (DFP).It contains 155,304 rows that can be used for a semantic similarity scoring task.The original data (without prompts) comes from the dataset stsb_multi_mt by May where only the French part has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/stsb_multi_mt_fr_prompt_sentence_similarity.semantic_sentence_similarity_ES
Dataset Card for "semantic_sentence_similarity_ES"
This dataset is based on https://huggingface.co/datasets/PlanTL-GOB-ES/sts-es, which includes the datasets presented at the SemEval 2014 and 2015 shared tasks on sentence similarity (see the link for more info about the citations). It also includes data from SemEval 2017.
Tamil-Sinhala-short-sentence-similarity-deep-learningThis research focuses on finding the best possible deep learning-based techniques to measure the short sentence similarity for low-resourced languages, focusing on Tamil and Sinhala sort sentences by utilizing existing unsupervised techniques for English. Original repo available on https://github.com/nlpcuom/Tamil-Sinhala-short-sentence-similarity-deep-learning
If you use this dataset, cite Nilaxan, S., & Ranathunga, S. (2021, July). Monolingual sentence similarity measurement using siamese… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Tamil-Sinhala-short-sentence-similarity-deep-learning.sentence_similaritysentence-transformer-parallel-En-Ko-with-Similarity
Preprocessing En-Ko subset of Parallel Sentences Datasets
해외에서 제작된 많은 대규모 번역 쌍 데이터들이 영어 텍스트와 한국어 텍스트를 문장 단위로 분리한 후 기계적으로 매핑시키고 있습니다.
이로 인해 전혀 엉뚱한 문장이 번역 쌍으로 매칭되어 있는 문제가 발생합니다.
임베딩 유사도 기반 전처리를 통해 이 문제를 해결할 수 있을 것 같아서 이 데이터 셋을 제작했습니다.
일부 데이터는 상업적 사용이 어려운 라이선스가 적용된 경우가 있습니다.
데이터 목록
parallel-sentences-global-voices
parallel-sentences-muse
parallel-sentences-jw300
parallel-sentences-tatoeba
parallel-sentences-wikititles
유사도 측정
BAAI/BGE-m3로 임베딩
영어 문장과 한국어 문장의… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/sentence-transformer-parallel-En-Ko-with-Similarity.sentence-similaritysentence-similarity-outputsentence-similarity-checkpoint-downloadsfor_Sentence_Similaritysentence-transformer-parallel-En-Hi-with-Similarity
