CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tum-nlp /span-similarity-dataset Span Similarity Dataset (SSD) Dataset Summary The Span Similarity Dataset (SSD) focuses on Explainable Textual Similarity. It consists of pairs of sentences with annotations pointing to both semantically equivalent and dissimilar spans. Languages The SSD includes exclusively texts in English. Dataset Structure The dataset is split into -train (800 samples), -eval (100 samples), and -test (100 samples), all of them provided as a .tsv… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/span-similarity-dataset.textsentence-similarity1K<n<10K2 likes88 downloads2mo agoHugging Face02lette2 /gvg-explainable-game-similarity GVG Explainable Game Similarity Dataset 2026 (v0.1.0) 50 human-reviewed pairs of similar PC games (82 games) from Game V Game. Each row says why the two games are similar, what differs, and who each suits, in English and Chinese, and cites the two official Steam store records the claims were checked against (with capture dates). Most similarity data says two games are alike and stops. This one is small on purpose: every row was read by a person against the current Steam records… See the full description on the dataset page: https://huggingface.co/datasets/lette2/gvg-explainable-game-similarity.tabulartext-classificationn<1K0 likes56 downloads23d agoHugging Face03mrp /Thai-Semantic-Textual-Similarity-BenchmarkSentence representation plays a crucial role in NLP downstream tasks such as NLI, text classification, and STS. Recent sentence representation training techniques require NLI or STS datasets. However, there are no equivalent Thai NLI or STS datasets for sentence representation training. To address this problem we provide the Thai sentence vector benchmark. We evaluate the Spearman correlation score of the sentence representations’ performance on Thai STS-B (translated version of STS-B).… See the full description on the dataset page: https://huggingface.co/datasets/mrp/Thai-Semantic-Textual-Similarity-Benchmark.tabular1K<n<10K2 likes45 downloads5y agoHugging Face04SaProtHub /Dataset-Structural_Similarity-ProteinShake Description Structure Similarity Prediction predicts the (aligned) Local Distance Difference Test (LDDT) of the structures given an unaligned pair of proteins. Target values are computed after alignment with TM-align for all pairs of 1000 randomly sampled single-chain proteins. Protein Format: SA sequence(PDB) Splits The dataset is from ProteinShake Building datasets and benchmarks for deep learning on protein structures. We use the splits based on 70% structure… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-Structural_Similarity-ProteinShake.text100K<n<1M2 likes38 downloads2y agoHugging Face05ClarusC64 /clinical-drift-based-drug-similarity-substitution-mapping-v0.1What this dataset tests Drug similarity based on systemic drift fingerprintsnot target class. Required outputs drift similarity score shared drift axes critical difference axes substitution recommendation class rationale trace monitoring plan Recommendation classes preferred substitute conditional substitute complement not substitute do not substitute avoid in frail profiles Use case Third layer of the Drug-Induced System Drift Library. texttext-classificationn<1K0 likes30 downloads8mo agoHugging Face06almogtavor /stanford-rare-word-similarity-dataset Stanford Rare Word (RW) Similarity Dataset Created by Minh-Thang Luong, Richard Socher, and Christopher D. Manning, Stanford University Computer Science Department. Available at: http://nlp.stanford.edu/~lmthang/morphoNLM Described in: Luong, M.-T., Socher, R., & Manning, C. D. (2013). Better Word Representations with Recursive Neural Networks for Morphology. CoNLL, Sofia, Bulgaria. Columns Word1 – First word in the pair Word2 – Second word in the pair… See the full description on the dataset page: https://huggingface.co/datasets/almogtavor/stanford-rare-word-similarity-dataset.tabular1K<n<10K1 likes29 downloads1y agoHugging Face07jaimevera1107 /similarity-sentences-spanish similarity-sentences-spanish (SSS) Dataset Summary This dataset comprises a collection of sentences generated using Chat GPT-3, covering various general topics. The dataset also includes sentences from two existing datasets, STS-ES and STSB-Multi-MT, as well as SICK, which were used as additional sources. The sentences in this dataset were generated to exhibit varying levels of similarity based on randomly divided prompts. Source Share (rows) Count (rows) Score… See the full description on the dataset page: https://huggingface.co/datasets/jaimevera1107/similarity-sentences-spanish.textsentence-similarity10K<n<100K7 likes26 downloads3y agoHugging Face08StephanAkkerman /english-words-human-similaritytabularn<1K0 likes21 downloads2y agoHugging Face09sinhala-nlp /semantic-textual-similaritytext1K<n<10K1 likes18 downloads2y agoHugging Face10math-similarity /class-zbmath-identifier class-zbmath-identifier This is a proxy dataset to model semantic similarity of short mathematical texts from zbMath.This proxy only contains zbMath.org identifiers (aka an) instead of full titles / abstracts. Columns an_a (string): zbMath.org identifier of work a MSC_a (string): primary MSC5 of work a MSC2_a (list(string)): secondary MSC5s of work a an_b (string): zbMath.org identifier of work b MSC_b (string): primary MSC5 of work b MSC2_b (list(string)): secondary… See the full description on the dataset page: https://huggingface.co/datasets/math-similarity/class-zbmath-identifier.tabulartext-classification100K<n<1M0 likes17 downloads2y agoHugging Face11ClarusC64 /clinical-structural-similarity-scoring-against-ground-truth-v0.1What this dataset tests Whether a model can match the later-discovered explanationby structural logic, not by diagnosis label. Input pre-explanation case summary and data predicted structure ground truth structure Required outputs structural_similarity_score_0_100 alignment_strengths divergence_points Representation format Predicted and ground truth structures use this schema text systems A B C nodes n1 n2 n3 edges n1->n2 n2->n3 phases p1 p2 p3 failure_modes f1 f2 Typical… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-structural-similarity-scoring-against-ground-truth-v0.1.texttext-classificationn<1K0 likes17 downloads8mo agoHugging Face12StephanAkkerman /orthographic-similarity-ratings Orthographic similarity ratings for English-Spanish cognates from the academic word list This dataset is derived from the paper "Orthographic similarity ratings for English-Spanish cognates from the academic word list" by Hout et al. You can find there paper here and the original dataset here. tabularn<1K0 likes14 downloads2y agoHugging Face13aliFurkan123 /semantic_textual_similarity_dataset Turkish Semantic Textual Similarity Dataset Bu veri seti, Türkçe cümle çiftleri arasındaki anlamsal benzerliği değerlendirmek amacıyla hazırlanmış 80 özgün örnekten oluşur. Veri yapısı sentence1: Birinci cümle sentence2: İkinci cümle score: İnsan değerlendirmesiyle belirlenen anlamsal benzerlik puanı Puanlama ölçeği Puanlar 0 ile 5 arasındadır: 5: Aynı anlamı ifade eden cümleler 4: Büyük ölçüde aynı, küçük ayrıntı farkları bulunan cümleler 3:… See the full description on the dataset page: https://huggingface.co/datasets/aliFurkan123/semantic_textual_similarity_dataset.textsentence-similarityn<1K0 likes10 downloads2mo agoHugging Face14ada-datadruids /bookfilm_summaries_with_similarity_and_sentimenttabularn<1K0 likes9 downloads2y agoHugging Face15StephanAkkerman /semantic-similarity English Word Semantic Similarity This dataset is a combination of the following datasets. Wordsim-353 Simlex-999 SimVerb-3500 The similarity score is scaled from 0 to 1, with 1 having the highest similarity. text1K<n<10K0 likes8 downloads2y agoHugging Face16Pushpendra817 /With-Similarity-Scorestabular1K<n<10K1 likes7 downloads3y agoHugging Face17elsheikhams /q2q_similarity_workshoptext10K<n<100K0 likes6 downloads3y agoHugging Face18luiseduardobrito /similarity-sentences-portuguese similarity-sentences-portuguese (SSP) Dataset Summary This dataset comprises a collection of sentences generated using Chat GPT-3, covering various general topics, originally in spanish by jaimevera1107. The sentences were translated to portuguese using seamless-m4t-medium. Languages Portuguese Dataset Structure Data Fields Sentence 1: The first sentence to be compared. Sentence 2: The second sentence to be compared. Score: A number… See the full description on the dataset page: https://huggingface.co/datasets/luiseduardobrito/similarity-sentences-portuguese.texttext-classification10K<n<100K4 likes6 downloads3y agoHugging Face19vedface /movie-similaritytabular1K<n<10K0 likes6 downloads1y agoHugging Face20soumakchak /similarity_dattext10K<n<100K0 likes5 downloads1y agoHugging Face21buiductai /Model_Semantic_Similarity_Analysis_BERT_VN_DATAtext1K<n<10K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.