datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
span-similarity-dataset
Span Similarity Dataset (SSD)
Dataset Summary
The Span Similarity Dataset (SSD) focuses on Explainable Textual Similarity. It consists
of pairs of sentences with annotations pointing to both semantically equivalent and
dissimilar spans.
Languages
The SSD includes exclusively texts in English.
Dataset Structure
The dataset is split into -train (800 samples), -eval (100 samples), and -test (100
samples), all of them provided as a .tsv… See the full description on the dataset page: https://huggingface.co/datasets/tum-nlp/span-similarity-dataset.gvg-explainable-game-similarity
GVG Explainable Game Similarity Dataset 2026 (v0.1.0)
50 human-reviewed pairs of similar PC games (82 games) from Game V Game. Each row says why the two games are similar, what differs, and who each suits, in English and Chinese, and cites the two official Steam store records the claims were checked against (with capture dates).
Most similarity data says two games are alike and stops. This one is small on purpose: every row was read by a person against the current Steam records… See the full description on the dataset page: https://huggingface.co/datasets/lette2/gvg-explainable-game-similarity.Thai-Semantic-Textual-Similarity-BenchmarkSentence representation plays a crucial role in NLP downstream tasks such as NLI, text classification, and STS. Recent sentence representation training techniques require NLI or STS datasets. However, there are no equivalent Thai NLI or STS datasets for sentence representation training.
To address this problem we provide the Thai sentence vector benchmark. We evaluate the Spearman correlation score of the sentence representations’ performance on Thai STS-B (translated version of STS-B).… See the full description on the dataset page: https://huggingface.co/datasets/mrp/Thai-Semantic-Textual-Similarity-Benchmark.Dataset-Structural_Similarity-ProteinShake
Description
Structure Similarity Prediction predicts the (aligned) Local Distance Difference Test (LDDT) of the structures given an unaligned pair of proteins. Target values are computed after alignment with TM-align for all pairs of 1000 randomly sampled single-chain proteins.
Protein Format: SA sequence(PDB)
Splits
The dataset is from ProteinShake Building datasets and benchmarks for deep learning on protein structures. We use the splits based on 70% structure… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-Structural_Similarity-ProteinShake.clinical-drift-based-drug-similarity-substitution-mapping-v0.1What this dataset tests
Drug similarity based on systemic drift fingerprintsnot target class.
Required outputs
drift similarity score
shared drift axes
critical difference axes
substitution recommendation class
rationale trace
monitoring plan
Recommendation classes
preferred substitute
conditional substitute
complement not substitute
do not substitute
avoid in frail profiles
Use case
Third layer of the Drug-Induced System Drift Library.
stanford-rare-word-similarity-dataset
Stanford Rare Word (RW) Similarity Dataset
Created by Minh-Thang Luong, Richard Socher, and Christopher D. Manning, Stanford University Computer Science Department.
Available at: http://nlp.stanford.edu/~lmthang/morphoNLM
Described in:
Luong, M.-T., Socher, R., & Manning, C. D. (2013). Better Word Representations with Recursive Neural Networks for Morphology. CoNLL, Sofia, Bulgaria.
Columns
Word1 – First word in the pair
Word2 – Second word in the pair… See the full description on the dataset page: https://huggingface.co/datasets/almogtavor/stanford-rare-word-similarity-dataset.similarity-sentences-spanish
similarity-sentences-spanish (SSS)
Dataset Summary
This dataset comprises a collection of sentences generated using Chat GPT-3, covering various general topics.
The dataset also includes sentences from two existing datasets, STS-ES and STSB-Multi-MT, as well as SICK, which were used as additional sources.
The sentences in this dataset were generated to exhibit varying levels of similarity based on randomly divided prompts.
Source
Share (rows)
Count (rows)
Score… See the full description on the dataset page: https://huggingface.co/datasets/jaimevera1107/similarity-sentences-spanish.english-words-human-similaritysemantic-textual-similarityclass-zbmath-identifier
class-zbmath-identifier
This is a proxy dataset to model semantic similarity of short mathematical texts from zbMath.This proxy only contains zbMath.org identifiers (aka an) instead of full titles / abstracts.
Columns
an_a (string): zbMath.org identifier of work a
MSC_a (string): primary MSC5 of work a
MSC2_a (list(string)): secondary MSC5s of work a
an_b (string): zbMath.org identifier of work b
MSC_b (string): primary MSC5 of work b
MSC2_b (list(string)): secondary… See the full description on the dataset page: https://huggingface.co/datasets/math-similarity/class-zbmath-identifier.clinical-structural-similarity-scoring-against-ground-truth-v0.1What this dataset tests
Whether a model can match the later-discovered explanationby structural logic, not by diagnosis label.
Input
pre-explanation case summary and data
predicted structure
ground truth structure
Required outputs
structural_similarity_score_0_100
alignment_strengths
divergence_points
Representation format
Predicted and ground truth structures use this schema text
systems A B C
nodes n1 n2 n3
edges n1->n2 n2->n3
phases p1 p2 p3
failure_modes f1 f2
Typical… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-structural-similarity-scoring-against-ground-truth-v0.1.orthographic-similarity-ratings
Orthographic similarity ratings for English-Spanish cognates from the academic word list
This dataset is derived from the paper "Orthographic similarity ratings for English-Spanish cognates from the academic word list" by Hout et al.
You can find there paper here and the original dataset here.
semantic_textual_similarity_dataset
Turkish Semantic Textual Similarity Dataset
Bu veri seti, Türkçe cümle çiftleri arasındaki anlamsal benzerliği değerlendirmek amacıyla hazırlanmış 80 özgün örnekten oluşur.
Veri yapısı
sentence1: Birinci cümle
sentence2: İkinci cümle
score: İnsan değerlendirmesiyle belirlenen anlamsal benzerlik puanı
Puanlama ölçeği
Puanlar 0 ile 5 arasındadır:
5: Aynı anlamı ifade eden cümleler
4: Büyük ölçüde aynı, küçük ayrıntı farkları bulunan cümleler
3:… See the full description on the dataset page: https://huggingface.co/datasets/aliFurkan123/semantic_textual_similarity_dataset.bookfilm_summaries_with_similarity_and_sentimentsemantic-similarity
English Word Semantic Similarity
This dataset is a combination of the following datasets.
Wordsim-353
Simlex-999
SimVerb-3500
The similarity score is scaled from 0 to 1, with 1 having the highest similarity.
With-Similarity-Scoresq2q_similarity_workshopsimilarity-sentences-portuguese
similarity-sentences-portuguese (SSP)
Dataset Summary
This dataset comprises a collection of sentences generated using Chat GPT-3, covering various general topics, originally in spanish by jaimevera1107.
The sentences were translated to portuguese using seamless-m4t-medium.
Languages
Portuguese
Dataset Structure
Data Fields
Sentence 1: The first sentence to be compared.
Sentence 2: The second sentence to be compared.
Score: A number… See the full description on the dataset page: https://huggingface.co/datasets/luiseduardobrito/similarity-sentences-portuguese.movie-similaritysimilarity_datModel_Semantic_Similarity_Analysis_BERT_VN_DATA
