datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gvg-explainable-game-similarity
GVG Explainable Game Similarity Dataset 2026 (v0.1.0)
50 human-reviewed pairs of similar PC games (82 games) from Game V Game. Each row says why the two games are similar, what differs, and who each suits, in English and Chinese, and cites the two official Steam store records the claims were checked against (with capture dates).
Most similarity data says two games are alike and stops. This one is small on purpose: every row was read by a person against the current Steam records… See the full description on the dataset page: https://huggingface.co/datasets/lette2/gvg-explainable-game-similarity.Thai-Semantic-Textual-Similarity-BenchmarkSentence representation plays a crucial role in NLP downstream tasks such as NLI, text classification, and STS. Recent sentence representation training techniques require NLI or STS datasets. However, there are no equivalent Thai NLI or STS datasets for sentence representation training.
To address this problem we provide the Thai sentence vector benchmark. We evaluate the Spearman correlation score of the sentence representations’ performance on Thai STS-B (translated version of STS-B).… See the full description on the dataset page: https://huggingface.co/datasets/mrp/Thai-Semantic-Textual-Similarity-Benchmark.stanford-rare-word-similarity-dataset
Stanford Rare Word (RW) Similarity Dataset
Created by Minh-Thang Luong, Richard Socher, and Christopher D. Manning, Stanford University Computer Science Department.
Available at: http://nlp.stanford.edu/~lmthang/morphoNLM
Described in:
Luong, M.-T., Socher, R., & Manning, C. D. (2013). Better Word Representations with Recursive Neural Networks for Morphology. CoNLL, Sofia, Bulgaria.
Columns
Word1 – First word in the pair
Word2 – Second word in the pair… See the full description on the dataset page: https://huggingface.co/datasets/almogtavor/stanford-rare-word-similarity-dataset.english-words-human-similarityclass-zbmath-identifier
class-zbmath-identifier
This is a proxy dataset to model semantic similarity of short mathematical texts from zbMath.This proxy only contains zbMath.org identifiers (aka an) instead of full titles / abstracts.
Columns
an_a (string): zbMath.org identifier of work a
MSC_a (string): primary MSC5 of work a
MSC2_a (list(string)): secondary MSC5s of work a
an_b (string): zbMath.org identifier of work b
MSC_b (string): primary MSC5 of work b
MSC2_b (list(string)): secondary… See the full description on the dataset page: https://huggingface.co/datasets/math-similarity/class-zbmath-identifier.orthographic-similarity-ratings
Orthographic similarity ratings for English-Spanish cognates from the academic word list
This dataset is derived from the paper "Orthographic similarity ratings for English-Spanish cognates from the academic word list" by Hout et al.
You can find there paper here and the original dataset here.
With-Similarity-Scoresbookfilm_summaries_with_similarity_and_sentimentmovie-similarity
