datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
linguistic-similaritylm-similarity
Great Models Think Alike and this Undermines AI Oversight
This is the data collection for the publication "Great Models Think Alike and this Undermines AI Oversight."
judge_scores_mmlu_pro_free_filtered: Judge scores of nine judges without access to the reference answers on the filtered, open-style MMLU-Pro dataset.
judge_w_gt_mmlu_pro_free_filtered: Ensemble judge scores of five judges with access to the reference options and ground-truth information on the filtered OSQ MMLU-Pro.… See the full description on the dataset page: https://huggingface.co/datasets/bethgelab/lm-similarity.pubchem_10m_genmol_similarity
PubChem 10M GenMol Fingerprint Similarity Dataset
This dataset is an augmented version of the PubChem 10M dataset, enhanced with molecular similarity data generated using GenMol, as described in the paper GenMol: A Drug Discovery Generalist with Discrete Diffusion.
The dataset contains molecular structures represented as SMILES strings along with their corresponding molecular fingerprints, similarity scores, and various molecular properties.
This dataset is for training a Chem-MRL… See the full description on the dataset page: https://huggingface.co/datasets/Derify/pubchem_10m_genmol_similarity.gvg-explainable-game-similarity
GVG Explainable Game Similarity Dataset 2026 (v0.1.0)
50 human-reviewed pairs of similar PC games (82 games) from Game V Game. Each row says why the two games are similar, what differs, and who each suits, in English and Chinese, and cites the two official Steam store records the claims were checked against (with capture dates).
Most similarity data says two games are alike and stops. This one is small on purpose: every row was read by a person against the current Steam records… See the full description on the dataset page: https://huggingface.co/datasets/lette2/gvg-explainable-game-similarity.flickr30k_CLIP_ViT-B-32_subset_pairs_SimCSE_similarity_copyThai-Semantic-Textual-Similarity-BenchmarkSentence representation plays a crucial role in NLP downstream tasks such as NLI, text classification, and STS. Recent sentence representation training techniques require NLI or STS datasets. However, there are no equivalent Thai NLI or STS datasets for sentence representation training.
To address this problem we provide the Thai sentence vector benchmark. We evaluate the Spearman correlation score of the sentence representations’ performance on Thai STS-B (translated version of STS-B).… See the full description on the dataset page: https://huggingface.co/datasets/mrp/Thai-Semantic-Textual-Similarity-Benchmark.flickr30k_CLIP_ViT-B-32_subset_pairs_SimCSE_similaritydataset-biased-1.4M-pairwise-similaritystanford-rare-word-similarity-dataset
Stanford Rare Word (RW) Similarity Dataset
Created by Minh-Thang Luong, Richard Socher, and Christopher D. Manning, Stanford University Computer Science Department.
Available at: http://nlp.stanford.edu/~lmthang/morphoNLM
Described in:
Luong, M.-T., Socher, R., & Manning, C. D. (2013). Better Word Representations with Recursive Neural Networks for Morphology. CoNLL, Sofia, Bulgaria.
Columns
Word1 – First word in the pair
Word2 – Second word in the pair… See the full description on the dataset page: https://huggingface.co/datasets/almogtavor/stanford-rare-word-similarity-dataset.rag-embeddings-relevance-similarity
Dataset Card for rag-embeddings-relevance-similarity
This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into your Argilla server as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Using this dataset with Argilla
To load with Argilla, you'll just need to install Argilla as pip install argilla --upgrade and then use the following code:
import argilla as rg
ds =… See the full description on the dataset page: https://huggingface.co/datasets/argilla/rag-embeddings-relevance-similarity.english-words-human-similarityclass-zbmath-identifier
class-zbmath-identifier
This is a proxy dataset to model semantic similarity of short mathematical texts from zbMath.This proxy only contains zbMath.org identifiers (aka an) instead of full titles / abstracts.
Columns
an_a (string): zbMath.org identifier of work a
MSC_a (string): primary MSC5 of work a
MSC2_a (list(string)): secondary MSC5s of work a
an_b (string): zbMath.org identifier of work b
MSC_b (string): primary MSC5 of work b
MSC2_b (list(string)): secondary… See the full description on the dataset page: https://huggingface.co/datasets/math-similarity/class-zbmath-identifier.sentence-transformer-parallel-En-Ko-with-Similarity
Preprocessing En-Ko subset of Parallel Sentences Datasets
해외에서 제작된 많은 대규모 번역 쌍 데이터들이 영어 텍스트와 한국어 텍스트를 문장 단위로 분리한 후 기계적으로 매핑시키고 있습니다.
이로 인해 전혀 엉뚱한 문장이 번역 쌍으로 매칭되어 있는 문제가 발생합니다.
임베딩 유사도 기반 전처리를 통해 이 문제를 해결할 수 있을 것 같아서 이 데이터 셋을 제작했습니다.
일부 데이터는 상업적 사용이 어려운 라이선스가 적용된 경우가 있습니다.
데이터 목록
parallel-sentences-global-voices
parallel-sentences-muse
parallel-sentences-jw300
parallel-sentences-tatoeba
parallel-sentences-wikititles
유사도 측정
BAAI/BGE-m3로 임베딩
영어 문장과 한국어 문장의… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/sentence-transformer-parallel-En-Ko-with-Similarity.web-cache-similarity-benchmark
Web Cache Semantic Similarity Benchmark
This dataset was created by Alete to benchmark the effectiveness of modern semantic vector caching against traditional cryptographic string hashing (SHA-256) directly on the live web, evaluating temporal cache decay across multiple time frames in a uniform environment.
Naïve string hashing fails on the modern web because minor changes in page layouts—such as dynamic trending articles list, dynamic comments count, ad units, and modified… See the full description on the dataset page: https://huggingface.co/datasets/alete-ai/web-cache-similarity-benchmark.Finding_new_crystal_compounds_using_chemical_similarity
Cite this dataset Wang, H., Botti, S., and Marques, M. A. L. Finding new crystal compounds using chemical similarity. ColabFit, 2025. https://doi.org/10.60732/b9e7eedf
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_sk8zwvk3qxur_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Finding_new_crystal_compounds_using_chemical_similarity.DeepSeek-7B_aime24_similarity_0902orthographic-similarity-ratings
Orthographic similarity ratings for English-Spanish cognates from the academic word list
This dataset is derived from the paper "Orthographic similarity ratings for English-Spanish cognates from the academic word list" by Hout et al.
You can find there paper here and the original dataset here.
DeepSeek-1.5B_dapo2k_similarity_0907Qwen3-1.7B_aime24_similarity_0902With-Similarity-Scoresrag-embeddings-relevance-similaritybookfilm_summaries_with_similarity_and_sentimentmashqa-cosine-similarity-bgeNuminaMath-CoT-2048-cosine-similarity-rankETD_Detoxification_Dataset_similarity_metricscrs-2014-2023-housing-similaritycat_names_similarity_pairs_three_scores_with_cat_scores_fullarxiv_similarity_300arc-barc-processed-diagnostic-insights-similarity-rubriclaion400m_top10000_similarity
