CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01chendelong /linguistic-similaritytabularn<1K1 likes706 downloads2y agoHugging Face02bethgelab /lm-similarity Great Models Think Alike and this Undermines AI Oversight This is the data collection for the publication "Great Models Think Alike and this Undermines AI Oversight." judge_scores_mmlu_pro_free_filtered: Judge scores of nine judges without access to the reference answers on the filtered, open-style MMLU-Pro dataset. judge_w_gt_mmlu_pro_free_filtered: Ensemble judge scores of five judges with access to the reference options and ground-truth information on the filtered OSQ MMLU-Pro.… See the full description on the dataset page: https://huggingface.co/datasets/bethgelab/lm-similarity.tabularquestion-answering10K<n<100K5 likes266 downloads2y agoHugging Face03Derify /pubchem_10m_genmol_similarity PubChem 10M GenMol Fingerprint Similarity Dataset This dataset is an augmented version of the PubChem 10M dataset, enhanced with molecular similarity data generated using GenMol, as described in the paper GenMol: A Drug Discovery Generalist with Discrete Diffusion. The dataset contains molecular structures represented as SMILES strings along with their corresponding molecular fingerprints, similarity scores, and various molecular properties. This dataset is for training a Chem-MRL… See the full description on the dataset page: https://huggingface.co/datasets/Derify/pubchem_10m_genmol_similarity.tabularother10M<n<100M1 likes75 downloads1y agoHugging Face04lette2 /gvg-explainable-game-similarity GVG Explainable Game Similarity Dataset 2026 (v0.1.0) 50 human-reviewed pairs of similar PC games (82 games) from Game V Game. Each row says why the two games are similar, what differs, and who each suits, in English and Chinese, and cites the two official Steam store records the claims were checked against (with capture dates). Most similarity data says two games are alike and stops. This one is small on purpose: every row was read by a person against the current Steam records… See the full description on the dataset page: https://huggingface.co/datasets/lette2/gvg-explainable-game-similarity.tabulartext-classificationn<1K0 likes56 downloads23d agoHugging Face05closji /flickr30k_CLIP_ViT-B-32_subset_pairs_SimCSE_similarity_copytabular10M<n<100M0 likes55 downloads4y agoHugging Face06mrp /Thai-Semantic-Textual-Similarity-BenchmarkSentence representation plays a crucial role in NLP downstream tasks such as NLI, text classification, and STS. Recent sentence representation training techniques require NLI or STS datasets. However, there are no equivalent Thai NLI or STS datasets for sentence representation training. To address this problem we provide the Thai sentence vector benchmark. We evaluate the Spearman correlation score of the sentence representations’ performance on Thai STS-B (translated version of STS-B).… See the full description on the dataset page: https://huggingface.co/datasets/mrp/Thai-Semantic-Textual-Similarity-Benchmark.tabular1K<n<10K2 likes41 downloads5y agoHugging Face07closji /flickr30k_CLIP_ViT-B-32_subset_pairs_SimCSE_similaritytabular10M<n<100M1 likes35 downloads4y agoHugging Face08stage-babylm /dataset-biased-1.4M-pairwise-similaritytabular1M<n<10M0 likes30 downloads1mo agoHugging Face09almogtavor /stanford-rare-word-similarity-dataset Stanford Rare Word (RW) Similarity Dataset Created by Minh-Thang Luong, Richard Socher, and Christopher D. Manning, Stanford University Computer Science Department. Available at: http://nlp.stanford.edu/~lmthang/morphoNLM Described in: Luong, M.-T., Socher, R., & Manning, C. D. (2013). Better Word Representations with Recursive Neural Networks for Morphology. CoNLL, Sofia, Bulgaria. Columns Word1 – First word in the pair Word2 – Second word in the pair… See the full description on the dataset page: https://huggingface.co/datasets/almogtavor/stanford-rare-word-similarity-dataset.tabular1K<n<10K1 likes26 downloads1y agoHugging Face10argilla /rag-embeddings-relevance-similarity Dataset Card for rag-embeddings-relevance-similarity This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into your Argilla server as explained in Load with Argilla, or used directly with the datasets library in Load with datasets. Using this dataset with Argilla To load with Argilla, you'll just need to install Argilla as pip install argilla --upgrade and then use the following code: import argilla as rg ds =… See the full description on the dataset page: https://huggingface.co/datasets/argilla/rag-embeddings-relevance-similarity.tabular1K<n<10K1 likes24 downloads2y agoHugging Face11StephanAkkerman /english-words-human-similaritytabularn<1K0 likes21 downloads2y agoHugging Face12math-similarity /class-zbmath-identifier class-zbmath-identifier This is a proxy dataset to model semantic similarity of short mathematical texts from zbMath.This proxy only contains zbMath.org identifiers (aka an) instead of full titles / abstracts. Columns an_a (string): zbMath.org identifier of work a MSC_a (string): primary MSC5 of work a MSC2_a (list(string)): secondary MSC5s of work a an_b (string): zbMath.org identifier of work b MSC_b (string): primary MSC5 of work b MSC2_b (list(string)): secondary… See the full description on the dataset page: https://huggingface.co/datasets/math-similarity/class-zbmath-identifier.tabulartext-classification100K<n<1M0 likes18 downloads2y agoHugging Face13werty1248 /sentence-transformer-parallel-En-Ko-with-Similarity Preprocessing En-Ko subset of Parallel Sentences Datasets 해외에서 제작된 많은 대규모 번역 쌍 데이터들이 영어 텍스트와 한국어 텍스트를 문장 단위로 분리한 후 기계적으로 매핑시키고 있습니다. 이로 인해 전혀 엉뚱한 문장이 번역 쌍으로 매칭되어 있는 문제가 발생합니다. 임베딩 유사도 기반 전처리를 통해 이 문제를 해결할 수 있을 것 같아서 이 데이터 셋을 제작했습니다. 일부 데이터는 상업적 사용이 어려운 라이선스가 적용된 경우가 있습니다. 데이터 목록 parallel-sentences-global-voices parallel-sentences-muse parallel-sentences-jw300 parallel-sentences-tatoeba parallel-sentences-wikititles 유사도 측정 BAAI/BGE-m3로 임베딩 영어 문장과 한국어 문장의… See the full description on the dataset page: https://huggingface.co/datasets/werty1248/sentence-transformer-parallel-En-Ko-with-Similarity.tabulartranslation1M<n<10M1 likes18 downloads2y agoHugging Face14alete-ai /web-cache-similarity-benchmark Web Cache Semantic Similarity Benchmark This dataset was created by Alete to benchmark the effectiveness of modern semantic vector caching against traditional cryptographic string hashing (SHA-256) directly on the live web, evaluating temporal cache decay across multiple time frames in a uniform environment. Naïve string hashing fails on the modern web because minor changes in page layouts—such as dynamic trending articles list, dynamic comments count, ad units, and modified… See the full description on the dataset page: https://huggingface.co/datasets/alete-ai/web-cache-similarity-benchmark.tabulartext-classificationn<1K1 likes18 downloads2mo agoHugging Face15colabfit /Finding_new_crystal_compounds_using_chemical_similarity Cite this dataset Wang, H., Botti, S., and Marques, M. A. L. Finding new crystal compounds using chemical similarity. ColabFit, 2025. https://doi.org/10.60732/b9e7eedf This dataset has been curated and formatted for the ColabFit Exchange This dataset is also available on the ColabFit Exchange: https://materials.colabfit.org/id/DS_sk8zwvk3qxur_0 Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Finding_new_crystal_compounds_using_chemical_similarity.tabular100K<n<1M0 likes17 downloads11mo agoHugging Face16guanning-ai /DeepSeek-7B_aime24_similarity_0902tabular1K<n<10K0 likes15 downloads1y agoHugging Face17StephanAkkerman /orthographic-similarity-ratings Orthographic similarity ratings for English-Spanish cognates from the academic word list This dataset is derived from the paper "Orthographic similarity ratings for English-Spanish cognates from the academic word list" by Hout et al. You can find there paper here and the original dataset here. tabularn<1K0 likes14 downloads2y agoHugging Face18guanning-ai /DeepSeek-1.5B_dapo2k_similarity_0907tabular10K<n<100K0 likes14 downloads1y agoHugging Face19guanning-ai /Qwen3-1.7B_aime24_similarity_0902tabular1K<n<10K0 likes11 downloads1y agoHugging Face20Pushpendra817 /With-Similarity-Scorestabular1K<n<10K1 likes10 downloads3y agoHugging Face21davidberenstein1957 /rag-embeddings-relevance-similaritytabular1K<n<10K0 likes10 downloads2y agoHugging Face22ada-datadruids /bookfilm_summaries_with_similarity_and_sentimenttabularn<1K0 likes9 downloads2y agoHugging Face23316usman /mashqa-cosine-similarity-bgetabularn<1K0 likes9 downloads2y agoHugging Face24ChenWu98 /NuminaMath-CoT-2048-cosine-similarity-ranktabular100K<n<1M0 likes9 downloads1y agoHugging Face25TheMrguiller /ETD_Detoxification_Dataset_similarity_metricstabular100K<n<1M0 likes9 downloads8mo agoHugging Face26alex-miller /crs-2014-2023-housing-similaritytabular100K<n<1M0 likes8 downloads2y agoHugging Face27youssefkhalil320 /cat_names_similarity_pairs_three_scores_with_cat_scores_fulltabular100K<n<1M0 likes8 downloads1y agoHugging Face28Adanato /arxiv_similarity_300tabular10K<n<100K0 likes8 downloads11mo agoHugging Face29asingh15 /arc-barc-processed-diagnostic-insights-similarity-rubrictabularn<1K0 likes8 downloads8mo agoHugging Face30quasar529 /laion400m_top10000_similarityimage10K<n<100K0 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.