CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01chendelong /linguistic-similaritytabularn<1K1 likes706 downloads2y agoHugging Face02dell-research-harvard /headlines-semantic-similarity Dataset Card for HEADLINES Dataset Summary HEADLINES is a massive English-language semantic similarity dataset, containing 396,001,930 pairs of different headlines for the same newspaper article, taken from historical U.S. newspapers, covering the period 1920-1989. Languages The text in the dataset is in English. Dataset Structure Each year in the dataset is divided into a distinct file (eg. 1952_headlines.json), giving a total of 70 files. The… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/headlines-semantic-similarity.textsentence-similarity10M<n<100M12 likes520 downloads2y agoHugging Face03ZurichNLP /x-span-similarity X-Span-Similarity (X-SSD) Expanding Lozano et al.'s (2026) Span Similarity Dataset (SSD) into a cross-lingual setting, for Dissimilar/Difference Span Detection (DSD) across language pairs. Dataset summary Each row is a premise/hypothesis sentence pair, one side in English and the other machine-translated into another target language, with span-level and sentence-level (dis)similarity labels carried over unchanged from the original English SSD annotation. Spans… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/x-span-similarity.texttext-classification100K<n<1M1 likes366 downloads2mo agoHugging Face04bethgelab /lm-similarity Great Models Think Alike and this Undermines AI Oversight This is the data collection for the publication "Great Models Think Alike and this Undermines AI Oversight." judge_scores_mmlu_pro_free_filtered: Judge scores of nine judges without access to the reference answers on the filtered, open-style MMLU-Pro dataset. judge_w_gt_mmlu_pro_free_filtered: Ensemble judge scores of five judges with access to the reference options and ground-truth information on the filtered OSQ MMLU-Pro.… See the full description on the dataset page: https://huggingface.co/datasets/bethgelab/lm-similarity.tabularquestion-answering10K<n<100K5 likes266 downloads2y agoHugging Face05duke-trust-lab /human-aligned-similarity-benchmark Human Aligned Similarity Benchmark You are welcome to go to alignedmachine.com to contribute. Overview This dataset contains human-aligned similarity judgments for embedding text and multimodal AI model evaluation. The benchmark is designed to assess how well AI models align with human cognitive preferences in similarity perception across text and image modalities. Dataset Structure Concept Files This dataset contains human preference judgments for… See the full description on the dataset page: https://huggingface.co/datasets/duke-trust-lab/human-aligned-similarity-benchmark.textn<1K0 likes179 downloads9mo agoHugging Face06math-similarity /arXiv-metadata-oai-snapshot-111text1M<n<10M0 likes133 downloads2y agoHugging Face07mrcuddle /gpt4-instruct-similarity-0.9-dataset_yourgpt Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/mrcuddle/gpt4-instruct-similarity-0.9-dataset_yourgpt.text10K<n<100K0 likes46 downloads2y agoHugging Face08MedSwin /MedSwin-Passage-SimilarityCleaned and curated dataset specifically used for medical semantic similarity and passage comparison (pos/neg passages). Ideal for finetuning (long-context) on: Biomedical Reranker Biomedical Embedding Data Collection 1) BioASQ (Generated Queries) Used as: (query, document) positives; negatives sampled from rolling buffer. Specialised to handle the complex terminology and high precision required for Task B (Biomedical Semantic QA). The reranker acts as a critical… See the full description on the dataset page: https://huggingface.co/datasets/MedSwin/MedSwin-Passage-Similarity.textquestion-answering100K<n<1M1 likes45 downloads8mo agoHugging Face09Antix5 /Product_Similarity_DatasetThis following dataset is a rich dataset of product similarity. The dataset has been design to be challenging to train on by having quite a lot of hard negatives This dataset is especially targeted toward fine-tuning usecase, especially to finetune reranker or embedding model. The data are especially adapted for listwise loss like LambdaLoss or ListNetLoss. The data are in JSONL and each line follow the same format as here below : A "query", the anchor product label "docs", the potential… See the full description on the dataset page: https://huggingface.co/datasets/Antix5/Product_Similarity_Dataset.texttext-ranking10K<n<100K0 likes35 downloads1y agoHugging Face10whalning /Similarity-matching-of-Chinese-disease-questions本数据来自平安医疗科技疾病问答 Given question pairs from five different disease types, we are required to determine whether the semantics of the two sentences are the same or similar. 骆迅,倪渊,汤步洲,雷健波. 基于竞赛视角探讨文本语义匹配技术在中文医学文本领域中的应用 [J]. 中国数字医学. 2021 (11) text10K<n<100K0 likes29 downloads2y agoHugging Face11mohitskaushal /bigbird-legal-dataset-after-similarity-score-addingtext10K<n<100K0 likes27 downloads6mo agoHugging Face12Ibisbill /Semantic_similarity_deduplicated_reasoning_data_english Semantic_similarity_deduplicated_reasoning_data_english 数据集描述 Semantic similarity deduplicated reasoning data filtered from OpenThoughts2-1M, 77662 examples in total, 10000 examples for each category 文件结构 semantic_similarity_deduplicated_reasoning_data_english.jsonl: 主数据文件(JSONL格式) 数据格式 数据集包含以下字段: question: str quality: int difficulty: int topic: str validity: int 使用方法 方法1: 使用datasets库 from datasets import load_dataset #… See the full description on the dataset page: https://huggingface.co/datasets/Ibisbill/Semantic_similarity_deduplicated_reasoning_data_english.texttext-generation10K<n<100K0 likes23 downloads1y agoHugging Face13avinot /Champion-Similarity-v6-c2cp_tftext1K<n<10K1 likes23 downloads4mo agoHugging Face14avinot /Champion-Similarity-v6-ogtext1K<n<10K0 likes17 downloads4mo agoHugging Face15locuslab /string-similaritytextn<1K0 likes14 downloads2y agoHugging Face16mohitskaushal /inlegal-laysum-for-t5-with-similarity-scoretext10K<n<100K0 likes11 downloads6mo agoHugging Face17AAAndyZ /AGCI_Similarity_Datasetimage10K<n<100K0 likes9 downloads1y agoHugging Face18napronald /AI-Paper-Similarity-Searchtext10K<n<100K2 likes7 downloads2y agoHugging Face19avinot /Champion-Similarity-v5text10K<n<100K0 likes6 downloads11mo agoHugging Face20avinot /Champion-Similarity-v6text10K<n<100K0 likes5 downloads4mo agoHugging Face21avinot /Champion-Similarity-v6-og_c2cptext1K<n<10K0 likes2 downloads4mo agoHugging Face22lucabolzonello /ONC_Question_Similarity_And_Matchingtabularn<1K0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.