datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
linguistic-similarityheadlines-semantic-similarity
Dataset Card for HEADLINES
Dataset Summary
HEADLINES is a massive English-language semantic similarity dataset, containing 396,001,930 pairs of different headlines for the same newspaper article, taken from historical U.S. newspapers, covering the period 1920-1989.
Languages
The text in the dataset is in English.
Dataset Structure
Each year in the dataset is divided into a distinct file (eg. 1952_headlines.json), giving a total of 70 files.
The… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/headlines-semantic-similarity.x-span-similarity
X-Span-Similarity (X-SSD)
Expanding Lozano et al.'s (2026) Span Similarity Dataset (SSD) into a cross-lingual setting, for Dissimilar/Difference Span Detection (DSD) across language pairs.
Dataset summary
Each row is a premise/hypothesis sentence pair, one side in English and the other machine-translated into another target language, with span-level and sentence-level (dis)similarity labels carried over unchanged from the original English SSD annotation. Spans… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/x-span-similarity.lm-similarity
Great Models Think Alike and this Undermines AI Oversight
This is the data collection for the publication "Great Models Think Alike and this Undermines AI Oversight."
judge_scores_mmlu_pro_free_filtered: Judge scores of nine judges without access to the reference answers on the filtered, open-style MMLU-Pro dataset.
judge_w_gt_mmlu_pro_free_filtered: Ensemble judge scores of five judges with access to the reference options and ground-truth information on the filtered OSQ MMLU-Pro.… See the full description on the dataset page: https://huggingface.co/datasets/bethgelab/lm-similarity.human-aligned-similarity-benchmark
Human Aligned Similarity Benchmark
You are welcome to go to alignedmachine.com to contribute.
Overview
This dataset contains human-aligned similarity judgments for embedding text and multimodal AI model evaluation. The benchmark is designed to assess how well AI models align with human cognitive preferences in similarity perception across text and image modalities.
Dataset Structure
Concept Files
This dataset contains human preference judgments for… See the full description on the dataset page: https://huggingface.co/datasets/duke-trust-lab/human-aligned-similarity-benchmark.arXiv-metadata-oai-snapshot-111gpt4-instruct-similarity-0.9-dataset_yourgpt
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/mrcuddle/gpt4-instruct-similarity-0.9-dataset_yourgpt.MedSwin-Passage-SimilarityCleaned and curated dataset specifically used for medical semantic similarity and passage comparison (pos/neg passages). Ideal for finetuning (long-context) on:
Biomedical Reranker
Biomedical Embedding
Data Collection
1) BioASQ (Generated Queries)
Used as: (query, document) positives; negatives sampled from rolling buffer.
Specialised to handle the complex terminology and high precision required for Task B (Biomedical Semantic QA). The reranker acts as a critical… See the full description on the dataset page: https://huggingface.co/datasets/MedSwin/MedSwin-Passage-Similarity.Product_Similarity_DatasetThis following dataset is a rich dataset of product similarity. The dataset has been design to be challenging to train on by having quite a lot of hard negatives
This dataset is especially targeted toward fine-tuning usecase, especially to finetune reranker or embedding model.
The data are especially adapted for listwise loss like LambdaLoss or ListNetLoss.
The data are in JSONL and each line follow the same format as here below :
A "query", the anchor product label
"docs", the potential… See the full description on the dataset page: https://huggingface.co/datasets/Antix5/Product_Similarity_Dataset.Similarity-matching-of-Chinese-disease-questions本数据来自平安医疗科技疾病问答
Given question pairs from five different disease types, we are required to determine whether the semantics of the two sentences are the same or similar.
骆迅,倪渊,汤步洲,雷健波. 基于竞赛视角探讨文本语义匹配技术在中文医学文本领域中的应用 [J]. 中国数字医学. 2021 (11)
bigbird-legal-dataset-after-similarity-score-addingSemantic_similarity_deduplicated_reasoning_data_english
Semantic_similarity_deduplicated_reasoning_data_english
数据集描述
Semantic similarity deduplicated reasoning data filtered from OpenThoughts2-1M, 77662 examples in total, 10000 examples for each category
文件结构
semantic_similarity_deduplicated_reasoning_data_english.jsonl: 主数据文件(JSONL格式)
数据格式
数据集包含以下字段:
question: str
quality: int
difficulty: int
topic: str
validity: int
使用方法
方法1: 使用datasets库
from datasets import load_dataset
#… See the full description on the dataset page: https://huggingface.co/datasets/Ibisbill/Semantic_similarity_deduplicated_reasoning_data_english.Champion-Similarity-v6-c2cp_tfChampion-Similarity-v6-ogstring-similarityinlegal-laysum-for-t5-with-similarity-scoreAGCI_Similarity_DatasetAI-Paper-Similarity-SearchChampion-Similarity-v5Champion-Similarity-v6Champion-Similarity-v6-og_c2cpONC_Question_Similarity_And_Matching
