datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Test_Semantic_Searchdoorkey-semantic-reasoning-labels
GPT-OSS-20B DoorKey semantic reasoning labels
This dataset contains automatic sentence-level semantic-function annotations
for 7,038 reasoning sentences produced by openai/gpt-oss-20b on 46 fixed
DoorKey environment states. Each target sentence is paired with its preceding
reasoning context and assigned one or more human-readable discourse labels.
The annotation run produced 7,036 valid rows and two schema failures. These are
model-generated exploratory annotations, not human… See the full description on the dataset page: https://huggingface.co/datasets/project-telos/doorkey-semantic-reasoning-labels.Thai-Semantic-Textual-Similarity-BenchmarkSentence representation plays a crucial role in NLP downstream tasks such as NLI, text classification, and STS. Recent sentence representation training techniques require NLI or STS datasets. However, there are no equivalent Thai NLI or STS datasets for sentence representation training.
To address this problem we provide the Thai sentence vector benchmark. We evaluate the Spearman correlation score of the sentence representations’ performance on Thai STS-B (translated version of STS-B).… See the full description on the dataset page: https://huggingface.co/datasets/mrp/Thai-Semantic-Textual-Similarity-Benchmark.semantic-montecarlo-benchmark
Semantic Monte Carlo Benchmark
A synthetic benchmark of numeric research and forecasting questions for
evaluating the
semantic-montecarlo
pipeline.
This release contains only benchmark inputs. Cached experiments, individual
run artifacts, and aggregate results are intentionally excluded.
At a glance
Questions
Language
Splits
License
300
English
Validation and test
CC0 1.0
Dataset structure
The dataset has no training split:… See the full description on the dataset page: https://huggingface.co/datasets/cynosural/semantic-montecarlo-benchmark.Semantic-KG
Dataset Card for Semantic-KG
This is a dataset for evaluating semantic similarity containing semantically similar and dissimilar natural language statement pairs generated from knowledge graphs across four domains: general knowledge, biomedicine, finance, and biology.
Dataset Details
Dataset Description
This dataset is designed to evaluate semantic-textual similarity (STS) methods across diverse domains. It is generated using the Semantic-KG framework, an… See the full description on the dataset page: https://huggingface.co/datasets/QiyaoWei/Semantic-KG.lab05-semantic-searchSemanticTextualSimilarityDataset
Semantic Textual Similarity (STS) Dataset (Turkish)
This repository contains a Turkish Semantic Textual Similarity (STS) dataset created as part of a university assignment on semantic similarity, sentence embeddings, and vector representations in Natural Language Processing (NLP).
Authors
Muhammet Enes Nas
Salih Dede
About
The purpose of this project was to gain practical experience with:
Semantic Textual Similarity (STS)
Sentence Embeddings… See the full description on the dataset page: https://huggingface.co/datasets/menesnas/SemanticTextualSimilarityDataset.AI_Articles_Scraped_from_arXiv-Semantic_Scholar
📘 AI Articles Scraped from arXiv & Semantic Scholar
🧩 Description
This dataset contains information on articles related to major AI conferences such as AAAI, NeurIPS, IJCAI, ICML, ICLR, collected through scraping from ArXiv and Semantic Scholar.It is intended to be used as a training dataset for various model training tasks and other desired uses.
📂 File Structure
File
Description
AI_Titles_v2025.csv
Main dataset
README.md
This file… See the full description on the dataset page: https://huggingface.co/datasets/d-e-c-d/AI_Articles_Scraped_from_arXiv-Semantic_Scholar.semantic-consistency-base-predsemantic-memessemantic-feature-production-normssemantic-memeslab05-semantic-searchsemantic-scholar-datasciencecustomerservice_qa_semantic_lexical_resultssemantic-collapse-mitigation-pilot
Pilot: Semantic Collapse Mitigation via Context Scaffolding (n=2)
1. Abstract
This repository archives the raw data and empirical findings from an exploratory pilot (n=2) testing the efficacy of "Context Scaffolding" (specifically, the systematic injection of Identity and Context constraints) against AI-induced semantic collapse.
When foundational models operate without specific directorial constraints, they statistically converge toward the "automated average"… See the full description on the dataset page: https://huggingface.co/datasets/danicollada/semantic-collapse-mitigation-pilot.semantic_prime_T5_predictions_and_metricslab05-semantic-searchlab05-semantic-search
