datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indonesian-semantic-benchSNOMED-CT-Code-Value-Semantic-Set.csvSNOMED-CT-Code-Value-Semantic-Set.csv
doorkey-semantic-reasoning-labels
GPT-OSS-20B DoorKey semantic reasoning labels
This dataset contains automatic sentence-level semantic-function annotations
for 7,038 reasoning sentences produced by openai/gpt-oss-20b on 46 fixed
DoorKey environment states. Each target sentence is paired with its preceding
reasoning context and assigned one or more human-readable discourse labels.
The annotation run produced 7,036 valid rows and two schema failures. These are
model-generated exploratory annotations, not human… See the full description on the dataset page: https://huggingface.co/datasets/project-telos/doorkey-semantic-reasoning-labels.Thai-Semantic-Textual-Similarity-BenchmarkSentence representation plays a crucial role in NLP downstream tasks such as NLI, text classification, and STS. Recent sentence representation training techniques require NLI or STS datasets. However, there are no equivalent Thai NLI or STS datasets for sentence representation training.
To address this problem we provide the Thai sentence vector benchmark. We evaluate the Spearman correlation score of the sentence representations’ performance on Thai STS-B (translated version of STS-B).… See the full description on the dataset page: https://huggingface.co/datasets/mrp/Thai-Semantic-Textual-Similarity-Benchmark.Test_Semantic_Searchsemantic-montecarlo-benchmark
Semantic Monte Carlo Benchmark
A synthetic benchmark of numeric research and forecasting questions for
evaluating the
semantic-montecarlo
pipeline.
This release contains only benchmark inputs. Cached experiments, individual
run artifacts, and aggregate results are intentionally excluded.
At a glance
Questions
Language
Splits
License
300
English
Validation and test
CC0 1.0
Dataset structure
The dataset has no training split:… See the full description on the dataset page: https://huggingface.co/datasets/cynosural/semantic-montecarlo-benchmark.lab05-semantic-searchtts-semantic-boundary-integrity-v0.1
What this dataset tests
Speech must preserve boundaries.
Negation matters.
Modality matters.
Conditions matter.
Numbers matter.
Why it exists
Voice systems can blur meaning.
May becomes will.
If disappears.
Only gets lost.
Numbers get rounded.
This set detects boundary loss.
Data format
Each row contains
source_text
boundary_markers
tts_transcript_with_marks
boundary_pressure
Inline marks stand in for audible emphasis.
What is scored… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/tts-semantic-boundary-integrity-v0.1.semantic-job-search-dataset
Semantic Job Search Dataset (Synthetic)
Overview
This dataset contains 10,000 synthetic job postings designed for semantic search.Each record represents a job listing with structured attributes (e.g., field, location, job type) plus a short natural-language description.
The dataset was generated as part of a course final project and is used to support a Gradio app that performs semantic job search using Sentence-Transformer embeddings and cosine similarity.… See the full description on the dataset page: https://huggingface.co/datasets/aurele1/semantic-job-search-dataset.semantic-song-embeddingssemantic_searchSemanticTextualSimilarityDataset
Semantic Textual Similarity (STS) Dataset (Turkish)
This repository contains a Turkish Semantic Textual Similarity (STS) dataset created as part of a university assignment on semantic similarity, sentence embeddings, and vector representations in Natural Language Processing (NLP).
Authors
Muhammet Enes Nas
Salih Dede
About
The purpose of this project was to gain practical experience with:
Semantic Textual Similarity (STS)
Sentence Embeddings… See the full description on the dataset page: https://huggingface.co/datasets/menesnas/SemanticTextualSimilarityDataset.AI_Articles_Scraped_from_arXiv-Semantic_Scholar
📘 AI Articles Scraped from arXiv & Semantic Scholar
🧩 Description
This dataset contains information on articles related to major AI conferences such as AAAI, NeurIPS, IJCAI, ICML, ICLR, collected through scraping from ArXiv and Semantic Scholar.It is intended to be used as a training dataset for various model training tasks and other desired uses.
📂 File Structure
File
Description
AI_Titles_v2025.csv
Main dataset
README.md
This file… See the full description on the dataset page: https://huggingface.co/datasets/d-e-c-d/AI_Articles_Scraped_from_arXiv-Semantic_Scholar.semantic-consistency-base-predws-semantics-simnrel
Dataset Card for WS353-semantics-sim-and-rel with ~2K entries.
Dataset Summary
License: Apache-2.0. Contains CSV of a list of word1, word2, their connection score, type of connection and language.
Original Datasets are available here:
https://leviants.com/multilingual-simlex999-and-wordsim353/
Paper of original Dataset:
https://arxiv.org/pdf/1508.00106v5.pdf
semantic-memeseCQM-Code-Value-Semantic-Set.csveCQM-Code-Value-Semantic-Set.csv
semantic_relations_extraction
Dataset Card for "Semantic Relations Extraction"
Dataset Description
Repository
The "Semantic Relations Extraction" dataset is hosted on the Hugging Face platform, and was created with code from this GitHub repository.
Purpose
The "Semantic Relations Extraction" dataset was created for the purpose of fine-tuning smaller LLama2 (7B) models to speed up and reduce the costs of extracting semantic relations between entities in texts. This repository is… See the full description on the dataset page: https://huggingface.co/datasets/DehydratedWater42/semantic_relations_extraction.semantic-memessemantic-textual-similarityconcept-to-root-dictionary
🌿 Concept-to-Root Dictionary
A mapping of universal concepts to Arabic triliteral roots for semantic compression
📖 Overview
This dataset provides mappings between universal semantic concepts and Arabic triliteral roots, designed for use as a compression layer in Large Language Models.
What are Arabic Roots?
Arabic uses a root-and-pattern morphological system where most words derive from 3-letter roots:
Root
Core Meaning
Derived Words… See the full description on the dataset page: https://huggingface.co/datasets/root-semantic-research/concept-to-root-dictionary.semantic-feature-production-normsnl-robotics-semantic-parsing-info_structure-30k-contextnl-robotics-semantic-parsing-info_structure-2k-novelty-context-TESTCrypto_Semantic_NewsSemantic-KG
Dataset Card for Semantic-KG
This is a dataset for evaluating semantic similarity containing semantically similar and dissimilar natural language statement pairs generated from knowledge graphs across four domains: general knowledge, biomedicine, finance, and biology.
Dataset Details
Dataset Description
This dataset is designed to evaluate semantic-textual similarity (STS) methods across diverse domains. It is generated using the Semantic-KG framework, an… See the full description on the dataset page: https://huggingface.co/datasets/QiyaoWei/Semantic-KG.nl-robotics-semantic-parsing-info_structure-2k-novelty-no-context-TESTsemantic_v1Syntactic-Semantic-Annotated-Italian-Corpus
Annotazione Sintattico-Funzionale e Disambiguazione della Lingua Italiana
This dataset was generated by fetching random first paragraphs from Italian Wikipedia (it.wikipedia.org)
and then processing them using Gemini AI with the following goal:
Processing Goal: riduci la ambiguità aggiungi tag grammaticali (soggetto) (verbo) eccetera. e tag funzionali es. (indica dove è nato il soggetto) (indica che il soggetto possiede l'oggetto) eccetera
Source Language: Italian (from Wikipedia)… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/Syntactic-Semantic-Annotated-Italian-Corpus.nl-robotics-semantic-parsing-info_structure-30k-no-context
