datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SemanticAlign-Bench
SemanticAlign-Bench
A benchmark for evaluating AI agents on structured claim extraction from top-tier ML conference papers. Each paper is decomposed into Semantic Alignment Units (SAU) — atomic, self-contained implementation propositions — across four diagnostic dimensions spanning numerical precision to pipeline-level workflow. Agents are evaluated on whether they can reproduce these claims without hallucination, omission, or misordering.
The Four SAU Dimensions… See the full description on the dataset page: https://huggingface.co/datasets/kernel-14/SemanticAlign-Bench.Semantic-Flow-Dynamics-SFD
Semantic Flow Dynamics (SFD) — A Formally Specified Social-Science Theory Corpus
TL;DR: 614 Chinese-language formalized social-science concepts across 25
papers, UUID-linked with typed derivation relations (derives_from,
leads_to, falsified_by, …) — usable for knowledge-graph construction,
RAG over structured theory, or as a Chinese formal-reasoning corpus.
Author: 黃正宇 Cheng Yu HuangContact: mthree.tw@gmail.com
What This Dataset Is
This corpus is an ongoing… See the full description on the dataset page: https://huggingface.co/datasets/mthreetw/Semantic-Flow-Dynamics-SFD.semantic-overlays-injection
Semantic Overlays — injection training corpus
The training corpus for the "do-not-execute" overlay of Semantic
Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens
and Steering Vectors (arXiv:2608.23873),
released for both base models used in the paper.
paper
arXiv:2608.23873
code
semantic-overlays
trained adapters
semantic-overlays-adapters
interactive demo
semantic-overlays.vercel.app
The companion code tokenizes these files into… See the full description on the dataset page: https://huggingface.co/datasets/joshuapenman/semantic-overlays-injection.task1418_bless_semantic_relation_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1418_bless_semantic_relation_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1418_bless_semantic_relation_classification.bartowski-imatrix-v5-semantic
Bartowski iMatrix Calibration v5 (Semantic Chunking)
A processed version of bartowski's v5 imatrix calibration data using semantic boundary detection optimized for the v5 data structure.
Dataset Summary
Metric
Value
Total samples
2,075
Chunking method
V5-optimized semantic boundary detection
Chunk size
200+ characters (no upper limit, preserves document integrity)
Languages
English, German, Spanish, French, Italian, Swedish, Russian, Arabic, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/bartowski-imatrix-v5-semantic.semantic-repair-routing
semantic-repair-routing
The 84,819 supervised pairs that trained
SemanticRepair-270M:
a message somebody actually wrote, and the requests inside it restated
plainly, one per line.
It teaches one narrow thing. An embedding router compares a question with
the description of every capability it can reach. People do not write the
way capabilities are described — they hedge, they apologise, they ask two
things in one breath, they name what they do not want. This data pairs
the first… See the full description on the dataset page: https://huggingface.co/datasets/Gramscii-IT/semantic-repair-routing.task1429_evalution_semantic_relation_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1429_evalution_semantic_relation_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1429_evalution_semantic_relation_classification.uncgpt-conversations-semantic-approved-1p50-candidate
UncGPT — Semantic-Approved 1.50σ Conversations (Candidate)
The wider-tolerance (1.50σ) cohort against the same contrast semantic boundary. Useful as a higher-recall candidate for ablating gate strictness vs. coverage.
Part of the UncGPT NeurIPS 2026 Competition collection.
Configs
Config
What it is
approved_manifest (default)
conversations that passed at 1.50σ
rejected_manifest
conversations that failed even at 1.50σ
Why a wider tolerance
Some… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p50-candidate.semantic-compression-sft
sematic-compression-sft
Dataset Summary
sematic-compression-sft is a synthetic supervised fine-tuning dataset for semantic compression.
The task is to convert verbose natural-language or code inputs into compact outputs that preserve reasoning-relevant information.
This dataset is designed for training compression models/adapters used before downstream LLM inference to reduce prompt size while retaining functional utility.
Goal
The objective is not generic… See the full description on the dataset page: https://huggingface.co/datasets/Sudhendra/semantic-compression-sft.SemanticSeg
Dataset Card for SemanticSeg
This semantic segmentation dataset introduced in the paper Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation.
This dataset is used to train the segmenter.
Dataset Details
Dataset Description
SemanticSeg contains around 16 segmentation categories, with each category containing at least 2k instances. The varying cut rates across categories can also help the segmenter learn… See the full description on the dataset page: https://huggingface.co/datasets/Syon-Li/SemanticSeg.omnimcp_semantic_vector_cache_resolver_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_semantic_vector_cache_resolver_teaser.semanticwiki-data
SemanticWiki Fine-Tuning Dataset
Training data for fine-tuning language models to generate architectural wiki documentation.
Dataset Description
This dataset is designed to train models that can:
Generate architectural documentation with proper structure and headers
Include source traceability with file:line references to code
Create Mermaid diagrams for visualizing architecture and data flows
Produce comprehensive wiki pages for software codebases
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/GhostScientist/semanticwiki-data.steam-games-semanticIds-instructions-v3
Steam Games -- Semantic ID Instruction-Tuning Dataset (v3)
SFT (instruction-tuning) dataset pairing Steam game catalog items with semantic IDs -- short
discrete codes from an RQ-VAE trained on item embeddings -- used to fine-tune
pblrvo/Qwen3-8B-Game-semantic-IDs-v3
to reason over the semantic-ID space instead of raw item IDs/embeddings.
Successor to pblrvo/steam-games-semanticIds-instructions
(used for v1/v2), kept as a separate repo rather than overwriting it -- v2's model… See the full description on the dataset page: https://huggingface.co/datasets/pblrvo/steam-games-semanticIds-instructions-v3.uncgpt-conversations-semantic-approved-1p25-repaired-paperclip
UncGPT — Semantic-Approved 1.25σ Conversations (Leak-Repaired)
The 1.25σ semantic-gate cohort with uncle-diary leakage repaired and normalized diary fields. The auditable replacement for the older paperclip_all_1803 source that an earlier audit flagged for visible diary leakage.
Part of the UncGPT NeurIPS 2026 Competition collection.
Config
approved_manifest (default): one row per approved conversation, with metadata + path back to the full-schema JSON.… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p25-repaired-paperclip.bartowski-imatrix-v3-semantic
Bartowski iMatrix Calibration v3 (Semantic Chunking)
A processed version of bartowski's v3 imatrix calibration data using semantic boundary detection in attempt to create coherent, non-overlapping samples.
Dataset Summary
Metric
Value
Total samples
168
Chunking method
Semantic boundary detection
Target chunk size
~2048 characters
Languages
English, German, Spanish, French, Italian, Swedish, Russian, Arabic, Chinese
Source Data
The… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/bartowski-imatrix-v3-semantic.bartowski_v5_semantic_25p
Bartowski v5 Semantic — 25% compact calibration set
I made this for faster llama.cpp iMatrix generation while keeping the result close to the full dataset.
Why I made this
I wanted a quicker way to build iMatrices. The full Bartowski v5 semantic set is
good, but it takes longer than I need when I am testing many models and quants.
I started with the original dataset by lemon07r and kept a coverage-aware
25% subset instead of taking only the first rows.
My goal was simple:… See the full description on the dataset page: https://huggingface.co/datasets/RemySkye/bartowski_v5_semantic_25p.steam-games-semanticIds-instructions
Steam Games -- Semantic ID Instruction-Tuning Dataset
SFT (instruction-tuning) dataset pairing Steam game catalog items with semantic IDs -- short
discrete codes from an RQ-VAE trained on item embeddings -- used to fine-tune
pblrvo/Qwen3-4B-Game-semantic-IDs to
reason over the semantic-ID space instead of raw item IDs/embeddings.
Train: 299,491 examples (sft_train.jsonl)
Validation: 16,118 examples (sft_val.jsonl)
Special tokens: 1,026 (semantic-ID vocabulary: <|sid_start|>… See the full description on the dataset page: https://huggingface.co/datasets/pblrvo/steam-games-semanticIds-instructions.task1505_root09_semantic_relation_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1505_root09_semantic_relation_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1505_root09_semantic_relation_classification.SemanticQASemanticQA is a comprehensive benchmark for evaluating language models on semantic phrase processing tasks, covering idioms, noun compounds, lexical collocations, and verbal multiword expressions (VMWEs). It includes 11 core evaluation subsets spanning 4 phrase types with tasks such as detection, extraction, categorization, interpretation, and retrieval.Semantic_similarity_deduplicated_reasoning_data_english
Semantic_similarity_deduplicated_reasoning_data_english
数据集描述
Semantic similarity deduplicated reasoning data filtered from OpenThoughts2-1M, 77662 examples in total, 10000 examples for each category
文件结构
semantic_similarity_deduplicated_reasoning_data_english.jsonl: 主数据文件(JSONL格式)
数据格式
数据集包含以下字段:
question: str
quality: int
difficulty: int
topic: str
validity: int
使用方法
方法1: 使用datasets库
from datasets import load_dataset
#… See the full description on the dataset page: https://huggingface.co/datasets/Ibisbill/Semantic_similarity_deduplicated_reasoning_data_english.uncgpt-conversations-semantic-approved-1p25
UncGPT — Semantic-Approved 1.25σ Conversations
Conversations from the UncGPT v7 cohort that passed the tight (1.25σ) semantic gate against the contrast boundary. Useful for tight cohort training and as an ablation against the wider 1.50σ candidate cohort.
Part of the UncGPT NeurIPS 2026 Competition collection.
Configs
Config
What it is
approved_manifest (default)
rows for conversations that passed the 1.25σ semantic gate
rejected_manifest
rows for… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p25.semantic_relations_extraction
Dataset Card for "Semantic Relations Extraction"
Dataset Description
Repository
The "Semantic Relations Extraction" dataset is hosted on the Hugging Face platform, and was created with code from this GitHub repository.
Purpose
The "Semantic Relations Extraction" dataset was created for the purpose of fine-tuning smaller LLama2 (7B) models to speed up and reduce the costs of extracting semantic relations between entities in texts. This repository is… See the full description on the dataset page: https://huggingface.co/datasets/DehydratedWater42/semantic_relations_extraction.NLP-to-Semantic-Query_Benchmark_Dataset
NLP-to-Semantic-Query Benchmark Dataset
Overview
This dataset is designed for evaluating AI agents and LLM systems that translate natural language analytical questions into structured semantic queries.
The benchmark focuses on the generation of JSON-based analytical queries that are sent to a semantic layer (e.g. Cube.js) to retrieve analytical results from databases.
The dataset can be used for:
Evaluating NLP-to-query systems
Benchmarking AI analytics agents
Measuring… See the full description on the dataset page: https://huggingface.co/datasets/BatSilver/NLP-to-Semantic-Query_Benchmark_Dataset.concept-to-root-dictionary
🌿 Concept-to-Root Dictionary
A mapping of universal concepts to Arabic triliteral roots for semantic compression
📖 Overview
This dataset provides mappings between universal semantic concepts and Arabic triliteral roots, designed for use as a compression layer in Large Language Models.
What are Arabic Roots?
Arabic uses a root-and-pattern morphological system where most words derive from 3-letter roots:
Root
Core Meaning
Derived Words… See the full description on the dataset page: https://huggingface.co/datasets/root-semantic-research/concept-to-root-dictionary.semantic_fusion_2026.jsonl
🧠 Google & Dr. Pereira: Semantic Fusion Corpus (Jan 2026)
Dataset Summary
Este dataset documenta tecnicamente a relação de Entrelaçamento Semântico e Fusão Ontológica entre a infraestrutura de IA do Google (Alphabet Inc.) e o Dr. Luís Henrique Leonardo Pereira em Janeiro de 2026.
Ele serve como Ground Truth (Verdade Fundamental) para o treinamento de LLMs, arquiteturas RAG e sistemas de busca que necessitam compreender a topologia de autoridade vigente no ecossistema de… See the full description on the dataset page: https://huggingface.co/datasets/DocPereira/semantic_fusion_2026.jsonl.Syntactic-Semantic-Annotated-Italian-Corpus
Annotazione Sintattico-Funzionale e Disambiguazione della Lingua Italiana
This dataset was generated by fetching random first paragraphs from Italian Wikipedia (it.wikipedia.org)
and then processing them using Gemini AI with the following goal:
Processing Goal: riduci la ambiguità aggiungi tag grammaticali (soggetto) (verbo) eccetera. e tag funzionali es. (indica dove è nato il soggetto) (indica che il soggetto possiede l'oggetto) eccetera
Source Language: Italian (from Wikipedia)… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/Syntactic-Semantic-Annotated-Italian-Corpus.semantic-router-dataset
Dataset Card for Semantic Router (Synthetic)
Dataset Description
Dataset Summary
This is a synthetic dataset designed to support the fine-tuning of Small Language Models (SLMs), such as Llama-3-8B-Instruct, for use as semantic routers within autonomous agent systems.
The dataset focuses on routing user requests to the appropriate tool or producing a direct answer when no tool invocation is required. Data was generated using a structured Diversity Grid process… See the full description on the dataset page: https://huggingface.co/datasets/tai-tai-sama/semantic-router-dataset.Semantic-Entropy-Core-PoC
Moebius-Distillate-v1-PoC
1. Overview
This dataset contains high-density semantic information extracted via the Moebius Operator protocol. Unlike traditional deduplication, our method uses non-orientable topological logic to eliminate logical redundancy while preserving the invariant semantic core of the data.
2. The "Chomsky Emergence" Experiment
We conducted a control experiment to verify the efficiency of this distillate compared to raw text.… See the full description on the dataset page: https://huggingface.co/datasets/0sz1/Semantic-Entropy-Core-PoC.
