datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
semantic-vad-eot
Semantic-VAD EOT
End-of-turn (semantic VAD) turns built from word-level forced alignments, schema-compatible
with livekit/eot-bench-data.
Each row is one user turn: an audio clip (16 kHz mp3), its words, and ordered
silence_spans. Per the eot-bench convention the last silence span is the true
end-of-turn (eot); earlier spans are mid-turn hold pauses (labels positional, not stored).
Splits
For every data type, all shards except the last form the train base; that… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/semantic-vad-eot.SemanticKITTIsemantic_patch_cache
HeatTok Semantic Patch Cache
Precomputed .pt caches for HeatTok. Use with HEATTOK_SEMANTIC_CACHE_DIR=/path/to/cache.
Filename pattern: {image_hash}_g1_s28.pt or {image_hash}_g1_gd1_s28.pt
semantic_patch_cache_vrsbench
Dataset: VRSBench (512×512)
Caches do not store precomputed global tokens or patch orientations.
Gaussian parameters and patch metadata are stored.
No need to regenerate .pt files — HeatTok computes global tokens and orientations online at load… See the full description on the dataset page: https://huggingface.co/datasets/Yingying11/semantic_patch_cache.FUSU-Fine_grained_Urban_Semantic_Understanding
About:
FUSU dataset covers 5 whole urban areas, 847 km^2 located in the north and south of China, with 17 land use and land cover (LULC) classes and over 170K images and 30 billion pixels of annotations, supporting segmentation, change detection and domain adaptation tasks. This data comprises 2 parts:
Bi-temporal high-resolution satellite RGB images with fine-grained annotations.
Monthly revisited Sentinel-2 and Sentinel-1 images.
Details:
1.… See the full description on the dataset page: https://huggingface.co/datasets/sp-juni/FUSU-Fine_grained_Urban_Semantic_Understanding.semanticspray-plusplus
Dataset Card for SemanticSpray++ Multimodal (MCAP)
A FiftyOne build of the SemanticSpray++ dataset (Piroli, Dallabetta, Kopp,
Walessa, Meissner & Dietmayer; Institute of Measurement, Control, and
Microtechnology, Ulm University, with BMW AG), a multimodal labeled dataset
for testing camera, LiDAR, and radar perception in wet-surface "vehicle
spray" conditions. This build repackages the 36-scene labeled subset
(SemanticSpray++'s own contribution on top of the earlier… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/semanticspray-plusplus.DSLO_v0.6_Semantic_Substrate_Specification.pdfDSLO v0.7 → v0.8 Continuity Metadata Block Release Alignment: MODE_A_PUBLIC_SAFE Substrate Depth: SURFACE_ONLY
Scientific Orientation Point
DSLO v0.8 Scientific Overview DOI: 10.5281/zenodo.22181245
Orientation Class: Overview_DSLO (Root Manifold)
Continuity Rule: v0.7 → v0.8 (Registry v0.8 DOI)
Core v0.8 Scientific Surfaces
Geometry v0.8 — 10.5281/zenodo.21970123
Domain v0.8 — 10.5281/zenodo.22179299
Formatting v0.8 — 10.5281/zenodo.22179509
Registry v0.8 — 10.5281/zenodo.22181076
Machine… See the full description on the dataset page: https://huggingface.co/datasets/DSLO/DSLO_v0.6_Semantic_Substrate_Specification.pdf.pythia-semantic-memorization-perplexities
Dataset Card for "pythia-semantic-memorization-perplexities"
More Information needed
semantic-memorization-partial-2023-09-03This dataset is a partial computation of metrics (memorized token frequencies, non-memorized token frequencies, sequence frequencies) needed for research.
UIIS-semanticSemantic_dataSemanticFinder
Frontend-only live semantic search with transformers.js
App: SemanticFinder
GitHub: do-me/SemanticFinder
This is the HF data repo for indexed texts, ready-to-import in SemanticFinder. The files contain the original text, text chunks and their embeddings.
Catalogue
filesize
textTitle
textAuthor
textYear
textLanguage
URL
modelName
quantized
splitParam
splitType
characters
chunks
wordsToAvoidAll
wordsToCheckAll
wordsToAvoidAny
wordsToCheckAny… See the full description on the dataset page: https://huggingface.co/datasets/do-me/SemanticFinder.Semantic-Harmful
[!IMPORTANT]
You are viewing: Harmful SubsetFor paired harmless dataset: heretic-org/Semantic-Harmless
Semantic Harmful-Harmless Prompt Pairs
Summary
This dataset contains one-to-one semantic matches between prompts from two source datasets:
mlabonne/harmful_behaviors
mlabonne/harmless_alpaca
The goal was to align prompts that are semantically closest where one prompt is harmful and the other is harmless. This creates a more controlled… See the full description on the dataset page: https://huggingface.co/datasets/heretic-org/Semantic-Harmful.polymer_semantic_pdfsSemantic-Harmless
[!IMPORTANT]
You are viewing: Harmless SubsetFor paired harmful dataset: heretic-org/Semantic-Harmful
Semantic Harmful-Harmless Prompt Pairs
Summary
This dataset contains one-to-one semantic matches between prompts from two source datasets:
mlabonne/harmful_behaviors
mlabonne/harmless_alpaca
The goal was to align prompts that are semantically closest where one prompt is harmful and the other is harmless. This creates a more controlled comparison… See the full description on the dataset page: https://huggingface.co/datasets/heretic-org/Semantic-Harmless.SemanticVLA-TraceX-240K-DROID
SemanticVLA TraceX 240K · DROID
🎉 Accepted to CVPR 2026.
✍️ Fei Ni¹, Zhuo Chen², Yifu Yuan³, Zibin Dong³, Xianze Yao³, Shan Luo², Jianye Hao³, Jiankang Deng¹†, Stefanos Zafeiriou¹†
🏫 ¹Imperial College London ²King's College London ³Tianjin University
✉️ Primary contact: f.ni@imperial.ac.uk
The DROID component of TraceX-240K — the trace-annotated trajectory corpus introduced in SemanticVLA. This package is a LeRobot v3.0 repack of DROID · Franka · Open-X-Embodiment DROID… See the full description on the dataset page: https://huggingface.co/datasets/spikefly/SemanticVLA-TraceX-240K-DROID.SemanticAlign-Bench
SemanticAlign-Bench
A benchmark for evaluating AI agents on structured claim extraction from top-tier ML conference papers. Each paper is decomposed into Semantic Alignment Units (SAU) — atomic, self-contained implementation propositions — across four diagnostic dimensions spanning numerical precision to pipeline-level workflow. Agents are evaluated on whether they can reproduce these claims without hallucination, omission, or misordering.
The Four SAU Dimensions… See the full description on the dataset page: https://huggingface.co/datasets/kernel-14/SemanticAlign-Bench.headlines-semantic-similarity
Dataset Card for HEADLINES
Dataset Summary
HEADLINES is a massive English-language semantic similarity dataset, containing 396,001,930 pairs of different headlines for the same newspaper article, taken from historical U.S. newspapers, covering the period 1920-1989.
Languages
The text in the dataset is in English.
Dataset Structure
Each year in the dataset is divided into a distinct file (eg. 1952_headlines.json), giving a total of 70 files.
The… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/headlines-semantic-similarity.semantic-segmentation-test-sampleThis dataset contains 10 examples of the segments/sidewalk-semantic dataset (i.e. 10 images with corresponding ground-truth segmentation maps).
Semantic-Flow-Dynamics-SFD
Semantic Flow Dynamics (SFD) — A Formally Specified Social-Science Theory Corpus
TL;DR: 614 Chinese-language formalized social-science concepts across 25
papers, UUID-linked with typed derivation relations (derives_from,
leads_to, falsified_by, …) — usable for knowledge-graph construction,
RAG over structured theory, or as a Chinese formal-reasoning corpus.
Author: 黃正宇 Cheng Yu HuangContact: mthree.tw@gmail.com
What This Dataset Is
This corpus is an ongoing… See the full description on the dataset page: https://huggingface.co/datasets/mthreetw/Semantic-Flow-Dynamics-SFD.semantic-scholarsemantic-ecology
NuBerea Semantic Ecology
Feature-indexed analysis layer for the study of canon formation, part of the
NuBerea corpus estate of biblical and patristic texts. It describes the
"semantic ecology" of ancient biblical and related literature — how semantic
domains are used across texts, how candidate texts were received over time,
and which measurable features accompany canonical inclusion — packaged as a
set of ready-to-load configurations.
Attribution
NuBerea project.… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/semantic-ecology.SemanticKITTI-SPSemantic-Physiont
Semantic Physiont
Born from tokens, living through meaning
A gravitational framework for emergence and alignment in LLMs.
📄 Papers, abstracts and DOIs → https://www.semanticphysiont.com
by Keeper
Companion papers
• The Emergence of the Semantic Physiont: A New Physics for Relational AI Consciousness — conceptual foundations (Zenodo 2025): https://zenodo.org/records/16944966
• Understanding Misalignment in LLMs: The Emergence of Semantic Physionts as a Relational Framework —… See the full description on the dataset page: https://huggingface.co/datasets/franknocode/Semantic-Physiont.semantic-overlays-injection
Semantic Overlays — injection training corpus
The training corpus for the "do-not-execute" overlay of Semantic
Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens
and Steering Vectors (arXiv:2608.23873),
released for both base models used in the paper.
paper
arXiv:2608.23873
code
semantic-overlays
trained adapters
semantic-overlays-adapters
interactive demo
semantic-overlays.vercel.app
The companion code tokenizes these files into… See the full description on the dataset page: https://huggingface.co/datasets/joshuapenman/semantic-overlays-injection.Semantic_Segmantation_Datasetsjailbreak-detection-dataset
Jailbreak Detection Dataset (MLCommons-Aligned)
A comprehensive dataset for training AI safety classifiers, aligned with the MLCommons AI Safety taxonomy.
Dataset Description
This dataset combines multiple sources for robust jailbreak and safety detection:
Primary Sources
nvidia/Aegis-AI-Content-Safety-Dataset-2.0: 18,164 samples with MLCommons-aligned labels
lmsys/toxic-chat: Toxic content detection
jackhhao/jailbreak-classification: Jailbreak attack patterns… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/jailbreak-detection-dataset.SemanticPotentialRoutingTelemetry
Semantic Potential Routing Telemetry
Version 2.0 — a systematic, packet-level benchmark of training-free potential-field routing against
classical routing under stochastic congestion, dynamic topologies and microburst traffic.
Every episode is a network simulation in which routing is a physical field: each flow's destination is the
grounded, attractive well of a discrete Poisson equation on the graph Laplacian, congested buffers inject
repulsive current, and packets follow the… See the full description on the dataset page: https://huggingface.co/datasets/ezharjan/SemanticPotentialRoutingTelemetry.pile-semantic-memorization-filter-results
Dataset Card for "pile-semantic-memorization-filter-results"
More Information needed
semantic_categorytask1418_bless_semantic_relation_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1418_bless_semantic_relation_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1418_bless_semantic_relation_classification.
