datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
semantic-vad-eot
Semantic-VAD EOT
End-of-turn (semantic VAD) turns built from word-level forced alignments, schema-compatible
with livekit/eot-bench-data.
Each row is one user turn: an audio clip (16 kHz mp3), its words, and ordered
silence_spans. Per the eot-bench convention the last silence span is the true
end-of-turn (eot); earlier spans are mid-turn hold pauses (labels positional, not stored).
Splits
For every data type, all shards except the last form the train base; that… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/semantic-vad-eot.Semantic-Harmful
[!IMPORTANT]
You are viewing: Harmful SubsetFor paired harmless dataset: heretic-org/Semantic-Harmless
Semantic Harmful-Harmless Prompt Pairs
Summary
This dataset contains one-to-one semantic matches between prompts from two source datasets:
mlabonne/harmful_behaviors
mlabonne/harmless_alpaca
The goal was to align prompts that are semantically closest where one prompt is harmful and the other is harmless. This creates a more controlled… See the full description on the dataset page: https://huggingface.co/datasets/heretic-org/Semantic-Harmful.Semantic-Harmless
[!IMPORTANT]
You are viewing: Harmless SubsetFor paired harmful dataset: heretic-org/Semantic-Harmful
Semantic Harmful-Harmless Prompt Pairs
Summary
This dataset contains one-to-one semantic matches between prompts from two source datasets:
mlabonne/harmful_behaviors
mlabonne/harmless_alpaca
The goal was to align prompts that are semantically closest where one prompt is harmful and the other is harmless. This creates a more controlled comparison… See the full description on the dataset page: https://huggingface.co/datasets/heretic-org/Semantic-Harmless.headlines-semantic-similarity
Dataset Card for HEADLINES
Dataset Summary
HEADLINES is a massive English-language semantic similarity dataset, containing 396,001,930 pairs of different headlines for the same newspaper article, taken from historical U.S. newspapers, covering the period 1920-1989.
Languages
The text in the dataset is in English.
Dataset Structure
Each year in the dataset is divided into a distinct file (eg. 1952_headlines.json), giving a total of 70 files.
The… See the full description on the dataset page: https://huggingface.co/datasets/dell-research-harvard/headlines-semantic-similarity.SemanticVLA-TraceX-240K-DROID
SemanticVLA TraceX 240K · DROID
🎉 Accepted to CVPR 2026.
✍️ Fei Ni¹, Zhuo Chen², Yifu Yuan³, Zibin Dong³, Xianze Yao³, Shan Luo², Jianye Hao³, Jiankang Deng¹†, Stefanos Zafeiriou¹†
🏫 ¹Imperial College London ²King's College London ³Tianjin University
✉️ Primary contact: f.ni@imperial.ac.uk
The DROID component of TraceX-240K — the trace-annotated trajectory corpus introduced in SemanticVLA. This package is a LeRobot v3.0 repack of DROID · Franka · Open-X-Embodiment DROID… See the full description on the dataset page: https://huggingface.co/datasets/spikefly/SemanticVLA-TraceX-240K-DROID.semantic-scholarSemantic-Flow-Dynamics-SFD
Semantic Flow Dynamics (SFD) — A Formally Specified Social-Science Theory Corpus
TL;DR: 614 Chinese-language formalized social-science concepts across 25
papers, UUID-linked with typed derivation relations (derives_from,
leads_to, falsified_by, …) — usable for knowledge-graph construction,
RAG over structured theory, or as a Chinese formal-reasoning corpus.
Author: 黃正宇 Cheng Yu HuangContact: mthree.tw@gmail.com
What This Dataset Is
This corpus is an ongoing… See the full description on the dataset page: https://huggingface.co/datasets/mthreetw/Semantic-Flow-Dynamics-SFD.semantic-ecology
NuBerea Semantic Ecology
Feature-indexed analysis layer for the study of canon formation, part of the
NuBerea corpus estate of biblical and patristic texts. It describes the
"semantic ecology" of ancient biblical and related literature — how semantic
domains are used across texts, how candidate texts were received over time,
and which measurable features accompany canonical inclusion — packaged as a
set of ready-to-load configurations.
Attribution
NuBerea project.… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/semantic-ecology.semantic-overlays-injection
Semantic Overlays — injection training corpus
The training corpus for the "do-not-execute" overlay of Semantic
Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens
and Steering Vectors (arXiv:2608.23873),
released for both base models used in the paper.
paper
arXiv:2608.23873
code
semantic-overlays
trained adapters
semantic-overlays-adapters
interactive demo
semantic-overlays.vercel.app
The companion code tokenizes these files into… See the full description on the dataset page: https://huggingface.co/datasets/joshuapenman/semantic-overlays-injection.SemanticPotentialRoutingTelemetry
Semantic Potential Routing Telemetry
Version 2.0 — a systematic, packet-level benchmark of training-free potential-field routing against
classical routing under stochastic congestion, dynamic topologies and microburst traffic.
Every episode is a network simulation in which routing is a physical field: each flow's destination is the
grounded, attractive well of a discrete Poisson equation on the graph Laplacian, congested buffers inject
repulsive current, and packets follow the… See the full description on the dataset page: https://huggingface.co/datasets/ezharjan/SemanticPotentialRoutingTelemetry.pile-semantic-memorization-filter-results
Dataset Card for "pile-semantic-memorization-filter-results"
More Information needed
jailbreak-detection-dataset
Jailbreak Detection Dataset (MLCommons-Aligned)
A comprehensive dataset for training AI safety classifiers, aligned with the MLCommons AI Safety taxonomy.
Dataset Description
This dataset combines multiple sources for robust jailbreak and safety detection:
Primary Sources
nvidia/Aegis-AI-Content-Safety-Dataset-2.0: 18,164 samples with MLCommons-aligned labels
lmsys/toxic-chat: Toxic content detection
jackhhao/jailbreak-classification: Jailbreak attack patterns… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/jailbreak-detection-dataset.task1418_bless_semantic_relation_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1418_bless_semantic_relation_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1418_bless_semantic_relation_classification.semantic-image-search-assetsdataset_v1_rawyodas2-mm-semanticEEG-semantic-text-relevanceWe release a novel dataset containing 23,270 time-locked (0.7s) word-level EEG recordings acquired from participants who read both text that was semantically relevant and irrelevant to self-selected topics.
The raw EEG data and the datasheet are available at https://osf.io/xh3g5/.
See code repository for benchmark results.
EEG data acquisition:
Explanations of the variables:
event corresponds to a specific point in time during EEG data collection and represents the onset of an event… See the full description on the dataset page: https://huggingface.co/datasets/Quoron/EEG-semantic-text-relevance.fact-check-classification-dataset
Fact-Check Classification Dataset
🎯 Purpose: Binary classification dataset for determining whether a prompt needs external fact-checking.
Dataset Description
This dataset is designed to train classifiers that can route LLM requests based on whether they require external fact verification. It's part of the vLLM Semantic Router project.
Labels
FACT_CHECK_NEEDED (1): Information-seeking questions requiring external verification
Factual questions about dates… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/fact-check-classification-dataset.semantic-economyall things are now lawful to you in jack feist
EA-RHIZOME-SE-01 — semantic economy
The archive's political economy of meaning, seeded at the stolon the collapse body left open. Its roles are NOT the collapse roles: where that body asks what narrows, this one asks who gains by the narrowing.
These are symbola. They are for traversal.
A token broken in two, each half held by a different party, no half carrying complete authority, the fit of the fracture proving the… See the full description on the dataset page: https://huggingface.co/datasets/leesharks/semantic-economy.rehydra-semanticgeneration-semantic-memorization-filtersav_semantic_anomaliesSemantic-Scholar-Papersbartowski-imatrix-v5-semantic
Bartowski iMatrix Calibration v5 (Semantic Chunking)
A processed version of bartowski's v5 imatrix calibration data using semantic boundary detection optimized for the v5 data structure.
Dataset Summary
Metric
Value
Total samples
2,075
Chunking method
V5-optimized semantic boundary detection
Chunk size
200+ characters (no upper limit, preserves document integrity)
Languages
English, German, Spanish, French, Italian, Swedish, Russian, Arabic, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/bartowski-imatrix-v5-semantic.brain-tumor-image-dataset-semantic-segmentation
Dataset Card for "brain-tumor-image-dataset-semantic-segmentation"
Dataset Description
The Brain Tumor Image Dataset (BTID) for Semantic Segmentation contains MRI images and annotations aimed at training and evaluating segmentation models. This dataset was sourced from Kaggle and includes detailed segmentation masks indicating the presence and boundaries of brain tumors.
This dataset can be used for developing and benchmarking algorithms for medical image segmentation… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/brain-tumor-image-dataset-semantic-segmentation.github_semantic_searcharxiv-semantic-scholar
📚 arxiv-semantic-scholar
A paper metadata dataset covering every paper on arXiv, with title, authors, full abstract, download link, and submission history.
Each row is additionally enriched by a call to the Semantic Scholar API, adding citation counts and venue information.
🔗 Sources
arXiv metadata: https://www.kaggle.com/datasets/Cornell-University/arxiv (CC0)
Semantic Scholar Academic Graph: https://api.semanticscholar.org/ (ODC-BY)
Snapshot taken 2026-07-04.… See the full description on the dataset page: https://huggingface.co/datasets/dx2102/arxiv-semantic-scholar.memories-semantic-memorization-filter-results
Dataset Card for "memories-semantic-memorization-filter-results"
More Information needed
dataset_v2_classifiedsec-10k-semantic-deduped
