datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
semantic-memorization-partial-2023-09-03This dataset is a partial computation of metrics (memorized token frequencies, non-memorized token frequencies, sequence frequencies) needed for research.
pythia-semantic-memorization-perplexities
Dataset Card for "pythia-semantic-memorization-perplexities"
More Information needed
Semantic_dataSemantic-Harmful
[!IMPORTANT]
You are viewing: Harmful SubsetFor paired harmless dataset: heretic-org/Semantic-Harmless
Semantic Harmful-Harmless Prompt Pairs
Summary
This dataset contains one-to-one semantic matches between prompts from two source datasets:
mlabonne/harmful_behaviors
mlabonne/harmless_alpaca
The goal was to align prompts that are semantically closest where one prompt is harmful and the other is harmless. This creates a more controlled… See the full description on the dataset page: https://huggingface.co/datasets/heretic-org/Semantic-Harmful.Semantic-Harmless
[!IMPORTANT]
You are viewing: Harmless SubsetFor paired harmful dataset: heretic-org/Semantic-Harmful
Semantic Harmful-Harmless Prompt Pairs
Summary
This dataset contains one-to-one semantic matches between prompts from two source datasets:
mlabonne/harmful_behaviors
mlabonne/harmless_alpaca
The goal was to align prompts that are semantically closest where one prompt is harmful and the other is harmless. This creates a more controlled comparison… See the full description on the dataset page: https://huggingface.co/datasets/heretic-org/Semantic-Harmless.semantic-segmentation-test-sampleThis dataset contains 10 examples of the segments/sidewalk-semantic dataset (i.e. 10 images with corresponding ground-truth segmentation maps).
semantic-ecology
NuBerea Semantic Ecology
Feature-indexed analysis layer for the study of canon formation, part of the
NuBerea corpus estate of biblical and patristic texts. It describes the
"semantic ecology" of ancient biblical and related literature — how semantic
domains are used across texts, how candidate texts were received over time,
and which measurable features accompany canonical inclusion — packaged as a
set of ready-to-load configurations.
Attribution
NuBerea project.… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/semantic-ecology.SemanticVLA-TraceX-240K-DROID
SemanticVLA TraceX 240K · DROID
🎉 Accepted to CVPR 2026.
✍️ Fei Ni¹, Zhuo Chen², Yifu Yuan³, Zibin Dong³, Xianze Yao³, Shan Luo², Jianye Hao³, Jiankang Deng¹†, Stefanos Zafeiriou¹†
🏫 ¹Imperial College London ²King's College London ³Tianjin University
✉️ Primary contact: f.ni@imperial.ac.uk
The DROID component of TraceX-240K — the trace-annotated trajectory corpus introduced in SemanticVLA. This package is a LeRobot v3.0 repack of DROID · Franka · Open-X-Embodiment DROID… See the full description on the dataset page: https://huggingface.co/datasets/spikefly/SemanticVLA-TraceX-240K-DROID.SemanticPotentialRoutingTelemetry
Semantic Potential Routing Telemetry
Version 2.0 — a systematic, packet-level benchmark of training-free potential-field routing against
classical routing under stochastic congestion, dynamic topologies and microburst traffic.
Every episode is a network simulation in which routing is a physical field: each flow's destination is the
grounded, attractive well of a discrete Poisson equation on the graph Laplacian, congested buffers inject
repulsive current, and packets follow the… See the full description on the dataset page: https://huggingface.co/datasets/ezharjan/SemanticPotentialRoutingTelemetry.pile-semantic-memorization-filter-results
Dataset Card for "pile-semantic-memorization-filter-results"
More Information needed
jailbreak-detection-dataset
Jailbreak Detection Dataset (MLCommons-Aligned)
A comprehensive dataset for training AI safety classifiers, aligned with the MLCommons AI Safety taxonomy.
Dataset Description
This dataset combines multiple sources for robust jailbreak and safety detection:
Primary Sources
nvidia/Aegis-AI-Content-Safety-Dataset-2.0: 18,164 samples with MLCommons-aligned labels
lmsys/toxic-chat: Toxic content detection
jackhhao/jailbreak-classification: Jailbreak attack patterns… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/jailbreak-detection-dataset.semantic-duplicatesSemantic-Scholar-Papersyodas2-mm-semanticgeneration-semantic-memorization-filtersdexmg_dataset_semantic_goalsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "robomimic",
"total_episodes": 1014,
"total_frames": 326707,
"total_tasks": 1,
"total_videos": 3042,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:1014"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ishika/dexmg_dataset_semantic_goals.EEG-semantic-text-relevanceWe release a novel dataset containing 23,270 time-locked (0.7s) word-level EEG recordings acquired from participants who read both text that was semantically relevant and irrelevant to self-selected topics.
The raw EEG data and the datasheet are available at https://osf.io/xh3g5/.
See code repository for benchmark results.
EEG data acquisition:
Explanations of the variables:
event corresponds to a specific point in time during EEG data collection and represents the onset of an event… See the full description on the dataset page: https://huggingface.co/datasets/Quoron/EEG-semantic-text-relevance.fact-check-classification-dataset
Fact-Check Classification Dataset
🎯 Purpose: Binary classification dataset for determining whether a prompt needs external fact-checking.
Dataset Description
This dataset is designed to train classifiers that can route LLM requests based on whether they require external fact verification. It's part of the vLLM Semantic Router project.
Labels
FACT_CHECK_NEEDED (1): Information-seeking questions requiring external verification
Factual questions about dates… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/fact-check-classification-dataset.task1418_bless_semantic_relation_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1418_bless_semantic_relation_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1418_bless_semantic_relation_classification.dataset_v2_classifiedbrain-tumor-image-dataset-semantic-segmentation
Dataset Card for "brain-tumor-image-dataset-semantic-segmentation"
Dataset Description
The Brain Tumor Image Dataset (BTID) for Semantic Segmentation contains MRI images and annotations aimed at training and evaluating segmentation models. This dataset was sourced from Kaggle and includes detailed segmentation masks indicating the presence and boundaries of brain tumors.
This dataset can be used for developing and benchmarking algorithms for medical image segmentation… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/brain-tumor-image-dataset-semantic-segmentation.memories-semantic-memorization-filter-results
Dataset Card for "memories-semantic-memorization-filter-results"
More Information needed
longcontext-haldetect
Long-Context Hallucination Detection Benchmark
A synthetic benchmark dataset for evaluating hallucination detection models on long documents (8K-24K tokens). This dataset is specifically designed to test models that can handle contexts beyond the typical 8K token limit.
Dataset Summary
Property
Value
Total samples
3,366
Token range
8,005 - 23,998
Average tokens
17,852
Hallucinated
1,681 (49.9%)
Supported
1,685 (50.1%)
Splits… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/longcontext-haldetect.sec-10k-semantic-dedupedgeneration-semantic-intermediate-filtersdataset_v1_rawgeneration-semantic-filtersfeedback-detector-dataset
Feedback Detector Dataset
A large-scale multilingual dataset for 4-class user feedback classification, labeled using GPT-OSS-120B on AMD MI300X GPU.
Dataset Description
This dataset contains 51,694 examples of user feedback classified into 4 categories:
Label
Description
Count
%
SAT
User is satisfied
8,649
17%
NEED_CLARIFICATION
User needs more information
16,179
31%
WRONG_ANSWER
System gave incorrect response
19,919
39%
WANT_DIFFERENT
User wants… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/feedback-detector-dataset.sidewalk-semantic
Dataset Card for sidewalk-semantic
Dataset Summary
A dataset of sidewalk images gathered in Belgium in the summer of 2021. Label your own semantic segmentation datasets on segments.ai
Supported Tasks and Leaderboards
semantic-segmentation: The dataset can be used to train a semantic segmentation model, where each pixel is classified. The model performance is measured by how high its mean IoU (intersection over union) to the reference is.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/segments/sidewalk-semantic.Semantic-SVG-Benchmark
Semantic SVG Benchmark
A benchmark of 203 SVG files annotated with human-written semantic object-decomposition trees:
every rendered shape (<path>, <rect>, <circle>, …) in each SVG is assigned to a named semantic
object (e.g. judge, gavel), and objects may be further decomposed into parts
(e.g. Bamboo planter → pot, bamboo). It is the evaluation benchmark of
Compositional SVG Generation via VLM-Driven Hierarchical Semantic Parsing
(EMNLP 2026). The annotations are ours; the SVGs… See the full description on the dataset page: https://huggingface.co/datasets/KU-MIIL/Semantic-SVG-Benchmark.
