datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
semantic-memorization-partial-2023-09-03This dataset is a partial computation of metrics (memorized token frequencies, non-memorized token frequencies, sequence frequencies) needed for research.
pythia-semantic-memorization-perplexities
Dataset Card for "pythia-semantic-memorization-perplexities"
More Information needed
SemanticVLA-TraceX-240K-DROID
SemanticVLA TraceX 240K · DROID
🎉 Accepted to CVPR 2026.
✍️ Fei Ni¹, Zhuo Chen², Yifu Yuan³, Zibin Dong³, Xianze Yao³, Shan Luo², Jianye Hao³, Jiankang Deng¹†, Stefanos Zafeiriou¹†
🏫 ¹Imperial College London ²King's College London ³Tianjin University
✉️ Primary contact: f.ni@imperial.ac.uk
The DROID component of TraceX-240K — the trace-annotated trajectory corpus introduced in SemanticVLA. This package is a LeRobot v3.0 repack of DROID · Franka · Open-X-Embodiment DROID… See the full description on the dataset page: https://huggingface.co/datasets/spikefly/SemanticVLA-TraceX-240K-DROID.semantic-ecology
NuBerea Semantic Ecology
Feature-indexed analysis layer for the study of canon formation, part of the
NuBerea corpus estate of biblical and patristic texts. It describes the
"semantic ecology" of ancient biblical and related literature — how semantic
domains are used across texts, how candidate texts were received over time,
and which measurable features accompany canonical inclusion — packaged as a
set of ready-to-load configurations.
Attribution
NuBerea project.… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/semantic-ecology.SemanticPotentialRoutingTelemetry
Semantic Potential Routing Telemetry
Version 2.0 — a systematic, packet-level benchmark of training-free potential-field routing against
classical routing under stochastic congestion, dynamic topologies and microburst traffic.
Every episode is a network simulation in which routing is a physical field: each flow's destination is the
grounded, attractive well of a discrete Poisson equation on the graph Laplacian, congested buffers inject
repulsive current, and packets follow the… See the full description on the dataset page: https://huggingface.co/datasets/ezharjan/SemanticPotentialRoutingTelemetry.pile-semantic-memorization-filter-results
Dataset Card for "pile-semantic-memorization-filter-results"
More Information needed
semantic-duplicatesEEG-semantic-text-relevanceWe release a novel dataset containing 23,270 time-locked (0.7s) word-level EEG recordings acquired from participants who read both text that was semantically relevant and irrelevant to self-selected topics.
The raw EEG data and the datasheet are available at https://osf.io/xh3g5/.
See code repository for benchmark results.
EEG data acquisition:
Explanations of the variables:
event corresponds to a specific point in time during EEG data collection and represents the onset of an event… See the full description on the dataset page: https://huggingface.co/datasets/Quoron/EEG-semantic-text-relevance.dexmg_dataset_semantic_goalsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "robomimic",
"total_episodes": 1014,
"total_frames": 326707,
"total_tasks": 1,
"total_videos": 3042,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:1014"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ishika/dexmg_dataset_semantic_goals.generation-semantic-memorization-filtersmemories-semantic-memorization-filter-results
Dataset Card for "memories-semantic-memorization-filter-results"
More Information needed
sec-10k-semantic-dedupedgeneration-semantic-filtersgeneration-semantic-intermediate-filtersSemanticVLA-TraceX-240K-Bridge
SemanticVLA TraceX 240K · Bridge
🎉 Accepted to CVPR 2026.
✍️ Fei Ni¹, Zhuo Chen², Yifu Yuan³, Zibin Dong³, Xianze Yao³, Shan Luo², Jianye Hao³, Jiankang Deng¹†, Stefanos Zafeiriou¹†
🏫 ¹Imperial College London ²King's College London ³Tianjin University
✉️ Primary contact: f.ni@imperial.ac.uk
The Bridge component of TraceX-240K — the trace-annotated trajectory corpus introduced in SemanticVLA. This package is a LeRobot v3.0 repack of BridgeData V2 · WidowX with dense… See the full description on the dataset page: https://huggingface.co/datasets/spikefly/SemanticVLA-TraceX-240K-Bridge.SemanticVLA-TraceX-240K-Fractal
SemanticVLA TraceX 240K · Fractal (RT-1)
🎉 Accepted to CVPR 2026.
✍️ Fei Ni¹, Zhuo Chen², Yifu Yuan³, Zibin Dong³, Xianze Yao³, Shan Luo², Jianye Hao³, Jiankang Deng¹†, Stefanos Zafeiriou¹†
🏫 ¹Imperial College London ²King's College London ³Tianjin University
✉️ Primary contact: f.ni@imperial.ac.uk
The Fractal (RT-1) component of TraceX-240K — the trace-annotated trajectory corpus introduced in SemanticVLA. This package is a LeRobot v3.0 repack of Fractal · Google Robot ·… See the full description on the dataset page: https://huggingface.co/datasets/spikefly/SemanticVLA-TraceX-240K-Fractal.semantic-filters-intermediateSemantic-SVG-Benchmark
Semantic SVG Benchmark
A benchmark of 203 SVG files annotated with human-written semantic object-decomposition trees:
every rendered shape (<path>, <rect>, <circle>, …) in each SVG is assigned to a named semantic
object (e.g. judge, gavel), and objects may be further decomposed into parts
(e.g. Bamboo planter → pot, bamboo). It is the evaluation benchmark of
Compositional SVG Generation via VLM-Driven Hierarchical Semantic Parsing
(EMNLP 2026). The annotations are ours; the SVGs… See the full description on the dataset page: https://huggingface.co/datasets/KU-MIIL/Semantic-SVG-Benchmark.multieurlex21-pt-semantic-cache
MultiEURLEX-21 PT — frozen semantic chunk embeddings
Precomputed, unit-normalised chunk embeddings of the official MTEB
MultiEURLEXMultilabelClassification Portuguese data, produced once so that
downstream experiments (frozen MaleCNS connectome reservoir, label-free
controls, ablations) never pay the sentence encoders again. No benchmark labels
are stored or used here.
source dataset: mteb/eurlex-multilingual revision 2aea5a6dc8fdcfeca41d0fb963c0a338930bde5c, subset pt
splits:… See the full description on the dataset page: https://huggingface.co/datasets/franklinbaldo/multieurlex21-pt-semantic-cache.semantic-history-search
Semantic History (Synthetic)
Semantic History is a synthetic dataset for semantic history search research.It ships as three normalized Parquet tables:
Split
Description
docs
Search history records with url, title, description, frecency, last_visit_date, and tags.
queries
One row per query, tagged by profile and temporal/multi-label flags.
qrels
Relevance pairs linking queries ↔ docs with rank and relevance.
All content is synthetic (no real browsing logs).… See the full description on the dataset page: https://huggingface.co/datasets/frankjc2022/semantic-history-search.uncgpt-conversations-semantic-approved-1p50-candidate
UncGPT — Semantic-Approved 1.50σ Conversations (Candidate)
The wider-tolerance (1.50σ) cohort against the same contrast semantic boundary. Useful as a higher-recall candidate for ablating gate strictness vs. coverage.
Part of the UncGPT NeurIPS 2026 Competition collection.
Configs
Config
What it is
approved_manifest (default)
conversations that passed at 1.50σ
rejected_manifest
conversations that failed even at 1.50σ
Why a wider tolerance
Some… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p50-candidate.SemanticVLA-TraceX-240K-BC-Z
SemanticVLA TraceX 240K · BC-Z
🎉 Accepted to CVPR 2026.
✍️ Fei Ni¹, Zhuo Chen², Yifu Yuan³, Zibin Dong³, Xianze Yao³, Shan Luo², Jianye Hao³, Jiankang Deng¹†, Stefanos Zafeiriou¹†
🏫 ¹Imperial College London ²King's College London ³Tianjin University
✉️ Primary contact: f.ni@imperial.ac.uk
The BC-Z component of TraceX-240K — the trace-annotated trajectory corpus introduced in SemanticVLA. This package is a LeRobot v3.0 repack of BC-Z · Open-X-Embodiment BC-Z v0.1.0 with… See the full description on the dataset page: https://huggingface.co/datasets/spikefly/SemanticVLA-TraceX-240K-BC-Z.mini_kitti_semanticSemanticSeg
Dataset Card for SemanticSeg
This semantic segmentation dataset introduced in the paper Towards Generalization of Block Attention via Automatic Segmentation and Block Distillation.
This dataset is used to train the segmenter.
Dataset Details
Dataset Description
SemanticSeg contains around 16 segmentation categories, with each category containing at least 2k instances. The varying cut rates across categories can also help the segmenter learn… See the full description on the dataset page: https://huggingface.co/datasets/Syon-Li/SemanticSeg.doorkey-semantic-reasoning-labels
GPT-OSS-20B DoorKey semantic reasoning labels
This dataset contains automatic sentence-level semantic-function annotations
for 7,038 reasoning sentences produced by openai/gpt-oss-20b on 46 fixed
DoorKey environment states. Each target sentence is paired with its preceding
reasoning context and assigned one or more human-readable discourse labels.
The annotation run produced 7,036 valid rows and two schema failures. These are
model-generated exploratory annotations, not human… See the full description on the dataset page: https://huggingface.co/datasets/project-telos/doorkey-semantic-reasoning-labels.pg19-semantic-novelty
PG19 Semantic Novelty Dataset
Paragraph-by-paragraph semantic novelty curves for 28,535 books from the PG19 corpus (Project Gutenberg, pre-1920 English literature).
What is Semantic Novelty?
For each paragraph in a book, we compute:
novelty(p) = 1 - cosine_similarity(embedding(p), running_centroid)
where embedding() uses SBERT all-mpnet-base-v2 (768-dimensional) and running_centroid is the mean of all preceding paragraph embeddings. This measures how much new information… See the full description on the dataset page: https://huggingface.co/datasets/wfzimmerman/pg19-semantic-novelty.SemanticWikipediaArticleGraphThai-Semantic-Textual-Similarity-BenchmarkSentence representation plays a crucial role in NLP downstream tasks such as NLI, text classification, and STS. Recent sentence representation training techniques require NLI or STS datasets. However, there are no equivalent Thai NLI or STS datasets for sentence representation training.
To address this problem we provide the Thai sentence vector benchmark. We evaluate the Spearman correlation score of the sentence representations’ performance on Thai STS-B (translated version of STS-B).… See the full description on the dataset page: https://huggingface.co/datasets/mrp/Thai-Semantic-Textual-Similarity-Benchmark.Test_Semantic_Searchrav4-semantic-video-index
RAV4 Image to Video Semantic Retrieval Dataset
This dataset contains outputs from a system that searches a video using an image of a car part.
We give the system a picture (for example a hood or door), and it returns the moments in the video where that same part appears.
The system matches meaning (car parts), not exact pixels.
How the system works
The video is converted into frames (1 frame per second)
A trained detector finds car parts in every frame
All… See the full description on the dataset page: https://huggingface.co/datasets/shiniagarwal/rav4-semantic-video-index.
