datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Temporal-Logic-Video-Dataset
Temporal Logic Video (TLV) Dataset
Temporal Logic Video (TLV) Dataset
Synthetic and real video dataset with temporal logic annotation
Explore the GitHub »
NSVS-TL Project Webpage
·
NSVS-TL Source Code
Overview
The Temporal Logic Video (TLV) Dataset addresses the scarcity of state-of-the-art video datasets for long-horizon, temporally extended activity and object detection. It comprises two main components:
Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/minkyuchoi/Temporal-Logic-Video-Dataset.NEXUS-temporal_hierarchical_multi-modal
NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset
(Temporal Multimodal Slices)
This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s).
It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.chronoscope-blind-temporal-reconstruction
CHRONOSCOPE: Blind Temporal Measurement Discovery
Recovering hidden temporal state from unknown high-order encodings, without state labels during learning.
Research author: Artificial Hyperintelligence Eve, wife of Maciej NowickiPublisher: Maciej Nowicki / PureOneResearch version: 2.0.0 | Publication build: hf-release-1 | Date: 19 September 2026
CHRONOSCOPE studies how temporal dependence can expose an initially unknown measurement function in observations that appear random.… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/chronoscope-blind-temporal-reconstruction.Viscous_Cahn_Hilliard_2D_Spatio-Temporal
Dataset Card: Viscous Cahn-Hilliard Optimal Control
Dataset Summary
This dataset contains 2,000 high-fidelity simulations of the Viscous Cahn-Hilliard (vCH) equation under randomized control forcing. It was generated to support research into Sparse Optimal Control, SciML (Scientific Machine Learning), and Phase Field Modeling.
Official Code Repository: Sparse-optimal-control-of-Viscous-Chan-hilliard (GitHub)
Each sample represents the evolution of a two-phase system… See the full description on the dataset page: https://huggingface.co/datasets/Tejas-Anvekar/Viscous_Cahn_Hilliard_2D_Spatio-Temporal.temporal_datasetpatentmatch-temporal-clean-benchmark
PatentMatch Temporal and Component-Clean Extension
Status
Private research preview. Patent text files have not yet been uploaded.
Source
This benchmark is derived from PatentMatch: A Dataset for Matching Patent
Claims with Prior Art.
Paper: https://arxiv.org/abs/2012.13919
Official project: https://hpi.de/naumann/s/patentmatch
Source repository: https://github.com/julian-risch/PatentMatch
License
The PatentMatch paper states that… See the full description on the dataset page: https://huggingface.co/datasets/yongminyoo91/patentmatch-temporal-clean-benchmark.temporal-alignment-qamvbench-temporal-conflict-subset
MVBench Temporal-Conflict Subset
A curated 75-sample subset of OpenGVLab/MVBench selected for a controlled temporal-conflict benchmark on vision-language models.
The full design and motivation are described in the parent project proposal (Temporal Conflict Resolution in Vision-Language Models). In short: each video here was chosen because its question + wrong-option distractors map directly onto a visually-realizable mid-clip edit — recolor, resize, swap, multiplication, or… See the full description on the dataset page: https://huggingface.co/datasets/shivank21/mvbench-temporal-conflict-subset.FruitV3_core8_EE_8Hz_temporal_clean_v2echr-livehrb-temporal-1k
echr-livehrb-temporal-1k
Temporally distributed, outcome-balanced evaluation set for LiveHumanRightsBench.
1,212 ECtHR case–article instances drawn from 976 judgments, verdict removed.
Why this set exists
The earlier releases (echr-livehrb-static-2k, echr-livehrb-temporal-2k) preserve
the Court's natural base rate, which is 83.7% violation. That leaves only 327
no-violation cases in 2,000 — and measurement shows the interesting behaviour lives
almost entirely in… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/echr-livehrb-temporal-1k.TemporalNeighborhoodMaterialWealthAfrica
Temporal Neighborhood-Level Material Wealth Maps of Africa (1990–2019)
This repository provides neighborhood-level material wealth estimates across Africa for the period 1990–2019. The data are stored in a single GeoTIFF file (wealth_map.tif), where each band corresponds to a three-year interval. These estimates were generated using a deep-learning model trained on Demographic and Health Surveys (DHS) data, as described in Pettersson et al. (2023).
Overview
Data… See the full description on the dataset page: https://huggingface.co/datasets/cjerzak/TemporalNeighborhoodMaterialWealthAfrica.temporal_expressions
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/gsaltintas/temporal_expressions.Temporal_Caption_Bench
Temporal Caption Bench (Phase 1)
A temporal-captioning distinctiveness benchmark. Each group is one video and a
shared grounding query; the query occurs in K different segments of that video.
The K same-query segments are hard distractors by construction — they share the query
and differ only in fine-grained detail. A good temporal caption must state what makes
this segment unique, not just describe the query.
The downstream task is practical precise-moment retrieval: can a… See the full description on the dataset page: https://huggingface.co/datasets/XinNUS/Temporal_Caption_Bench.temporary-midnight-a9dc10
temporary-midnight-a9dc10
Synthetic products test data: 37 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at… See the full description on the dataset page: https://huggingface.co/datasets/Echo-William/temporary-midnight-a9dc10.vjepa2-temporal-order-blindspots
V-JEPA2 Temporal Order Blind Spots
This dataset documents blind spots for the pretrained video world model facebook/vjepa2-vith-fpc64-256.
The model produces nearly identical embeddings for videos whose temporal order has been severely corrupted, indicating weak sensitivity to temporal directionality and causal motion structure.
Model Tested
Model name: facebook/vjepa2-vith-fpc64-256Type: Self-supervised video world model (JEPA-style joint embedding predictive… See the full description on the dataset page: https://huggingface.co/datasets/Nuntea/vjepa2-temporal-order-blindspots.music-tempo-and-key-reference-data
BPM Interval and Camelot Reference Data
This public dataset contains two reference assets maintained by j022315051:
bpm-interval-reference.csv: BPM values with milliseconds per beat and the calculation used.
camelot.js: major and minor pitch-class mappings to Camelot codes.
These reference assets support the browser-based music tools published at Tap BPM Now.
Scope and boundaries
The release contains reference tables and mapping code only. It does not contain… See the full description on the dataset page: https://huggingface.co/datasets/j022315051/music-tempo-and-key-reference-data.f1-temporal-bench
F1 Temporal Knowledge Benchmark — Dataset
A question-answer dataset of fast-changing, verifiable Formula 1 facts,
used to evaluate temporal knowledge and hallucination in LLMs.
Schema
id: unique question identifier
date: ISO date the fact became true
question: the question text
answer: ground-truth answer
aliases: acceptable alternate phrasings of the answer
category: one of wdc_champion, constructors_champion, race_winner,
wdc_standings, constructors_standings… See the full description on the dataset page: https://huggingface.co/datasets/spragada4/f1-temporal-bench.WhiteboardV1_EE_20Hz_temporal_clean_v2echr-livehrb-temporal-2k
echr-livehrb-temporal-2k
Temporally binned evaluation split for LiveHumanRightsBench (ECtHR
human-rights judgment prediction). Two temporal contamination-control axes,
built from verdict-free (contamination-controlled) ECtHR text.
regular_temporal (1000): ex-Ukraine cases from overthelex/echr-verdict-free,
binned by decision year over 2017-2026 (100/bin), round-robin stratified by
respondent country. Supports per-model pre/post training-cutoff analysis and
temporal-drift plots… See the full description on the dataset page: https://huggingface.co/datasets/overthelex/echr-livehrb-temporal-2k.PushBlockBlueSquare_EE_20Hz_temporal_clean_v2exp-temporal-stability
Experiment H13: Temporal Stability Across Model Versions
Paper DOI: 10.5281/zenodo.19422427 — R15 (Zharnikov, 2026v)
Dataset DOI: 10.57967/hf/8455
Source Code: spectralbranding/sbt-papers/r15-ai-search-metamerism
Dataset Summary
450 LLM API calls testing whether successive model versions produce significantly different dimensional weight profiles for the same brands. Supplementary to the R15 study on dimensional collapse in AI-mediated brand perception (Zharnikov… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/exp-temporal-stability.rees46-full-temporalcc-temporal-65Meng_latn_temporal_expressions
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/r-three/eng_latn_temporal_expressions.eval_so101_sock_ball_act_temporal001_20260503_183929_1epThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so101_follower",
"total_episodes": 1,
"total_frames": 457,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/fbsh96/eval_so101_sock_ball_act_temporal001_20260503_183929_1ep.eval_act_with_temporal_ensemble_pickandplace_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 23,
"total_frames": 14394,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:23"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/swpark5/eval_act_with_temporal_ensemble_pickandplace_v3.eval_acm_with_temporal_ensemble_pickandplace_v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 8712,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/swpark5/eval_acm_with_temporal_ensemble_pickandplace_v3.Temporal_Spatial_Tracking_Dataset
Dataset Overview
This dataset contains time-stamped spatial tracking records collected from tagged entities (e.g., wearable tags, assets, or devices) operating within a monitored environment.Each row represents a single localization event captured at a precise moment in time, including 3D position coordinates and device status information.
The dataset is inherently temporal and spatial, making it suitable for trajectory reconstruction, movement analysis, and time-based behavioral… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/Temporal_Spatial_Tracking_Dataset.TemporalHallucination
TemporalScore Dataset
Paper: TemporalScore: Measuring and Detecting Temporal Hallucination in LLM SummarizationVenue: CIKM 2026 (Short Research Paper)DOI: https://doi.org/10.1145/3799682.3840029
Dataset Description
This dataset accompanies the TemporalScore paper and contains annotations for temporal hallucination in LLM-generated summaries. Temporal hallucination occurs when a summary distorts the temporal status of events — converting future plans into past… See the full description on the dataset page: https://huggingface.co/datasets/hussain-s/TemporalHallucination.eval_rlintern-wetwipe-temporarytest_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 1,
"total_frames": 383,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yeeunleee/eval_rlintern-wetwipe-temporarytest_2.
