datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
EmbodiedGenDatahttps://huggingface.co/spaces/HorizonRobotics/EmbodiedGen-Gallery-Explorer
Long-Horizon-Terminal-Bench
Long-Horizon Terminal-Bench (LHTB)
LHTB is a 46-task benchmark for measuring how well LLM agents sustain useful
work in a containerized terminal over hundreds of steps. Unlike short-horizon
coding benchmarks where an agent writes one artifact and stops, LHTB drops the agent
into a stateful environment and grades it with hidden, rebuild-from-artifact
verifiers — self-reported progress does not count.
📝 Blog: https://zli12321.github.io/LHTB/
🏆 Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/Long-Horizon-Terminal-Bench.AIGIBench
Is Artificial Intelligence Generated Image Detection a Solved Problem?
Ziqiang Li1, Jiazhen Yan1, Ziwen He1, Kai Zeng2, Weiwei Jiang1, Lizhi Xiong1, Zhangjie Fu1‡
‡Corresponding author
1Nanjing University of Information Science and Technology 2University of Siena
Paper | GitHub Repository
This repository is the official dataset of the AIGIBench.
AIGIBench dataset contains two types of training and 25 test subsets. This dataset has the following advantages:
Comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/HorizonTEL/AIGIBench.Long-Horizon-GUI-Datasetrobo_orchard_sim_assetssequence_recovery_centered_200_horizontally
RoboMME — SequenceRecoveryHorizontally (Robot HDF5 Demonstrations)
Raw robot demonstration data for the SequenceRecoveryHorizontally task from
RoboMME, a ManiSkill/SAPIEN benchmark for
memory-augmented robotic manipulation. Each HDF5 file is one recorded episode
containing observations (RGB/state), actions, and metadata for imitation
learning.
Layout
train/ 100 episodes
val/ 50 episodes
test/ 50 episodes
Files are named episode_<idx>_seed_<seed>.h5.… See the full description on the dataset page: https://huggingface.co/datasets/Hiesh/sequence_recovery_centered_200_horizontally.HorizonBench
HorizonBench
Long-Horizon Personalization with Evolving Preferences
HorizonBench evaluates whether language models can track user preferences as they evolve across months of interaction. Each benchmark item is a 5-option multiple-choice question embedded within a conversation history averaging ~163K tokens. Pre-evolution preference values serve as hard-negative distractors, enabling diagnosis of belief-update failure: models retrieve the user's originally stated preference but fail… See the full description on the dataset page: https://huggingface.co/datasets/stellalisy/HorizonBench.horizon-zero-dawn-gameplay-data
地平线零之曙光
This public dataset repository contains local gameplay data uploaded from F:\地平线零之曙光.
Contents
Files: 182
Total local size: 125.65 GB
Generated: 2026-06-05 14:52:20 UTC
File Types
.jsonl: 56
.json: 42
.png: 42
.txt: 14
.parquet: 14
.mkv: 14
Notes
This repository may contain gameplay video, images, Parquet files, JSON/JSONL metadata, and keyboard/mouse event logs.
The license is marked as other; review game footage, audio… See the full description on the dataset page: https://huggingface.co/datasets/xiaoluo11/horizon-zero-dawn-gameplay-data.300k-horizonsso101-long-horizon-datasetombo
so101-long-horizon-datasetombo
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
latent-programming-horizons-trajs
latent-programming-horizons-trajs
Agent trajectories and per-edit correctness labels from the
program-probes project, which
measures whether a language model's internal hidden states linearly predict
properties of its own agentic output (e.g. "does the code currently compile?")
before those properties are realised.
Each trajectory is a run of a coding agent (mini-SWE-agent) attempting a
SWE-bench (Verified or Pro) instance. This dataset contains the raw
transcripts and labels… See the full description on the dataset page: https://huggingface.co/datasets/ASSERT-KTH/latent-programming-horizons-trajs.rl-game-traces-horizon-forbidden-west
地平线之西之绝境
This public dataset repository contains gameplay trace data uploaded from F:\地平线之西之绝境.
Contents
Files: 442
Total local size: 375.37 GB
Generated: 2026-06-07T09:25:23+00:00
File Types
.jsonl: 132
.json: 100
.png: 79
.parquet: 33
.mkv: 33
.txt: 33
.jpg: 32
Notes
This repository may contain gameplay video, Parquet files, JSON/JSONL metadata, and input event logs.
The license is marked as other; review game footage, audio… See the full description on the dataset page: https://huggingface.co/datasets/yinhuankuang/rl-game-traces-horizon-forbidden-west.RoboTwin_shake_bottle_horizontally_randomizedHorizontalInequalityConflictAfricafsrvln_datasetsOpenVid-1M
Summary
This is the dataset proposed in our paper "OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation".
OpenVid-1M is a high-quality text-to-video dataset designed for research institutions to enhance video quality, featuring high aesthetics, clarity, and resolution. It can be used for direct training or as a quality tuning complement to other video datasets.
All videos in the OpenVid-1M dataset have resolutions of at least 512×512. Furthermore, we… See the full description on the dataset page: https://huggingface.co/datasets/lodestone-horizon/OpenVid-1M.omx_gelsight_env1_multitask_horizontal_vertical_line_nogelThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "omx_follower",
"total_episodes": 120,
"total_frames": 105619,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:120"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/WoojongKim/omx_gelsight_env1_multitask_horizontal_vertical_line_nogel.eval_smolvla_policy_omx_gelsight_env1_multitask_horizontal_vertical_line_nogel_20260820_192704This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "omx_follower",
"total_episodes": 10,
"total_frames": 7847,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/WoojongKim/eval_smolvla_policy_omx_gelsight_env1_multitask_horizontal_vertical_line_nogel_20260820_192704.HorizonGSso101-block-horizontal-layComb12
so101-block-horizontal-layComb12
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
R-HORIZON-AMC23
R-HORIZON
How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
📃 Paper • 🌐 Project Page • 🤗 Dataset
R-HORIZON is a novel method designed to stimulate long-horizon reasoning behaviors in Large Reasoning Models (LRMs) through query composition. We transform isolated problems into complex multi-step reasoning scenarios, revealing that even the most advanced LRMs suffer significant performance degradation when facing interdependent problems that span… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/R-HORIZON-AMC23.grab-the-black-box-horizontalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 100,
"total_frames": 32355,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Tank-123/grab-the-black-box-horizontal.gsplat-training-frames_horizontobservatorium
Dataset Card for Gaussian Splatting Drone Frames — Horizontobservatorium (DE)
Direct Use
Training and evaluation of Gaussian Splatting (3DGS/gsplat) and NeRF variants.
3D reconstruction with SfM/MVS (e.g., COLMAP) and validation of photogrammetry pipelines.
Benchmarks/ablations (PSNR/SSIM/LPIPS), pose estimation, approximate intrinsic calibration, metric scaling with GPS.
Out-of-Scope Use
Person/vehicle recognition or surveillance: the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/naruone90/gsplat-training-frames_horizontobservatorium.long-horizon-coding-trajectories
Long Horizon Coding Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/long-horizon-coding-trajectories.ARSG-110K
ARSG-110K
Project Page | Paper | GitHub
ARSG-110K is a large-scale scene-level dataset comprising over 110K diverse scenes and 3M annotated images with high-fidelity 3D ground truth. It is designed to support the training and evaluation of compositional 3D scene generation and in-place completion models. The dataset provides accurate 3D object-level ground-truth, layout, and annotations.
This dataset was introduced as part of the paper: 3D-Fixer: Coarse-to-Fine In-place Completion… See the full description on the dataset page: https://huggingface.co/datasets/HorizonRobotics/ARSG-110K.kaggle-animal-crossing-new-horizons-nookplaza-dataset
Dataset Card for Animal Crossing New Horizons Catalog
Dataset Summary
Context
This dataset comes from this spreadsheet, a comprehensive Item Catalog for Animal Crossing New Horizons (ACNH). As described by Wikipedia,
> ACNH is a life simulation game released by Nintendo for Nintendo Switch on March 20, 2020. It is the fifth main series title in the Animal Crossing series and, with 5 million digital copies sold, has broken the record for Switch title with most… See the full description on the dataset page: https://huggingface.co/datasets/osanseviero/kaggle-animal-crossing-new-horizons-nookplaza-dataset.long-horizon-traj
long-horizon-traj
Long-horizon agent trajectories with multi-step planning and constraints
Dataset Description
This dataset contains agent trajectories for multi-turn tool use tasks, including reasoning traces, tool calls, and responses.
Dataset Structure
The dataset is organized by model name, with each model having separate JSONL files for different experimental passes.
long-horizon-traj/
├── model-1/
│ ├── pass@1.jsonl
│ ├── pass@2.jsonl
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/ToolGym/long-horizon-traj.Long-Horizon-Execution
Long Horizon Execution
This project contains the dataset accompanying the paper "The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs"
Abstract
Does continued scaling of large language models (LLMs) yield diminishing returns? Real-world value often stems from the length of task an agent can complete. We start this work by observing the simple but counterintuitive fact that marginal gains in single-step accuracy can compound into exponential… See the full description on the dataset page: https://huggingface.co/datasets/arvindh75/Long-Horizon-Execution.natural-horizons
Horizons
64 episodes per horizon, padded to 2048 Qwen3 tokens by the provided tokenizer utility.
Preparation contract
Sources and counts are recorded in each configuration's manifest.json. Smoke manifests are small validation runs, not full releases. Columns episode_00 through episode_63 contain lists of complete conversations; num_tokens columns count each conversation before right padding. OOD/test split original latent families. The first eight episodes each… See the full description on the dataset page: https://huggingface.co/datasets/aklein4/natural-horizons.300k-horizons-single
