datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
action-atlas-rollout-videos
Action Atlas — VLA Rollout Videos
Local rollout/ablation videos for Pi0.5, OpenVLA-OFT, X-VLA, GR00T, SmolVLA, and ACT/ALOHA,
organized by model. Companion to action-atlas-{pi05,oft,xvla,groot,smolvla} (SAEs + activations + concepts).
368283 unique mp4 clips, 61.8 GB. Per-model: {'act_aloha': 990, 'groot': 163891, 'oft': 24284, 'pi05': 62468, 'smolvla': 56844, 'xvla': 59806}
manifest.jsonl: one row per clip (model, env, experiment, sha256, bytes, hf_path).
action-worldmodel-benchaction3daction-atlas-viz
Action Atlas visualization bundle
The precomputed data the Action Atlas web frontend reads (https://action-atlas.com). This is all you
need to run the site locally; the SAE weights and raw activations are not required at runtime.
Contents
processed/ per-layer clustering layouts and feature scatter data the frontend renders.
feature_embeddings/ embeddings used for semantic feature search.
descriptions/ generated natural-language feature and concept descriptions.… See the full description on the dataset page: https://huggingface.co/datasets/bag100/action-atlas-viz.minecraft-text-action-datasetaction-atlas-xvla-saetrainaction-atlas-oft
Action Atlas: OpenVLA-OFT sparse autoencoders and concepts
Sparse autoencoders (SAEs) and identified concepts for OpenVLA-OFT, part of the Action Atlas
release accompanying the paper on cross-task activation injection in vision-language-action models.
The interactive explorer is at https://action-atlas.com.
What is here
saes/ TopK SAEs (k=64, 8x expansion) over the OpenVLA-OFT (Llama-2 7B backbone, continuous L1 action head), residual stream 4096-dim, 32 layers.… See the full description on the dataset page: https://huggingface.co/datasets/bag100/action-atlas-oft.Temporal_Action_Detectionaction-atlas-groot-activationsaction100m_tiny_subset
Dataset Card for action100m
This is a FiftyOne dataset with 1144 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/action100m_tiny_subset")
# Launch the App
session = fo.launch_app(dataset)
Dataset Card for… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/action100m_tiny_subset.actionbench
🎬 ActionBench: Paired Video-3D Synthetic Benchmark
📖 Overview
ActionBench is a benchmark dataset of 128 paired video ↔ animated point-cloud samples for evaluating animated 3D mesh generation from video.
The dataset consists of synthetic scenes of animated objects from ObjaverseXL, rendered using Blender 3.5.1.
Each sample contains:
Video: 16 RGBA frames with alpha mask
Camera (camera.json): Camera parameters using Blender convention (X_cam = X @ R^T + T, camera looks… See the full description on the dataset page: https://huggingface.co/datasets/facebook/actionbench.xiaoluo-gaming-action3000-20260910-media
Action-boundary review examples
Media for 300 selected examples from Xiaoluo (Cyberpunk 2077 and Rise of the Tomb Raider) and Gaming 500 Hours, 150 examples per dataset.
Includes 15-second review videos, observed boundary frames, and available action clips. These are visual model estimates; boundaries require human review. Source game and dataset rights remain with their respective owners.
Gallery and annotation manifests:… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/xiaoluo-gaming-action3000-20260910-media.Breakfast-Actions
🍳 Breakfast Actions Dataset (HF + WebDataset Ready)
This repository hosts the Breakfast Actions dataset metadata and videos, organized for modern deep learning workflows.It provides:
4 evaluation splits (s1, s2, s3, s4)
JSONL metadata describing each video, participant, camera, and frame-level action segments
Raw AVI videos stored directly on HuggingFace
Optional WebDataset shards for streaming training
📁 Folder Layout
Breakfast-Actions/
│
├──… See the full description on the dataset page: https://huggingface.co/datasets/CVML-TueAI/Breakfast-Actions.action100m-preview
Action100M: A Large-scale Video Action Dataset
Paper | GitHub
Action100M is a large-scale dataset constructed from 1.2M Internet instructional videos (14.6 years of duration), yielding ~100 million temporally localized segments with open-vocabulary action supervision and rich captions. It serves as a foundation for scalable research in video understanding and world modeling.
Load Action100M Annotations
Our data can be loaded from the 🤗 huggingface repo at… See the full description on the dataset page: https://huggingface.co/datasets/facebook/action100m-preview.ActionNet
Fourier ActionNet Dataset
Introduction
The data collected from two primary sources: the robot side and the camera side. The HDF5 file contains the robot-side data, while the camera-side data is stored in the corresponding episode folder. While there is also a metadata.json file in the dataset, which contains all episodes id, and it's prompt.
Download the dataset
First, you can easily download the dataset online, which will be in a .tar file. After downloading… See the full description on the dataset page: https://huggingface.co/datasets/FourierIntelligence/ActionNet.action-atlas-smolvla-activations4RC-ActionESK_action_segmentationPaper | GitHub
🍳 EPFL-Smart-Kitchen: Action segmentation benchmark
📚 Introduction
Given an untrimmed video, action segmentation requires the model to predict one or multiple action classes for every frame.Given the absence of popular (and comprehensive) action segmentation benchmarks from 3D pose we built an action segmentation benchmark that compares the impact of different input data (body,hand,eyes, videofeatures).One might expect that actions such as moving through… See the full description on the dataset page: https://huggingface.co/datasets/amathislab/ESK_action_segmentation.mobile-actions
Mobile Actions: A Dataset for On-Device Function Calling
The dataset contains conversational traces designed to train lightweight models (such as FunctionGemma 270M) to translate natural language instructions into executable function calls for Android OS system tools.
Dataset Format
The dataset is provided in JSONL format. Each line represents a data sample. The
dataset is pre-split into training and evaluation sets. This distinction is
denoted by the metadata field… See the full description on the dataset page: https://huggingface.co/datasets/google/mobile-actions.actionnet-subset100-gtdepth
ActionNet subset100 with ground-truth depth
This is a 100-episode subset redistribution of a third-party dataset, plus our derived
artifacts. Read the attribution before using it.
Attribution and license
Upstream dataset
FourierIntelligence/ActionNet (Fourier Intelligence)
What is redistributed
100 episodes out of 30,121, byte-identical to the upstream tars: rgb.mp4 (1280x800 fisheye), depth.mkv (lossless 16-bit), timestamps.json, <ULID>.hdf5… See the full description on the dataset page: https://huggingface.co/datasets/glory-hyeok/actionnet-subset100-gtdepth.HRDexDB-DTW-Action-Synced-224
HRDexDB DTW Action-Synced 224
This dataset contains temporally aligned human–robot episode pairs derived from HRDexDB.
Contents
Human and robot videos aligned with object-trajectory Dynamic Time Warping (DTW)
224×224 RGB video encoded as H.264 at 30 FPS
Robot arm and hand actions synchronized to every output frame
Human MANO vertices, shared topology, parameters, and joints synchronized to every output frame
Per-pair manifests preserving source frame IDs and… See the full description on the dataset page: https://huggingface.co/datasets/HRDexDB/HRDexDB-DTW-Action-Synced-224.action-world-model-atlas-1500-media-20260914
Action World Model Atlas
Public browsing previews for 1,500 unique action clips from the completed
6,033-video bundle. OpenPixel2Play, Gaming 500 Hours, and Xiaoluo each contribute
500 examples. All 46 games in the completed bundle are represented.
Videos preserve the full five-second duration and 81 frames. They are existing
browser previews and can be smaller than the native training videos. Video and
poster checksums are verified against the source media manifests.… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/action-world-model-atlas-1500-media-20260914.human_assisted_action_preference_optimizationgfmc_hyworld1.5_processed_160latents_16fps_actionaction-atlas-oft-activationsminecraft-motion-action-datasetKTH-Action-Recognition-mirrorrobot-action-prediction-dataset
Robotic Action Prediction Dataset
Dataset Description
This dataset contains triplets of (current observation, action instruction, future observation) for training models to predict future frames of robotic actions.
Dataset Structure
Data Fields
current_frame: Input image (RGB) of the current observation
instruction: Textual description of the action to perform
future_frame: Target image (RGB) showing the expected outcome 50 frames later… See the full description on the dataset page: https://huggingface.co/datasets/bryandts/robot-action-prediction-dataset.fire_actioncam
Fire Actioncam
This dataset is a collection of several real-world fire scenes, introduced by the ECCV paper "Gaussians on Fire: High-Frequency Reconstruction of Flames".
Overview
The dataset consists of 17 real-world scenes of burning paper, cardboard, wood, gasoline, ethanol, and propane. We captured each scene with three regular actioncams, synchronizing them with µs precision using a custom LED pattern.
Property
Value
Scenes
17 (two outdoor… See the full description on the dataset page: https://huggingface.co/datasets/jna-358/fire_actioncam.dual_needle_concat_action_staticThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unity",
"total_episodes": 250,
"total_frames": 97953,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:250"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/inaas/dual_needle_concat_action_static.
