datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xiaoluo-gaming-action3000-20260910-media
Action-boundary review examples
Media for 300 selected examples from Xiaoluo (Cyberpunk 2077 and Rise of the Tomb Raider) and Gaming 500 Hours, 150 examples per dataset.
Includes 15-second review videos, observed boundary frames, and available action clips. These are visual model estimates; boundaries require human review. Source game and dataset rights remain with their respective owners.
Gallery and annotation manifests:… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/xiaoluo-gaming-action3000-20260910-media.actionbench
🎬 ActionBench: Paired Video-3D Synthetic Benchmark
📖 Overview
ActionBench is a benchmark dataset of 128 paired video ↔ animated point-cloud samples for evaluating animated 3D mesh generation from video.
The dataset consists of synthetic scenes of animated objects from ObjaverseXL, rendered using Blender 3.5.1.
Each sample contains:
Video: 16 RGBA frames with alpha mask
Camera (camera.json): Camera parameters using Blender convention (X_cam = X @ R^T + T, camera looks… See the full description on the dataset page: https://huggingface.co/datasets/facebook/actionbench.fire_actioncam
Fire Actioncam
This dataset is a collection of several real-world fire scenes, introduced by the ECCV paper "Gaussians on Fire: High-Frequency Reconstruction of Flames".
Overview
The dataset consists of 17 real-world scenes of burning paper, cardboard, wood, gasoline, ethanol, and propane. We captured each scene with three regular actioncams, synchronizing them with µs precision using a custom LED pattern.
Property
Value
Scenes
17 (two outdoor… See the full description on the dataset page: https://huggingface.co/datasets/jna-358/fire_actioncam.action-world-model-atlas-1500-media-20260914
Action World Model Atlas
Public browsing previews for 1,500 unique action clips from the completed
6,033-video bundle. OpenPixel2Play, Gaming 500 Hours, and Xiaoluo each contribute
500 examples. All 46 games in the completed bundle are represented.
Videos preserve the full five-second duration and 81 frames. They are existing
browser previews and can be smaller than the native training videos. Video and
poster checksums are verified against the source media manifests.… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/action-world-model-atlas-1500-media-20260914.rlbench_joint_vel_action_lerobot_trainThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "panda",
"total_episodes": 1800,
"total_frames": 375567,
"total_tasks": 202,
"total_videos": 0,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:1800"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/daixianjie/rlbench_joint_vel_action_lerobot_train.robot-action-prediction-dataset
Robotic Action Prediction Dataset
Dataset Description
This dataset contains triplets of (current observation, action instruction, future observation) for training models to predict future frames of robotic actions.
Dataset Structure
Data Fields
current_frame: Input image (RGB) of the current observation
instruction: Textual description of the action to perform
future_frame: Target image (RGB) showing the expected outcome 50 frames later… See the full description on the dataset page: https://huggingface.co/datasets/bryandts/robot-action-prediction-dataset.Human_Action_Recognition
Dataset Summary
A dataset from kaggle. origin: https://dphi.tech/challenges/data-sprint-76-human-activity-recognition/233/data
Introduction
The dataset features 15 different classes of Human Activities.
The dataset contains about 12k+ labelled images including the validation images.
Each image has only one human activity category and are saved in separate folders of the labelled classes
PROBLEM STATEMENT
Human Action Recognition (HAR) aims to understand… See the full description on the dataset page: https://huggingface.co/datasets/Bingsu/Human_Action_Recognition.peg_04_16_cam0and1_cam_action_no_cropThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 65,
"total_frames": 31241,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:65"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/radoolonto/peg_04_16_cam0and1_cam_action_no_crop.dual_needle_concat_action_staticThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unity",
"total_episodes": 250,
"total_frames": 97953,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:250"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/inaas/dual_needle_concat_action_static.ActionEQA
ActionEQA: Action Interface for Embodied Question Answering
Tianwei Bao1* · Qineng Wang1* · Kangrui Wang1 · Mingkai Deng2 · Guangyi Liu5 · Jiayuan Mao3
Larry Birnbaum1 · Zhiting Hu4 · Eric P. Xing2,5 · Zhaoran Wang1 · Manling Li1
1 Northwestern University 2 Carnegie Mellon University 3 UPenn
4 UC San Diego 5 MBZUAI
* Equal contribution
ActionEQA is the first action-centric Embodied Question Answering (EQA) benchmark designed to systematically evaluate… See the full description on the dataset page: https://huggingface.co/datasets/TianweiBao/ActionEQA.openp2p-action-clips-media-3000-20260911drive-actionactionnet_3k_og
actionnet_3k_og — training subset, archived
Everything actionnet_grasp_2b_480 / _720 opens, plus the LeRobot data/ and meta/
that describe the same episodes: ~35.0 GB in 9 archives instead of ~14,900 loose files.
A 13th archive set, videos_15fps_768x432/, comes from a different store — read the
note below before using it.
This is still a subset of a larger tree. The source also holds videos/,
first_frames/, grasp_frames/, source_mask_merge/, rendering_videos/, gtdepth_s0/
and… See the full description on the dataset page: https://huggingface.co/datasets/seungkukim/actionnet_3k_og.video-web-actionours_joint_actionjam-actions-v0
Dataset Card for jam-actions-v0 (public subset)
Version: 0.5.1 — a documentation-only revision of the 0.5.0 record cut. No record, split or eval artifact changed; the card gained the fine-tuning evaluation banner and the "What's in a record" walkthrough, which had been added on Hugging Face and lived nowhere else.
Records built: 2026-07-11 Source tag: jam-actions-v0-0.5.0-cut-2026-07-11 (record-content correction release — Bach BWV 846 errata 001 + 002; see RELEASE_NOTES.md… See the full description on the dataset page: https://huggingface.co/datasets/mcp-tool-shop/jam-actions-v0.dual_needle_concat_action_two_wristsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unity",
"total_episodes": 250,
"total_frames": 97953,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:250"},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/inaas/dual_needle_concat_action_two_wrists.shotpath-action-diagnostic-venuslike-eval-20260709# ShotPath Action Diagnostic Venus-like Eval 20260709
This bundle contains the LLM-audited pure-operation diagnostic set for Venus-like evaluation.
Files:
action_diagnostic_pure_operation.jsonl: 908 examples after leakage audit.
images/: image files referenced by the jsonl.
scripts/eval_action_diagnostic_qwen25vl.py: Qwen2.5-VL base/LoRA evaluator.
scripts/run_action_diagnostic_venuslike_eval_server.sh: server runner for base7b, stage1 step200, stage2 step200.
Default server paths in the… See the full description on the dataset page: https://huggingface.co/datasets/purefall/shotpath-action-diagnostic-venuslike-eval-20260709.action_1_AUGcs2-action-inference-test
CS2 战术 Action 推理测试集
本测试集用于 WAN I2V 的战术动作定性测试。每个小类只保留 1 张真实比赛 POV 第一帧,以及两种英文文本条件;本版不提供 GT 视频。第一帧来源依据 parse-dem 的 events.csv、game_events.csv 或逐 tick 状态对齐到 opencs2_matches* 视频。
数据约定
共 45 个 case、9 个大类。
每个 case 只有一张 832x480 的 first_frame.png,作为 WAN I2V 条件图;不裁剪或复制 GT clip。首帧优先选择正常持械、水平视角、无遮挡且较开阔的画面。
prompt.txt 是完整英文 prompt,包含首帧可见环境、初始持械状态、画面保持要求和整段唯一动作变化。
chunk_prompts.json 固定包含 5 个英文 prompt,依次描述期望生成视频的 0-1、1-2、2-3、3-4、4-5 秒。
metadata.json… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/cs2-action-inference-test.dual_needle_concat_action_two_wrists_staticThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unity",
"total_episodes": 250,
"total_frames": 97953,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:250"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/inaas/dual_needle_concat_action_two_wrists_static.robotwin-action-prediction-dataset
Robotic Action Prediction Dataset
Dataset Description
This dataset contains triplets of (current observation, action instruction, future observation) for training models to predict future frames of robotic actions.
Dataset Structure
Data Fields
current_frame: Input image (RGB) of the current observation
instruction: Textual description of the action to perform
future_frame: Target image (RGB) showing the expected outcome 50 frames later… See the full description on the dataset page: https://huggingface.co/datasets/bryandts/robotwin-action-prediction-dataset.lpc-action-pixel-art-diffusion
LPC Action Pixel Art Diffusion Dataset
This dataset provides LPC-style action spritesheets for training
image-conditional diffusion models, where a 4-view character image is used
as the conditioning input and an action spritesheet is generated as the output.
It is developed as part of an undergraduate Third Year Project and is intended
for research and educational use.
The accompanying training code, preprocessing scripts, and experiments are
available in the project GitHub… See the full description on the dataset page: https://huggingface.co/datasets/carlosuperb/lpc-action-pixel-art-diffusion.drive-actionaction-prompt-checkpoint-video-eval-20260828
Action-Prompt checkpoint-10000 video evaluation
This dataset contains fresh end-to-end Wan video rollouts used by the companion gallery Space. It contains 24 generated MP4s (four experiments, six samples each), 24 ground-truth clips, input/generated frame PNGs, and per-sample JSON receipts. Every MP4 is 81 frames at 832x480 and 16 fps.
The global-env experiments use the 2026-08-27 environment-prompt manifests and chunk_causal=true. The self-attn-full experiments use the… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/action-prompt-checkpoint-video-eval-20260828.thought_action_bid_only_version_3
Dataset Card for "thought_action_bid_only_version_3"
More Information needed
peg_04_16_cam1_cam_action_no_cropThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 65,
"total_frames": 31241,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:65"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/radoolonto/peg_04_16_cam1_cam_action_no_crop.VideoGUI-Action
Dataset Card for VideoGUI-Action
This is the Action execution part of VideoGUI benchmark.
We uploaded the full metadata, recording at Google Drive.
Data Fields
"app" (str): The name of the software being used.
"task_id" (str): A unique identifier for each full task.
"screenshot_start" (image): The screenshot image representing the state at the beginning of the action.
"action_type" (str): The type of action performed e.g., (right) click, drag, type, scroll… See the full description on the dataset page: https://huggingface.co/datasets/VideoGUI/VideoGUI-Action.Vegetable_actionThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "Songling",
"total_episodes": 172,
"total_frames": 28475,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:172"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/owenchan2001/Vegetable_action.actionnet-3k-og-lerobot
ActionNet 3k — original resolution, corrected state mapping
3,027 Fourier GR1-T1 episodes as LeRobot v2.1: the camera at its native 1280x800, and the robot
state resampled onto the video's own clock with the offset measured rather than guessed.
This is a re-derivation of an existing subset, not new data. It exists because the two conversions
people were using each lose something this one keeps — and because the state/video alignment in
both of them is wrong in a way that is… See the full description on the dataset page: https://huggingface.co/datasets/glory-hyeok/actionnet-3k-og-lerobot.
