CoolFace
Datasetpublic

AxonData/human-movement-pov-dataset-for-robotics

VLA Egocentric Video Dataset — 100+ Hours of 4K Human Manipulation 100+ hours of first-person (egocentric POV) video of real humans performing real-life hand tasks for training vision-language-action (VLA) models, imitation learning policies, and embodied AI systems Key Highlights 100+ hours of real-world egocentric video Continuous, uncut - full task arcs (preparation → execution → result) Head-mounted POV - matches robot wrist/head sensor geometry better than… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/human-movement-pov-dataset-for-robotics.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes102downloads
Dataset Card

VLA Egocentric Video Dataset — 100+ Hours of 4K Human Manipulation

100+ hours of first-person (egocentric POV) video of real humans performing real-life hand tasks for training vision-language-action (VLA) models, imitation learning policies, and embodied AI systems

Key Highlights

  • —100+ hours of real-world egocentric video
  • —Continuous, uncut - full task arcs (preparation → execution → result)
  • —Head-mounted POV - matches robot wrist/head sensor geometry better than third-person
  • —Diverse tasks - repair, assembly, sewing, household, outdoor work, electronics, gardening, bike maintenance
  • —Manually verified - every clip reviewed for hand visibility, task continuity, signal clarity
  • —Commercial license available - full version cleared for production ML training, VLA pretraining, foundation model training

Use This Dataset For

  • —VLA (vision-language-action) model pretraining - large-scale aligned egocentric corpus
  • —Imitation learning / behavior cloning for robot manipulation policies
  • —Hand-object interaction (HOI) research - grasping, tool use, dexterous manipulation
  • —Embodied AI / physical AI foundation model training
  • —Humanoid robot manipulation - household and outdoor task generalization
  • —Action recognition at sub-action granularity (reach, grasp, lift, transport, release)
  • —Video understanding - long-form continuous task structure

How This Compares to Public Egocentric Datasets

DatasetDurationLicenseResolutionContinuous?Best for
Axon Labs VLA Egocentric (full)100+ hoursCommercial4K @ 30 FPSYes (5–60 min uncut)Production VLA, imitation learning, embodied AI
Ego4D (Meta)3,670 hoursResearch onlyMixedMixedDaily activity recognition, perception
EgoDex (Apple)829 hoursResearch onlyApple Vision Pro specShort tabletop clipsDexterous tabletop manipulation
EgoExo4D (Meta)1,286 hoursResearch onlyMixedSkilled activitiesEgo-exo synchronized learning
EPIC-Kitchens100 hoursResearch only1080pKitchen workflowsKitchen action recognition

Ego4D, EgoDex, and EgoExo4D remain leading academic resources for their respective tasks. This dataset complements them by providing the commercial license, 4K resolution, and continuous task arcs required for production ML training and VLA pretraining at scale

Why VLA / Physical AI Teams Need This Data

Recent work establishes human egocentric video as a first-class training input for robot manipulation, not a fallback:

  • —EgoMimic (CoRL 2024) showed that 1 hour of human egocentric data contributed more to policy performance than 1 additional hour of robot teleoperation
  • —EgoDex (Apple, 2025) defined the state of the art for dexterous tabletop manipulation training
  • —EgoScale (NVIDIA, 2026) demonstrated a log-linear scaling law - every doubling of human egocentric hours produces a predictable downstream task improvement

The bottleneck is no longer whether human data transfers - it's whether you can get commercially licensed, high-resolution, manipulation-relevant egocentric data at scale. That's what this dataset provides

Full version of dataset is available for commercial usage — leave a request on our website [Axonlabs](https://axonlab.ai/?utm_source=hugging-face&utm_medium=cpc&utm_campaign=profile&utm_content=profile_link) to purchase the dataset 💰

What Makes This Dataset Unique

  • —Commercial license - none of the leading academic alternatives offer this
  • —Continuous uncut footage - full task arcs, not curated short clips
  • —Real-life hand tasks - household, outdoor, repair, maintenance — broader than tabletop-only datasets like EgoDex

FAQ

Q: Is this dataset suitable for VLA (vision-language-action) model training? Yes, this is one of the primary intended use cases. Each video is paired with a natural-language task description (task_name), making the dataset directly aligned with the VLA training paradigm popularized by RT-2 and π0. The 4K resolution, continuous task arcs, and diverse manipulation contexts make it well-suited as a pretraining corpus for vision-language-action foundation models

Q: How does this compare to Ego4D and EgoDex? Ego4D is larger (3,670 hours) but research-only and not designed for manipulation specifically. EgoDex is research-only and limited to tabletop tasks. Our dataset is smaller (200+ hours) but commercially licensed, captured at 4K, includes diverse real-life household and outdoor tasks (not just tabletop), and uses head-mounted POV that's closer to robot deployment sensor geometry. For production manipulation training, this matters more than raw size

Q: Can I use this dataset for imitation learning / behavior cloning? Yes. The continuous full-task footage (5–60 minutes per clip, no cuts) captures complete demonstration arcs: preparation, execution, result, which is what behavior cloning policies need. The 4K resolution preserves fine-grained hand-object interaction detail. Pair with robot demonstration data (per the EgoMimic methodology) for best policy performance

Q: Is this dataset suitable for Physical AI / embodied AI foundation models? Yes. The diversity of tasks, environments, and tools makes it valuable as part of a large-scale corpus for foundation model pretraining. Following the EgoScale scaling-law results, additional commercially licensed egocentric hours produce predictable downstream gains

Q: What hardware was used for recording? Smile 5K camera at 4K, 30 FPS. Head-mounted (majority) and chest-mounted setups for first-person perspective. Original camera MP4 format preserved: no re-encoding, no compression loss. Full hardware specs and capture protocol available on request for the commercial version

keywords: VLA dataset, vision language action dataset, egocentric video dataset, POV video dataset, first person video dataset, robot manipulation dataset, imitation learning dataset, behavior cloning dataset, hand object interaction dataset, embodied AI dataset, physical AI dataset, humanoid robot dataset, robot training data, 4K video dataset, real world manipulation dataset, ego4d alternative, egodex alternative

Visit us at **Axonlabs** to request a full version of the dataset for commercial usage