datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xr-motion-dataset-catalogue
XR Motion Dataset Catalogue
Overview
The XR Motion Dataset Catalogue, accompanying our paper "Navigating the Kinematic Maze: A Comprehensive Guide to XR Motion Dataset Standards," standardizes and simplifies access to Extended Reality (XR) motion datasets. The catalogue represents our initiative to streamline the usage of kinematic data in XR research by aligning various datasets to a consistent format and structure.
Dataset Specifications
All datasets in this… See the full description on the dataset page: https://huggingface.co/datasets/cschell/xr-motion-dataset-catalogue.MotionDecode
🆕 Open-Source Release: Unitree G1 Retargeted Data
!!We are releasing 1000 hours of robot-ready motion trajectories retargeted to the Unitree G1 humanoid. All data is provided in CSV format under the samples/ directory. Please indicate the source of the data when using it: from Chingmu.
ChingMu 1000-Hour Embodied Motion Dataset
High-precision optical motion capture data for humanoid robots, dexterous hands, embodied AI, and virtual production.… See the full description on the dataset page: https://huggingface.co/datasets/CMRobot/MotionDecode.MotionBench
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
[🍎 Project Page] [📖 arXiv Paper] [📊 Dataset] [💻 GitHub] [🏆 Leaderboard] [🏆 HF Leaderboard]
MotionBench is a comprehensive evaluation benchmark designed to assess the fine-grained motion comprehension of video understanding models. It evaluates models' motion-level perception through six primary categories of motion-oriented question types and includes data… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/MotionBench.MotionHub
MotionHub
MotionHub is a curated multi-domain human-motion dataset collection released for training and evaluating generalist motion models. The released version contains motion, language, music, speech, and two-person interaction supervision in a unified MotionHub annotation format.
This public release is used by VersatileMotion (ECCV 2026). Every subset listed below has been visually inspected, converted to the repository SMPL-H convention, re-split where needed, and uploaded… See the full description on the dataset page: https://huggingface.co/datasets/ZeyuLing/MotionHub.Motion-o-MCoT-PLM-motion-keyframes
Motion-o-MCoT (PLM + motion keyframes)
Subset of STGR: STR_plm_rdcap rows with <motion in reasoning_process, plus sharded keyframes under videos/stgr/plm/kfs/.
Train split: 3,168 examples (see export_manifest.json in the repo for exact export stats).
Keyframes: JPEGs are stored under shard subfolders (e.g. videos/stgr/plm/kfs/plm_0150/…) so each directory stays under Hugging Face file-count limits. Each key_frames[].path in the JSON is relative to videos/stgr/plm/kfs/ (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/bishoygaloaa/Motion-o-MCoT-PLM-motion-keyframes.h2_retargeted_motions
H2 retargeted motions (full corpus)
129,785 Bones-SEED motions retargeted onto the Unitree H2 (31 DoF, 32
bodies) with SOMA Retargeter.
This is the corpus that trained
junsooki/h2_checkpoints.
A 10-clip sample for quick trials lives at
junsooki/h2_reference_motions.
Format
One joblib pickle per clip, keyed by motion name:
field
shape
meaning
dof
(T, 31)
joint angles, IsaacLab order
root_trans_offset
(T, 3)
root position, metres
root_rot
(T, 4)
root… See the full description on the dataset page: https://huggingface.co/datasets/junsooki/h2_retargeted_motions.motionatlas-data
MotionAtlas-Data
MotionAtlas-Data is a large-scale dataset for region-aware motion captioning. Instead of describing a whole clip globally, each sample pairs a video with a spatiotemporal region and a precise description of the motion inside that region, reducing visual clutter and motion entanglement.
159K high-quality region-level motion captioning samples
Built with a scalable pipeline using self-bootstrap refinement to suppress fine-grained hallucinations
Designed to… See the full description on the dataset page: https://huggingface.co/datasets/maxLWSv2/motionatlas-data.human-motion-tracking-deeplabcutThis dataset is used to adapt DeepLabCut for Human motion tracking.
Structure of the dataset
videos contains 100+ videos of 4 candidates recorded during a game of darts.
labeled-data contains labels on the corresponding frames of the videos. These labels are used to adapt DeepLabCut for human motion tracking. Under labeled-data there are 2 folders for every video.
video_name has all the relevant frames extracted from the video, xy coordinates of the labels in the csv file and the… See the full description on the dataset page: https://huggingface.co/datasets/GT-Neuronext/human-motion-tracking-deeplabcut.human-motion-tracking-deeplabcutThis dataset is used to adapt DeepLabCut for Human motion tracking.
Structure of the dataset
videos contains 100+ videos of 4 candidates recorded during a game of darts.
labeled-data contains labels on the corresponding frames of the videos. These labels are used to adapt DeepLabCut for human motion tracking. Under labeled-data there are 2 folders for every video.
video_name has all the relevant frames extracted from the video, xy coordinates of the labels in the csv file and the… See the full description on the dataset page: https://huggingface.co/datasets/pratikshapai/human-motion-tracking-deeplabcut.molmo-motion-1m
MolmoMotion-1M
MolmoMotion-1M is a dataset of 3D point-trajectory annotations curated across
seven video corpora — ego-centric manipulation, real-world robot teleoperation,
dynamic real-world scenes, and simulator renders. Each clip ships motion-filtered 3D
tracks (and, for most datasets, 2D pixel tracks), a short action caption, per-frame
camera, and a train/test split — all frame-aligned to the source video.
Datasets
We do not re-host the original videos, and… See the full description on the dataset page: https://huggingface.co/datasets/allenai/molmo-motion-1m.QtMeshEditor-motion-corpus
QtMeshEditor Motion Corpus
A permissively-licensed animated-humanoid corpus: rigged 3D characters
with skeletal animation clips, harvested for
QtMeshEditor's text-to-motion
v2 work (epic #837)
— the template clip library and the training set for a from-scratch
flow-matching motion model.
Every asset is CC0 or CC-BY — nothing here derives from Mixamo, LAFAN1,
Bandai-Namco, AMASS/HumanML3D, or game rips (all license-poisoned for
commercial redistribution). That makes this corpus —… See the full description on the dataset page: https://huggingface.co/datasets/fernandotonon/QtMeshEditor-motion-corpus.motion-capture-data
Motion Capture Data
Dataset Description
This dataset contains human motion capture data and expert demonstrations for humanoid robot control. The goal is training a base decoder-only transformer model to output
motion control instructions given a motion prefix, then further finetuning that base to specific anime characters. The latter part will need additional datasets not provided
here.
Overview
The dataset consists of expert demonstrations collected by… See the full description on the dataset page: https://huggingface.co/datasets/nekomata-project/motion-capture-data.MotionDecode
🆕 Open-Source Release: Unitree G1 Retargeted Data
!!We are releasing 1000 hours of robot-ready motion trajectories retargeted to the Unitree G1 humanoid. All data is provided in CSV format under the samples/ directory. Please indicate the source of the data when using it: from Chingmu.
ChingMu 1000-Hour Embodied Motion Dataset
High-precision optical motion capture data for humanoid robots, dexterous hands, embodied AI, and virtual production.… See the full description on the dataset page: https://huggingface.co/datasets/Linmove/MotionDecode.excavator-motion
Introduction
Excavator-motion is a large-scale dataset of excavator motion trajectories, encompassing three distinct excavator models. Each H5 file in the dataset records a complete motion trajectory from the initiation of digging to the full loading of a truck.
Compared to our another released dataset Excavator-video, excavator-motion not only contains RGB images from identical view but also includes elevation maps in global view and fine-grained excavator joint angle… See the full description on the dataset page: https://huggingface.co/datasets/fuxi-robot/excavator-motion.Motion-Xplusplus
Data for Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset
Here, we release our dataset, "Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset," which includes various motion modalities. It features 2D keypoints for mesh recovery and motion generation. Additionally, we provide SMPL-X annotations that differentiate between translations and orientations in camera and world coordinate systems. The dataset also includes action descriptions… See the full description on the dataset page: https://huggingface.co/datasets/YuhongZhang/Motion-Xplusplus.motion-smd-data
Motion-SMD Data
Data release for "Encoder-Free Human Motion Understanding via Structured Motion Descriptions".
🌐 Project page: https://yaozhang182.github.io/motion-smd/
💻 Code: https://github.com/yaozhang182/motion-smd
🤗 LoRA adapters: https://huggingface.co/zyyy12138/motion-smd-lora
📄 Paper (arXiv): https://arxiv.org/abs/2604.21668
What's here
Four subdirectories, each with its own README.md describing files, provenance, and license:
Subdir
Contents
Our… See the full description on the dataset page: https://huggingface.co/datasets/zyyy12138/motion-smd-data.motionxImageNet-C-motion_blur-severity_5waymo_open_dataset_motion_v_1_3_0MotionBlind
MotionBlind
A contrastive benchmark for physical-motion perception in Video-LLMs.
📄 Paper: MotionBlind: Probing the Illusion of Motion Understanding in Video-LLMs
A Video-LLM can watch two clips of the same person in the same room and name every
object in both, yet fail to tell you which one moves faster, which way a hand travels,
or how far a box slides. MotionBlind is a contrastive, minimal-pairs benchmark
built to expose exactly that gap: pairs of near-identical clips that… See the full description on the dataset page: https://huggingface.co/datasets/augmentedcognitionlab/MotionBlind.Kimodo-Motion-Gen-Benchmark
Kimodo Human Motion Generation Benchmark
Kimodo Codebase, Benchmark Documentation
Dataset Description:
This dataset provides the necessary metadata to construct the suite of test cases that make up the Kimodo human motion generation benchmark. This includes test cases that evaluate text-following for the text-to-motion task, along with constraint-following for constraint-conditioned motion generation.
The benchmark is constructed from the SOMA uniform version of the… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Kimodo-Motion-Gen-Benchmark.MotionEdit-Trainminecraft-motion-action-datasetanimacy-human-motion
squaredcuber/animacy-human-motion
Canonical human conversational motion (animacy.human.v1, 30 Hz) captured with
animacy — the open interaction layer for
expressive robots. Each clip is motion.parquet (28 channels: head 6-DoF, gaze,
brows, eyes, mouth, torso, a puppet arm chain, speaking/validity flags),
audio.wav (16 kHz mono, same clock) and meta.json.
Retarget any clip to a robot with animacy retarget --robot <name>; robots are
one ROBOT.md each (Autonomous Lamp and Reachy… See the full description on the dataset page: https://huggingface.co/datasets/squaredcuber/animacy-human-motion.MotionFix
MotionFix MotionHub Format
This dataset contains the processed MotionFix motion-editing data in the MotionHub format.
It stores paired source and target motions as SMPL-H 52-joint parameter files, plus editing instructions and split annotations.
Please also follow the license and terms of the original MotionFix dataset.
Structure
MotionFix/
├── smplh_52/
│ ├── train/
│ │ ├── 000000_000499/
│ │ ├── 000500_000999/
│ │ └── ...
│ ├── val/
│ └── test/… See the full description on the dataset page: https://huggingface.co/datasets/ZeyuLing/MotionFix.StarCraft-MotionMotionMillion
🔑 Key Features
Over 2000 hours of high-quality human motion captured from web-scale human video data, covering:
Martial Arts (23.7%)
Fitness (26.4%)
Performance (17.5%)
Dance (14.9%)
Non-Human (2.9%)
Sports (2.4%)
Over 20 detailed annotations per motion, including:
Age
Body Characteristics
Movement Styles
Emotions
Environments
👨🏫 Get Started
Download the Dataset
To download the full dataset, use the following code. If you encounter any… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/MotionMillion.MotionSightThis is the dataset proposed in our paper MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs.
We split the dataset into multiple small files, you can recover by cat:
cat MotionSightDataset_part* > MotionSightDataset.zip
unzip MotionSightDataset
Project Page | Github
Motion247
Unitree G1 Retargeted Motion Dataset
A comprehensive collection of human motion data retargeted to the Unitree G1 humanoid robot. This dataset contains 174 motion sequences spanning diverse categories including locomotion, dance, sports, and expressive movements.
Dataset Overview
This dataset provides motion data in a unified format suitable for training humanoid robot control policies. All motions have been retargeted from human motion capture data (SMPL format)… See the full description on the dataset page: https://huggingface.co/datasets/Motion247/Motion247.DOMINO_absolute_motion_v2
DOMINO Absolute Motion v2
DOMINO_absolute_motion_v2 is the complete packed training corpus for
DynamicWAM's exact-simulator-time motion pipeline. It is a training-ready
derivative of H-EmbodVis/DOMINO,
not a copy of the raw RGB dataset.
Every sample was packed as one aligned record containing video latents,
action/state targets, frame indices, a language-group identifier, and four
history intervals of absolute motion descriptors. The alignment is fixed at
conversion time;… See the full description on the dataset page: https://huggingface.co/datasets/KhalilGao/DOMINO_absolute_motion_v2.
