datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MotionDecode
🆕 Open-Source Release: Unitree G1 Retargeted Data
!!We are releasing 1000 hours of robot-ready motion trajectories retargeted to the Unitree G1 humanoid. All data is provided in CSV format under the samples/ directory. Please indicate the source of the data when using it: from Chingmu.
ChingMu 1000-Hour Embodied Motion Dataset
High-precision optical motion capture data for humanoid robots, dexterous hands, embodied AI, and virtual production.… See the full description on the dataset page: https://huggingface.co/datasets/CMRobot/MotionDecode.xr-motion-dataset-catalogue
XR Motion Dataset Catalogue
Overview
The XR Motion Dataset Catalogue, accompanying our paper "Navigating the Kinematic Maze: A Comprehensive Guide to XR Motion Dataset Standards," standardizes and simplifies access to Extended Reality (XR) motion datasets. The catalogue represents our initiative to streamline the usage of kinematic data in XR research by aligning various datasets to a consistent format and structure.
Dataset Specifications
All datasets in this… See the full description on the dataset page: https://huggingface.co/datasets/cschell/xr-motion-dataset-catalogue.MotionBench
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
[🍎 Project Page] [📖 arXiv Paper] [📊 Dataset] [💻 GitHub] [🏆 Leaderboard] [🏆 HF Leaderboard]
MotionBench is a comprehensive evaluation benchmark designed to assess the fine-grained motion comprehension of video understanding models. It evaluates models' motion-level perception through six primary categories of motion-oriented question types and includes data… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/MotionBench.Cauldron-JA
Dataset Card for The Cauldron-JA
Dataset description
The Cauldron-JA is a Vision Language Model dataset that translates 'The Cauldron' into Japanese using the DeepL API. The Cauldron is a massive collection of 50 vision-language datasets (training sets only) that were used for the fine-tuning of the vision-language model Idefics2.
To create a Japanese Vision Language Dataset, datasets related to OCR, coding, and graphs were excluded because translating them into Japanese… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Cauldron-JA.MotionHub
MotionHub
MotionHub is a curated multi-domain human-motion dataset collection released for training and evaluating generalist motion models. The released version contains motion, language, music, speech, and two-person interaction supervision in a unified MotionHub annotation format.
This public release is used by VersatileMotion (ECCV 2026). Every subset listed below has been visually inspected, converted to the repository SMPL-H convention, re-split where needed, and uploaded… See the full description on the dataset page: https://huggingface.co/datasets/ZeyuLing/MotionHub.h2_retargeted_motions
H2 retargeted motions (full corpus)
129,785 Bones-SEED motions retargeted onto the Unitree H2 (31 DoF, 32
bodies) with SOMA Retargeter.
This is the corpus that trained
junsooki/h2_checkpoints.
A 10-clip sample for quick trials lives at
junsooki/h2_reference_motions.
Format
One joblib pickle per clip, keyed by motion name:
field
shape
meaning
dof
(T, 31)
joint angles, IsaacLab order
root_trans_offset
(T, 3)
root position, metres
root_rot
(T, 4)
root… See the full description on the dataset page: https://huggingface.co/datasets/junsooki/h2_retargeted_motions.Motion-o-MCoT-PLM-motion-keyframes
Motion-o-MCoT (PLM + motion keyframes)
Subset of STGR: STR_plm_rdcap rows with <motion in reasoning_process, plus sharded keyframes under videos/stgr/plm/kfs/.
Train split: 3,168 examples (see export_manifest.json in the repo for exact export stats).
Keyframes: JPEGs are stored under shard subfolders (e.g. videos/stgr/plm/kfs/plm_0150/…) so each directory stays under Hugging Face file-count limits. Each key_frames[].path in the JSON is relative to videos/stgr/plm/kfs/ (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/bishoygaloaa/Motion-o-MCoT-PLM-motion-keyframes.yue2-minted-corpus
YuE2 minted corpus
Songs generated by m-a-p/YuE2-3B (with YuE2-Vae) from
MusicForge plan_50k prompts, queued genre-round-robin (355 genres, ~35% instrumental), seeds from the plan. Every track keeps the exact
intermediate the model produced, so the set is a labelled corpus for training an audio → semantic-token encoder (the piece YuE2 does not ship)
and a grammar regularizer for AR fine-tunes.
tracks/<pid>/
file
content
audio.flac
48 kHz stereo 24-bit
semantic.npy… See the full description on the dataset page: https://huggingface.co/datasets/Mothersuperior/yue2-minted-corpus.DatasetDemo
Motus Training Dataset Demo
Introduction
This repository serves as a demonstration dataset illustrating the required data format for training the Motus model. It provides a reference for structuring your data to ensure compatibility with the training pipeline. The demo data come from Robotwin-clean benchmark.
Directory Structure
Data is generally organized following a hierarchy of Dataset Name, Task Name (optional), and Data Type.
Standard Format:… See the full description on the dataset page: https://huggingface.co/datasets/motus-robotics/DatasetDemo.motMOT17
MOT17
MOT17 is a benchmark dataset for single-camera multi-object tracking (MOT), focused primarily on pedestrian tracking in real-world video sequences. This Hugging Face repository provides the MOT17 data in the original MOTChallenge-style structure for research, benchmarking, training, and evaluation of multi-object tracking systems.
MOT17 extends MOT16 with more accurate ground-truth annotations and provides each sequence with three public detection sets:
DPM
Faster R-CNN /… See the full description on the dataset page: https://huggingface.co/datasets/Lekim89/MOT17.Motius-Leaderboard-Cases
Motius Leaderboard Case Assets
This dataset stores compact browser assets for the all-case comparison pages in
the public Motius leaderboards. It is a
visualization companion, not a training or evaluation dataset.
Folder
Population
Comparison
m2t-humanml3d-smpl/
4,400
HumanML3D input motion with TM2T, MotionGPT, MotionGPT3, and VerMo captions
t2m-humanml3d-smpl/
4,042
HumanML3D selected captions with every released T2M output
babel-sequential-smpl/
1,295
BABEL GT… See the full description on the dataset page: https://huggingface.co/datasets/ZeyuLing/Motius-Leaderboard-Cases.MOT20
MOT20
MOT20 is a benchmark dataset for single-camera multi-object tracking (MOT) and pedestrian detection in very crowded real-world scenes. This Hugging Face repository provides MOT20 in the original MOTChallenge-style structure for research, benchmarking, training, and evaluation of multi-object tracking systems.
MOT20 was introduced to stress-test MOT methods in high-density pedestrian scenes, including crowded squares, indoor train stations, stadium exits, and pedestrian… See the full description on the dataset page: https://huggingface.co/datasets/Lekim89/MOT20.Japan-Open-Driving-Dataset-Sample
Japan Open Driving Dataset Sample
Overview
This repository contains a sample subset of the Japan Open Driving Dataset, a large-scale autonomous driving dataset comprising over 100 hours of driving data collected in Tokyo, Japan.
The data is stored in nuScenes format and can be loaded with the nuscenes-devkit.
In addition to sensor data and 3D annotations, this dataset includes virtual captioned data for training Vision-Language-Model (VLM) and Vision-Language-Action (VLA)… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Japan-Open-Driving-Dataset-Sample.visdrone-mot
Dataset Card for VisDrone2019-DET
This is a FiftyOne version of the VisDrone2019-DET dataset with 8629 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
# Load the dataset
# Note: other available arguments include 'max_samples', 'persistent`, 'overwrite' etc
dataset = fouh.load_from_hub("Voxel51/VisDrone2019-DET")
# Launch the App
session =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/visdrone-mot.DeepSea-MOT
DeepSea MOT
DeepSea MOT is a benchmark dataset for multi-object tracking on deep-sea video.
Dataset Description
DeepSea MOT consists of 4 video sequences (2 midwater, 2 benthic) with a total of 2,400 frames and 57,376 annotated objects comprising 188 tracks. The videos were captured by the Monterey Bay Aquarium Research Institute (MBARI) using remotely operated vehicles (ROVs) Doc Ricketts and Ventana in deep-sea environments, showcasing a variety of marine species and… See the full description on the dataset page: https://huggingface.co/datasets/MBARI-org/DeepSea-MOT.motionatlas-data
MotionAtlas-Data
MotionAtlas-Data is a large-scale dataset for region-aware motion captioning. Instead of describing a whole clip globally, each sample pairs a video with a spatiotemporal region and a precise description of the motion inside that region, reducing visual clutter and motion entanglement.
159K high-quality region-level motion captioning samples
Built with a scalable pipeline using self-bootstrap refinement to suppress fine-grained hallucinations
Designed to… See the full description on the dataset page: https://huggingface.co/datasets/maxLWSv2/motionatlas-data.motor_two_wheel_rider
Dataset Card for motor_two_wheel_rider
MOTOR (MOtorized TwO-wheeler Rider) is the first large-scale, multi-view, multimodal dataset dedicated to understanding two-wheeler rider behavior in dense, unstructured traffic conditions typical of the Global South. The full dataset comprises 1,629 annotated sequences (~25 hours) from 16 riders collected across diverse traffic scenarios in India.
This repository contains a subset of the MOTOR dataset imported into FiftyOne format for easy… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/motor_two_wheel_rider.CoVLA-Dataset
CoVLA-Dataset
WACV 2025 Oral
CoVLA-Dataset is a dataset comprising real-world driving videos spanning more than 80 hours. This dataset leverages a novel, scalable approach based on automated data processing and a caption generation pipeline to generate accurate driving trajectories paired with detailed natural language descriptions of driving environments and maneuvers. It includes 10,000 30-second video clips, paired with trajectory targets and language annotations generated from… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/CoVLA-Dataset.molmo-motion-1m
MolmoMotion-1M
MolmoMotion-1M is a dataset of 3D point-trajectory annotations curated across
seven video corpora — ego-centric manipulation, real-world robot teleoperation,
dynamic real-world scenes, and simulator renders. Each clip ships motion-filtered 3D
tracks (and, for most datasets, 2D pixel tracks), a short action caption, per-frame
camera, and a train/test split — all frame-aligned to the source video.
Datasets
We do not re-host the original videos, and… See the full description on the dataset page: https://huggingface.co/datasets/allenai/molmo-motion-1m.human-motion-tracking-deeplabcutThis dataset is used to adapt DeepLabCut for Human motion tracking.
Structure of the dataset
videos contains 100+ videos of 4 candidates recorded during a game of darts.
labeled-data contains labels on the corresponding frames of the videos. These labels are used to adapt DeepLabCut for human motion tracking. Under labeled-data there are 2 folders for every video.
video_name has all the relevant frames extracted from the video, xy coordinates of the labels in the csv file and the… See the full description on the dataset page: https://huggingface.co/datasets/GT-Neuronext/human-motion-tracking-deeplabcut.human-motion-tracking-deeplabcutThis dataset is used to adapt DeepLabCut for Human motion tracking.
Structure of the dataset
videos contains 100+ videos of 4 candidates recorded during a game of darts.
labeled-data contains labels on the corresponding frames of the videos. These labels are used to adapt DeepLabCut for human motion tracking. Under labeled-data there are 2 folders for every video.
video_name has all the relevant frames extracted from the video, xy coordinates of the labels in the csv file and the… See the full description on the dataset page: https://huggingface.co/datasets/pratikshapai/human-motion-tracking-deeplabcut.QtMeshEditor-motion-corpus
QtMeshEditor Motion Corpus
A permissively-licensed animated-humanoid corpus: rigged 3D characters
with skeletal animation clips, harvested for
QtMeshEditor's text-to-motion
v2 work (epic #837)
— the template clip library and the training set for a from-scratch
flow-matching motion model.
Every asset is CC0 or CC-BY — nothing here derives from Mixamo, LAFAN1,
Bandai-Namco, AMASS/HumanML3D, or game rips (all license-poisoned for
commercial redistribution). That makes this corpus —… See the full description on the dataset page: https://huggingface.co/datasets/fernandotonon/QtMeshEditor-motion-corpus.STRIDE-QA-Dataset
STRIDE-QA Dataset
📦 Dataset
STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks.
Category
Description
Object-centric Spatial QA
Spatial relations between two… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset.motif-1k
Dataset Card for MotIF-1K
MotIF-1K is a robotics motion dataset containing 1,022 demonstrations across 13 task categories, used to benchmark and fine-tune vision-language models (VLMs) for motion-based success detection. Each demonstration includes a video of the motion, multiple pre-rendered trajectory visualizations, task instructions, and motion descriptions.
The FiftyOne dataset is a grouped dataset where each group represents one trajectory and each group slice represents a… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/motif-1k.MotionMillion
🔑 Key Features
Over 2000 hours of high-quality human motion captured from web-scale human video data, covering:
Martial Arts (23.7%)
Fitness (26.4%)
Performance (17.5%)
Dance (14.9%)
Non-Human (2.9%)
Sports (2.4%)
Over 20 detailed annotations per motion, including:
Age
Body Characteristics
Movement Styles
Emotions
Environments
👨🏫 Get Started
Download the Dataset
To download the full dataset, use the following code. If you encounter any… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/MotionMillion.MOT17
MOT17
MOT17 is a benchmark dataset for single-camera multi-object tracking (MOT), focused primarily on pedestrian tracking in real-world video sequences. This Hugging Face repository provides the MOT17 data in the original MOTChallenge-style structure for research, benchmarking, training, and evaluation of multi-object tracking systems.
MOT17 extends MOT16 with more accurate ground-truth annotations and provides each sequence with three public detection sets:
DPM
Faster R-CNN /… See the full description on the dataset page: https://huggingface.co/datasets/Morrison1025/MOT17.motion-capture-data
Motion Capture Data
Dataset Description
This dataset contains human motion capture data and expert demonstrations for humanoid robot control. The goal is training a base decoder-only transformer model to output
motion control instructions given a motion prefix, then further finetuning that base to specific anime characters. The latter part will need additional datasets not provided
here.
Overview
The dataset consists of expert demonstrations collected by… See the full description on the dataset page: https://huggingface.co/datasets/nekomata-project/motion-capture-data.synT1CE-motfm-OOD-peds
BraTS-PEDs — qualitative export
24 slices over 8 subjects. Four files per slice:
*_input_t1n.png non-contrast T1, the model input
*_generated.png synthesized T1CE
*_real_t1ce.png ground-truth T1CE (this is what earlier exports lacked)
*_mask.png expert segmentation, where available
Intensities are min-max normalized per slice for display only; all quantitative
figures come from the metric pipeline, not from these PNGs.
MotionDecode
🆕 Open-Source Release: Unitree G1 Retargeted Data
!!We are releasing 1000 hours of robot-ready motion trajectories retargeted to the Unitree G1 humanoid. All data is provided in CSV format under the samples/ directory. Please indicate the source of the data when using it: from Chingmu.
ChingMu 1000-Hour Embodied Motion Dataset
High-precision optical motion capture data for humanoid robots, dexterous hands, embodied AI, and virtual production.… See the full description on the dataset page: https://huggingface.co/datasets/Linmove/MotionDecode.
