datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
human-motion-tracking-deeplabcutThis dataset is used to adapt DeepLabCut for Human motion tracking.
Structure of the dataset
videos contains 100+ videos of 4 candidates recorded during a game of darts.
labeled-data contains labels on the corresponding frames of the videos. These labels are used to adapt DeepLabCut for human motion tracking. Under labeled-data there are 2 folders for every video.
video_name has all the relevant frames extracted from the video, xy coordinates of the labels in the csv file and the… See the full description on the dataset page: https://huggingface.co/datasets/GT-Neuronext/human-motion-tracking-deeplabcut.MUOT_3M-A_3_Million_Frame_Underwater_Object_Tracking_Dataset
🌊 MUOT-3M: The Largest Multimodal Underwater Object Tracking Dataset
Official repository for MUOT-3M📄 MUOT-3M: The Largest Multimodal Underwater Object Tracking Dataset and MUTrack Tracking Method
🚀 Overview
MUOT-3M is currently the largest underwater object tracking dataset, containing over 3 million annotated frames across 3,030 underwater videos with synchronized multimodal annotations.
The benchmark is designed to advance research in:
Underwater object tracking… See the full description on the dataset page: https://huggingface.co/datasets/AhsanBB/MUOT_3M-A_3_Million_Frame_Underwater_Object_Tracking_Dataset.human-motion-tracking-deeplabcutThis dataset is used to adapt DeepLabCut for Human motion tracking.
Structure of the dataset
videos contains 100+ videos of 4 candidates recorded during a game of darts.
labeled-data contains labels on the corresponding frames of the videos. These labels are used to adapt DeepLabCut for human motion tracking. Under labeled-data there are 2 folders for every video.
video_name has all the relevant frames extracted from the video, xy coordinates of the labels in the csv file and the… See the full description on the dataset page: https://huggingface.co/datasets/pratikshapai/human-motion-tracking-deeplabcut.VISEM-Tracking
Dataset Card for VISEM-Tracking
Dataset Summary
VISEM-Tracking is a dataset of human spermatozoa tracking, featuring 20 video recordings of 30 seconds each, captured from wet semen preparations. The dataset contains 29,196 frames with manually annotated bounding-box coordinates and a set of sperm characteristics analyzed by domain experts. In addition to the labeled dataset, unlabeled videos are provided to enable research in self-supervised learning. This dataset serves… See the full description on the dataset page: https://huggingface.co/datasets/sperm-net/VISEM-Tracking.pig-detection-and-tracking
🐷 Dataset for the paper: Benchmarking pig detection and tracking under diverse and challenging conditions
Note: This dataset does not provide a load_dataset interface.
More detailed descriptions about the dataset can be found in the corresponding paper.
⬇️ Downloading the dataset
You can download the dataset using the Hugging Face CLI. First, install the following package:
pip install huggingface_hub
The entire dataset (roughly 25 GB) can then be downloaded as follows:… See the full description on the dataset page: https://huggingface.co/datasets/jonaden/pig-detection-and-tracking.VISEM-Tracking
Dataset Card for VISEM-Tracking
Dataset Summary
VISEM-Tracking is a dataset of human spermatozoa tracking, featuring 20 video recordings of 30 seconds each, captured from wet semen preparations. The dataset contains 29,196 frames with manually annotated bounding-box coordinates and a set of sperm characteristics analyzed by domain experts. In addition to the labeled dataset, unlabeled videos are provided to enable research in self-supervised learning. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/LandyJJ/VISEM-Tracking.Multi-Camera-Multi-Vehicle-Tracking-System
Multi-Camera Multi-Vehicle Tracking System (UAV Dataset)
This dataset contains synchronized multi-UAV video footage and comprehensive tracking annotations, accompanying the paper A Topology-Aware Spatiotemporal Handover Framework for Continuous Multi-UAV Tracking.
It is designed for evaluating real-time multi-camera multi-vehicle tracking (MCMT) systems, focusing on solving trajectory fragmentation and maintaining global identity persistence across isolated UAV fields of view.… See the full description on the dataset page: https://huggingface.co/datasets/jye9/Multi-Camera-Multi-Vehicle-Tracking-System.astribot_tracking_redballSoccerNet-Tracking-RAW-Video
SoccerNet Tracking — Raw Single-Camera Video
Raw single-camera broadcast footage used in the SoccerNet Player Tracking benchmark, covering the 12 single-camera games.
Gated — request access on this page.
File
Description
single_camera_games_public.zip
Raw single-camera video for all 12 tracking games
Download
Using the SoccerNet pip package (recommended):
from SoccerNet.Downloader import SoccerNetDownloader
d =… See the full description on the dataset page: https://huggingface.co/datasets/SoccerNet/SoccerNet-Tracking-RAW-Video.pexels-object-tracking-test-videosgripper_tracking_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 11,
"total_frames": 40060,
"total_tasks": 1,
"total_videos": 22,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:11"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Flova/gripper_tracking_test.XMR_Demo_Object_Tracking
Demo for Hyperspectral Object Tracking
Video spectroscopy beyond the visible spectrum, applied to object tracking in a crowded bus-station scene. Captured with a Cubert Ultris XMR camera — 61 bands per pixel, 430–910 nm, 1080 × 1000 pixels at 15 Hz.
Object tracking is a general computer-vision problem — the "object" could be a vehicle, a container, a piece of equipment, or an animal. In this demo the target class is humans: a crowded public scene with people in… See the full description on the dataset page: https://huggingface.co/datasets/cubert-gmbh/XMR_Demo_Object_Tracking.v14-real-tracking-any-granularity-videosTTA_Tracking
TTA_Tracking
Dataset Description
TTA_Tracking is a ball tracking dataset for table tennis, recorded under realistic conditions at a professional Paralympics level. The dataset is intended to support research in sports video understanding, object detection, and ball tracking under challenging real-world conditions.
Dataset Summary
Task: Ball tracking / object detection
Sport: Table Tennis (Paralympics)
Recording conditions: Realistic professional match… See the full description on the dataset page: https://huggingface.co/datasets/AugustRushG123/TTA_Tracking.tracking_basket_4point_0425This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 37,
"total_frames": 35808,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:37"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/chen4803/tracking_basket_4point_0425.tracking_basket_0425This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/chen4803/tracking_basket_0425.hand_tracking_pv_carton_dual_view
hand_tracking_pv_carton_dual_view
The raw PressureVision carton recordings. 54 episodes, 13,369 frames, 10 Hz,
front and side RGB, six-joint state and action. LeRobot format v3.0.
This is the source the rest of the carton line derives from, which is why it
is here: hand_tracking_pv_carton_phase_b (30 reviewed episodes) and its
_train24 / _train24_onset_qreplace splits are all built from these
recordings by
build_phase_b_training_dataset.py.
A derived set can be rebuilt from this… See the full description on the dataset page: https://huggingface.co/datasets/stevenzenith/hand_tracking_pv_carton_dual_view.hand_tracking_pv_carton_middle_standard
hand_tracking_pv_carton_middle_standard
13 episodes, 2,902 frames, 10 Hz, front and side RGB, six-joint state and
action. LeRobot format v3.0.
Recorded with the middle-pose convention, and used together with
hand_tracking_pv_carton_phase_b
to fine-tune
act_carton_middle_labels_5k
from a 50k ACT base. The exact episode selection and the hand-made recovery
labels live in
training/train_middle_labeled_act.py:
episodes 0-7, 10 and 11 of this dataset, plus 1, 9, 14, 20 and 26 of… See the full description on the dataset page: https://huggingface.co/datasets/stevenzenith/hand_tracking_pv_carton_middle_standard.human-motion-tracking-deeplabcuttracking_dataset
Converted Robotics Dataset
This dataset stores authoritative timestamps and metadata in Parquet tables and
stores frame payloads in video containers. Video presentation timestamps are not
used for synchronization; use tables/frames.parquet and video_frame_idx.
Core loading rule:
decoded video frame k == frames.parquet row where video_frame_idx == k
Useful files:
dataset_manifest.parquet: converted sessions in this dataset root
sessions/<session_id>/metadata/streams.parquet: stream… See the full description on the dataset page: https://huggingface.co/datasets/DAIRLab/tracking_dataset.hand_tracking_pick_placeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 10,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/stevenzenith/hand_tracking_pick_place.hand_tracking_pv_carton_phase_bThis dataset was created using LeRobot.
Dataset Description
Private Phase B training dataset for adaptive 250 g paper-carton grasping with an SO-101 arm. It contains 30 reviewed formal episodes at 10 Hz with front and side RGB video, six-joint observations, and six-joint actions. Audit-only PressureVision/teacher fields and constant grip-context dimensions are intentionally excluded. See source_episode_map.csv and phase_b_training_manifest.json for provenance and coverage… See the full description on the dataset page: https://huggingface.co/datasets/stevenzenith/hand_tracking_pv_carton_phase_b.bicycle-vibration-dataset-3d-point-trackingtraffic-tracking-sample-data
