datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
artem-fold-towelThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi1_gripper"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/artem-fold-towel.time-lapse-artifacts
Time-Lapse Artifacts
873 indexed video files document one artist's traditional drawing practice.
The recorded finish dates span September 17, 2024 through September 20, 2026;
nine Pre-Standard dates remain unknown. Standardized acquisition began July 13,
2025. The current indexes contain 2,196,054,134,482 indexed video bytes
(approximately 2.20 TB).
The recordings began as personal practice documentation and a durable record of
manual work. The archive was initially organized as… See the full description on the dataset page: https://huggingface.co/datasets/maxwellinked/time-lapse-artifacts.artem-pour-waterThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
14
],
"names": [
"umi1_x",
"umi1_y",
"umi1_z",
"umi1_rx",
"umi1_ry",
"umi1_rz",
"umi1_gripper"… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/artem-pour-water.veo3-video-prompts
Veo 3 Video Generation Dataset
English | Português do Brasil
English
Summary
A collection of AI-generated videos created with Google's Veo 3 family of models. Each record contains the original text prompt, the model variant used, the generated video, and (when applicable) the input reference image. Videos are organized into one configuration per model variant.
Videos: 5,811
Input images: 1,354
Configurations: 6
Language of prompts: multilingual… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/veo3-video-prompts.SoftVTBench
SoftVTBench
Visuo-tactile manipulation demonstrations for rigid and soft/deformable
LIBERO-style pick-and-place tasks, collected with a tactile-sensing Franka arm
in Isaac Lab (Tabero simulation stack).
Mirrored on both hubs:
Hugging Face — Arthur12137/SoftVTBench
ModelScope — Arthur12137/SoftVTBench
Earlier releases (evaluation USD assets and the first partial data drops) were
moved to Arthur12137/SoftVTBench-archive on both hubs.
Contents
Four subsets, each 10… See the full description on the dataset page: https://huggingface.co/datasets/Arthur12137/SoftVTBench.artem-fold-towel-filtered
Artem fold-towel filtered trajectories
Observation-only LeRobot v3 derivative of brandonyang/artem-fold-towel. It contains 781 demonstrations (1048134 frames) accepted by the continuous bimanual YAM replayability pipeline.
The 14-D observation.state contains the smoothed, trajectory-optimized YAM-achievable UMI1 pose, normalized UMI1 gripper, UMI2 pose, and normalized UMI2 gripper. The two original UMI videos, timestamps, frame cadence, and task are preserved; action is… See the full description on the dataset page: https://huggingface.co/datasets/brandonyang/artem-fold-towel-filtered.MUSDB18
MUSDB 2018 Dataset
The musdb18 consists of 150 songs of different styles along with the images of their constitutive objects.
musdb18 contains two folders, a folder with a training set: "train", composed of 100 songs, and a folder with a test set: "test", composed of 50 songs. Supervised approaches should be trained on the training set and tested on both sets.
All files from the musdb18 dataset are encoded in the Native Instruments stems format (.mp4). It is a multitrack format… See the full description on the dataset page: https://huggingface.co/datasets/artemisweb/MUSDB18.WorldRover-art_nouveau
WorldRover — art_nouveau
Art Nouveau mansion, indoor. 47 clips per view, 30.2 min each of panoramic and first-person video,
30 fps, with lossless per-frame depth, camera pose and action labels.
pano/<clip_id> and fp/<clip_id> share the same camera path: the first-person clip was rendered
from the panoramic clip's per-frame trajectory, so frame k of one is frame k of the other and
the two pose files match exactly.
Clips
47 per view (94 total)
Duration
30.2 min per… See the full description on the dataset page: https://huggingface.co/datasets/AlayaLab/WorldRover-art_nouveau.Artifact-Bench
Artifact-Bench
Paper | Github
Artifact-Bench is a comprehensive benchmark for evaluating whether Multimodal Large Language Models (MLLMs) can truly detect and reason about the artifacts of AI-generated videos. Instead of focusing only on semantic understanding, Artifact-Bench emphasizes artifact-aware realism perception and fine-grained video analysis across photorealistic, animated, and CG-style domains.
Tasks
Artifact-Bench defines three complementary tasks:
Task… See the full description on the dataset page: https://huggingface.co/datasets/DogNeverSleep/Artifact-Bench.SoftVTBench-archive
SoftVTBench — archive
Frozen snapshot of everything that lived in Arthur12137/SoftVTBench before the
2026-08-13 re-release: the evaluation USD assets, the soft-body assets, and the
first partial data drops.
This repo is not maintained. The current dataset is at
Arthur12137/SoftVTBench.
Contents: eval-assets/, soft-assets/, object-rigid/, object-soft/,
spatial-rigid/, spatial-soft/ (9319 files, 2.3 GB).
Original dataset card (kept verbatim)
SoftVTBench dataset… See the full description on the dataset page: https://huggingface.co/datasets/Arthur12137/SoftVTBench-archive.sushi_atelier_artifacts
sushi_atelier_artifacts
Images, videos and other artefacts needed to run the website
shape-sorting-so101-30fpsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
6
],
"names": [
"Rotation",
"Pitch",
"Elbow",
"Wrist_Pitch",
"Wrist_Roll",
"Jaw"
]
}… See the full description on the dataset page: https://huggingface.co/datasets/Artefacts/shape-sorting-so101-30fps.eval-artifactsrobusto-2
Dataset: Robusto-2
Paper Link on ArXiv: https://arxiv.org/abs/2606.20980
Description
This dataset contains 20 videos, which were specifically used in this paper. These videos were selected from a larger set of 200 dashcam videos recorded in various cities across Peru (Lima) and New York City (NYC), available as an extended dataset. They are split evenly by region — 10 from Lima/Peru and 10 from NYC — so model and human behavior can be compared across a familiar… See the full description on the dataset page: https://huggingface.co/datasets/Artificio/robusto-2.artem_screwdriver_100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "mcx",
"total_episodes": 10,
"total_frames": 11525,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/antwoor/artem_screwdriver_100.martial-arts-video-v1
Martial Arts Action Video Preview
This is a small preview of five short, staged-looking martial-arts action clips. The clips show guarded stances, upper-body strikes, close combat, footwork, falls/recovery, and apparent weapon exchanges.
Contents
videos/: five short MP4 clips
metadata.csv: one row per video with broad scene and action labels
segments.csv: coarse time windows with action descriptions
metadata.jsonl: the same metadata in JSON Lines format
frames/:… See the full description on the dataset page: https://huggingface.co/datasets/thordata/martial-arts-video-v1.time-lapse-artifacts-derived
Time-Lapse Artifacts: Derived Access Layer
This repository contains reproducible access derivatives for
maxwellinked/time-lapse-artifacts.
The archival masters remain in the source dataset and remain authoritative.
Related interfaces
Public browser (pinned snapshot)
Browser source and revision history
The browser is a static presentation layer and may lag the current Hugging Face
dataset. Hugging Face remains authoritative for media, record identities, and… See the full description on the dataset page: https://huggingface.co/datasets/maxwellinked/time-lapse-artifacts-derived.artem_screwdriver_50This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "mcx",
"total_episodes": 10,
"total_frames": 11525,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/antwoor/artem_screwdriver_50.ur5_lab_test_tube_camera_shiftsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5",
"total_episodes": 5,
"total_frames": 2153,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/arturaah/ur5_lab_test_tube_camera_shifts.mva-repro-artifactsltx2-optims-artifactsContains the artifacts for https://github.com/sayakpaul/ltx2-simple-optims.
Videos generated during the benchmarks are in videos.
Logs are available in results.
Claude's trace and memory are available in claude_stuff.
draw_pixel_artThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_follower",
"total_episodes": 50,
"total_frames": 26066,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Beegbrain/draw_pixel_art.wipe3
wipe3
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
ur5_place_and_pour_nuts_camera_shiftsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5",
"total_episodes": 118,
"total_frames": 97790,
"total_tasks": 1,
"total_videos": 708,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:118"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/arturaah/ur5_place_and_pour_nuts_camera_shifts.lab_test_tube_100episodes_deptheval_so100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 50,
"total_frames": 22605,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ArturFrost/eval_so100.shape-sorting-so101_split_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
6
],
"names": [
"Rotation",
"Pitch",
"Elbow",
"Wrist_Pitch",
"Wrist_Roll",
"Jaw"
]
}… See the full description on the dataset page: https://huggingface.co/datasets/Artefacts/shape-sorting-so101_split_test.so100_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 2,
"total_frames": 1171,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Artdelia/so100_test.art_screw_sasha_tape_5050This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 101,
"total_frames": 117139,
"total_tasks": 2,
"total_videos": 202,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:101"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/antwoor/art_screw_sasha_tape_5050.trial1
trial1
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
