datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.BridgeData-V2-Scripted-Images
BridgeData V2 Image Triplets Dataset
This dataset contains image triplets from BridgeData V2 trajectories in ImageFolder format.
Derived From
This dataset is a derivative of the 30 GB scripted subset of BridgeData V2 from RAIL-Berkeley. All rights and original licensing apply.
Dataset Structure
initial_images/: Contains first frame images (initial state)
intermediate_images/: Contains intermediate frame images (frame 38)
final_images/: Contains final frame… See the full description on the dataset page: https://huggingface.co/datasets/VyoJ/BridgeData-V2-Scripted-Images.scripted_atomic_train_frac_0.3_large_goal_annotationThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 664,
"total_frames": 116214,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:664"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_train_frac_0.3_large_goal_annotation.anv_data_ke_kikuyu_scriptedscripted_atomic_step_pose_0.6This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 955,
"total_frames": 159935,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:955"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_step_pose_0.6.scripted_atomic_step_train_frac0.3_large_imageThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 664,
"total_frames": 116214,
"total_tasks": 1,
"total_videos": 1328,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:664"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_step_train_frac0.3_large_image.scripted_atomic_step_train_frac0.3_large_blindThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 664,
"total_frames": 116214,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:664"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_step_train_frac0.3_large_blind.scripted_atomic_step_train_frac0.2_largeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "ur5_wsg50_lego_atomic_step",
"total_episodes": 512,
"total_frames": 107566,
"total_tasks": 1,
"total_videos": 1024,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:512"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/windfromthenorth/scripted_atomic_step_train_frac0.2_large.LGRM_BridgeDataV2_Scriptedxarm_scripted_1BridgeData-V2-Scripted-Videos
BridgeData V2 Video Dataset
This dataset contains videos from BridgeData V2 trajectories in VideoFolder format.
Derived From
This dataset is a derivative of the 30 GB scripted subset of BridgeData V2 from RAIL-Berkeley. All rights and original licensing apply.
Citation
If you use this dataset, please cite the original BridgeData V2 paper:
@inproceedings{walke2023bridgedata,
title={BridgeData V2: A Dataset for Robot Learning at Scale},
author={Walke… See the full description on the dataset page: https://huggingface.co/datasets/VyoJ/BridgeData-V2-Scripted-Videos.dynamic_robot_bench_dr_scripted_10k
dynamic_robot_bench_dr_scripted_10k
10,000 scripted-expert demonstrations across all 100 dynamic task families of
dynamic-robot-bench — a conveyor-belt dynamic-manipulation benchmark (Franka
Panda + wrist camera, ManiSkill 3 / SAPIEN GPU sim). One LeRobot v2.1 dataset:
100 episodes per family, success-filtered, language-prompted per episode.
Collection configuration (identical for every family)
Scripted expert with per-step auto-derived speed caps, recorded as… See the full description on the dataset page: https://huggingface.co/datasets/Damin3927/dynamic_robot_bench_dr_scripted_10k.openapps-scripted-navigation-3024ep
OpenApps Scripted Navigation (3,024 episodes, 20 routes)
Inter-app navigation trajectories in OpenApps, collected with a scripted
policy using only real UI actions (click "Return to List of Apps", click the
target app icon) — no goto() teleports, so transitions are learnable and
plannable from the action space alone.
20 routes: 5 source apps (todo, calendar, messages, codeeditor, map) × 4 targets
3,024 episodes after filtering (from 4,000 collected), 20 steps each
Episode… See the full description on the dataset page: https://huggingface.co/datasets/FruitPunchSamuraiG/openapps-scripted-navigation-3024ep.common-voice-scripted-speech-kab-26-huge
Common Voice Scripted Speech 26.0 - Kabyle (Huge, Cleaned)
Full cleaned dataset of Mozilla Common Voice 26.0 for Kabyle (Taqbaylit) ASR. No speaker cap, no splits — all validated, cleaned, GlotLID-filtered clips.
Source
Original: Mozilla Common Voice 26.0 (cv-corpus-26.0-2026-06-12)
Dataset ID: cmqim4fux00tynq07ljtyhzfh (Mozilla Data Collective)
License: CC0-1.0
Generated: 2026-07-12
Cleaning Pipeline
Quality filter: ≥2 upvotes, 0 downvotes… See the full description on the dataset page: https://huggingface.co/datasets/boffire/common-voice-scripted-speech-kab-26-huge.UR7e_Scripted_DrawerOpen_100epi_10fpsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur7e",
"total_episodes": 100,
"total_frames": 19315,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Cache-SCA/UR7e_Scripted_DrawerOpen_100epi_10fps.common-voice-scripted-speech-quebec
common-voice-scripted-speech-quebec
This dataset is a filtered subset of the Mozilla Common Voice Scripted Speech 25.0 - French dataset. It exclusively contains audio clips from speakers with Canadian and Québécois accents.
Dataset Summary
Language: French (fr)
Total Clips: 25,198
Total Duration: 36.83 hours (132,578.64 seconds)
License: CC-0
Filtering Criteria
This subset was generated by extracting rows from the original cv-corpus-25.0-2026-03-09 dataset… See the full description on the dataset page: https://huggingface.co/datasets/thomasgauthier/common-voice-scripted-speech-quebec.LGIE_BridgeData_V2_Scripted_V2_Centroidpiperx-scripted-stack-real2sim-v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 250,
"total_frames": 147831,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:250"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/axiboai/piperx-scripted-stack-real2sim-v1.UR7e_Scripted_DrawerOpen_TestThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur7e",
"total_episodes": 3,
"total_frames": 578,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:3"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Cache-SCA/UR7e_Scripted_DrawerOpen_Test.piper-stack-scripted-smolvla-v4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 200,
"total_frames": 125458,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/axiboai/piper-stack-scripted-smolvla-v4.ttt-scripted-smoke
OpenEnv rollouts
Collected with OpenEnv (openenv collect).
Episodes: 20
Run metadata:
key
value
env
openspiel:tic_tac_toe
env_base_url
https://sergiopaniego-openspiel-ttt-env.hf.space
provider
scripted
model
None
num_episodes_requested
20
temperature
0.2
keep_losses
False
Schema
Each line of results.jsonl is one episode:
episode_id (string)
messages (chat transcript; TRL SFTTrainer-compatible)
reward (float)
done (bool)
tool_trace (list of… See the full description on the dataset page: https://huggingface.co/datasets/sergiopaniego/ttt-scripted-smoke.cv-en-scripted-test-500
Common Voice English Scripted Test Set — 500 clips
n = 500 utterances · private eval set for ASR benchmarking
Source
Derived from Mozilla Common Voice Scripted Speech 25.0 — English (test split), downloaded via the Mozilla Data Collective API (dataset ID cmndapwry02jnmh07dyo46mot, 94 GB tarball).
Construction
Starting from the full CV 25.0 English test split (16,398 rows), a stratified 500-clip subset was produced using the same recipe as… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/cv-en-scripted-test-500.piper-stack-scripted-v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 400,
"total_frames": 218735,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:400"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/axiboai/piper-stack-scripted-v2.common-voice-scripted-speech-24.0-mongolianpiper-stack-scripted-pi0fast-v4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 200,
"total_frames": 125458,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/axiboai/piper-stack-scripted-pi0fast-v4.piper-stack-scripted-smolvla-v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": null,
"total_episodes": 200,
"total_frames": 125272,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:200"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/axiboai/piper-stack-scripted-smolvla-v3.anv_scripted_multilingual-1common-voice-scripted-speech-kab-26-tiny
Common Voice Scripted Speech 26.0 - Kabyle (Cleaned)
This is a cleaned, speaker-disjoint subset of Mozilla Common Voice 26.0 for Kabyle (Taqbaylit) ASR.
Source
Original: Mozilla Common Voice 26.0 (cv-corpus-26.0-2026-06-12)
Dataset ID: cmqim4fux00tynq07ljtyhzfh (Mozilla Data Collective)
License: CC0-1.0
Generated: 2026-07-12
Cleaning Pipeline
Step
Input
Output
Filter
Quality filter
609,940
573,073
≥2 upvotes, 0 downvotes
Character… See the full description on the dataset page: https://huggingface.co/datasets/boffire/common-voice-scripted-speech-kab-26-tiny.so-100-pick-place-scripted-v3piper-stack-scripted-pi0fast-v4-success
