datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bc_z_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "google_robot",
"total_episodes": 39350,
"total_frames": 5471693,
"total_tasks": 104,
"total_videos": 39350,
"total_chunks": 40,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:39350"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/bc_z_lerobot.bc_z_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "google_robot",
"total_episodes": 39350,
"total_frames": 5471693,
"total_tasks": 104,
"total_videos": 39350,
"total_chunks": 40,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:39350"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ygtxr1997/bc_z_lerobot.thinking_bc_z_lerobot_output_qwen3vlbc_z_rawLOKIThis repository contains the data of the paper LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models.
UrBenchbc_z-bootstapir_checkpoint_v2MolmoAct2-BC-Z-Dataset
MolmoAct2-BC-Z Dataset
This dataset was created using LeRobot.
Language Annotations
This dataset includes annotated language instructions in meta/tasks_annotated.parquet. The file is indexed by episode_index and has a task column containing our per-episode annotated instruction.
The standard LeRobot loader resolves a frame's language instruction through task_index: each data row stores a task_index, which is looked up in meta/tasks.parquet. When you use these… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MolmoAct2-BC-Z-Dataset.bc_z_interleaveConverted bc_z for training Interleave-VLA.
Please refer to link for more details on this RLDS dataset format and conversion.
bc_zThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "google_robot",
"total_episodes": 39350,
"total_frames": 5471693,
"total_tasks": 104,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:39350"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/typoverflow/bc_z.bc_z_lerobotOXE_bc_z_embeddingsLanguage Table (LeRobot) — Embedding-Only Release
(DINOv3 + SigLIP2 image features; EmbeddingGemma task-text features)
This repository packages a re-encoded variant of IPEC-COMMUNITY/bc_z_lerobot where raw videos are replaced by fixed-length image embeddings, and task strings are augmented with text embeddings. All indices, splits, and semantics remain consistent with the source dataset while storage and I/O are substantially lighter. To make the dataset practical to upload/download and stream… See the full description on the dataset page: https://huggingface.co/datasets/saaduddinM/OXE_bc_z_embeddings.bc_z_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "google_robot",
"total_episodes": 39350,
"total_frames": 5471693,
"total_tasks": 104,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:39350"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/FedorX8/bc_z_lerobot.bc_z_lerobot_v30This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "google_robot",
"total_episodes": 39350,
"total_frames": 5471693,
"total_tasks": 104,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:39350"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/tailong-wu/bc_z_lerobot_v30.SemanticVLA-TraceX-240K-BC-Z
SemanticVLA TraceX 240K · BC-Z
🎉 Accepted to CVPR 2026.
✍️ Fei Ni¹, Zhuo Chen², Yifu Yuan³, Zibin Dong³, Xianze Yao³, Shan Luo², Jianye Hao³, Jiankang Deng¹†, Stefanos Zafeiriou¹†
🏫 ¹Imperial College London ²King's College London ³Tianjin University
✉️ Primary contact: f.ni@imperial.ac.uk
The BC-Z component of TraceX-240K — the trace-annotated trajectory corpus introduced in SemanticVLA. This package is a LeRobot v3.0 repack of BC-Z · Open-X-Embodiment BC-Z v0.1.0 with… See the full description on the dataset page: https://huggingface.co/datasets/spikefly/SemanticVLA-TraceX-240K-BC-Z.drifting-vla-v2-bc_zdysarthric_german
Dataset Card for "dysarthric_german"
More Information needed
bc_zrewards_bc_z_1_perCityBench-v0.3CityVQA-v0.2CityBench-v0.2CityBench-SubTasksSyntheticBench-Videosocto_bc_zKate1
