datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pii-masking-openpii-1.5m
OpenPII 1.5M: Multilingual PII Masking Dataset (Asia Pacific Extension)
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Overview
The OpenPII 1.5M dataset extends OpenPII 1M
with a new Asia Pacific corpus, bringing global coverage to 30 languages
across Europe, Americas, and Asia Pacific.
This is the flagship release of the PII-Masking-3M family, the world's
largest open multilingual PII masking corpus. Built to advance open… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1.5m.VLA_Arena_L0_L_lerobot_openpi
VLA-Arena Dataset (L0 - Large Variant)
About VLA-Arena
VLA-Arena is an open-source benchmark designed for the systematic evaluation of Vision-Language-Action (VLA) models. It provides a complete and unified toolchain covering scene modeling, demonstration collection, model training, and evaluation. Featuring 150+ tasks across 11 specialized suites, VLA-Arena assesses models through hierarchical difficulty levels (L0-L2) to ensure comprehensive metrics for safety… See the full description on the dataset page: https://huggingface.co/datasets/VLA-Arena/VLA_Arena_L0_L_lerobot_openpi.libero_gen_goal_chain_train_openpipii-masking-openpii-1m
OpenPII 1M — Multilingual PII Masking Dataset
Overview
The OpenPII 1M dataset is a large-scale, multilingual collection of 1,428,143 synthetic text examples with fine-grained PII (Personally Identifiable Information) annotations, spanning 23 European languages and 19 entity types.
Built to advance open research in privacy-preserving NLP, this dataset enables the development and benchmarking of Named Entity Recognition (NER) models, token classification pipelines… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1m.openpi_datasets
openpi_datasets
version
folder
data
v0.1.0
put_the_box
image_higt,image_left,state,action
v0.2.0
pi0_put_the_yellow_mango_on_the_blue_plate
image_higt,image_left,state,action
openpi
Real-World Franka Panda Task Data (224px Square Center Crop)
Tasks
Task
Description
Train Demos
Val Demos
task_1
put the pink noodle in the bowl
90
10
task_2
put the bread in the bowl
90
10
task_3
pour the pasta into the pan
89
10
task_4
close the right cabinet door
90
10
Data Format
Each demo is an HDF5 file with:
root/
actions: (T, 10) float64 — [pos(3), rot6d(6), gripper(1)]
agentview/video: (1, T, 3… See the full description on the dataset page: https://huggingface.co/datasets/zhicao/openpi.open-pii-masking-500k-ai4privacy
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy.Isaaclab-so101_11task_openpi_v21libero_gen_spatial_combination_train_openpihacker-news
Hacker News posts and comments
This is a dataset of all HN posts and comments, current as of November 1, 2023.
lego_assemblies_openpi_trainVLA_Arena_L1_L_lerobot_openpi
VLA-Arena Dataset (L1 - Large Variant)
About VLA-Arena
VLA-Arena is an open-source benchmark designed for the systematic evaluation of Vision-Language-Action (VLA) models. It provides a complete and unified toolchain covering scene modeling, demonstration collection, model training, and evaluation. Featuring 150+ tasks across 11 specialized suites, VLA-Arena assesses models through hierarchical difficulty levels (L0-L2) to ensure comprehensive metrics for safety… See the full description on the dataset page: https://huggingface.co/datasets/VLA-Arena/VLA_Arena_L1_L_lerobot_openpi.pour_foam_25_08_28_openpi_lerobotVLA_Arena_L0_M_lerobot_openpi
VLA-Arena Dataset (L0 - Medium Variant)
About VLA-Arena
VLA-Arena is an open-source benchmark designed for the systematic evaluation of Vision-Language-Action (VLA) models. It provides a complete and unified toolchain covering scene modeling, demonstration collection, model training, and evaluation. Featuring 150+ tasks across 11 specialized suites, VLA-Arena assesses models through hierarchical difficulty levels (L0-L2) to ensure comprehensive metrics for safety… See the full description on the dataset page: https://huggingface.co/datasets/VLA-Arena/VLA_Arena_L0_M_lerobot_openpi.libero_gen_goal_chain_no_secondstep_train_openpiThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 6767,
"total_frames": 1885353,
"total_tasks": 165,
"total_videos": 0,
"total_chunks": 7,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:6767"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/austinpatel/libero_gen_goal_chain_no_secondstep_train_openpi.privacy-filter-openpii-masking
1. Overview
privacy-filter-openpii-masking is a Korean and English entity-detection dataset for fine-tuning token-classification models. It is derived from ai4privacy/pii-masking-openpii-1.5m, relabeled to a 29-label taxonomy, and supplemented with statically authored or contextualized financial, customer-service/VOC, security, identity, and infrastructure scenarios.
The dataset provides entity annotations rather than application-specific redaction output. masked_text replaces… See the full description on the dataset page: https://huggingface.co/datasets/BCCard/privacy-filter-openpii-masking.franka-record-red-block-0730-openpiThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 51,
"total_frames": 59280,
"total_tasks": 1,
"total_videos": 102,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/zjushining/franka-record-red-block-0730-openpi.VLA_Arena_L0_S_lerobot_openpi
VLA-Arena Dataset (L0 - Small Variant)
About VLA-Arena
VLA-Arena is an open-source benchmark designed for the systematic evaluation of Vision-Language-Action (VLA) models. It provides a complete and unified toolchain covering scene modeling, demonstration collection, model training, and evaluation. Featuring 150+ tasks across 11 specialized suites, VLA-Arena assesses models through hierarchical difficulty levels (L0-L2) to ensure comprehensive metrics for safety… See the full description on the dataset page: https://huggingface.co/datasets/VLA-Arena/VLA_Arena_L0_S_lerobot_openpi.openpi-libero-lerobotThis dataset was created using LeRobot.
Dataset Description
This dataset combines four individual Libero datasets: Libero-Spatial, Libero-Object, Libero-Goal and Libero-10.
All datasets were taken from here and converted into LeRobot format.
Homepage: https://libero-project.github.io
Paper: https://arxiv.org/abs/2306.03310
License: CC-BY 4.0
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "panda",
"total_episodes": 1693… See the full description on the dataset page: https://huggingface.co/datasets/XinY0201/openpi-libero-lerobot.pour_foam_25_08_27_openpi_lerobotput_can_on_shelf_25_08_25_openpi_lerobotopenpi_sim_pick_place
SO-101 Pick and Place Dataset (OpenPi Format)
This dataset contains 40 episodes of a simulated SO-101 robot performing pick-and-place tasks, converted to OpenPi/RLDS format for use with Physical Intelligence's Pi0/Pi0.5 models.
Source
Converted from LeRobot dataset: danbhf/sim_pick_place_merged_40ep
Format
Each episode is stored as an NPZ file containing:
Key
Shape
Type
Description
observation/state
(N, 6)
float32
Joint positions (6 DoF)… See the full description on the dataset page: https://huggingface.co/datasets/danbhf/openpi_sim_pick_place.sortletters_parallel_121_v21_openpi
sortletters_parallel_121_v21_openpi
STOP — rejected for behavior training. This public revision is preserved only as a
reproducible engineering fixture for loader, padding-mask, and training-throughput tests.
It is not a valid real-robot policy-training dataset and must not be used to claim
sortletters behavior quality or real-robot success rate.
警告:该数据已被行为训练资格门拒绝。 仅允许用于 loader、mask 和训练吞吐工程验证;
不得用于声称有效真机策略训练、字母排序行为质量或真机成功率。
Qualification evidence
Independent… See the full description on the dataset page: https://huggingface.co/datasets/Shiki42/sortletters_parallel_121_v21_openpi.sanitizer-pick-place-openpi-eef-v21
Tuesday-Lab/sanitizer-pick-place-openpi-eef-v21
LeRobot-format robot dataset uploaded from a local dataset root.
Metadata
codebase_version: v2.1
fps: 50
robot_type: tuesday_arm
Load Example
from lerobot.datasets.lerobot_dataset import LeRobotDataset
dataset = LeRobotDataset(
repo_id="Tuesday-Lab/sanitizer-pick-place-openpi-eef-v21",
root=None, # Cache path; keep None to use default HF cache
video_backend="pyav",
)
print(len(dataset))
openpi-interpretability-data
openpi-interpretability-data
Interpretability artifacts (activations, conceptors, linear steering vectors, sparse autoencoder vectors and checkpoints) extracted from open vision-language-action (VLA) policy models on the LIBERO, MetaWorld, and RoboCasa benchmarks.
This dataset accompanies an anonymous submission and is shared for double-blind peer review.
Models and benchmarks
Model
Family
Benchmarks
pi0_5 (pi05)
π-series VLA
LIBERO
pi0_fast (pi0fast)… See the full description on the dataset page: https://huggingface.co/datasets/NeurIPsMay1234/openpi-interpretability-data.Open-Pixel-1T
🌌 Open-Pixel-1T (Visual Atlas)
A Large-Scale, High-Entropy Synthetic Image Dataset for Foundational Pre-Training
📑 Dataset Summary
Open-Pixel-1T is a monumental open-source initiative designed to create a "Visual Atlas" of stochastic imagery. Unlike traditional datasets scraped from social media which contain inherent human bias, Open-Pixel-1T is constructed using high-entropy random seeds to generate unique, diverse visual signals.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/LAYEK-143/Open-Pixel-1T.openpi_PiPER_demo_train_3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "piper",
"total_episodes": 98,
"total_frames": 37649,
"total_tasks": 1,
"total_videos": 294,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 200,
"splits": {
"train": "0:98"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aromanus/openpi_PiPER_demo_train_3.pickup-carrot-openpi-20251029This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "bi_widowxai_follower",
"total_episodes": 28,
"total_frames": 13130,
"total_tasks": 1,
"total_videos": 112,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:28"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/argus-systems/pickup-carrot-openpi-20251029.franka-insert-marker-v2-openpiThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 50,
"total_frames": 9828,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ankile/franka-insert-marker-v2-openpi.sam_openpi_solder2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 5,
"total_frames": 1495,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/1g0rrr/sam_openpi_solder2.
