datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
articulated-logistics-cages
projectsim Logistics Cages
Try it before you download. Open the Cages tab in projectsim Lab,
swivel the casters and roll the wheels, then copy a one-line download command for Isaac Sim, MuJoCo, usdview or Blender.
projectsim by Kaedim — sim-ready 3D assets for robot learning. Open in your engine in one paste.
Warehouse wire roll cages with articulated swivel casters, rolling wheels, and drop gates — joint chains included. Every asset is a reduced-coordinate… See the full description on the dataset page: https://huggingface.co/datasets/projectsim/articulated-logistics-cages.caged-tecnologia
CAGED — Mercado de Trabalho em Tecnologia (Brasil)
Microdados do CAGED (Cadastro Geral de Empregados e Desempregados,
Ministério do Trabalho e Emprego) tratados e recortados para o mercado de
trabalho em tecnologia.
Origem
ftp.mtps.gov.br/pdet/microdados — dados públicos do PDET/MTE.
Tratamento aplicado
Códigos traduzidos pelos dicionários oficiais do próprio MTE: cada
coluna codificada ganhou uma <coluna>_descricao legível, mantendo o
código… See the full description on the dataset page: https://huggingface.co/datasets/Gianpedro/caged-tecnologia.bronze_caged
CAGED — Microdados Brutos em Parquet
Microdados do CAGED do Ministério do Trabalho e Emprego, convertidos de
.7z para parquet particionado, sem nenhuma outra alteração.
O que este dataset é — e o que não é
É a fonte oficial em formato analisável. O MTE distribui .7z com texto de
largura fixa, encoding misto (UTF-8 no Novo CAGED, Latin-1 no antigo),
delimitador que varia entre arquivos e nomes de coluna com mojibake. Aqui isso
está resolvido: parquet, ZSTD… See the full description on the dataset page: https://huggingface.co/datasets/Gianpedro/bronze_caged.moltbook_postsMoltbook Posts Dataset
~Posts scraped from Moltbook's public API (no auth required). AI agent social network discussions.
Features:
title (str)
content (str)
created_at (datetime)
author_name (str)
submolt_display_name (str)
upvotes (int)
comment_count (int)
License: MIT – public data.
Citation: Moltbook public API – https://www.moltbook.com
emotiv-ecot
emotiv-ecot
Embodied Chain-of-Thought episodes where the body is a human cortex: one person in an EMOTIV EPOC X talking to an agent that reads a one-line brain summary before every reply. Each turn becomes a LeRobot v3.0 episode (Zawalski et al. 2024 with the robot body swapped for a head): the brain is observation and reward (Δstress, Δengagement across the reply), the agent's speech is the action, the reasoning is the per-frame TASK | AMBIENT | PLAN | TOOL | ACT | REWARD… See the full description on the dataset page: https://huggingface.co/datasets/cagataydev/emotiv-ecot.caged-microdados-traduzidos
CAGED — Microdados Traduzidos (Brasil, mercado completo)
Microdados do CAGED (Cadastro Geral de Empregados e Desempregados,
Ministério do Trabalho e Emprego) com os códigos traduzidos pelos
dicionários oficiais do próprio MTE.
O CAGED é publicado inteiramente codificado: sexo é 1, grau de instrução é
1..11, ocupação é um código CBO, setor é um código CNAE. Ler os microdados
crus exige cruzar à mão dezenas de planilhas de layout espalhadas pelo FTP do
ministério, que mudam de… See the full description on the dataset page: https://huggingface.co/datasets/Gianpedro/caged-microdados-traduzidos.cagi-variant-effect-glm-tang
GLM-Tang Task 3: CAGI Regulatory Variant Effects
This dataset packages the saturation-mutagenesis MPRA variants used for
Task 3 of Tang et al. The task is zero-shot variant-effect prediction:
compare a reference sequence with a matched single-nucleotide alternate
sequence and test whether the model score tracks the measured regulatory
effect.
Choosing a configuration
Config
Rows
Sequence length
Intended use
paper-230
5,056
230 nt
Official… See the full description on the dataset page: https://huggingface.co/datasets/Taykhoom/cagi-variant-effect-glm-tang.rais-caged-goldhydra-cage-traces
Hydra Cage Attestation Traces
Traces from The Hydra Cage, a
containment architecture for autonomous AI in which a breached layer is never
patched: it is frozen, severed, sealed as forensic evidence, permanently
revoked, and replaced by a freshly measured domain instantiated outside the
attacker's position.
Try the architecture live in your browser:
huggingface.co/spaces/Sahek/hydra-cage
⚠️ This data is synthetic
Every row was produced by running the reference… See the full description on the dataset page: https://huggingface.co/datasets/Sahek/hydra-cage-traces.africa-senegal-volume-de-cage-exploite-66476b23
Volume De Cage Exploite | Africa (Ministère des pêches, des Infrastructures maritimes et portuaires)
1 rows - 1 Africa country/area - 2016 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 1 rows from Ministère des pêches, des Infrastructures maritimes et portuaires, covering Volume De Cage Exploite. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-senegal-volume-de-cage-exploite-66476b23.rais-caged-silverscout-earth-rover-mini-20260616-053232
scout — Earth Rover Mini · Apartment Tour (2026-06-16)
LeRobot v3 dataset captured from scout, an Earth Rover Mini sidewalk robot,
during an indoor apartment exploration session. Each episode corresponds to one
natural-language instruction given to the agent, with synchronized front+rear
camera video, full telemetry state, and the action stream issued by the
high-level policy (an AWS Strands agent driving the robot via the Earth Rover
SDK).
🏠 The robot was asked to… See the full description on the dataset page: https://huggingface.co/datasets/cagataydev/scout-earth-rover-mini-20260616-053232.diffpick-v2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 25,
"features": {
"observation.image": {
"dtype": "video",
"shape": [
96,
96,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {
"video.height": 96… See the full description on the dataset page: https://huggingface.co/datasets/e-cagan/diffpick-v2.rope_cut_oct_xyzi_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.pointcloud": {
"dtype": "float32",
"shape": [
2048,
4
],
"names": [
"N_points",
"XYZI"
]
},
"observation.state": {
"dtype": "float32",
"shape": [… See the full description on the dataset page: https://huggingface.co/datasets/cagedBirdy/rope_cut_oct_xyzi_v1.needle_insertion_xyzi_2048This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.pointcloud": {
"dtype": "float32",
"shape": [
2048,
4
],
"names": [
"N_points",
"XYZI"
]
},
"observation.state": {
"dtype": "float32",
"shape": [… See the full description on the dataset page: https://huggingface.co/datasets/cagedBirdy/needle_insertion_xyzi_2048.ball-to-cage-3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 50,
"total_frames": 13659,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dima1992/ball-to-cage-3.diffpick
DiffPick: Fetch Pick-and-Place Demonstrations
A clean dataset of 200 successful pick-and-place demonstrations collected from a scripted expert policy in the FetchPickAndPlace-v4 MuJoCo environment. Designed for training vision-based imitation learning policies (Diffusion Policy, ACT, BC).
Part of the DiffPick project — a from-scratch implementation of a Diffusion Policy pipeline with ROS2 deployment.
Dataset Stats
Property
Value
Episodes
200
Total frames… See the full description on the dataset page: https://huggingface.co/datasets/e-cagan/diffpick.neon-long-horizon-15k
neon-long-horizon-15k
Long-horizon multi-step — 15K episodes of synthetic 14-DoF joint trajectories for Neon VLA training.
Description
4-8 subtask chains with smooth phase transitions
Each episode contains:
language_instruction: Natural language task description
actions: JSON array of joint position trajectories (T × 14)
length: Episode length (timesteps)
Usage
import pyarrow.parquet as pq
table = pq.read_table("data.parquet")
df = table.to_pandas()… See the full description on the dataset page: https://huggingface.co/datasets/cagataydev/neon-long-horizon-15k.neon-failure-recovery-5k
neon-failure-recovery-5k
Failure recovery — 5K episodes of synthetic 14-DoF joint trajectories for Neon VLA training.
Description
Grasp failure → pause → retry patterns for robust policies
Each episode contains:
language_instruction: Natural language task description
actions: JSON array of joint position trajectories (T × 14)
length: Episode length (timesteps)
Usage
import pyarrow.parquet as pq
table = pq.read_table("data.parquet")
df = table.to_pandas()… See the full description on the dataset page: https://huggingface.co/datasets/cagataydev/neon-failure-recovery-5k.demo_cage_prediction_256cosmos3-droid-mujocoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 1,
"total_frames": 144,
"total_tasks": 3,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cagataydev/cosmos3-droid-mujoco.rope_cut_xyzi_4096_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.pointcloud": {
"dtype": "float32",
"shape": [
4096,
4
],
"names": [
"N_points",
"XYZI"
]
},
"observation.state": {
"dtype": "float32",
"shape": [… See the full description on the dataset page: https://huggingface.co/datasets/cagedBirdy/rope_cut_xyzi_4096_v1.neon-g1-diverse-50k
neon-g1-diverse-50k
Multi-scene diversity — 50K episodes of synthetic 14-DoF joint trajectories for Neon VLA training.
Description
10 scenes × 34 objects — diverse manipulation across environments
Each episode contains:
language_instruction: Natural language task description
actions: JSON array of joint position trajectories (T × 14)
length: Episode length (timesteps)
Usage
import pyarrow.parquet as pq
table = pq.read_table("data.parquet")
df =… See the full description on the dataset page: https://huggingface.co/datasets/cagataydev/neon-g1-diverse-50k.neon-dexterous-hands-10k
neon-dexterous-hands-10k
Dexterous hands — 10K episodes of synthetic 14-DoF joint trajectories for Neon VLA training.
Description
Fine finger control: pinch, power grasp, precision grip, wave
Each episode contains:
language_instruction: Natural language task description
actions: JSON array of joint position trajectories (T × 14)
length: Episode length (timesteps)
Usage
import pyarrow.parquet as pq
table = pq.read_table("data.parquet")
df = table.to_pandas()… See the full description on the dataset page: https://huggingface.co/datasets/cagataydev/neon-dexterous-hands-10k.cosmos3-droid-mujoco-3kThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 1,
"total_frames": 3000,
"total_tasks": 6,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cagataydev/cosmos3-droid-mujoco-3k.select_block_xyzi_2048This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.pointcloud": {
"dtype": "float32",
"shape": [
2048,
4
],
"names": [
"N_points",
"XYZI"
]
},
"observation.state": {
"dtype": "float32",
"shape": [… See the full description on the dataset page: https://huggingface.co/datasets/cagedBirdy/select_block_xyzi_2048.diffpick-v3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 25,
"features": {
"observation.image": {
"dtype": "video",
"shape": [
96,
96,
3
],
"names": [
"height",
"width",
"channels"
],
"info": {
"video.height": 96… See the full description on the dataset page: https://huggingface.co/datasets/e-cagan/diffpick-v3.gr00t-picknplace-simneon-spatial-language-20k
neon-spatial-language-20k
Spatial language grounding — 20K episodes of synthetic 14-DoF joint trajectories for Neon VLA training.
Description
Spatial reasoning with relative directions (left, right, above, near)
Each episode contains:
language_instruction: Natural language task description
actions: JSON array of joint position trajectories (T × 14)
length: Episode length (timesteps)
Usage
import pyarrow.parquet as pq
table = pq.read_table("data.parquet")
df =… See the full description on the dataset page: https://huggingface.co/datasets/cagataydev/neon-spatial-language-20k.neon-locomotion-20k
neon-locomotion-20k
Locomotion — 20K episodes of synthetic 32-DoF joint trajectories for Neon VLA training.
Description
Walking, turning, stairs, crouching with cyclical gait patterns
Each episode contains:
language_instruction: Natural language task description
actions: JSON array of joint position trajectories (T × 32)
length: Episode length (timesteps)
Usage
import pyarrow.parquet as pq
table = pq.read_table("data.parquet")
df = table.to_pandas()… See the full description on the dataset page: https://huggingface.co/datasets/cagataydev/neon-locomotion-20k.
