bucket
Datasets
All datasets matching “bucket”bucketexp-bucket
LeWM multi-domain robot corpus
Multi-domain robot-manipulation data for latent world-model training and dataset-replay
evaluation (reset an env to any stored frame, roll out a planner, score against the stored
goal). Seven domains, unified where it matters, documented where it can't be.
Format: Lance. Each dataset is one <name>.lance table,
one row per frame, episode_idx + step_idx addressing. Column-level detail for every task is
in probe_manifest.json at the repo root… See the full description on the dataset page: https://huggingface.co/datasets/mh-hf/exp-bucket.viptrossen-sim-put-cap-in-bucket
Robot Manipulation Dataset: Put Cucumber in Bucket
Dataset Description
This dataset contains demonstration episodes for a bimanual robotic manipulation task: putting a cucumber in a bucket. The data was collected in MuJoCo simulation using Trossen AI bimanual robotic arms.
Dataset Summary
The dataset consists of robot manipulation demonstrations collected in simulation. Each episode contains multi-view camera observations, joint positions/velocities… See the full description on the dataset page: https://huggingface.co/datasets/stonesstones/trossen-sim-put-cap-in-bucket.emolia-thinking-balanced-buckets
Emolia-Thinking — Balanced Per-Dimension Bucket Subset
A balanced, per-dimension bucket subset of
VoiceNet/emolia-thinking,
derived from that dataset's zero-shot VoiceNet-dimension labels.
For every VoiceNet voice/prosody/timbre/style dimension, this subset draws a
roughly equal number of clips from each ordinal bucket (0–6), so that
downstream training / probing sees a balanced distribution along each axis
instead of the strongly skewed natural distribution.
How… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-thinking-balanced-buckets.cc12m-1mp-plus-realistic-bucketed-1024
CC12M 1MP+ Realistic Bucketed 1024
This dataset is a self-contained, recaptioned export of the train split from
opendiffusionai/cc12m-1mp_plus-realistic.
The upstream dataset provides metadata and image URLs; this release contains the
downloaded image bytes, so training does not require fetching images from the
original URLs.
It contains 573,052 image-caption pairs packaged as aspect-ratio-bucketed TAR
shards for text-to-image training. Images are assigned to buckets targeting a… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/cc12m-1mp-plus-realistic-bucketed-1024.
