azithromycin/williamdgomez_SNOW.SB
SNOW.SB TsFile This repository is an Apache TsFile conversion of the LeRobot dataset williamdgomez/SNOW.SB. The original Hugging Face repository and its commit history identify williamdgomez as the dataset author/uploader. The source dataset was created with LeRobot v3.0 for the lekiwi_client robot. Source Dataset Dataset: williamdgomez/SNOW.SB Author/uploader: williamdgomez License: Apache-2.0 Task: Pick up the block and place it in the bin Split: train (0:91)… See the full description on the dataset page: https://huggingface.co/datasets/azithromycin/williamdgomez_SNOW.SB.
SNOW.SB TsFile
This repository is an Apache TsFile conversion of the LeRobot dataset `williamdgomez/SNOW.SB`. The original Hugging Face repository and its commit history identify williamdgomez as the dataset author/uploader. The source dataset was created with LeRobot v3.0 for the lekiwi_client robot.
Source Dataset
- Dataset: `williamdgomez/SNOW.SB`
- Author/uploader: `williamdgomez`
- License: Apache-2.0
- Task: Pick up the block and place it in the bin
- Split:
train(0:91) - Sampling rate: 24 fps
- Scale: 91 episodes, 73,578 frames, 1 task
- Source data shards: 91 Parquet files under
data/chunk-000/file-*.parquet - Source video files: 273 MP4 files, 91 per camera stream
The source video streams are stored at `videos/`:
videos/observation.images.front/chunk-000/file-{file_index:03d}.mp4videos/observation.images.wrist/chunk-000/file-{file_index:03d}.mp4videos/observation.images.top/chunk-000/file-{file_index:03d}.mp4
Videos are not included in this TsFile repository. They remain available in the original Hugging Face dataset and can be aligned with rows using episode_index, frame_index, and the source meta/episodes offsets.
Converted Files
- TsFile:
data/williamdgomez_snow_sb.tsfile - Table:
williamdgomez_snow_sb - Rows: 73,578
- Devices: 91 (
episode_indexxtask_index) - Time precision: milliseconds
- Metadata:
meta/mirrors the source metadata;meta/info.jsonrecords the TsFile path and conversion mapping.
Schema
Time is round(timestamp * 1000) and is stored in milliseconds. The source timestamp is dropped after this conversion because it is exactly the seconds representation of Time; Time restarts at zero for each episode.
Vector names preserve their source prefixes (. becomes _) and append a zero-based element index. The three video features (observation.images.front, .wrist, .top) are omitted from the TsFile because TsFile stores the numeric time-series table while the original MP4 files remain at the source URL.
Encoding and Compression
The local conversion used the type-aware TsFile writer profile requested for this dataset:
- FLOAT/DOUBLE:
GORILLAwithLZ4 - INT32/INT64 and
Time:TS_2DIFFwithLZ4 - BOOLEAN (if present):
RLEwithLZ4 - TAG columns: TsFile table/device TAG mechanism
The generated TsFile is 1,456,157 bytes. The merged staged Parquet is 1,235,578 bytes, and the 91 source Parquet shards total 3,152,836 bytes.
Validation
Apache TsFile Python SDK readback succeeded. The TsFile metadata row count and query readback both equal the staged Parquet count: 73,578 rows. All 91 episode devices have monotonic Time values with no duplicate (episode_index, task_index, Time) keys.
Usage
from tsfile import TsFileReader
reader = TsFileReader("data/williamdgomez_snow_sb.tsfile")
table = reader.get_all_table_schemas()["williamdgomez_snow_sb"]
columns = [c.get_column_name() for c in table.get_columns() if c.get_column_name() != "Time"]
with reader.query_table("williamdgomez_snow_sb", columns, batch_size=65536) as result:
batch = result.read_arrow_batch()Citation
The original dataset card does not provide a paper or BibTeX citation.
