datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
austin_sailor_dataset_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "franka",
"total_episodes": 240,
"total_frames": 353094,
"total_tasks": 4,
"total_videos": 480,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:240"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/austin_sailor_dataset_lerobot.austin_sailor_datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 240,
"total_frames": 353094,
"total_tasks": 4,
"total_videos": 480,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:240"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lerobot/austin_sailor_dataset.sailor2-sft-stage1austin_sailor_dataset_augmented
austin_sailor_dataset_augmented
Overview
Codebase version: v3.0
Robots: google_robot, images, jaco, kinova3, kuka_iiwa, sawyer, ur5e, widowX, xarm7
FPS: 20
Episodes: 240
Frames: 353,094
Splits:
train: 0:240
Data Layout
data_path : data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet
video_path: videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Features
Feature
dtype
shape
description… See the full description on the dataset page: https://huggingface.co/datasets/oxe-auge/austin_sailor_dataset_augmented.austin-sailor
Extra pre-processing*:
smooth zero-state samples (ee pose & joints, below 1e-5) by interpolating from neighbors with window size of 5
filter out first 10 frames where robot state (ee pose, joints) jumps from 0
mapped gripper value in action from [-1, 1] to [0, 1] where range is [close, open]. gripper value in state is distance in meters in range [0.0, 0.08]
Note: empirically chosen number 5 for smooting window and 10 for trimming leading frames are median, and there might be still… See the full description on the dataset page: https://huggingface.co/datasets/sanjar-normuradov-agile-robots/austin-sailor.OXE_austin_sailor_dataset_embeddingsLanguage Table (LeRobot) — Embedding-Only Release
(DINOv3 + SigLIP2 image features; EmbeddingGemma task-text features)
This repository packages a re-encoded variant of IPEC-COMMUNITY/austin_sailor_dataset_lerobot where raw videos are replaced by fixed-length image embeddings, and task strings are augmented with text embeddings. All indices, splits, and semantics remain consistent with the source dataset while storage and I/O are substantially lighter. To make the dataset practical to… See the full description on the dataset page: https://huggingface.co/datasets/saaduddinM/OXE_austin_sailor_dataset_embeddings.sea-wildbenchSEA-WildBench Code: GPT-4 evaluation for Chat Model
austin_sailor_dataset_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 240,
"total_frames": 353094,
"total_tasks": 4,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:240"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"… See the full description on the dataset page: https://huggingface.co/datasets/FedorX8/austin_sailor_dataset_lerobot.sailor2-sft-stage2old-sailor-2-ichigo-tokenssailor-instruct-merge-shortsailor-instruct-merge-longsailor2-sft-stage1-thaiSakalti__Sailor-japanese-details
Dataset Card for Evaluation run of Sakalti/Sailor-japanese
Dataset automatically created during the evaluation run of model Sakalti/Sailor-japanese
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Sakalti__Sailor-japanese-details.sailor2-instruct-stage1-vie
