datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Cobot_Magic_make_hamburger
Cobot_Magic_make_hamburger
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
restaurant
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_make_hamburger.agent-task-recursive-task-synthesis
Recursive-Task-Synthesis for tmax
Images require building: the complete dataset and build contexts are included. Image builds are deferred; run the resumable script below before using these environments.
All 37,484 task directories from Zhongzhi1228/Recursive-Task-Synthesis, pinned to be44f96808d5a9b599d5cb024341ff00091adeb7, converted to tmax's swerl_vanillux_sandbox format.
The train split uses the same messages, ground_truth, dataset, env_config, and source schema as the… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/agent-task-recursive-task-synthesis.agent-task-calibforge
CalibForge for tmax
Images available: this conversion reuses verified public upstream images pinned by digest.
All 5,431 task directories from AweAI-Team/CalibForge, pinned to fb1e75441a94b8bb0ced08acd6b59e711704d70a, converted to tmax's swerl_vanillux_sandbox format.
The train split uses the same messages, ground_truth, dataset, env_config, and source schema as the other hamishivi/agent-task-* datasets. Messages use the tmax Vanillux templates; dataset is passthrough. Task IDs… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/agent-task-calibforge.agent-task-litecoder-terminal-rl-preview
LiteCoder-Terminal-RL-preview for tmax
All 602 task directories from Lite-Coder/LiteCoder-Terminal-RL-preview, pinned to 6fe7e994ff12d678de9b803da5c9907c8394a89c, converted to tmax's swerl_vanillux_sandbox format.
The train split uses the same messages, ground_truth, dataset, env_config, and source schema as the other hamishivi/agent-task-* datasets. Messages use the tmax Vanillux templates; dataset is passthrough. Task IDs are prefixed with litecoder_terminal_rl__ to avoid… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/agent-task-litecoder-terminal-rl-preview.PMC-VQA-1
PMC-VQA-1
This dataset is a streaming-friendly version of the PMC-VQA dataset, specifically containing the "Compounded Images" version (version-1). It is designed to facilitate efficient training and evaluation of Visual Question Answering (VQA) models in the medical domain, straight from the repository
Dataset Description
The original PMC-VQA dataset, available at https://huggingface.co/datasets/xmcmic/PMC-VQA, comprises Visual Question Answering pairs derived from… See the full description on the dataset page: https://huggingface.co/datasets/hamzamooraj99/PMC-VQA-1.RoboTwin_beat_block_hammer_randomizedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "aloha",
"total_episodes": 500,
"total_frames": 54582,
"total_tasks": 365,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/suz22/RoboTwin_beat_block_hammer_randomized.agent-task-facet-terminal-6k
FACET-Terminal-Tasks-6k for tmax
Images require building: the complete dataset and build contexts are included. Image builds are deferred; run the resumable script below before using these environments.
All 6,020 task directories from FACET-Terminal/FACET-Terminal-Tasks-6k, pinned to b2d02645932e3989c8332a41e57d0f0855de7002, converted to tmax's swerl_vanillux_sandbox format.
The train split uses the same messages, ground_truth, dataset, env_config, and source schema as the other… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/agent-task-facet-terminal-6k.DAPO-Math-17k-Processed_filteredvisual_distracting_metaworld_with_masksagent-task-terminal-lego-15k
Terminal-Lego-15k for tmax
Images require building: the dataset conversion and complete task archive are published. 47 task image(s) were published during preparation; the remaining images are intentionally left for the included resumable build script. No build job is running for this release. Only tasks with available images can run immediately. See image-build-status.json for the publication-time snapshot.
All 15,048 task directories from Lego-X/Terminal-Lego-15k, pinned to… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/agent-task-terminal-lego-15k.INDiC-BPCC-hq
Description
General and Technical Domain Data with LaBSE Score >= 0.8 (Bidirectional only with Indic Languages). Includes Samanantar by AI4Bharat and NPTEL Data by Speech Lab.
HAM10000
Dataset Card for "HAM10000"
More Information needed
hamela_books_text_full_okswerl-tmax-15kvisual_distracting_control_suite
Visual Distracting Control Suite Benchmark
This dataset contains expert trajectories generated by a Proximal Policy Optimization (PPO) reinforcement learning agent trained on 4 environments of the Distracting Control Suite. For each environment we collect data with different levels of distraction, which we define below, and masks for the agent.
Levels of distraction:
None: Vanilla DeepMind Control Suite without visual distractions. The environment uses the default static background… See the full description on the dataset page: https://huggingface.co/datasets/hamza-adnan/visual_distracting_control_suite.aihub_math_donutFrench-PD-diverse43,085,129,931 words
dcs_mujoco_with_masksagent-task-seta-env
SETA-Env for tmax
Images require building: the dataset conversion and complete task archive are published. 0 task image(s) were published during preparation; the remaining images are intentionally left for the included resumable build script. No build job is running for this release. Only tasks with available images can run immediately. See image-build-status.json for the publication-time snapshot.
All 4,567 task directories from camel-ai/SETA-Env, pinned to… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/agent-task-seta-env.captcha-dataOpenThoughts2-1Mspeech-accent-archive-v2
Dataset Card for "speech-accent-archive-v2"
More Information needed
Recorrected_Classification_Data_filtered_traindataset_ashSS_tieragent-task-termigenso100_bi_pl1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100_bimanual",
"total_episodes": 9,
"total_frames": 3147,
"total_tasks": 1,
"total_videos": 27,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:9"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hamidkaloorazi/so100_bi_pl1.GPQA-train-RLVRagent-task-combinedtmax-sft-full-20260403
